Device having a data processing engine array

The integration of a data processing engine array with shared memory and interconnects within programmable ICs addresses inefficiencies in existing ICs, achieving efficient and low-power data processing for various applications.

JP7701788B2Active Publication Date: 2025-07-02XILINX INC
View PDF 6 Cites 0 Cited by

Patent Information

Application Number
JP2020554171
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Priority Date
2018-04-03
Filing Date
2019-04-01
Publication Date
2025-07-02
Estimated Expiration
2039-04-01

AI Technical Summary

Technical Problem

Existing programmable integrated circuits (ICs) face challenges in efficiently integrating data processing engines and memory modules within a programmable circuitry framework, leading to suboptimal performance in terms of power consumption, area usage, and data throughput.

Method used

The integration of a data processing engine array (DPE) with a memory module accessible by multiple cores, both within the same DPE and across neighboring DPEs, facilitated by a DPE interconnect that supports communication and configuration, along with a system-on-chip interface block for data exchange.

Benefits of technology

This configuration enables efficient data processing with reduced power consumption and area usage, while achieving predictable data throughput and latency, suitable for applications like wireless radio, 5G, and machine learning operations.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007701788000001
    Figure 0007701788000001
  • Figure 0007701788000002
    Figure 0007701788000002
  • Figure 0007701788000003
    Figure 0007701788000003
Patent Text Reader

Abstract

The device may include multiple data processing engines (304), each of which may include a core (602) and a memory module (604). Each core (602) may be configured to access a memory module (604) in the same data processing engine (304) and a memory module (604) in at least one other of the multiple data processing engines (304).
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] Technical Field The present disclosure relates to integrated circuit devices (devices), and more particularly, to devices including data processing engines and / or arrays of data processing engines.

Background Art

[0002] Background A programmable integrated circuit (IC) refers to a type of IC that includes programmable circuitry. An example of a programmable IC is a field programmable gate array (FPGA). An FPGA is characterized by including programmable circuit blocks. Examples of programmable circuit blocks include, but are not limited to, input / output blocks (IOBs), configurable logic blocks (CLBs), dedicated random access memory blocks (BRAMs), digital signal processing blocks (DSPs), processors, clock managers, and delay lock loops (DLLs).

[0003] Circuit design can be physically realized within the programmable circuitry of a programmable IC by loading configuration data, sometimes referred to as a configuration bitstream, into the device. The configuration data can be loaded into the internal configuration memory cells of the device. The collective state of the individual configuration memory cells determines the functionality of the programmable IC. For example, the specific operations performed by various programmable circuit blocks and the connectivity between the programmable circuit blocks of the programmable IC are defined by the collective state of the configuration memory cells when the configuration data is loaded.

Summary of the Invention

Means for Solving the Problems

[0004] Summary In one or more embodiments, a device may include a plurality of data processing engines. Each data processing engine may include a core and a memory module. Each core may be configured to access a memory module within the same data processing engine and a memory module within at least one other data processing engine of the plurality of data processing engines.

[0005] In one or more embodiments, a method may include a first core of a first data processing engine that generates data, the first core writing the data to a first memory module within the first data processing engine, and the method may further include a second core of a second data processing engine that reads the data from the first memory module.

[0006] In one or more embodiments, a device may include a plurality of data processing engines, a subsystem, and a system-on-chip (SoC) interface block coupled to the plurality of data processing engines and the subsystem. The SoC interface block may be configured to exchange data between the subsystem and the plurality of data processing engines.

[0007] In one or more embodiments, a tile for a SoC interface block may include a memory-mapped switch configured to provide a first portion of configuration data to neighboring tiles and a second portion of configuration data to a certain data processing engine among a plurality of data processing engines. The tile may include a stream switch configured to provide first data to at least one neighboring tile and second data to the certain data processing engine among the plurality of data processing engines. The tile may include an event broadcast circuit configured to receive events generated within the tile and events from circuitry external to the tile, the event broadcast circuit being programmable to provide selected ones of the events to selected destinations. The tile may include an interface circuit coupling the memory-mapped switch, the stream switch, and the event broadcast circuit to a subsystem of a device including the tile.

[0008] In one or more embodiments, a device may include a plurality of data processing engines. Each data processing engine may include a core and a memory module. The plurality of data processing engines may be arranged in a plurality of columns. Each core may be configured to communicate with other neighboring data processing engines among the plurality of data processing engines by way of shared access to the memory modules of the other neighboring data processing engines.

[0009] In one or more embodiments, a device may include a plurality of data processing engines. Each data processing engine may include a memory pool having a plurality of memory banks, a plurality of cores each coupled to the memory pool and configured to access the plurality of memory banks, a memory-mapped switch coupled to the memory pool and a memory-mapped switch of at least one neighboring data processing engine, and a stream switch coupled to each of the plurality of cores and to a stream switch of at least one neighboring data processing engine.

[0010] This summary section is provided merely to introduce certain concepts and is not provided to identify any important or essential features of the claimed subject matter. Other features of the present invention will become apparent from the accompanying drawings and the following detailed description.

[0011] The configuration of the present invention is shown, by way of example, in the accompanying drawings. However, the drawings should not be construed as limiting the configuration of the present invention to the specific examples shown. Various aspects and advantages will become apparent upon reviewing the following detailed description and referring to the drawings.

Brief Description of the Drawings

[0012]

Figure 1

Figure 2A

Figure 2B

Figure 2C

Figure 2D

Figure 3

Figure 4A

Figure 4B

Figure 5A

Figure 5B

Figure 5C

Figure 5D

Figure 5E

Figure 5F

Figure 5G

Figure 6

Figure 7

Figure 8

Figure 9

Figure 10A

Figure 10B

Figure 10C

Figure 10D

Figure 10E

Figure 11

Figure 12

Figure 13

Figure 14A

Figure 14B

Figure 14C

Figure 15

Figure 16

Figure 17

Figure 18

Figure 19

Figure 20

Figure 21

Figure 22

Figure 23

Figure 24

DETAILED DESCRIPTION OF THE INVENTION

[0013] Detailed Description The present disclosure defines the claims that specify novel features, but it is believed that the various features described within the present disclosure will be better understood by considering the description in conjunction with the drawings. The processes, machines, manufactures, and any variations thereof described herein are provided for purposes of illustration. The specific structural and functional details described within the present disclosure should not be construed as limiting, but rather as a representative basis for teaching one of ordinary skill in the art to variously employ the features described in substantially any appropriately detailed structure as a basis for the claims. Further, the terms and expressions used within the present disclosure are not intended to be limiting; rather, they are intended to provide an understandable description of the features described.

[0014] The present disclosure relates to an integrated circuit device (device) that includes one or more data processing engines (DPEs) and / or an array of DPEs. The array of DPEs refers to a plurality of hardwired circuit blocks. The plurality of circuit blocks may be programmable. The array of DPEs may include a plurality of DPEs and system-on-chip (SoC) interface blocks. Generally, a DPE includes a core that can provide data processing capabilities. The DPE further includes a memory module accessible by one or more cores within the DPE. In certain embodiments, the memory module of a DPE may also be accessed by one or more other cores in different DPEs of the array of DPEs.

[0015] A DPE can further include a DPE interconnect. The DPE interconnect refers to a circuit that can enable communication between other DPEs in the array of DPEs and / or communication between different subsystems of the device that includes the array of DPEs. The DPE interconnect may further support the configuration of the DPE. In certain embodiments, the DPE interconnect can carry control data and / or debug data.

[0016] The DPE array can be configured using any of a variety of different architectures. In one or more embodiments, the DPE array can be configured into one or more rows and one or more columns. Optionally, the columns and / or rows of DPEs are aligned. In some embodiments, each DPE can include a single core coupled to a memory module. In other embodiments, one or more DPEs or each DPE of the DPE array can be implemented to include more than two cores coupled to a memory module.

[0017] In one or more embodiments, the DPE array is implemented as a homogeneous structure where each DPE is the same as each other DPE. In other embodiments, the DPE array is implemented as a heterogeneous structure and the DPE array includes more than two different types of DPEs. For example, the DPE array can include DPEs with a single core, DPEs with multiple cores, DPEs that include different types of cores therein, and / or DPEs with different physical architectures.

[0018] The DPE array can be implemented in various sizes. For example, the DPE array can be implemented to span the entire width and / or length of the die of the device. In another example, the DPE array can be implemented to span a portion of the entire width and / or length of such a die. In further embodiments, more than one DPE array can be implemented within the die, and the different DPE arrays are distributed in different regions on the die, have different sizes, have different shapes, and / or have different architectures (e.g., aligned rows and / or columns, homogeneous and / or heterogeneous) described herein. Further, the DPE array can include a different number of rows of DPEs and / or a different number of columns of DPEs.

[0019] The DPE array can be utilized with and coupled to any of a variety of different subsystems within a device. Such subsystems can include, but are not limited to, a processor and / or a processor system, programmable logic, and / or a network-on-chip (NoC). In certain embodiments, the NoC may be programmable. Further examples of subsystems that can be included in a device and coupled to the DPE array can include, but are not limited to, an application-specific integrated circuit (ASIC), a hardwired circuit block, analog and / or mixed-signal circuits, a graphics processing unit (GPU), and / or a general-purpose processor (e.g., a central processing unit or CPU). An example of a CPU is a processor having an x86-type architecture. As used herein, the term “ASIC” can refer to an IC, die, and / or a portion of a die that includes an application-specific circuit combined with one or more other types of circuits, and / or an IC and / or die that is entirely formed of an application-specific circuit.

[0020] In certain embodiments, a device that includes one or more DPE arrays can be implemented using a single die architecture. In that case, the DPE array and any other subsystems utilized with the DPE array are implemented on the same die of the device. In other embodiments, a device that includes one or more DPE arrays can be implemented as a multi-die device that includes two or more dies. In some multi-die devices, one or more DPE arrays can be implemented on one die, and one or more other subsystems can be implemented on one or more other dies. In other multi-die devices, one or more DPE arrays can be implemented in combination with one or more other subsystems of the multi-die device in one or more dies (e.g., the DPE array is implemented within the same die as at least one subsystem).

[0021] The DPE arrays described in this disclosure can implement an optimized digital signal processing (DSP) architecture. The DSP architecture can efficiently execute any of a variety of different operations. Examples of types of operations that can be performed by the architecture include, but are not limited to, operations related to wireless radio, decision feedback equalization (DFE), 5G / baseband, wireless backhaul, machine learning, automotive driver assistance, embedded vision, cable access, and / or radar. The DPE arrays described herein can perform such operations while consuming less power than other solutions that utilize conventional programmable (e.g., FPGA type) circuits. Further, a DPE array-based solution can be implemented using less die area than other solutions that utilize conventional programmable circuits. The DPE arrays can further perform operations as described herein while meeting predictable and guaranteed data throughput and latency metrics.

[0022] Further aspects of the configuration of the present invention are described in more detail below with reference to the drawings. For simplicity and clarity of explanation, the elements shown in the figures are not necessarily drawn to scale. For example, the dimensions of some elements may be emphasized relative to other elements for clarity. Further, where appropriate, reference numerals are repeated between the figures to indicate corresponding, similar, or like features.

[0023] FIG. 1 shows an example of a device 100 including a DPE array 102. In the example of FIG. 1, the DPE array 102 includes a SoC interface block 104. Also, the device 100 includes one or more subsystems 106-1 to 106-N. In one or more embodiments, the device 100 is implemented as a system-on-chip (SoC) type device. Generally, an SoC refers to an IC including two or more subsystems capable of interacting with each other. As an example, an SoC may include a processor that executes program code and one or more other circuits. The other circuits may be implemented as hardwired circuits, programmable circuits, other subsystems, and / or any combination thereof. The circuits can operate in cooperation with each other and / or with the processor.

[0024] The DPE array 102 is formed of a plurality of interconnected DPEs. Each DPE is a hardwired circuit block. Each DPE may be programmable. The SoC interface block 104 may include one or more tiles. Each of the tiles of the SoC interface block 104 may be hardwired. Each tile of the SoC interface block 104 may be programmable. The SoC interface block 104 provides an interface between the DPE array 102, e.g., the DPEs, and other parts of the SoC such as the subsystems 106 of the device 100. The subsystems 106-1 to 106-N may represent, for example, one or more or any combination of a processor and / or a processor system (e.g., a CPU, a general-purpose processor, and / or a GPU), programmable logic, a NoC, an ASIC, analog and / or mixed-signal circuits, and / or hardwired circuit blocks.

[0025] In one or more embodiments, device 100 is implemented using a single die architecture. In that case, DPE array 102 and at least one subsystem 106 may be included in or implemented on a single die. In one or more other embodiments, device 100 is implemented using a multi-die architecture. In that case, DPE array 102 and subsystem 106 may be implemented across two or more dies. For example, DPE array 102 may be implemented on one die and subsystem 106 may be implemented on one or more other dies. In another example, SoC interface block 104 may be implemented on a die different from the DPEs of DPE array 102. In yet another example, DPE array 102 and at least one subsystem 106 may be implemented on the same die, and other subsystems and / or other DPE arrays may be implemented on other dies. Further examples of single die and multi-die architectures are described in more detail below in connection with FIGS. 2, 3, 4, and 5.

[0026] FIGS. 2A, 2B, 2C, and 2D (collectively referred to as "FIG. 2") show an exemplary architecture for a device that includes one or more DPE arrays 102. More specifically, FIG. 2 shows an example of a single die architecture for device 100. For illustrative purposes, SoC interface block 104 is not shown in FIG. 2.

[0027] Figure 2A shows an exemplary architecture of a device 100 that includes a single DPE array. In the example of Figure 2A, the DPE array 102 is implemented in the device 100 together with the subsystem 106-1. The DPE array 102 and the subsystem 106-1 are implemented within the same die. The DPE array 102 may extend across the entire width of the die of the device 100, or may extend only partially across the die of the device 100. As shown, the DPE array 102 is implemented in the upper region of the device 100. However, it should be understood that the DPE array 102 may be implemented in another region of the device 100. Therefore, the arrangement and / or size of the DPE array 102 in Figure 2A are not intended to be limiting. The DPE array 102 may be coupled to the subsystem 106-1 by a SoC interface block 104 (not shown).

[0028] Figure 2B shows an exemplary architecture of a device 100 that includes multiple DPE arrays. In the example of Figure 2B, multiple DPE arrays are implemented and shown as DPE array 102-1 and DPE array 102-2. Figure 2B shows that multiple DPE arrays may be implemented within the same die of the device 100 together with the subsystem 106-1. The DPE array 102-1 and / or the DPE array 102-2 may extend across the entire width of the die of the device 100, or may extend only partially across the die of the device 100. As shown, the DPE array 102-1 is implemented in the upper region of the device 100, and the DPE array 102-2 is implemented in the bottom region of the device 100. As described above, the arrangement and / or size of the DPE arrays 102-1 and 102-2 in Figure 2B are not intended to be limiting.

[0029] In one or more embodiments, DPE arrays 102-1 and 102-2 may be substantially similar or the same. For example, DPE array 102-1 may be the same as DPE array 102-2 with respect to the size, shape, number of DPEs, and whether the DPEs are of the same kind or of a similar type and arrangement in each respective DPE array. In one or more other embodiments, DPE array 102-1 may be different from DPE array 102-2. For example, DPE array 102-1 may be different from DPE array 102-2 with respect to the size, shape, number, core type of DPEs, and whether the DPEs are of the same kind or of different types and / or arrangements in each respective DPE array.

[0030] In one or more embodiments, each of DPE arrays 102-1 and 102-2 is coupled to subsystem 106-1 via its own SoC interface block (not shown). For example, a first SoC interface block may be included and used to couple DPE array 102-1 to subsystem 106-1, and a second SoC interface block may be included and used to couple DPE array 102-2 to subsystem 106-1. In another embodiment, a single SoC interface block can be used to couple both DPE array 102-1 and DPE array 102-2 to subsystem 106-1. In the latter case, for example, one of the DPE arrays may not include an SoC interface block. The DPEs within that array may be coupled to subsystem 106-1 using the SoC interface block of the other DPE array.

[0031] FIG. 2C shows an exemplary architecture of a device 100 that includes a plurality of DPE arrays and a plurality of subsystems. In the example of FIG. 2C, a plurality of DPE arrays are implemented and shown as DPE array 102-1 and DPE array 102-2. FIG. 2C shows that the plurality of DPE arrays can be implemented within the same die of device 100 and that the placement or location of DPE arrays 102 can vary. Further, DPE arrays 102-1 and 102-2 are implemented within the same die as subsystems 106-1 and 106-2.

[0032] In the example of FIG. 2C, DPE array 102-1 and DPE array 102-2 do not extend across the entire width of the die of device 100. Rather, each of DPE arrays 102-1 and 102-2 extends partially across the die of device 100 and is thus implemented within a region that is a portion of the width of the die of device 100. As in the example of FIG. 2B, DPE array 102-1 and DPE array 102-2 of FIG. 2C may be substantially similar or the same, or may be different.

[0033] In one or more embodiments, each of DPE array 102-1 and DPE array 102-2 is coupled to subsystem 106-1 and / or subsystem 106-2 via its own SoC interface block (not shown). By way of illustrative and non-limiting example, a first SoC interface block may be included and used to couple DPE array 102-1 to subsystem 106-1, and a second SoC interface block may be included and used to couple DPE array 102-2 to subsystem 106-2. In that case, each DPE array communicates with a subset of the available subsystems of device 100. In another example, a first SoC interface block may be included and used to couple DPE array 102-1 to subsystems 106-1 and 106-2, and a second SoC interface block may be included and used to couple DPE array 102-2 to subsystems 106-1 and 106-2. In yet another example, a single SoC interface block can be used to couple both DPE array 102-1 and DPE array 102-2 to subsystem 106-1 and / or subsystem 106-2. As described above, the arrangement and / or size of DPE arrays 102-1 and 102-2 in FIG. 2C are not intended to be limiting.

[0034] Figure 2D shows another exemplary architecture of device 100 that includes a plurality of DPE arrays and a plurality of subsystems. In the example of Figure 2D, the plurality of DPE arrays are implemented and shown as DPE array 102-1 and DPE array 102-2. Figure 2D also shows that the plurality of DPE arrays can be implemented within the same die of device 100, and that the placement and / or location of DPE array 102 can vary. In the example of Figure 2D, DPE array 102-1 and DPE array 102-2 do not extend across the full width of the die of device 100. Rather, each of DPE arrays 102-1 and 102-2 is implemented in a region that is a portion of the width of the die of device 100. Further, device 100 of Figure 2D includes subsystems 106-1, 106-2, 106-3, and 106-4 within the same die as DPE arrays 102-1 and 102-2. As in the example of Figure 2B, DPE array 102-1 and DPE array 102-2 of Figure 2D may be substantially similar or the same, or may be different.

[0035] The connectivity between the DPE arrays and the subsystems in the example of Figure 2D can vary. In some cases, a DPE array can be coupled to only a subset of the subsystems available in device 100. In other cases, a DPE array can be coupled to more than one subsystem or to each subsystem within device 100.

[0036] The example of Figure 2 is provided for purposes of illustration and not limitation. A device having a single die can include one or more different DPE arrays located in different regions of the die. The number, placement, and / or size of the DPE arrays can vary. Further, the DPE arrays can be the same or different. One or more DPE arrays may be implemented in combination with one or more and / or any combination of different types of subsystems described within this disclosure.

[0037] In one or more embodiments, two or more DPE arrays may be configured to communicate directly with each other. For example, DPE array 102-1 may be able to communicate directly with DPE array 102-2 and / or additional DPE arrays. In certain embodiments, DPE array 102-1 may communicate with DPE array 102-2 and / or other DPE arrays via one or more SoC interface blocks.

[0038] FIG. 3 shows another exemplary architecture of device 100. In the example of FIG. 3, DPE array 102 is implemented as a two-dimensional array of DPEs 304 that includes an SoC interface block 104. DPE array 102 may be implemented using any of a variety of different architectures that are described in more detail below. By way of example and not limitation, FIG. 3 shows DPEs 304 arranged in aligned rows and aligned columns as will be described in more detail in connection with FIG. 19. However, in other embodiments, the DPEs 304 may be arranged such that the DPEs within a selected row and / or column are horizontally inverted or flipped with respect to the DPEs in adjacent rows and / or columns. An example of a horizontal inversion of DPEs is described in connection with FIG. 18. In one or more other embodiments, the rows and / or columns of DPEs may be offset with respect to adjacent rows and / or columns. One or more or all of the DPEs 304 may be implemented to include a single core as generally described in connection with FIGS. 6 and 8, or to include two or more cores as generally described in connection with FIG. 12.

[0039] The SoC interface block 104 can couple the DPE 304 to one or more other subsystems of the device 100. In one or more embodiments, the SoC interface block 104 is coupled to an adjacent DPE 304. For example, the SoC interface block 104 can be directly coupled to each DPE 304 in the bottom row of DPEs in the DPE array 102. By way of illustration, the SoC interface block 104 can be directly connected to DPEs 304-1, 304-2, 304-3, 304-4, 304-5, 304-6, 304-7, 304-8, 304-9, and 304-10.

[0040] FIG. 3 is shown for illustrative purposes. In other embodiments, the SoC interface block 104 can be located at the top of the DPE array 102, on the left side of the DPE array 102 (e.g., as a column), on the right side of the DPE array 102 (e.g., as a column), or at multiple locations within and around the DPE array 102 (e.g., as one or more intervening rows and / or columns within the DPE array 102). Depending on the layout and location of the SoC interface block 104, the particular DPEs coupled to the SoC interface block 104 can vary.

[0041] By way of illustration and not limitation, if the SoC interface block 104 is located on the left side of the DPE 304, the SoC interface block 104 can be directly coupled to the left column of DPEs including DPE 304-1, DPE 304-11, DPE 304-21, and DPE 304-31. If the SoC interface block 104 is located on the right side of the DPE 304, the SoC interface block 104 can be directly coupled to the right column of DPEs including DPE 304-10, DPE 304-20, DPE 304-30, and DPE 304-40. If the SoC interface block 104 is located above the DPE 304, the SoC interface block 104 can be coupled to the top row of DPEs including DPE 304-31, DPE 304-32, DPE 304-33, DPE 304-34, DPE 304-35, DPE 304-36, DPE 304-37, DPE 304-38, DPE 304-39, and DPE 304-40. If the SoC interface block 104 is located in multiple locations, the specific DPEs directly connected to the SoC interface block 104 can vary. For example, if the SoC interface block is implemented as a row and / or column within the DPE array 102, the DPEs directly coupled to the SoC interface block 104 can be those adjacent to the SoC interface block 104 on one or more or each side of the SoC interface block 104.

[0042] The DPEs 304 are interconnected by DPE interconnects (not shown) that form a DPE interconnect network when taken collectively. Thus, the SoC interface block 104 can communicate with one or more selected DPEs 304 of the DPE array 102 directly connected to the SoC interface block 104 and communicate with any DPE 304 of the DPE array 102 by utilizing the DPE interconnect network formed by the DPE interconnects implemented within each respective DPE 304.

[0043] The SoC interface block 104 can couple each DPE 304 within the DPE array 102 to one or more other subsystems of the device 100. For illustration, the device 100 includes subsystems (e.g., subsystem 106) such as the NoC 308, programmable logic (PL) 310, processor system (PS) 312, and / or any of the hardwired circuit blocks 314, 316, 318, 320, and / or 322. For example, the SoC interface block 104 can establish a connection between a selected DPE 304 and the PL 310. The SoC interface block 104 can also establish a connection between a selected DPE 304 and the NoC 308. Through the NoC 308, a selected DPE 304 can communicate with the PS 312 and / or the hardwired circuit blocks 320 and 322. A selected DPE 304 can communicate with the hardwired circuit blocks 314 - 318 via the SoC interface block 104 and the PL 310. In certain embodiments, the SoC interface block 104 can be directly coupled to one or more subsystems of the device 100. For example, the SoC interface block 104 can be directly coupled to the PS 312 and / or other hardwired circuit blocks. In certain embodiments, the hardwired circuit blocks 314 - 322 can be regarded as examples of ASICs.

[0044] In one or more embodiments, the DPE array 102 includes a single clock domain. Other subsystems such as the NoC 308, PL 310, PS 312, and various hardwired circuit blocks can be in one or more separate or different clock domains. Further, the DPE array 102 can include additional clocks that can be used to interface with other subsystems among the subsystems. In certain embodiments, the SoC interface block 104 includes a clock signal generator that can generate one or more clock signals that can be provided or distributed to the DPEs 304 of the DPE array 102.

[0045] The DPE array 102 can be programmed by loading configuration data (also referred to herein as "configuration registers") into internal configuration memory cells that define the connectivity between the DPE 304 and the SoC interface block 104 and how the DPE 304 and the SoC interface block 104 operate. For example, the DPE 304 and the SoC interface block 104 are programmed in such a way for a particular DPE 304 or a group of DPE 304s to communicate with a subsystem. Similarly, the DPEs are programmed in such a way for one or more particular DPE 304s to communicate with one or more other DPE 304s. The DPE 304 and the SoC interface block 104 can be programmed by loading configuration data into the configuration registers within the DPE 304 and the SoC interface block 104, respectively. In another example, a clock signal generator, which is part of the SoC interface block 104, may be programmable using configuration data to vary the clock frequency provided to the DPE array 102.

[0046] The NoC 308 provides connections to selected circuit blocks (such as circuit blocks 320 and 322) among the PL310, the PS312, and the hardwired circuit blocks. In the example of FIG. 3, the NoC 308 is programmable. In the case of a programmable NoC used with other programmable circuits, the nets to be routed through the NoC 308 are unknown until the user circuit design is created for realization within the device 100. The NoC 308 can be programmed by loading configuration data into internal configuration registers that define how elements within the NoC 308, such as switches and interfaces, are configured and how data is passed from switch to switch and between NoC interfaces.

[0047] The NoC 308 is fabricated as part of the device 100 and is not physically modifiable, but can be programmed to establish connectivity between different master circuits and different slave circuits of the user circuit design. In this regard, the NoC 308 can adapt to different circuit designs, and each different circuit design has different combinations of master circuits and slave circuits realized at different locations within the device 100 that can be coupled by the NoC 308. The NoC 308 can be programmed to route data, such as application data and / or configuration data, between the master circuit and the slave circuit of the user circuit design. For example, the NoC 308 can be programmed to couple different user-specified circuits realized within the PL310 to different DPEs among the DPEs 304 via the PS 312 and the SoC interface block 104, to different hardwired circuit blocks, and / or to different circuits and / or systems external to the device 100.

[0048] The PL310 is a circuit that can be programmed to perform a specified function. As an example, the PL310 can be realized as a field programmable gate array (FPGA) circuit. The PL310 can include an array of programmable circuit blocks. Examples of programmable circuit blocks within the PL310 include, but are not limited to, input / output blocks (IOBs), configurable logic blocks (CLBs), dedicated random access memory blocks (BRAMs), digital signal processing blocks (DSPs), clock managers, and / or delay lock loops (DLLs).

[0049] Each programmable circuit block within the PL310 typically includes both a programmable interconnect circuit and a programmable logic circuit. The programmable interconnect circuit typically includes a number of interconnect wires of various lengths interconnected by programmable interconnect points (PIPs). Typically, the interconnect wires are configured (e.g., on a wire-by-wire basis) to provide connectivity on a bit-by-bit basis (e.g., each wire conveys 1 bit of information). The programmable logic circuit uses programmable elements, which can include, for example, look-up tables, registers, arithmetic logic, etc., to implement the logic of a user design. The programmable interconnect circuit and the programmable logic circuit can be programmed by loading configuration data defining how the programmable elements are configured and operate into internal configuration memory cells.

[0050] In the example of FIG. 3, the PL310 is shown in two separate sections. In another example, the PL310 can be implemented as an integrated area of programmable circuits. In yet another example, the PL310 can be implemented as three or more different areas of programmable circuits. The specific organization of the PL310 is not intended to be limiting.

[0051] In the example of FIG. 3, PS312 is implemented as a hardwired circuit fabricated as part of device 100. PS312 can be implemented as, or include, any of a variety of different processor types. For example, PS312 can be implemented as an individual processor, e.g., a single core that can execute program code. In another example, PS312 can be implemented as a multi-core processor. In yet another example, PS312 can include one or more cores, modules, coprocessors, interfaces, and / or other resources. PS312 can be implemented using any of a variety of different types of architectures. Exemplary architectures that can be used to implement PS312 can include, but are not limited to, the ARM processor architecture, the x86 processor architecture, the GPU architecture, the mobile processor architecture, the DSP architecture, or other suitable architectures that can execute computer-readable instructions or program code.

[0052] Circuit blocks 314 - 322 can be implemented as any of a variety of different hardwired circuit blocks. The hardwired circuit blocks 314 - 322 can be customized to perform a dedicated function. Examples of circuit blocks 314 - 322 can include, but are not limited to, input / output blocks (IOBs), transceivers, or other specialized circuit blocks. As described above, circuit blocks 314 - 322 can be considered an example of an ASIC.

[0053] The example of FIG. 3 shows an architecture that can be implemented in a device that includes a single die. Although DPE array 102 is shown as occupying the full width of device 100, in other embodiments, DPE array 102 can occupy less than the full width of device 100 and / or be located in different regions of device 100. Further, the number of DPEs 304 included can vary. Thus, the specific number of columns and / or rows of DPEs 304 can be different from that shown in FIG. 3.

[0054] In one or more other embodiments, a device such as device 100 may include two or more DPE arrays 102 located in different regions of device 100. For example, additional DPE arrays may be located under circuit blocks 320 and 322.

[0055] As described above, FIGS. 2-3 illustrate an exemplary architecture of a device that includes a single die. In one or more other embodiments, device 100 may be implemented as a multi-die device that includes one or more DPE arrays 102.

[0056] FIGS. 4A and 4B (collectively referred to as "FIG. 4") illustrate an example of a multi-die implementation of device 100. A multi-die device is a device or IC that includes two or more dies within a single package.

[0057] FIG. 4A shows a topography diagram of device 100. In the example of FIG. 4A, device 100 is implemented as a "stacked die" type of device formed by stacking a plurality of dies. Device 100 includes an interposer 402, a die 404, a die 406, and a substrate 408. Each of dies 404 and 406 is attached to a surface of interposer 402, e.g., the top surface. In one aspect, dies 404 and 406 are attached to interposer 402 using flip-chip technology. Interposer 402 is attached to the top surface of substrate 408.

[0058] In the example of FIG. 4A, interposer 402 is a die that has a plane on which dies 404 and 406 are horizontally stacked. As shown, dies 404 and 406 are arranged side by side on the plane of interposer 402. The number of dies shown for interposer 402 in FIG. 4A is for illustration purposes only and not a limitation. In other embodiments, three or more dies may be attached to interposer 402.

[0059] The interposer 402 provides a common mounting surface and electrical coupling for each of dies 404 and 406. The manufacture of the interposer 402 can include one or more process steps that enable the deposition of one or more conductive layers patterned to form wires. These conductive layers may be formed from aluminum, gold, copper, nickel, various silicides, and / or other suitable materials. The interposer 402 can be manufactured using one or more additional process steps that enable the deposition of one or more dielectric or insulating layers, such as silicon dioxide for example. The interposer 402 can also include vias and through vias (TVs). The TVs can be through-silicon vias (TSVs), through-glass vias (TGVs), or other via structures depending on the particular materials used to implement the interposer 402 and its substrate. When the interposer 402 is implemented as a passive die, the interposer 402 can have only various types of solder bumps, vias, wires, TVs, and under-bump metallization (UBM). When implemented as an active die, the interposer 402 can include additional process layers that form one or more active devices with respect to electrical devices such as transistors, diodes, etc. that include P-N junctions.

[0060] Each of dies 404 and 406 can be implemented as a passive die or an active die that includes one or more active devices. For example, one or more DPE arrays can be implemented on one or both of dies 404 and / or 406 when implemented as an active die. In one or more embodiments, die 404 can include one or more DPE arrays and die 406 can implement any of the different subsystems described herein. The examples provided herein are for illustrative purposes and are not intended to be limiting. For example, device 100 can include three or more dies, which can be of different types and / or functions.

[0061] Figure 4B is a side cross-sectional view of the device 100 of Figure 4A. Figure 4B shows a view of the device 100 from Figure 4A taken along the cut line 4B-4B. Each of the dies 404 and 406 is electrically and mechanically coupled to a first plane of the interposer 402 via solder bumps 410. In one example, the solder bumps 410 are implemented as microbumps. Additionally, any of a variety of other techniques may be used to attach the dies 404 and 406 to the interposer 402. For example, bond wires or edge wires may be used to mechanically and electrically attach the dies 404 and 406 to the interposer 402. In another example, an adhesive material may be used to mechanically attach the dies 404 and 406 to the interposer 402. As shown in Figure 4B, using the solder bumps 410 to attach the dies 404 and 406 to the interposer 402 is provided for illustrative purposes and is not intended as a limitation.

[0062] The interposer 402 includes one or more conductive layers 412 shown as dashed or dotted lines within the interposer 402. The conductive layers 412 are implemented using any of a variety of metal layers as described above. The conductive layers 412 are processed to form a patterned metal layer that realizes the wires 414 of the interposer 402. Wires realized within the interposer 402 that couple at least two different dies, such as dies 404 and 406, are referred to as inter-die wires. Figure 4B shows the wires 414 considered to be inter-die wires for illustrative purposes. The wires 414 carry inter-die signals between the die 404 and the die 406. For example, each of the wires 414 couples a solder bump 410 under the die 404 to a solder bump 410 under the die 406, thereby enabling the exchange of inter-die signals between the dies 404 and 406. The wires 414 can be data lines or power lines. A power line may be a wire that carries a voltage potential or a wire that has a ground potential or a reference voltage potential.

[0063] The different conductive layers 412 can be coupled using vias 416. Generally, via structures are used to implement vertical conductive paths (e.g., conductive paths perpendicular to the process layers of a device). In this regard, the vertical portions of the wires 414 that contact the solder bumps 410 are implemented as vias 416. By using multiple conductive layers to implement interconnects within the interposer 402, a greater number of signals can be routed and more complex signal routing can be achieved within the interposer 402.

[0064] Solder bumps 418 can be used to mechanically and electrically couple a second plane of the interposer 402 to the substrate 408. In certain embodiments, the solder bumps 418 are implemented as controlled collapse chip connection (C4) balls. The substrate 408 includes conductive paths (not shown) that couple the different solder bumps 418 to one or more nodes below the substrate 408. Thus, one or more of the solder bumps 418 couple a circuit within the interposer 402 to an external node of the device 100 via a circuit or wiring within the substrate 408.

[0065] The TV 420 is a via that forms an electrical connection that traverses the interposer 402 vertically, e.g., extending through a substantial portion, if not all, of the interposer 402. The TV 420 can be formed from any of a variety of different conductive materials including, but not limited to, copper, aluminum, gold, nickel, various silicides, and / or other suitable materials, like wires and vias. As shown, each of the TV 420s extends from the bottom surface of the interposer 402 to the conductive layer 412 of the interposer 402. The TV 420 can further be coupled to the solder bumps 410 through one or more of the conductive layers 412 in combination with one or more vias 416.

[0066] Figures 5A, 5B, 5C, 5D, 5E, 5F, and 5G (collectively referred to as "Figure 5") show exemplary multi-die implementations of device 100. The example of Figure 5 can be implemented as described in connection with Figure 4.

[0067] Referring to Figure 5A, die 404 includes one or more DPE arrays 102, and die 406 implements PS312.

[0068] Referring to Figure 5B, die 404 includes one or more DPE arrays 102, and die 406 implements ASIC 504. ASIC 504 can be implemented as any of a variety of different customized circuits suitable for performing specific or specialized operations.

[0069] Referring to Figure 5C, die 404 includes one or more DPE arrays 102, and die 406 implements PL310.

[0070] Referring to Figure 5D, die 404 includes one or more DPE arrays 102, and die 406 implements analog and / or mixed (analog / mixed) signal circuitry 508. Analog / mixed signal circuitry 508 can include one or more radio receivers, radio transmitters, amplifiers, analog-to-digital converters, digital-to-analog converters, or other analog and / or digital circuitry.

[0071] Figures 5E, 5F, and 5G show an example of device 100 having three dies 404, 406, and 510. Referring to Figure 5E, device 100 includes dies 404, 406, and 510. Die 404 includes one or more DPE arrays 102. Die 406 includes PL310. Die 510 includes ASIC 504.

[0072] Referring to Figure 5F, die 404 includes one or more DPE arrays 102. Die 406 includes PL310. Die 510 includes analog / mixed signal circuitry 508.

[0073] Referring to FIG. 5G, die 404 includes one or more DPE arrays 102. Die 406 includes ASIC 504. Die 510 includes analog / mixed signal circuitry 508. In one or more embodiments, the PS (e.g., PS 312) is an example of an ASIC.

[0074] In the example of FIG. 5, each of die 406 and / or 510 is shown as including a particular type of subsystem. In other embodiments, die 404, 406, and / or 510 can include one or more subsystems in combination with one or more DPE arrays 102. Further, die 404, 406, and / or 510 can include two or more different types of subsystems. Thus, any one or more of die 404, 406, and / or 510 can include one or more DPE arrays 102 in any combination with one or more subsystems.

[0075] In one or more embodiments, the interposer 402, as well as the dies 404, 406, and / or 510, may be implemented using the same IC manufacturing technology (e.g., feature size). In one or more other embodiments, the interposer 402 may be implemented using a particular IC manufacturing technology, and the dies 404, 406, and / or 510 may be implemented using different IC manufacturing technologies. In still other embodiments, the dies 404, 406, and / or 510 may be implemented using different IC manufacturing technologies that may be the same as or different from the IC manufacturing technology used to implement the interposer 402. By using different IC manufacturing technologies for different dies and / or interposers, a lower cost and / or more reliable IC manufacturing technology can be used for a particular die, while another IC manufacturing technology capable of forming smaller feature sizes can be used for other dies. For example, a more mature manufacturing technology may be used to implement the interposer 402, and another technology capable of forming smaller feature sizes may be used to implement the active die and / or the die including the DPE array 102.

[0076] The example of FIG. 5 shows an example of a multi-die implementation of the device 100 including two or more dies mounted on an interposer. The number of dies shown is for illustrative purposes only and not limiting. In other embodiments, the device 100 may include more than three dies mounted on the interposer 402.

[0077] In one or more other embodiments, the multi-divergence of device 100 can be realized using an architecture other than the stacked die architecture of FIG. 4. For example, device 100 can be realized as a multi-chip module (MCM). An MCM implementation example of device 100 may be realized using one or more pre-packaged ICs mounted on a circuit board having a form factor and / or footprint intended to mimic an existing chip package. In another example, an MCM implementation example of device 100 can be realized by integrating two or more dies on a high-density interconnect substrate. In yet another example, an MCM implementation example of device 100 can be realized as a "chip stack" package.

[0078] Using a DPE array as described herein in combination with one or more other subsystems, whether realized in a single die device or a multi-die device, increases the processing power of the device while keeping the area usage and power consumption low. For example, one or more DPE arrays can be used to accelerate specific operations in hardware and / or perform functions offloaded from one or more of the subsystems of the device described herein. For example, when used with a PS, the DPE array can be used as a hardware accelerator. The PS can offload operations to be performed by the DPE array or a portion thereof. In other examples, the DPE array can be used to perform computationally resource-intensive operations such as generating digital predistortion to be applied to an analog / mixed-signal circuit.

[0079] It should be understood that any of the various combinations of the DPE arrays and / or other subsystems described herein in connection with FIGS. 1, 2, 3, 4, and / or 5 can be realized in either a single die type of device or a multi-die type of device.

[0080] In various examples described herein, the SoC interface block is implemented within the DPE array. In one or more other embodiments, the SoC interface block may be implemented external to the DPE array. For example, the SoC interface block may be implemented as a circuit block separate from the circuit block that implements the plurality of DPEs, such as a stand-alone circuit block.

[0081] FIG. 6 shows an exemplary architecture of DPE 304 of DPE array 102. In the example of FIG. 6, DPE 304 includes a core 602, a memory module 604, and a DPE interconnect 606.

[0082] Core 602 provides the data processing capabilities of DPE 304. Core 602 may be implemented as any of a variety of different processing circuits. In the example of FIG. 6, core 602 includes an optional program memory 608. In one or more embodiments, core 602 is implemented as a processor capable of executing program code, such as computer-readable instructions. In that case, program memory 608 is included and can store the instructions executed by core 602. Core 602 may be implemented as, for example, a CPU, GPU, DSP, vector processor, or other type of processor capable of executing instructions. The core may be implemented using any of the various CPU and / or processor architectures described herein. In another example, core 602 is implemented as a very long instruction word (VLIW) vector processor or DSP.

[0083] In certain embodiments, program memory 608 is implemented as a dedicated program memory that is private to core 602. Program memory 608 may be used only by the cores of the same DPE304. Thus, program memory 608 may be accessed only by core 602 and is not shared with any other DPE or components of another DPE. Program memory 608 may include a single port for read and write operations. Program memory 608 can support program compression and is addressable using the memory-mapped network portion of DPE interconnect 606, which is described in more detail below. For example, program memory 608 may be loaded with program code executable by core 602 via the memory-mapped network of DPE interconnect 606.

[0084] In one or more embodiments, program memory 608 can support one or more error detection and / or error correction mechanisms. For example, program memory 608 may be implemented to support parity checking by adding parity bits. In another example, program memory 608 can be an error correction code (ECC) memory that can detect and correct different types of data corruption. In another example, program memory 608 can support both ECC and parity checking. The different types of error detection and / or error correction described herein are provided for illustrative purposes and are not intended to limit the described embodiments. Other error detection and / or error correction techniques can be used with program memory 608 other than those enumerated.

[0085] In one or more embodiments, core 602 may have a customized architecture to support an application-specific instruction set. For example, core 602 may be customized for a wireless application and configured to execute wireless-specific instructions. In another example, core 602 may be customized for machine learning and configured to execute machine learning-specific instructions.

[0086] In one or more other embodiments, core 602 is implemented as a hardwired circuit, such as a hardware intellectual property (IP) core dedicated to performing a particular operation(s). In this case, core 602 may not execute program code. In embodiments where core 602 does not execute program code, program memory 608 may be omitted. By way of example and not limitation, core 602 may be implemented as a hardware forward error correction (FEC) engine or other circuit block.

[0087] Core 602 may include configuration register 624. Configuration register 624 can load configuration data to control the operation of core 602. In one or more embodiments, core 602 may be activated and / or deactivated based on the configuration data loaded into configuration register 624. In the example of FIG. 6, configuration register 624 is addressable (e.g., readable and / or writable) via a memory-mapped network of DPE interconnect 606 described in more detail below.

[0088] In one or more embodiments, the memory module 604 can store data used by the core 602 and / or data generated by the core 602. For example, the memory module 604 can store application data. The memory module 604 can include a read / write memory such as a random access memory. Thus, the memory module 604 is capable of storing data that can be read and consumed by the core 602. The memory module 604 can also store data (e.g., results) written by the core 602.

[0089] In one or more other embodiments, the memory module 604 can store data, such as application data, that can be used and / or generated by one or more other cores of other DPEs within the DPE array. One or more other cores of the DPE can also read from and / or write to the memory module 604. In certain embodiments, the other cores that can read from and / or write to the memory module 604 can be cores of one or more neighboring DPEs. Another DPE that shares (e.g., is adjacent to) a boundary line or boundary with the DPE 304 is referred to as a "neighboring" DPE to the DPE 304. By enabling the core 602 and one or more other cores from neighboring DPEs to read from and / or write to the memory module 604, the memory module 604 implements a shared memory that supports communication between different DPEs and / or cores that can access the memory module 604.

[0090] Referring to FIG. 3, for example, DPE304-14, 304-16, 304-5, and 304-25 are considered to be DPEs in the vicinity of DPE304-15. In one example, the cores within each of DPE304-16, 304-5, and 304-25 can read from and write to the memory module within DPE304-15. In certain embodiments, only the DPEs in the vicinity adjacent to the memory module can access the memory module of DPE304-15. For example, DPE304-14 is adjacent to DPE304-15, but may not be adjacent to the memory module of DPE304-15 because the core of DPE304-15 can be located between the core of DPE304-14 and the memory module of DPE304-15. Thus, in certain embodiments, the core of DPE304-14 may not access the memory module of DPE304-15.

[0091] In certain embodiments, whether the core of a DPE can access the memory module of another DPE depends on the number of memory interfaces included in the memory module and whether such a core is connected to the available ones of the memory interfaces of the memory module. In the above example, the memory module of DPE304-15 includes four memory interfaces, and the cores of each of DPE304-16, 304-5, and 304-25 are connected to such memory interfaces. The core 602 within DPE304-15 itself is connected to the fourth memory interface. Each memory interface may include one or more read and / or write channels. In certain embodiments, each memory interface includes a plurality of read channels and a plurality of write channels, and the specific core attached thereto can read from and / or write to multiple banks within the memory module 604 simultaneously.

[0092] In other examples, more than four memory interfaces may be available. Such other memory interfaces may be used to enable the DPEs on the diagonal of DPE304-15 to access the memory modules of DPE304-15. For example, cores within DPEs such as DPE304-14, 304-24, 304-26, 304-4, and / or 304-6 can also access the memory modules within DPE304-15 if they are coupled to the available memory interfaces of the memory modules within DPE304-15.

[0093] Memory module 604 may include a configuration register 636. The configuration register 636 can load configuration data to control the operation of the memory module 604. In the example of FIG. 6, the configuration register 636 (and 624) is addressable (e.g., can be read from and / or written to) via the memory-mapped network of DPE interconnect 606, which is described in more detail below.

[0094] In the example of FIG. 6, DPE interconnect 606 is specialized for DPE304. DPE interconnect 606 facilitates various operations, including communication between DPE304 and one or more other DPEs of DPE array 102 and / or communication with other subsystems of device 100. DPE interconnect 606 further enables the configuration, control, and debugging of DPE304.

[0095] In certain embodiments, DPE interconnect 606 is implemented as an on-chip interconnect. Examples of on-chip interconnects are Advanced Microcontroller Bus Architecture (AMBA) eXtensible Interface (AXI) buses (e.g., or switches). The AMBA AXI bus is a built-in microcontroller bus interface for use in establishing on-chip connections between circuit blocks and / or systems. The AXI bus is provided herein as an example of an interconnect circuit that can be used with the configurations of the invention described within the present disclosure and is thus not intended as a limitation. Other examples of interconnect circuits can include other types of buses, crossbars, and / or other types of switches.

[0096] In one or more embodiments, DPE interconnect 606 includes two different networks. The first network can exchange data with other DPEs of DPE array 102 and / or other subsystems of device 100. For example, the first network can exchange application data. The second network can exchange data such as DPE configuration, control, and / or debug data.

[0097] In the example of FIG. 6, the first network of the DPE interconnect 606 is formed by a stream switch 626 and one or more stream interfaces. As shown, the stream switch 626 includes a plurality of stream interfaces (abbreviated as "SI" in FIG. 6). In one or more embodiments, each stream interface may include one or more masters (e.g., master interfaces or outputs) and / or one or more slaves (e.g., slave interfaces or inputs). Each master can be an independent output with a specific bit width. For example, each master included in the stream interface may be an independent AXI master. Each slave can be an independent input with a specific bit width. For example, each slave included in the stream interface may be an independent AXI slave.

[0098] The stream interfaces 610-616 are used to communicate with other DPEs within the DPE array 102 and / or the SoC interface block 104. For example, each of the stream interfaces 610, 612, 614, and 616 can communicate in different basic compass directions. In the example of FIG. 6, the stream interface 610 communicates with the DPE on the left (west). The stream interface 612 communicates with the DPE on the top (north). The stream interface 614 communicates with the DPE on the right (east). The stream interface 616 communicates with the DPE on the bottom (south) or the SoC interface block 104.

[0099] The stream interface 628 is used to communicate with the core 602. The core 602 includes, for example, a stream interface 638 that connects to the stream interface 628, thereby enabling the core 602 to communicate directly with other DPEs 304 via the DPE interconnect 606. For example, the core 602 may include instructions or hardwired circuitry that enables the core 602 to directly transmit and / or receive data via the stream interface 638. The stream interface 638 can be blocking or non-blocking. In one or more embodiments, when the core 602 attempts to read from an empty stream or write to a full stream, the core 602 may stall. In other embodiments, attempting to read from an empty stream or write to a full stream may not cause the core 602 to stall. Rather, the core 602 may continue execution or operation.

[0100] The stream interface 630 is used to communicate with the memory module 604. The memory module 604 includes, for example, a stream interface 640 that connects to the stream interface 630, thereby enabling other DPEs 304 to communicate with the memory module 604 via the DPE interconnect 606. The stream switch 626 enables DPEs that are not coupled to the memory interfaces of non-proximate DPEs and / or the memory module 604 to communicate with the core 602 and / or the memory module 604 via a DPE interconnect network formed by the DPE interconnects of each DPE 304 in the DPE array 102.

[0101] Referring back to FIG. 3, using DPE304-15 as a reference point, stream interface 610 can be coupled to and communicate with another stream interface located at the DPE interconnect of DPE304-14. Stream interface 612 can be coupled to and communicate with another stream interface located within the DPE interconnect of DPE304-25. Stream interface 614 can be coupled to and communicate with another stream interface located at the DPE interconnect of DPE304-16. Stream interface 616 can be coupled to and communicate with another stream interface located at the DPE interconnect of DPE304-5. Thus, core 602 and / or memory module 604 can also communicate with any DPE within DPE array 102 via the DPE interconnects within the DPEs.

[0102] Stream switch 626 can also be used to interface to subsystems such as PL310 and / or NoC308. In general, stream switch 626 can be programmed to operate as a circuit-switched stream interconnect or a packet-switched stream interconnect. A circuit-switched stream interconnect can implement a point-to-point dedicated stream suitable for high-bandwidth communication between DPEs. A packet-switched stream interconnect enables sharing of streams and time-division multiplexing of multiple logic streams onto one physical stream for medium-bandwidth communication.

[0103] The stream switch 626 may include a configuration register (abbreviated as "CR" in FIG. 6) 634. Configuration data may be written to the configuration register 634 via the memory-mapped network of the DPE interconnect 606. The configuration data loaded into the configuration register 634 determines which other DPEs and / or subsystems (e.g., NoC 308, PL 310, and / or PS 312) the DPE 304 communicates with, and whether such communication is established as a circuit-switched point-to-point connection or as a packet-switched connection.

[0104] It should be understood that the number of stream interfaces shown in FIG. 6 is for illustrative purposes only and not limiting. In other embodiments, the stream switch 626 may include fewer stream interfaces. In certain embodiments, the stream switch 626 may include more stream interfaces to facilitate connections to other components and / or subsystems within the device. For example, additional stream interfaces may be coupled to other non-proximate DPEs such as DPE 304-24, 304-26, 304-4, and / or 304-6. In one or more other embodiments, the stream switch may include stream interfaces to couple a DPE such as DPE 304-15 to other DPEs located separately by one or more DPEs. For example, one or more stream interfaces may be included that allow DPE 304-15 to directly couple to a stream interface within DPE 304-13, DPE 304-16, or other non-proximate DPEs.

[0105] The second network of DPE interconnect 606 is formed by a memory - mapped switch 632. The memory - mapped switch 632 includes a plurality of memory - mapped interfaces (abbreviated as "MMI" in FIG. 6). In one or more embodiments, each memory - mapped interface may include one or more masters (e.g., master interface or output) and / or one or more slaves (e.g., slave interface or input). Each master can be an independent output having a specific bit width. For example, each master included in a memory - mapped interface may be an independent AXI master. Each slave can be an independent input having a specific bit width. For example, each slave included in a memory - mapped interface may be an independent AXI slave.

[0106] In the example of FIG. 6, the memory - mapped switch 632 includes memory - mapped interfaces 620, 622, 642, 644, and 646. It should be understood that the memory - mapped switch 632 may include additional or fewer memory - mapped interfaces. For example, for each component of the DPE that can be read and / or written using the memory - mapped switch 632, the memory - mapped switch 632 may include a memory - mapped interface coupled to such a component. Further, to facilitate reading and / or writing of memory addresses, the component itself may include a memory - mapped interface that is coupled to the corresponding memory - mapped interface within the memory - mapped switch 632.

[0107] The memory-mapped interfaces 620 and 622 can be used to exchange configuration, control, and debug data for the DPE304. In the example of FIG. 6, the memory-mapped interface 620 can receive configuration data used to configure the DPE304. The memory-mapped interface 620 can receive configuration data from a DPE located below the DPE304 and / or from the SoC interface block 104. The memory-mapped interface 622 can transfer the configuration data received by the memory-mapped interface 620 to one or more other DPEs above the DPE304, the core 602 (e.g., the program memory 608 and / or the configuration register 624), the memory module 604 (e.g., the memory and / or the configuration register 636 within the memory module 604), and / or the configuration register 634 within the stream switch 626.

[0108] In certain embodiments, the memory-mapped interface 620 communicates with the tile of the lower DPE or the SoC interface block 104 as described herein. The memory-mapped interface 622 communicates with the upper DPE. Referring again to FIG. 3 and using the DPE304-15 as a reference point, the memory-mapped interface 620 can be coupled to and communicate with another memory-mapped interface located at the DPE interconnect of the DPE304-5. The memory-mapped interface 622 can be coupled to and communicate with another memory-mapped interface located at the DPE interconnect of the DPE304-25. In one or more embodiments, the memory-mapped switch 632 transfers control and / or debug data from south to north. In other embodiments, the memory-mapped switch 632 can also pass data from north to south.

[0109] The memory-mapped interface 646 can be coupled to a memory-mapped interface (not shown) within the memory module 604 to facilitate reading and / or writing of the memory within the configuration register 636 and / or the memory module 604. The memory-mapped interface 644 can be coupled to a memory-mapped interface (not shown) within the core 602 to facilitate reading and / or writing of the program memory 608 and / or the configuration register 624. The memory-mapped interface 642 can be coupled to the configuration register 634 to read and / or write to the configuration register 634.

[0110] In the example of FIG. 6, the memory-mapped switch 632 can communicate with the upper (e.g., north) and lower (e.g., south) circuits. In one or more other embodiments, the memory-mapped switch 632 includes additional memory-mapped interfaces coupled to the memory-mapped interfaces of the memory-mapped switches of the left and / or right DPEs. Using DPE304-15 as a reference point, such additional memory-mapped interfaces can be connected to the memory-mapped switches located in DPE304-14 and / or DPE304-16, thereby facilitating the communication of configuration, control, and debug data between DPEs in both the horizontal and vertical directions.

[0111] In other embodiments, the memory-mapped switch 632 may include additional memory-mapped interfaces connected to memory-mapped switches within the DPE that are diagonal to the DPE 304. For example, using DPE 304-15 as a reference point, such additional memory-mapped interfaces are coupled to memory-mapped switches located in DPEs 304-24, 304-26, 304-4, and / or 304-6, thereby facilitating diagonal communication of configuration, control, and debug information between DPEs.

[0112] The DPE interconnect 606 is coupled to the DPE interconnect and / or the SoC interface block 104 of each neighboring DPE depending on the location of the DPE 304. In summary, the DPE interconnects of the DPEs 304 form a DPE interconnect network, which may include a stream network and / or a memory-mapped network. The configuration registers of the stream switch of each DPE can be programmed by loading configuration data through the memory-mapped switch. Through configuration, the stream switch and / or the stream interface are programmed to establish connections, either packet-switched or circuit-switched, with other endpoints in any of one or more other DPEs 304 and / or the SoC interface block 104.

[0113] In one or more embodiments, the DPE array 102 is mapped to the address space of a processor system such as the PS 312. Accordingly, any configuration register and / or memory within the DPE 304 can be accessed via the memory-mapped interface. For example, the memory in the memory module 604, the program memory 608, the configuration register 624 in the core 602, the configuration register 636 in the memory module 604, and / or the configuration register 634 can be read and / or written via the memory-mapped switch 632.

[0114] In the example of FIG. 6, the memory-mapped interface can receive configuration data for the DPE 304. The configuration data can include program code to be loaded into program memory 608 (if included), configuration data for loading into configuration registers 624, 634, and / or 636, and / or data to be loaded into the memory of memory module 604 (e.g., memory banks). In the example of FIG. 6, configuration registers 624, 634, and 636 are shown as being located within the specific circuitry that they are intended to control, such as core 602, stream switch 626, and memory module 604. The example of FIG. 6 is for illustrative purposes only and shows that elements within core 602, memory module 604, and / or stream switch 626 can be programmed by loading configuration data into the corresponding configuration registers. In other embodiments, the configuration registers can be integrated within a specific region of the DPE 304, even though they control the operation of components that are distributed throughout the DPE 304.

[0115] Thus, the stream switch 626 can be programmed by loading configuration data into configuration register 634. The configuration data programs the stream switch 626 and / or stream interfaces 610 - 616 and / or 628 - 630 to operate as a circuit-switched stream interface between two different DPEs and / or other subsystems, or as a packet-switched stream interface coupled to a selected DPE and / or other subsystem. Thus, the connections established by the stream switch 626 to other stream interfaces are programmed within the DPE 304 by loading appropriate configuration data into configuration register 634 to establish an actual connection or application data path with other DPEs and / or other subsystems of device 100.

[0116] FIG. 7 shows an example of connectivity among a plurality of DPE304s. In the example of FIG. 7, the architecture shown in FIG. 6 is used to implement each of DPE304-14, 304-15, 304-24, and 304-25. FIG. 7 shows an embodiment in which stream interfaces are interconnected (on each side and vertically and horizontally) among neighboring DPEs, and memory-mapped interfaces are connected to the upper and lower DPEs. For the sake of explanation, stream switches and memory-mapped switches are not shown.

[0117] As noted, in other embodiments, additional memory-mapped interfaces can be included to couple DPEs in a vertical and horizontal direction as shown. Further, the memory-mapped interfaces can support two-way communication in the vertical and / or horizontal direction.

[0118] Memory-mapped interfaces 620 and 622 can implement a shared transaction exchange network through which transactions propagate from a memory-mapped switch to a memory-mapped switch. Each of the memory-mapped switches can dynamically route transactions, for example, based on an address. A transaction can be stalled at any given memory-mapped switch. Memory-mapped interfaces 620 and 622 enable other subsystems of device 100 to access the resources (e.g., components) of DPE304.

[0119] In certain embodiments, subsystems of device 100 can read the internal state of any register and / or memory element of a DPE via memory-mapped interfaces 620 and / or 622. Via memory-mapped interfaces 620 and / or 622, subsystems of device 100 can read and / or write to program memory 608 and any configuration register within DPE304.

[0120] Stream interfaces 610-616 (such as stream switch 626) can provide deterministic throughput with guaranteed fixed latency from source to destination. In one or more embodiments, stream interfaces 610 and 614 can receive four 32-bit streams and output four 32-bit streams. In one or more embodiments, stream interface 614 can receive four 32-bit streams and output six 32-bit streams. In certain embodiments, stream interface 616 can receive four 32-bit streams and output four 32-bit streams. The number of streams and the size of the streams for each stream interface are provided for illustrative purposes and are not intended to be limiting.

[0121] FIG. 8 illustrates further aspects of the exemplary architecture of FIG. 6. In the example of FIG. 8, details regarding the DPE interconnect 606 are not shown. FIG. 8 shows the connectivity of core 602 with other DPEs via shared memory. FIG. 8 also shows additional aspects of memory module 604. For purposes of explanation, FIG. 8 refers to DPE 304-15.

[0122] As shown, memory module 604 includes a plurality of memory interfaces 802, 804, 806, and 808. In FIG. 8, memory interfaces 802 and 808 are abbreviated as "MI". Memory module 604 further includes a plurality of memory banks 812-1 to 812-N. In certain embodiments, memory module 604 includes eight memory banks. In other embodiments, memory module 604 may include fewer or more memory banks 812. In one or more embodiments, each memory bank 812 is single port, thereby allowing a maximum of one access to each memory bank per clock cycle. If memory module 604 includes eight memory banks 812, such a configuration supports eight parallel accesses per clock cycle. In other embodiments, each memory bank 812 is dual port or multi port, thereby allowing more parallel accesses per clock cycle.

[0123] In one or more embodiments, memory module 604 can support one or more error detection and / or error correction mechanisms. For example, memory bank 812 can be implemented to support parity checking by adding parity bits. In another example, memory bank 812 can be ECC memory that can detect and correct various types of data corruption. In another example, memory bank 812 can support both ECC and parity checking. The different types of error detection and / or error correction described herein are provided for illustrative purposes and are not intended to limit the described embodiments. Other error detection and / or error correction techniques may be used with memory module 604 other than those enumerated.

[0124] In one or more other embodiments, the error detection and / or error correction mechanism may be implemented for each memory bank 812. For example, one or more of the memory banks 812 may include parity checking, and one or more others of the memory banks 812 may be implemented as ECC memory. Further, others of the memory banks 812 may support both ECC and parity checking. Thus, different combinations of error detection and / or error correction may be supported by different memory banks 812 and / or combinations of memory banks 812.

[0125] In the example of FIG. 8, each of the memory banks 812-1 to 812-N has an arbiter 814-1 to 814-N, respectively. Each of the arbiters 814 is capable of generating a stall signal in response to detecting a conflict. Each arbiter 814 can include arbitration logic. Further, each arbiter 814 can include a crossbar. Thus, any master can write to any particular one or more of the memory banks 812. As described in connection with FIG. 6, the memory module 604 can include a memory-mapped interface (not shown) that communicates with the memory-mapped interface 646 of the memory-mapped switch 632. The memory-mapped interface within the memory module 604 can be connected to the communication lines within the memory module 604 that couple the DMA engine 816, the memory interfaces 802, 804, 806, and 808, and the arbiter 814, for reading and / or writing to the memory bank 812.

[0126] Memory module 604 further includes a direct memory access (DMA) engine 816. In one or more embodiments, DMA engine 816 includes at least two interfaces. For example, one or more interfaces can receive an input data stream from DPE interconnect 606 and write the received data to memory bank 812. One or more other interfaces can read data from memory bank 812 and transmit that data via the stream interface of DPE interconnect 606. For example, DMA engine 816 can include stream interface 640 of FIG. 6.

[0127] Memory module 604 can operate as shared memory that can be accessed by multiple different DPEs. In the example of FIG. 8, memory interface 802 is coupled to core 602 via core interface 828 included in core 602. Memory interface 802 provides core 602 access to memory bank 812 via arbiter 814. Memory interface 804 is coupled to the core of DPE304-25. Memory interface 804 provides the core of DPE304-25 access to memory bank 812. Memory interface 806 is coupled to the core of DPE304-16. Memory interface 806 provides the core of DPE304-16 access to memory bank 812. Memory interface 808 is coupled to the core of DPE304-5. Memory interface 808 provides the core of DPE304-5 access to memory bank 812. Thus, in the example of FIG. 8, each DPE having a shared boundary with memory module 604 of DPE304-15 can read and write to memory bank 812. In the example of FIG. 8, the core of DPE304-14 does not have direct access to memory module 604 of DPE304-15.

[0128] The memory-mapped switch 632 can write data to the memory bank 812. For example, the memory-mapped switch 632 can be located in the memory module 604 and then coupled to a memory-mapped interface (not shown) that is coupled to the arbiter 814. Thus, specific data stored in the memory module 604 may be controlled, for example, written as part of a configuration, control, and / or debug process.

[0129] The core 602 can access the memory modules of other neighboring DPEs via the core interfaces 830, 832, and 834. In the example of FIG. 8, the core interface 834 is coupled to the memory interface of DPE304-25. Thus, the core 602 can access the memory module of DPE304-25 via the core interface 834 and the memory interface included within the memory module of DPE304-25. The core interface 832 is coupled to the memory interface of DPE304-14. Thus, the core 602 can access the memory module of DPE304-14 via the core interface 832 and the memory interface included within the memory module of DPE304-14. The core interface 830 is coupled to the memory interface within DPE304-5. Thus, the core 602 can access the memory module of DPE304-5 via the core interface 830 and the memory interface included within the memory module of DPE304-5. As described above, the core 602 can access the memory module 604 within DPE304-15 via the core interface 828 and the memory interface 802.

[0130] In the example of FIG. 8, core 602 can read and write to any of the memory modules of DPEs (e.g., DPEs 304-25, 304-14, and 304-5) that share a boundary with core 602 in DPE 304-15. In one or more embodiments, core 602 can view the memory modules within DPEs 304-25, 304-15, 304-14, and 304-5 as a single contiguous memory. Based on this contiguous memory model, core 602 can generate addresses for reading and writing. Core 602 can direct read and / or write requests to the appropriate core interfaces 828, 830, 832, and / or 834 based on the generated addresses.

[0131] In one or more other embodiments, memory module 604 includes additional memory interfaces that can be coupled to other DPEs. For example, memory module 604 can include memory interfaces that are coupled to the cores of DPEs 304-24, 304-26, 304-4, and / or 304-5. In one or more other embodiments, memory module 604 can include one or more memory interfaces that are used to connect to cores of DPEs that are not neighboring DPEs. For example, such additional memory interfaces can be connected to cores of DPEs that are separated from DPE 304-15 by one or more other DPEs in the same row, the same column, or a diagonal direction. Thus, the number of memory interfaces within memory module 604, and the particular DPEs to which such memory interfaces are connected, as shown in FIG. 8, are for illustration and not limitation.

[0132] As described above, the core 602 can map read and / or write operations in the correct direction via the core interfaces 828, 830, 832, and / or 834 based on the addresses of such operations. When the core 602 generates an address for a memory access, the core 602 can decode the address to determine the direction (e.g., the specific DPE to be accessed), and transfer the memory operation to the correct core interface in the determined direction.

[0133] Accordingly, the core 602 can communicate with the cores of DPE304-25 via shared memory that can be the memory module within DPE304-25 and / or the memory module 604 of DPE304-15. The core 602 can communicate with the core of DPE304-14 via shared memory that is the memory module within DPE304-14. The core 602 can communicate with the cores of DPE304-5 via shared memory that can be the memory module within DPE304-5 and / or the memory module 604 of DPE304-15. Further, the core 602 can communicate with the core of DPE304-16 via shared memory that is the memory module 604 within DPE304-15.

[0134] As discussed, the DMA engine 816 may include one or more stream - memory interfaces (e.g., stream interface 640). Via the DMA engine 816, application data can be received from other sources within the device 100 and stored in the memory module 604. For example, the data can be received by the stream switch 626 from other DPEs that share and / or do not share a boundary with the DPE304 - 15. The data can also be received by the SoC interface block 104 from other subsystems of the device 100 (e.g., NoC308, hard - wired circuit blocks, PL310, and / or PS312) via the stream switch of the DPE. The DMA engine 816 can receive such data from the stream switch and write the data to one or more appropriate memory banks 812 within the memory module 604.

[0135] The DMA engine 816 may include one or more memory - stream interfaces (e.g., stream interface 630). Through the DMA engine 816, data can be read from one or more memory banks 812 of the memory module 604 and transmitted to other destinations via the stream interface. For example, the DMA engine 816 can read data from the memory module 604 and transmit such data to other DPEs that share and / or do not share a boundary with the DPE304 - 15 by the stream switch. The DMA engine 816 can also transmit such data to other subsystems (e.g., NoC308, hard - wired circuit block PL310, and / or PS312) via the stream switch and the SoC interface block 104.

[0136] In one or more embodiments, the DMA engine 816 can be programmed by the memory-mapped switch 632 within the DPE 304-15. For example, the DMA engine 816 can be controlled by the configuration register 636. The configuration register 636 can be written using the memory-mapped switch 632 of the DPE interconnect 606. In certain embodiments, the DMA engine 816 can be controlled by the stream switch 626 within the DPE 304-15. For example, the DMA engine 816 can include a control register that can be written by the stream switch 626 to which it is connected (e.g., via the stream interface 640). Streams received via the stream switch 626 within the DPE interconnect 606 can be connected to the DMA engine 816 within the memory module 604 and / or directly to the core 602, depending on the configuration data loaded into the configuration registers 624, 634, and / or 636. Streams can be transmitted from the DMA engine 816 (e.g., the memory module 604) and / or the core 602, depending on the configuration data loaded into the configuration registers 624, 634, and / or 636.

[0137] The memory module 604 can further include a hardware synchronization circuit 820 (abbreviated as "HSC" in FIG. 8). Generally, the hardware synchronization circuit 820 can synchronize the operations of different cores (e.g., cores of neighboring DPEs), the core 602 of FIG. 8, the DMA engine 816, and other external masters (e.g., the PS 312) that can communicate via the DPE interconnect 606. By way of example and not limitation, the hardware synchronization circuit 820 can synchronize two different cores within different DPEs that access the same, e.g., shared buffer within the memory module 604.

[0138] In one or more embodiments, the hardware synchronization circuit 820 may include a plurality of different locks. The specific number of locks included in the hardware synchronization circuit 820 may depend on the number of entities that can access the memory module, but is not intended to be limiting. In certain embodiments, each different hardware lock can have an arbiter that can handle simultaneous requests. Further, each hardware lock can handle new requests every clock cycle. The hardware synchronization circuit 820 can have a plurality of requesting sides, such as cores from each of cores 602, DPEs 304-25, 304-16, and 304-5, DMA engines 816, and / or masters communicating via the DPE interconnect 606. The requesting side acquires a lock for a particular portion of memory from the local hardware synchronization circuit, for example, before accessing that particular portion of memory within the memory module. The requesting side may release the lock so that another requesting side may acquire the lock before accessing the same portion of memory.

[0139] In one or more embodiments, the hardware synchronization circuit 820 can synchronize access by multiple cores to the memory module 604, more specifically, to the memory banks 812. For example, the hardware synchronization circuit 820 can synchronize access by the cores 602 shown in FIG. 8, the cores of DPE304-25, the cores of DPE304-16, and the cores of DPE304-5 to the memory module 604 of FIG. 8. In certain embodiments, the hardware synchronization circuit 820 can synchronize access to the memory banks 812 for any core that can directly access the memory module 604 via the memory interfaces 802, 804, 806, and / or 808. Each core that can access the memory module 604 (e.g., the core 602 of FIG. 8 and the cores of one or more of the neighboring DPEs) can, for example, access the hardware synchronization circuit 820 to request and acquire a lock before accessing a specific portion of the memory within the memory module 604, and then release the lock to allow another core to access that portion of the memory when the other core acquires the lock. Similarly, the core 602 can access the hardware synchronization circuit 820, the hardware synchronization circuit within DPE304-14, the hardware synchronization circuit within DPE304-25, and the hardware synchronization circuit within DPE304-5 to request and acquire a lock to access a portion of the memory within the memory module of each respective DPE, and then release the lock. The hardware synchronization circuit 820 effectively manages the operation of the shared memory among the DPEs by coordinating and synchronizing access to the memory modules of the DPEs.

[0140] The hardware synchronization circuit 820 can also be accessed via the memory-mapped switch 632 of the DPE interconnect 606. In one or more embodiments, lock transactions are implemented as atomic acquire (e.g., test whether unlocked and set a lock) and release (e.g., unlock) operations on a resource. The locks of the hardware synchronization circuit 820 provide a way to efficiently transfer ownership of a resource between two participants. The resource can be any of a variety of circuit components, such as a buffer in local memory (e.g., a buffer in the memory module 604).

[0141] While the hardware synchronization circuit 820 can synchronize access to memory to support communication via shared memory, the hardware synchronization circuit 820 can also synchronize any of a variety of other resources and / or agents, including other DPEs and / or other cores. For example, since the hardware synchronization circuit 820 provides a shared pool of locks, the locks can be used by a DPE, e.g., a core of a DPE, to initiate and / or stop the operation of another DPE or core. The locks of the hardware synchronization circuit 820 can be assigned for different purposes, e.g., based on configuration data, such as to synchronize different agents and / or resources that may be required depending on the particular application implemented by the DPE array 102.

[0142] In certain embodiments, DPE access and DMA access to the locks of the hardware synchronization circuit 820 are blocking. Such access can stall the requesting core or DMA engine if the lock cannot be acquired immediately. When the hardware lock becomes available, the requesting core or DMA engine acquires the lock and is automatically unstalled.

[0143] In one embodiment, memory-mapped access can be made non-blocking so that a memory-mapped master can poll the status of the lock of the hardware synchronization circuit 820. For example, a memory-mapped switch can send a lock "acquire" request to the hardware synchronization circuit 820 as a normal memory read operation. The read address may encode the lock identifier and other request data. The read data, e.g., the response to the read request, can signal the success of the acquire request operation. The "acquire" sent as a memory read may be sent in a loop until successful. In another example, the hardware synchronization circuit 820 can issue an event so that a memory-mapped master receives an interrupt when the status of the requested lock changes.

[0144] Thus, when two neighboring DPEs share a data buffer via the memory module 604, the hardware synchronization circuit 820 in the specific memory module 604 containing that buffer synchronizes the access. Although not necessarily required, typically, the memory block can be double-buffered to improve throughput.

[0145] If the two DPEs are not neighboring DPEs, the two DPEs do not have access to a common memory module. In that case, application data can be transferred via a data stream (the terms "data stream" and "stream" can be used interchangeably within the present disclosure). Thus, the local DMA engine can convert that transfer from a local memory-based transfer to a stream-based transfer. In that case, the core 602 and the DMA engine 816 can synchronize using the hardware synchronization circuit 820.

[0146] The core 602 can further access the hardware synchronization circuits of nearby DPEs, such as the locks of the hardware synchronization circuits, to facilitate communication via shared memory. Thus, such other or nearby hardware synchronization circuits within the DPE can synchronize access to resources, such as memory, among the cores of the nearby DPE.

[0147] The PS312 can communicate with the core 602 via the memory-mapped switch 632. For example, the PS312 can access the memory module 604 and the hardware synchronization circuit 820 by initiating memory reads and writes. In another embodiment, the hardware synchronization circuit 820 may send an interrupt to the PS312 when the status of the lock changes to avoid polling of the hardware synchronization circuit 820 by the PS312. The PS312 can also communicate with the DPE304-15 via a stream interface.

[0148] The examples provided herein regarding entities that send memory-mapped requests and / or transfers are for illustration purposes only and not limiting. In certain embodiments, any entity external to the DPE array 102 can send memory-mapped requests and / or transfers. For example, a circuit block implemented in the PL310, an ASIC, or other circuitry external to the DPE array 102 described herein can send memory-mapped requests and / or transfers to the DPE304 and access the hardware synchronization circuits of the memory modules within such DPEs.

[0149] In addition to communicating with neighboring DPEs via a shared memory module and with neighboring and / or non-neighboring DPEs via DPE interconnect 606, core 602 may include a cascade interface. In the example of FIG. 8, core 602 includes cascade interfaces 822 and 824 (abbreviated as "CI" in FIG. 8). Cascade interfaces 822 and 824 can provide direct communication with other cores. As shown, cascade interface 822 of core 602 directly receives an input data stream from the core of DPE 304-14. The data stream received via cascade interface 822 can be provided to data processing circuitry within core 602. Cascade interface 824 of core 602 can directly transmit an output data stream to the core of DPE 304-16.

[0150] In the example of FIG. 8, each of cascade interface 822 and cascade interface 824 may include a first-in first-out (FIFO) interface for buffering. In certain embodiments, cascade interfaces 822 and 824 are capable of transmitting data streams that can be hundreds of bits wide. The particular bit widths of cascade interfaces 822 and 824 are not intended to be limiting. In the example of FIG. 8, cascade interface 824 is coupled to accumulator register 836 (abbreviated as "AC" in FIG. 8) within core 602. Cascade interface 824 can output the contents of accumulator register 836 and can do so on a per clock cycle basis. Accumulator register 836 can store data generated and / or processed by data processing circuitry within core 602.

[0151] In the example of FIG. 8, the cascade interfaces 822 and 824 can be programmed based on the configuration data loaded into the configuration register 624. For example, based on the configuration register 624, the cascade interface 822 can be activated or deactivated. Similarly, based on the configuration register 624, the cascade interface 824 can be activated or deactivated. The cascade interface 822 can be activated and / or deactivated independently of the cascade interface 824.

[0152] In one or more other embodiments, the cascade interfaces 822 and 824 are controlled by the core 602. For example, the core 602 can include instructions for reading and writing to the cascade interfaces 822 and / or 824. In another example, the core 602 can include a hardwired circuit that enables reading and / or writing to the cascade interfaces 822 and / or 824. In certain embodiments, the cascade interfaces 822 and 824 can be controlled by an entity external to the core 602.

[0153] In the embodiments described within the present disclosure, the DPE 304 does not include a cache memory. By omitting the cache memory, the DPE array 102 can achieve predictable, for example deterministic, performance. Further, since there is no need to maintain coherence between cache memories located in different DPEs, a large processing overhead is avoided.

[0154] According to one or more embodiments, the core 602 of the DPE 304 does not have input interrupts. Thus, the core 602 of the DPE 304 can operate without being interrupted. By omitting the input interrupts to the core 602 of the DPE 304, the DPE array 102 can also achieve predictable, for example deterministic, performance.

[0155] When one or more DPE304 communicate with an external agent implemented in PS312, PL310, a hardwired circuit block, and / or another subsystem of device 100 (such as an ASIC) via a shared buffer in an external read-write (such as DDR) memory, a coherence mechanism may be implemented using the coherence interconnect in PS312. In these scenarios, the application data transfer between the DPE array 102 and the external agent may cross both the NoC308 and / or the PL310.

[0156] In one or more embodiments, the DPE array 102 may be functionally separated into a plurality of groups consisting of one or more DPEs. For example, a specific memory interface may be enabled and / or disabled via configuration data to create a group of one or more DPEs, and each group may include one or more (such as a subset) of the DPEs of the DPE array 102. In another example, the stream interface may be configured independently for each group to communicate with other cores of the DPEs within the group and / or a specified input source and / or output destination.

[0157] In one or more embodiments, the core 602 can support debug functions via a memory-mapped interface. As discussed, the program memory 608, memory module 604, core 602, DMA engine 816, stream switch 626, and other components of the DPE are memory-mapped. The memory-mapped registers may be read and / or written by any source capable of generating memory-mapped requests, such as PS312, PL310, and / or a platform management controller within the IC. The requests may proceed through the SoC interface block 104 to the intended or targeted DPE within the DPE array 102.

[0158] Functions such as core suspension, core resumption, core single-stepping, and / or core reset can be executed through the memory-mapped switches within the DPE. Further, such operations can be initiated for multiple different DPEs. Other exemplary debug operations that can be performed include, for example, reading the status of the hardware synchronization circuit 820 and / or the DMA engine 816 via the memory-mapped interface described herein, and / or setting its state.

[0159] In one or more embodiments, the stream interface of the DPE can generate trace information that can be output from the DPE array 102. For example, the stream interface may be configured to extract trace information from the DPE array 102. The trace information can be generated as a packet-switched stream that includes timestamped data marking event occurrences and / or a limited branch trace of the execution flow. In one aspect, the trace generated by the DPE can be pushed to the local trace buffer implemented in the PL310 or to external RAM using the SoC interface block 104 and the NoC 308. In another aspect, the trace generated by the DPE can be sent to an on-chip implemented debug subsystem.

[0160] In certain embodiments, each core 602 and memory module 604 of each DPE may include an additional stream interface that can directly output trace data to the stream switch 626. The stream interface for trace data may be in addition to those already described. The stream switch 626 may be configured to direct trace data onto a packet-switched stream such that trace information from multiple cores and memory modules of different DPEs can move on a single data stream. As noted, the stream portion of the DPE interconnect network can be configured to send trace data directly to the on-chip debug system via the PL310, to external memory via the SoC interface block 104, or to a gigabit transceiver via the NoC 308. Examples of different types of trace streams that can be generated include a PC trace stream that generates PC values at branch instructions as opposed to each change in the program counter (PC), as well as an application data trace stream that includes intermediate results within the DPE (e.g., via respective trace data streams from the core and / or memory module).

[0161] FIG. 9 is a diagram showing exemplary connectivity of a cascade interface of cores in multiple DPEs. In the example of FIG. 9, only the cores 602 of the DPEs are shown. Other parts of the DPE such as the DPE interconnect and memory modules are omitted for purposes of illustration.

[0162] As shown, the cores are connected in series via the cascade interface described in connection with FIG. 8. Core 602-1 is coupled to core 602-2, core 602-2 is coupled to core 602-3, and core 602-3 is coupled to core 602-4. Thus, application data can propagate directly from core 602-1 to cores 602-2, 602-3, and 602-4. Core 602-4 is coupled to core 602-8 in the row above. Core 602-8 is coupled to core 602-7, core 602-7 is coupled to core 602-6, and core 602-6 is coupled to core 602-5. Thus, application data can propagate directly from core 602-4 to cores 608-8, 602-7, 602-6, and 602-5. Core 602-5 is coupled to core 602-9 in the row above. Core 602-9 is coupled to core 602-10, core 602-10 is coupled to core 602-11, and core 602-11 is coupled to core 602-12. Thus, application data can propagate directly from core 602-5 to cores 602-9, 602-10, 602-11, and 602-12. Core 602-12 is coupled to core 602-16 in the row above. Core 602-16 is coupled to core 602-15, core 602-15 is coupled to core 602-14, and core 602-14 is coupled to core 602-13. Thus, application data can propagate directly from core 602-12 to cores 608-16, 602-15, 602-14, and 602-13.

[0163] FIG. 9 is intended to illustrate how the cascade interface of the cores of the DPE can be coupled from one row of DPEs to another row of DPEs within the DPE array. The specific number of columns and / or rows of cores (e.g., DPEs) shown is not intended to be limiting. FIG. 9 shows that the connections between cores using the cascade interface can be made in an "S" or zigzag pattern at alternating ends of the rows of DPEs.

[0164] In an embodiment where the DPE array 102 realizes two or more different clusters consisting of DPEs 304, the first cluster of DPEs may not be coupled to the second cluster of DPEs via a cascade and / or stream interface. For example, if the first two rows of DPEs form the first cluster and the second two rows of DPEs form the second cluster, the cascade interface of core 602-5 can be programmed to be disabled so as not to pass data to the cascade input of core 602-9.

[0165] In the examples described in connection with FIGS. 8 and 9, each core is shown as having a cascade interface operating as an input and a cascade interface operating as an output. In one or more other embodiments, the cascade interface can be implemented as a bidirectional interface. In certain embodiments, the core may include additional cascade interfaces so that the core can communicate directly with other cores above, below, to the left, and / or to the right via the cascade interface. As noted, such interfaces may be unidirectional or bidirectional.

[0166] Figures 10A, 10B, 10C, 10D, and 10E illustrate examples of connectivity between DPEs. Figure 10A illustrates an example of connectivity between DPEs using shared memory. In the example of Figure 10A, a function or kernel implemented within core 602-15 (e.g., a user circuit design implemented in a DPE and / or DPE array) processes data 1005, e.g., application data, and places it in memory module 604-15 using the core interface and memory interface within DPE 304-15. DPE 304-15 and DPE 304-16 are neighboring DPEs. Thus, core 602-16 can access data 1005 from memory module 604-15 based on obtaining a lock from a hardware synchronization circuit (not shown) within memory module 604-15 for a buffer containing data 1005. The shared access to memory module 604-15 by cores 602-15 and 602-16 facilitates high-speed transaction processing because there is no need for data to be physically transferred from one memory to another for core 602-16 to process application data.

[0167] FIG. 10B shows an example of connectivity between DPEs using a stream switch. In the example of FIG. 10B, DPEs 304-15 and 304-17 are non-neighboring DPEs and are thus separated by one or more intervening DPEs. The function or kernel implemented within core 602-15 processes data 1005 and places it in memory module 604-15. The DMA engine 816-15 of memory module 604-15 retrieves data 1005 based on acquiring a lock for the buffer used to store data 1005 within memory module 604-15. The DMA engine 816-15 transmits data 1005 to DPE 304-17 via the stream switch of the DPE interconnect. The DMA engine 816-17 within memory module 604-17 can retrieve data 1005 from the stream switch within DPE 304-17 and store data 1005 within the buffer of memory module 604-17 after acquiring a lock from the hardware synchronization circuit within memory module 604-17 for the buffer within memory module 604-17. The connectivity shown in FIG. 10B can be programmed by loading configuration data so that the respective stream switches within DPEs 304-15 and 304-17 as well as DMA engines 816-15 and 816-17 are configured to operate as described.

[0168] Figure 10C shows another example of connectivity between DPEs using a stream switch. In the example of Figure 10C, DPEs 304-15 and 304-17 are non-neighboring DPEs and are thus separated by one or more intervening DPEs. Figure 10C shows that data 1005 can be provided directly from DMA816-15 to the core of another DPE via the stream switch. As shown, DMA816-15 places data 1005 on the stream switch of DPE 304-15. Core 602-17 can receive data 1005 directly from the stream switch within DPE 304-17 using the stream interface contained therein, and data 1005 does not traverse to memory module 604-17. The connectivity shown in Figure 10C can be programmed by loading configuration data so that the respective stream switches of DPEs 304-15 and 304-17 as well as DMA816-15 are configured to operate as described.

[0169] Generally, Figure 10C shows an example of DMA-core transfer of data. It should be understood that core-DMA transfer of data can also be realized. For example, core 602-17 can send data to DPE 304-15 via the stream interface contained therein and the stream switch of DPE 304-17. DMA engine 816-15 can retrieve data from the stream switch contained in DPE 304-15 and store it in memory module 604-15.

[0170] FIG. 10D shows another example of connectivity between DPEs using a stream switch. Referring to FIG. 10D, the cores 602-15, 602-17, and 602-19 of each different, non-neighboring DPE can communicate directly with each other via the respective DPE's stream interface. In the example of FIG. 10D, core 602-15 can broadcast the same data stream to core 602-17 and core 602-19. The broadcast function of the stream interface within each respective DPE, including cores 602-15, 602-17, and 602-19, can be programmed by loading configuration data to configure each respective stream switch and / or stream interface as described. In one or more other embodiments, core 602-15 can multicast data to the cores of other DPEs.

[0171] FIG. 10E shows an example of connectivity between DPEs using a stream switch and a cascade interface. Referring to FIG. 10E, DPE304-15 and DPE304-16 are neighboring DPEs. In some cases, a kernel can be split to operate on multiple cores. In that case, the intermediate cumulative results within one sub-kernel can be transferred to the sub-kernel of the next core via the cascade interface.

[0172] In the example of FIG. 10E, core 602-15 receives data 1005 via a stream switch and processes the data 1005. Core 602-15 generates intermediate result data 1010 and directly outputs it from the accumulator register in core 602-15 to core 602-16 via a cascade interface. In certain embodiments, the cascade interface of core 602-15 can transfer the accumulator value every clock cycle of DPE304-15. The data 1005 received by core 602-15 is further propagated to core 602-16 via a stream switch in the DPE interconnect, and core 602-16 can process both the data 1005 (e.g., the original data) and the intermediate result data 1010 generated by core 602-15.

[0173] In the example of FIG. 10, the transmission, broadcast, and / or multicast of data streams are shown horizontally. It should be understood that a data stream can be transmitted, broadcast, and / or multicast from one DPE to any other DPE within the DPE array. Thus, a data stream can be transmitted, broadcast, or multicast to a DPE left, right, up, down, and / or diagonally as required to reach each such DPE based on the configuration data loaded into the intended destination DPE.

[0174] FIG. 11 shows an example of an event processing circuit within a DPE. A DPE can include an event processing circuit interconnected to the event processing circuits of other DPEs. In the example of FIG. 11, the event processing circuit is implemented within core 602 and memory module 604. Core 602 can include an event broadcast circuit 1102 and event logic 1104. Memory module 604 can include a separate event processing circuit including an event broadcast circuit 1106 and event logic 1108.

[0175] The event broadcast circuit 1102 can be connected to the event broadcast circuits within each of the cores of the DPEs in the vicinity of the upper and lower sides of the exemplary DPE shown in FIG. 11. The event broadcast circuit 1102 can also be connected to the event broadcast circuits within the memory modules of the DPEs in the vicinity of the left side of the exemplary DPE shown in FIG. 11. As shown, the event broadcast circuit 1102 is connected to the event broadcast circuit 1106. The event broadcast circuit 1106 can be connected to the event broadcast circuits within each of the memory modules of the DPEs in the vicinity of the upper and lower sides of the exemplary DPE shown in FIG. 11. The event broadcast circuit 1106 can also be connected to the event broadcast circuits within the cores of the DPEs in the vicinity of the right side of the exemplary DPE shown in FIG. 11.

[0176] In this way, the event processing circuit of the DPE can form an independent event broadcast network within the DPE array. The event broadcast network within the DPE array can exist independently of the DPE interconnect network. Further, the event broadcast network can be individually configurable by loading suitable configuration data into the configuration registers 624 and / or 636.

[0177] In the example of FIG. 11, the event broadcast circuit 1102 and the event logic 1104 can be configured by the configuration register 624. The event broadcast circuit 1106 and the event logic 1108 can be configured by the configuration register 636. The configuration registers 624 and 636 can be written via the memory-mapped switch of the DPE interconnect 606. In the example of FIG. 11, the configuration register 624 programs the event logic 1104 to detect a specific type of event occurring within the core 602. The configuration data loaded into the configuration register 624 determines, for example, which of a plurality of different types of predetermined events are detected by the event logic 1104. Examples of events can include, but are not limited to, the start and / or end of a read operation by the core 602, the start and / or end of a write operation by the core 602, a stall, and the occurrence of other operations performed by the core 602. Similarly, the configuration register 636 programs the event logic 1108 to detect a specific type of event occurring within the memory module 604. Examples of events can include, but are not limited to, the start and / or end of a read operation by the DMA engine 816, the start and / or end of a write operation by the DMA engine 816, a stall, and the occurrence of other operations performed by the memory module 604. The configuration data loaded into the configuration register 636 determines, for example, which of a plurality of different types of predetermined events are detected by the event logic 1108. It should be understood that the event logic 1104 and / or the event logic 1108 can detect events arising from and / or related to the DMA engine 816, the memory-mapped switch 632, the stream switch 626, the memory interface of the memory module 604, the core interface of the core 602, the cascade interface of the core 602, and / or other components located within the DPE.

[0178] The configuration register 624 can further program the event broadcast circuit 1102, and the configuration register 636 can program the event broadcast circuit 1106. For example, the configuration data loaded into the configuration register 624 can determine which of the events received by the event broadcast circuit 1102 from other event broadcast circuits are further propagated to another event broadcast circuit and / or to the SoC interface block 104. The configuration data can also specify which of the events generated internally by the event logic 1104 are propagated to another event broadcast circuit and / or to the SoC interface block 104.

[0179] Similarly, the configuration data loaded into the configuration register 636 can determine which of the events received by the event broadcast circuit 1106 from other event broadcast circuits are further propagated to another event broadcast circuit and / or to the SoC interface block 104. The configuration data can also specify which of the events generated internally by the event logic 1108 are propagated to another event broadcast circuit and / or to the SoC interface block 104.

[0180] Therefore, the events generated by the event logic 1104 may be provided to the event broadcast circuit 1102 or broadcast to other DPEs. In the example of FIG. 11, the event broadcast circuit 1102 can broadcast events to the upper DPE, the left DPE, and the lower DPE or the SoC interface block 104, regardless of whether the events are generated internally or received from other DPEs. The event broadcast circuit 1102 can also broadcast events to the event broadcast circuit 1106 within the memory module 604.

[0181] The events generated by event logic 1108 may be provided to event broadcast circuit 1106 and may be broadcast to other DPEs. In the example of FIG. 11, the events can be broadcast to the upper DPE, the right DPE, and the lower DPE or SoC interface block 104, regardless of whether the events are generated internally or received from other DPEs. Event broadcast circuit 1106 can also broadcast events to event broadcast circuit 1102 within core 602.

[0182] In the example of FIG. 11, the event broadcast circuits located in the cores communicate vertically with the event broadcast circuits located in the cores of the DPEs in the upper and / or lower vicinity. If a DPE is directly above (or adjacent to) SoC interface block 104, the event broadcast circuit within the core of that DPE can communicate with SoC interface block 104. Similarly, the event broadcast circuits located in the memory modules communicate vertically with the event broadcast circuits located in the memory modules of the DPEs in the upper and / or lower vicinity. If a DPE is directly above (e.g., adjacent to) SoC interface block 104, the event broadcast circuit within the memory module of that DPE can communicate with SoC interface block 104. The event broadcast circuits can further communicate with the event broadcast circuits directly to the left and / or right, regardless of whether such event broadcast circuits are located in another DPE and / or within a core or memory module.

[0183] When configuration registers 624 and 636 are written, event logics 1104 and 1108 can operate in the background. In certain embodiments, event logic 1104 generates events only in response to detecting certain conditions within core 602; event logic 1108 generates events only in response to detecting certain conditions within memory module 604.

[0184] Figure 12 shows another exemplary architecture of DPE304. In the example of Figure 12, DPE304 includes a plurality of different cores and can be referred to as a "cluster" type DPE architecture. In Figure 12, DPE304 includes cores 1202, 1204, 1206, and 1208. Each of the cores 1202 - 1208 is connected to a memory pool 1220 via core interfaces 1210, 1212, 1214, 1216 (abbreviated as "core IF" in Figure 12), respectively. Each of the core interfaces 1210 - 1216 is coupled to a plurality of memory banks 1222-1 to 1222-N via a crossbar 1224. Through the crossbar 1224, any one of the cores 1202 - 1208 can access any one of the memory banks 1222-1 to 1222-N. Thus, within the exemplary architecture of Figure 12, the cores 1202 - 1208 can communicate with each other via the shared memory banks 1222 of the memory pool 1220.

[0185] In one or more embodiments, the memory pool 1220 may include 32 memory banks. The number of memory banks included in the memory pool 1220 is provided for illustrative purposes only and is not limiting. In other embodiments, the number of memory banks included in the memory pool 1220 may be more than 32 or less than 32.

[0186] In the example of FIG. 12, DPE 304 includes a memory-mapped switch 1226. The memory-mapped switch 1226 includes a plurality of memory-mapped interfaces (not shown) that can be coupled to memory-mapped switches in neighboring DPEs in each of the basic four compass directions (e.g., north, south, west, east) and to the memory pool 1220. Each memory-mapped interface can include one or more masters and one or more slaves. For example, the memory-mapped switch 1226 is coupled to the crossbar 1224 via a memory-mapped interface. The memory-mapped switch 1226 is capable of transmitting configuration, control, and debug data as described in connection with other exemplary DPEs within the present disclosure. Thus, the memory-mapped switch 1226 can load configuration registers (not shown) in DPE 304. In the example of FIG. 12, DPE 304 can include configuration registers for controlling the operation of the stream switch 1232, cores 1202-1208, and DMA engine 1234.

[0187] In the example of FIG. 12, the memory-mapped switch 1226 can communicate in each of the four basic compass directions. In other embodiments, the memory-mapped switch 1226 can communicate only in the north and south directions. In other embodiments, the memory-mapped switch 1226 can include additional memory-mapped interfaces that enable the memory-mapped switch 1226 to communicate with more than four other entities, thereby enabling communication with other diagonal DPEs and / or other non-neighboring DPEs.

[0188] DPE304 also includes stream switch 1232. Stream switch 1232 includes a plurality of stream interfaces (not shown) that can be coupled to stream switches in neighboring DPEs in each of the basic four compass directions (e.g., north, south, west, east), and to cores 1202 - 1208. Each stream interface can include one or more masters and one or more slaves. Stream switch 1232 further includes a stream interface coupled to DMA engine 1234.

[0189] DMA engine 1234 is coupled to crossbar 1224 via interface 1218. DMA engine 1234 can include two interfaces. For example, DMA engine 1234 can include a memory - stream interface that can read data from one or more of memory banks 1222 and transmit that data on stream switch 1232. DMA engine 1234 can also include a stream - memory interface that can receive data via stream switch 1232 and store that data in one or more of memory banks 1222. Each of the interfaces can support one input / output stream or multiple simultaneous input / output streams, regardless of whether it is a memory - stream or a stream - memory interface.

[0190] The exemplary architecture of FIG. 12 supports DPE - to - DPE communication via both memory - mapped switch 1226 and stream switch 1232. As shown, memory - mapped switch 1226 can communicate with memory - mapped switches in neighboring DPEs above, below, to the left, and to the right. Similarly, stream switch 1232 can communicate with stream switches in neighboring DPEs above, below, to the left, and to the right.

[0191] In one or more embodiments, both the memory-mapped switch 1226 and the stream switch 1232 can support data transfer between cores of other DPEs (both in-neighborhood and out-of-neighborhood) to share application data. The memory-mapped switch 1226 can further support the transfer of configuration, control, and debug data for the purpose of configuring the DPE 304. In certain embodiments, the stream switch 1232 supports the transfer of application data, and the memory-mapped switch 1226 supports only the transfer of configuration, control, and debug data.

[0192] In the example of FIG. 12, cores 1202-1208 are connected in series via a cascade interface as described above. Further, core 1202 is coupled to the cascade interface (e.g., output) of the rightmost core within the DPE in the in-neighborhood to the left of the DPE of FIG. 12, and core 1208 is coupled to the cascade interface (e.g., input) of the leftmost core within the DPE in the in-neighborhood to the right of the DPE of FIG. 12. The cascade interfaces of DPEs using a cluster architecture can be connected row-to-row as shown in FIG. 9. In one or more other embodiments, one or more of cores 1202-1208 can be connected via a cascade interface to cores within DPEs in the upper and / or lower in-neighborhoods instead of and / or in addition to a horizontal cascade connection.

[0193] The exemplary architecture of FIG. 12 can be used to implement a DPE and form a DPE array, as described herein. The exemplary architecture of FIG. 12 increases the amount of memory available to the cores as compared to other exemplary DPE architectures described within the present disclosure. Thus, in applications where the cores require access to larger amounts of memory, the architecture of FIG. 12 that clusters multiple cores together within a single DPE can be used. For illustration purposes, depending on the configuration of the DPE 304 of FIG. 12, it is not necessary to use all of the cores. Thus, one or more (e.g., less than all of the cores 1202 - 1208 of the DPE 304) can access the memory pool 1220 and have access to more memory than would otherwise be the case based on configuration data that would be loaded into a configuration register (not shown) in the example of FIG. 12.

[0194] FIG. 13 shows an exemplary architecture of the DPE array 102 of FIG. 1. In the example of FIG. 13, the SoC interface block 104 provides an interface between the DPE 304 and other subsystems of the device 100. The SoC interface block 104 integrates the DPE into the device. The SoC interface block 104 is capable of transmitting configuration data to the DPE 304, transmitting events from the DPE 304 to other subsystems, transmitting events from other subsystems to the DPE 304, generating interrupts and transmitting them to entities external to the DPE array 102, transmitting application data between other subsystems and the DPE 304, and / or transmitting trace and / or debug data between other subsystems and the DPE 304.

[0195] In the example of FIG. 13, the SoC interface block 104 includes a plurality of interconnected tiles. For example, the SoC interface block 104 includes tiles 1302, 1304, 1306, 1308, 1310, 1312, 1314, 1316, 1318, and 1320. In the example of FIG. 13, the tiles 1302-1320 are arranged in rows. In other embodiments, the tiles may be arranged in columns, in a grid, or in another layout. For example, the SoC interface block 104 may be implemented as a column of tiles such as to the left of the DPE 304, to the right of the DPE 304, or between columns of the DPE 304. In another embodiment, the SoC interface block 104 may be located on top of the DPE array 102. The SoC interface block 104 may be implemented such that the tiles are arranged in any combination below the DPE array 102, to the left of the DPE array 102, to the right of the DPE array 102, and / or on top of the DPE array 102. In this regard, FIG. 13 is shown for illustrative purposes and not for limitation.

[0196] In one or more embodiments, the tiles 1302-1320 have the same architecture. In one or more other embodiments, the tiles 1302-1320 may be implemented with two or more different architectures. In certain embodiments, different architectures are used to implement the tiles within the SoC interface block 104, and each different tile architecture can support communication with different types of subsystems or combinations of subsystems of the device 100.

[0197] In the example of FIG. 13, tiles 1302-1320 are coupled such that data can propagate from one tile to another. For example, data can propagate from tile 1302 through tiles 1304, 1306 and down the lines of tiles to tile 1320. Similarly, data can propagate in the reverse direction from tile 1320 to tile 1302. In one or more embodiments, each of tiles 1302-1320 can operate as an interface for a plurality of DPEs. For example, each of tiles 1302-1320 can operate as an interface for a subset of DPEs 304 of DPE array 102. The subset of DPEs for which each tile provides an interface may be mutually exclusive such that no DPE is provided an interface by more than two tiles of SoC interface block 104.

[0198] In one example, each of tiles 1302-1320 provides an interface for a column of DPEs 304. By way of example, tile 1302 provides an interface to the DPEs in column A. Tile 1304 provides an interface to the DPEs in column B. In each case, the tile includes a direct connection to an adjacent DPE within the column of DPEs, which in this example is the bottom DPE. Referring to column A, for example, tile 1302 is directly connected to DPE 304-1. Other DPEs within column A can communicate with tile 1302, however, through the DPE interconnects of the intervening DPEs within the same column.

[0199] For example, tile 1302 can receive data from PS312, PL310, and / or another source such as another hardwired circuit block, e.g., an ASIC block. Tile 1302 can provide a portion of the data addressed to the DPEs in column A to such DPEs and can send the data addressed to the DPEs in other columns (e.g., DPEs for which tile 1302 is not an interface) to tile 1304. Tile 1304 can perform the same or a similar process, providing the data addressed to the DPEs in column B received from tile 1302 to such DPEs while sending the data addressed to the DPEs in other columns to tile 1306.

[0200] In this way, data can propagate from tile to tile in the SoC interface block 104 until it reaches a tile operating as an interface for the DPE to which the data is addressed (e.g., the "target DPE"). A tile operating as an interface for the target DPE can direct the data to the target DPE using the memory-mapped switch of the DPE and / or the stream switch of the DPE.

[0201] As noted, the use of columns is an exemplary implementation. In other embodiments, each tile of the SoC interface block 104 can provide an interface to a row of DPEs in the DPE array 102. Such a configuration can be used when the SoC interface block 104 is implemented as a column of tiles, regardless of whether it is to the left, right, or between columns of DPEs 304. In other embodiments, the subset of DPEs for which each tile provides an interface can be any combination of fewer DPEs than all of the DPEs in the DPE array 102. For example, DPEs 304 can be distributed to the tiles of the SoC interface block 104. The particular physical layout of such DPEs can vary based on the connectivity of the DPEs established by the DPE interconnects. For example, tile 1302 may provide an interface to DPEs 304-1, 304-2, 304-11, and 304-12. Another tile of the SoC interface block 104 may provide an interface to four other DPEs, etc.

[0202] Figures 14A, 14B, and 14C show exemplary architectures for implementing tiles of the SoC interface block 104. Figure 14A shows an exemplary implementation of tile 1304. The architecture shown in Figure 14A can also be used to implement any other tile included in the SoC interface block 104.

[0203] Tile 1304 includes a memory-mapped switch 1402. The memory-mapped switch 1402 can include a plurality of memory-mapped interfaces for communicating in each of a plurality of different directions. By way of example and not limitation, the memory-mapped switch 1402 can include one or more memory-mapped interfaces, and the memory-mapped interface can have a master that connects perpendicularly to the memory-mapped interface of the directly above DPE. Thus, the memory-mapped switch 1402 can operate as a master with respect to the memory-mapped interfaces of one or more of the DPEs. In a particular example, the memory-mapped switch 1402 can operate as a master for a subset of the DPEs. For example, the memory-mapped switch 1402 can operate as a master for a column of DPEs on tile 1304, such as column B of FIG. 13. It should be understood that the memory-mapped switch 1402 can include additional memory-mapped interfaces for connecting to a plurality of different circuits (such as DPEs) within the DPE array 102. The memory-mapped interfaces of the memory-mapped switch 1402 can also include one or more slaves that can communicate with circuits (such as one or more DPEs) located above tile 1304.

[0204] In the example of FIG. 14A, the memory - mapped switch 1402 may include one or more memory - mapped interfaces that facilitate horizontal communication to memory - mapped switches within neighboring tiles (e.g., tiles 1302 and 1306). For illustration, the memory - mapped switch 1402 may be horizontally connected to neighboring tiles via the memory - mapped interfaces, and each such memory - mapped interface may include one or more masters and / or one or more slaves. Thus, the memory - mapped switch 1402 can move data (e.g., configuration, control, and / or debug data) from one tile to another tile, reach the correct DPE and / or the correct subset of multiple DPEs, and direct that data towards the target DPE, regardless of whether the data is in the column above tile 1304 or in another subset where another tile of the SoC interface block 104 operates as an interface. For example, when a memory - mapped transaction is received from the NoC 308, the memory - mapped switch 1402 can distribute the transaction horizontally, e.g., to other tiles within the SoC interface block 104.

[0205] The memory - mapped switch 1402 may also include a memory - mapped interface having one or more masters and / or slaves coupled to the configuration register 1436 within tile 1304. Through the memory - mapped switch 1402, configuration data can be loaded into the configuration register 1436 to control various functions and operations executed by components within tile 1304. FIGS. 14A, 14B, and 14C show the connections between the configuration register 1436 and one or more elements of tile 1304. However, it should be understood that the configuration register 1436 may control other elements of tile 1304 and thus may have connections to such other elements, but such connections are not shown in FIGS. 14A, 14B, and / or 14C.

[0206] The memory - mapped switch 1402 may include a memory - mapped interface coupled to the NoC interface 1426 via a bridge 1418. This memory - mapped interface can include one or more masters and / or slaves. The bridge 1418 can convert a memory - mapped data transfer from the NoC 308 (e.g., configuration, control, and / or debug data) into memory - mapped data that can be received by the memory - mapped switch 1402.

[0207] Tile 1304 may also include an event - processing circuit. For example, tile 1304 includes event logic 1432. The event logic 1432 can be configured by a configuration register 1436. In the example of FIG. 14A, the event logic 1432 is coupled to a control, debug, and trace (CDT) circuit 1420. The configuration data loaded into the configuration register 1436 defines specific events that can be locally detected within tile 1304. The event logic 1432 can detect various different events originating from and / or related to the DMA engine 1412, the memory - mapped switch 1402, the stream switch 1406, the FIFO memory located at the PL interface 1410, and / or the NoC stream interface 1414, according to the configuration register 1436. Examples of events can include, but are not limited to, the end of a DMA transfer, the release of a lock, the acquisition of a lock, the end of a PL transfer, or other events related to the start or end of a data flow through tile 1304. The event logic 1432 can provide such events to the event broadcast circuit 1404 and / or the CDT circuit 1420. For example, in another embodiment, the event logic 1432 may not have a direct connection to the CDT circuit 1420 and may be connected to the CDT circuit 1420 via the event broadcast circuit 1404.

[0208] Tile 1304 includes event broadcast circuit 1404 and event broadcast circuit 1430. Each of event broadcast circuit 1404 and event broadcast circuit 1430 provides an interface between the event broadcast network of DPE array 102, other tiles of SoC interface block 104, and PL310 of device 100. Event broadcast circuit 1404 is coupled to event broadcast circuits and event broadcast circuit 1430 within adjacent or neighboring tile 1302. Event broadcast circuit 1430 is coupled to event broadcast circuits within adjacent or neighboring tile 1306. In one or more other embodiments where the tiles of SoC interface block 104 are arranged in a grid or array, event broadcast circuit 1404 and / or event broadcast circuit 1430 may be connected to event broadcast circuits located in other tiles above and / or below tile 1304.

[0209] In the example of FIG. 14A, event broadcast circuit 1404 is coupled to the event broadcast circuit within the core of DPE adjacent to tile 1304, e.g., DPE304-2 immediately above tile 1304 in column B. Event broadcast circuit 1404 is also coupled to PL interface 1410. Event broadcast circuit 1430 is coupled to the event broadcast circuit in the memory module of DPE adjacent to tile 1304, e.g., DPE304-2 immediately above tile 1304 in column B. Although not shown, in other embodiments, event broadcast circuit 1430 may also be coupled to PL interface 1410.

[0210] Event broadcast circuits 1404 and 1430 can transmit events generated internally by event logic 1432, events received from other tiles of the SoC interface block 104, and / or events received from the DPE in column B (or other DPEs of the DPE array 102) to other tiles. Event broadcast circuit 1404 can further transmit such events to PL310 via PL interface 1410. In another example, an event can be transmitted from event broadcast circuit 1404 to other blocks and / or subsystems within device 100, such as ASIC and / or PL circuit blocks located outside of DPE array 102, using PL interface block 1410. Additionally, PL interface 1410 can receive events from PL310 and provide such events to event broadcast switch 1404 and / or stream switch 1406. In one aspect, event broadcast circuit 1404 can transmit any event received from PL310 to other tiles of the SoC interface block 104 via PL interface 1410, and / or to the DPE in column B and / or other DPEs of the DPE array 102. In another example, an event received from PL310 can be transmitted from event broadcast circuit 1404 to other blocks and / or subsystems within device 100, such as an ASIC. Since events can be broadcast between tiles within the SoC interface block 104, an event can reach any DPE within the DPE array 102 by traversing the tiles within the SoC interface block 104 and the event broadcast circuit to the target (e.g., intended) DPE. For example, an event broadcast circuit within a tile of the SoC interface block 104 under the column (or subset) of DPEs managed by the tile containing the target DPE can propagate the event to the target DPE.

[0211] In the example of FIG. 14A, the event broadcast circuit 1404 and the event logic 1432 are coupled to the CDT circuit 1420. The event broadcast circuit 1404 and the event logic 1432 can send events to the CDT circuit 1420. The CDT circuit 1420 can packetize the received events and send the events from the event broadcast circuit 1404 and / or the event logic 1432 to the stream switch 1406. In certain embodiments, the event broadcast circuit 1430 can also be connected to the stream switch 1406 and / or the CDT circuit 1420.

[0212] In one or more embodiments, the event broadcast circuit 1404 and the event broadcast circuit 1430 can collect broadcast events from one, more, or all directions (e.g., via any of the connections shown in FIG. 14A) as shown in FIG. 14A. In certain embodiments, the event broadcast circuit 1404 and / or the event broadcast circuit 1430 can perform a logical “OR” of the signals and transfer the result in one, more, or all directions (e.g., including the CDT circuit 1420). Each output from the event broadcast circuit 1404 and the event broadcast circuit 1430 can include a bitmask configurable by configuration data loaded into the configuration register 1436. The bitmask determines, on an individual basis, which events are broadcast in each direction. Such a bitmask can, for example, eliminate unwanted propagation or duplicate propagation of events.

[0213] Interrupt handler 1434 is coupled to event broadcast circuit 1404 and can receive events broadcast from event broadcast circuit 1404. In one or more embodiments, interrupt handler 1434 can be configured by configuration data loaded into configuration register 1436 to generate an interrupt in response to a selected event and / or a combination of events from event broadcast circuit 1404 (e.g., DPE generation events, events generated within tile 1304, and / or PL310 generation events). Interrupt handler 1434 can generate an interrupt to PS312 and / or other device-level management blocks within device 100 based on the configuration data. Thus, interrupt handler 1434 can notify PS312 and / or such other device-level management blocks of events occurring in DPE array 102, events occurring in the tiles of SoC interface block 104, and / or events occurring in PL310, based on the interrupt generated by interrupt handler 1434.

[0214] In certain embodiments, interrupt handler 1434 can be coupled to the interrupt handler or interrupt port of PS312 and / or other device-level management blocks by a direct connection. In one or more other embodiments, interrupt handler 1434 can be coupled to PS312 and / or other device-level management blocks by another interface.

[0215] The PL interface 1410 couples to the PL310 of the device 100 and provides an interface thereto. In one or more embodiments, the PL interface 1410 provides an asynchronous clock domain that crosses between the DPE array clock and the PL clock. The PL interface 1410 may also provide level shifters and / or isolation cells for integration with the PL power rails. In certain embodiments, the PL interface 1410 may be configured to provide FIFO support to 32-bit, 64-bit, and / or 128-bit interfaces to handle backpressure. The specific width of the PL interface 1410 may be controlled by configuration data loaded into the configuration register 1436. In the example of FIG. 14A, the PL interface 1410 couples directly to one or more PL interconnect blocks 1422. In certain embodiments, the PL interconnect block 1422 is implemented as a hardwired circuit block that couples to an interconnect circuit located in the PL310.

[0216] In one or more other embodiments, the PL interface 1410 is coupled to other types of circuit blocks and / or subsystems. For example, the PL interface 1410 may be coupled to an ASIC, an analog / mixed signal circuit, and / or other subsystems. Thus, the PL interface 1410 is capable of transferring data between the tile 1304 and such other subsystems and / or blocks.

[0217] In the example of FIG. 14A, tile 1304 includes stream switch 1406. Stream switch 1406 is coupled via one or more stream interfaces to stream switches in adjacent or neighboring tile 1302 and to stream switches in adjacent or neighboring tile 1306. Each stream interface can include one or more masters and / or one or more slaves. In certain embodiments, each pair of neighboring stream switches can exchange data via one or more streams in each direction. Stream switch 1406 is also coupled via one or more stream interfaces to the stream switch within DPE 304-2, which is directly above tile 1304 in column B. As discussed, a stream interface can include one or more stream slaves and / or stream masters. Stream switch 1406 is also coupled to PL interface 1410, DMA engine 1412, and / or NoC stream interface 1414 via stream multiplexer / demultiplexer 1408 (abbreviated as stream mux / demux in FIG. 14A). Stream switch 1406 can include one or more stream interfaces used to communicate with each of PL interface 1410, DMA engine 1412, and / or NoC stream interface 1414 via stream multiplexer / demultiplexer 1408, for example.

[0218] In one or more other embodiments, stream switch 1406 can be coupled to other circuit blocks in other directions and / or diagonal directions depending on the number of stream interfaces included and / or the arrangement of tiles and / or DPEs and / or other circuit blocks around tile 1304.

[0219] In one or more embodiments, the stream switch 1406 can be configured by configuration data loaded into the configuration register 1436. The stream switch 1406 can be configured, for example, to support packet - switching operations and / or circuit - switching operations based on the configuration data. Further, the configuration data defines specific DPEs and / or a plurality of DPEs within the DPE array 102 with which the stream switch 1406 communicates. In one or more embodiments, the configuration data defines specific DPEs and / or a subset of DPEs (e.g., DPEs within column B) of the DPE array 102 with which the stream switch 1406 communicates.

[0220] The stream multiplexer / demultiplexer 1408 can direct data received from the PL interface 1410, the DMA engine 1412, and / or the NoC stream interface 1414 to the stream switch 1406. Similarly, the stream multiplexer / demultiplexer 1408 can direct data received from the stream switch 1406 to the PL interface 1410, the DMA engine 1412, and / or the NoC stream interface 1414. For example, the stream multiplexer / demultiplexer 1408 can be programmed by the configuration data stored in the configuration register 1436 to route selected data to the DMA engine 1412, where such data is transmitted via the NoC 308 as a memory - mapped transaction, and / or to route selected data to the NoC stream interface 1414, where the data may be transmitted via the NoC 308 as one or more data streams.

[0221] The DMA engine 1412 can operate as a master for directing data through the selector block 1416 to the NoC interface 1426 and towards the NoC 308. The DMA engine 1412 can receive data from the DPE and provide such data to the NoC 308 as a memory-mapped data transaction. In one or more embodiments, the DMA engine 1412 can include a hardware synchronization circuit that can be used to synchronize a plurality of channels included in the DMA engine 1412 and / or a channel within the DMA engine 1412 with a master that polls and drives a lock request. For example, the master can be a device implemented within the PS 312 or the PL 310. The master can also receive an interrupt generated by the hardware synchronization circuit within the DMA engine 1412.

[0222] In one or more embodiments, the DMA engine 1412 can access an external memory. For example, the DMA engine 1412 can receive a data stream from the DPE and transmit the data stream to the external memory to a memory controller located within the SoC via the NoC 308. The memory controller then directs the data received as a data stream to the external memory (e.g., initiates a read and / or write of the external memory as requested by the DMA engine 1412). Similarly, the DMA engine 1412 can receive data from the external memory, and such data may be distributed to other tiles of the SoC interface block 104 and / or to the target DPE.

[0223] In certain embodiments, the DMA engine 1412 includes security bits that can be set using the DPE global control settings register (DPE GCS register) 1438. The external memory can be divided into different regions or partitions, and the DPE array 102 is permitted to access only specific regions of the external memory. The security bits within the DMA engine 1412 may be set such that the DPE array 102 can access only specific regions of the external memory permitted according to the security bits via the DMA engine 1412. For example, an application implemented by the DPE array 102 may be restricted to access only specific regions of the external memory, restricted to only read from specific regions of the external memory, and / or restricted from writing to the entire external memory using this mechanism.

[0224] The security bits within the DMA engine 1412 that control access to the external memory may be implemented to control the DPE array 102 as a whole, or may be implemented in a more granular manner, and access to the external memory may be specified and / or controlled for each DPE, for example for each core, or for a group of cores configured to operate cooperatively, for example a kernel and / or other applications that implement an application.

[0225] The NoC stream interface 1414 can receive data from the NoC 308 via the NoC interface 1426 and transfer the data as a stream to the multiplexer / demultiplexer 1408. The NoC stream interface 1414 can further receive data from the stream multiplexer / demultiplexer 1408 and transfer the data to the NoC interface 1426 via the selector block 1416. The selector block 1416 can be configured to pass data from the DMA engine 1412 or the NoC stream interface 1414 to the NoC interface 1426.

[0226] The CDT circuit 1420 can perform control, debugging, and trace operations within the tile 1304. Regarding debugging, each of the registers located in the tile 1304 is mapped onto a memory map accessible via the memory-mapped switch 1402. The CDT circuit 1420 may include circuits such as, for example, trace hardware, a trace buffer, performance counters, and / or stall logic. The trace hardware of the CDT circuit 1420 can collect trace data. The trace buffer of the CDT circuit 1420 can buffer the trace data. Further, the CDT circuit 1420 can output the trace data to the stream switch 1406.

[0227] In one or more embodiments, the CDT circuit 1420 can collect data, such as trace and / or debug data, packetize such data, and then output the packetized data through the stream switch 1406. For example, the CDT circuit 1420 can output the packetized data and provide such data to the stream switch 1406. Additionally, the configuration register 1436 or others can be read or written via memory-mapped transactions through the memory-mapped switch 1402 of each tile during debugging. Similarly, the performance counters within the CDT circuit 1420 can be read or written via memory-mapped transactions through the memory-mapped switch 1402 of each tile during profiling.

[0228] In one or more embodiments, the CDT circuit 1420 can receive any event propagated by the event broadcast circuit 1404 (or event broadcast circuit 1430), or an event selected according to a bitmask utilized by an interface of the event broadcast circuit 1404 coupled to the CDT circuit 1420. The CDT circuit 1420 can further receive an event generated by the event logic 1432. For example, the CDT circuit 1420 can receive a broadcast event from the PL310, from the DPE304, from the tile 1304 (e.g., the event logic 1432 and / or the event broadcast switch 1404), and / or from other tiles of the SoC interface block 104. The CDT circuit 1420 can pack a plurality of such events together into a packet, e.g., packetize, and can associate the packetized event with a timestamp. The CDT circuit 1420 can further transmit the packetized event to a destination external to the tile 1304 via the stream switch 1406. The event can be transmitted via the stream switch 1406 and the stream multiplexer / demultiplexer 1408 via the PL interface 1410, the DMA engine 1412, and / or the NoC stream interface 1414.

[0229] The DPE GCS register 1438 can store DPE global control settings / bits (also referred to herein as "security bits") used to enable or disable secure access to and / or from the DPE array 102. The DPE GCS register 1438 can be programmed via the SoC security / initialization interface as described in more detail below in connection with FIG. 14C. The security bits received from the SoC security / initialization interface can be propagated from one tile of the SoC interface block 104 to the next tile via a bus as shown in FIG. 14A.

[0230] In one or more embodiments, external memory-mapped data transfers to the DPE array 102 (e.g., using the NoC 308) are not secure or reliable. Without setting a security bit in the DPE GCS register 1438, any entity within the device 100 that can communicate via a memory-mapped data transfer (e.g., via the NoC 308) can communicate with the DPE array 102. By setting a security bit in the DPE GCS register 1438, the specific entities permitted to communicate with the DPE array 102 can be defined such that only the designated entities that can generate secure traffic can communicate with the DPE array 102.

[0231] For example, the memory-mapped interface of the memory-mapped switch 1402 can communicate with the NoC 308. Memory-mapped data transfers can include additional sideband signals, such as bits, that specify whether a transaction is secure or not. When the security bit in the DPE GCS register 1438 is set, the sideband signal must be set such that the memory-mapped transaction entering the SoC interface block 104 indicates that the memory-mapped transaction arriving at the SoC interface block 104 from the NoC 308 is secure. If the memory-mapped transaction arriving at the SoC interface block 104 has the sideband bit not set and the security bit is set in the DPE GCS register 1438, the SoC interface block 104 does not permit the transaction to enter or pass through the DPE 304.

[0232] In one or more embodiments, the SoC includes a secure agent (e.g., a circuit) that operates as a root of trust. The secure agent is permitted, when the security bit of the DPE GCS register 1438 is set, to configure different entities (e.g., circuits) within the SoC with the sideband bits set in a memory-mapped transaction to access the DPE array 102. When the SoC is configured, the secure agent grants permissions to different masters that may be implemented by the PL310 or PS312, thereby giving such masters the ability to issue (or not issue) secure transactions to the DPE array 102 via the NoC 308.

[0233] FIG. 14B shows another exemplary implementation of tile 1304. The exemplary architecture shown in FIG. 14B may also be used to implement any of the other tiles included in the SoC interface block 104. The example of FIG. 14B shows a simplified version of the architecture shown in FIG. 14A. The tile architecture of FIG. 14B provides connectivity between the DPEs within the device 100 as well as other subsystems and / or blocks. For example, the tile 1304 of FIG. 14B may provide an interface between the DPE and the PL310, analog / mixed signal circuit blocks, ASICs, or other subsystems described herein. The tile architecture of FIG. 14B does not provide connectivity to the NoC 308. Accordingly, the DMA engine 1412, NoC interface 1414, selector block 1416, bridge 1418, and stream multiplexer / demultiplexer 1408 are omitted. Accordingly, the tile 1304 of FIG. 14B may be implemented using less area of the SoC. Further, as shown, the stream switch 1406 is directly coupled to the PL interface 1410.

[0234] The exemplary architecture of FIG. 14B cannot receive memory-mapped data, such as configuration data, for the purpose of constructing a DPE from the NoC 308. Such configuration data can be received from neighboring tiles via the memory-mapped switch 1402 and directed to a subset of the DPEs managed by tile 1304 (e.g., until it enters the column of DPEs above tile 1304 in FIG. 14B).

[0235] FIG. 14C shows another exemplary implementation of tile 1304. In certain embodiments, the architecture shown in FIG. 14C can be used to implement just one tile within the SoC interface block 104. For example, the architecture shown in FIG. 14C can be used to implement tile 1302 within the SoC interface block 104. The architecture shown in FIG. 14C is similar to the architecture shown in FIG. 14B. FIG. 14C includes additional components such as the SoC secure / initialization interface 1440, the clock signal generator 1442, and the global timer 1444.

[0236] In the example of FIG. 14C, the SoC Secure / Initialization Interface 1440 provides an additional interface for the SoC Interface Block 104. In one or more embodiments, the SoC Secure / Initialization Interface 1440 is implemented as a NoC peripheral interconnect. The SoC Secure / Initialization Interface 1440 can provide access to a global reset register for the DPE array 102 (not shown) and the DPE GCS register 1438. In certain embodiments, the DPE GCS register 1438 includes a configuration register for the clock signal generator 1442. As shown, the SoC Secure / Initialization Interface 1440 can provide security bits to the DPE GCS register 1438 and propagate the security bits to other DPE GCS registers 1438 within other tiles of the SoC Interface Block 104. In certain embodiments, the SoC Secure / Initialization Interface 1440 implements a single slave endpoint for the SoC Interface Block 104.

[0237] In the example of FIG. 14C, the clock signal generator 1442 can generate one or more clock signals 1446 and / or one or more reset signals 1450. The clock signals 1446 and / or the reset signals 1450 can be distributed to each of the DPEs 304 and / or other tiles of the SoC Interface Block 104 of the DPE array 102. In one or more embodiments, the clock signal generator 1442 can include one or more phase-locked loop circuits (PLLs). As shown, the clock signal generator 1442 can receive a reference clock signal generated by another circuit that is external to the DPE array 102 and located on the SoC. The clock signal generator 1442 can generate the clock signal 1446 based on the received reference clock signal.

[0238] In the example of FIG. 14C, the clock signal generator 1442 is configured via the SoC secure / initialization interface 1440. For example, the clock signal generator 1442 may be configured by loading data into the DPE GCS register 1438. Thus, the generation of one or more clock frequencies and the reset signal 1450 of the DPE array 102 can be set by writing appropriate configuration data to the DPE GCS register 1438 via the SoC secure / initialization interface 1440. For test purposes, the clock signal 1446 and / or the reset signal 1450 may also be directly routed to the PL310.

[0239] The SoC secure / initialization interface 1440 may be coupled to an SoC control / debug (circuit) block (e.g., a control and / or debug subsystem of a device 100 not shown). In one or more embodiments, the SOC secure / initialization interface 1440 can provide status signals to the SOC control / debug block. By way of example and not limitation, the SoC secure / initialization interface 1440 can provide a "PLL lock" signal generated internally by the clock signal generator 1440 to the SoC control / debug block. The PLL lock signal can indicate when the PLL acquires lock on the reference clock signal.

[0240] The SoC secure / initialization interface 1440 can receive instructions and / or data via the interface 1448. The data can include security bits, clock signal generator configuration data, and / or other data that can be written to the DPE GCS register 1438 as described herein.

[0241] The global timer 1444 can interface with the CDT circuit 1420. For example, the global timer 1444 can be coupled to the CDT circuit 1420. The global timer 1444 can provide a signal used by the CDT circuit 1420 to timestamp events used for tracking. In one or more embodiments, the global timer 1444 can be coupled to a CDT circuit 1420 within another tile of the tiles of the SoC interface circuit 104. For example, the global timer 1444 can be coupled to the CDT circuit 1420 within the exemplary tiles of FIGS. 14A, 14B, and / or 14C. The global timer 1444 can also be coupled to the SoC control / debug block.

[0242] Referring collectively to the architectures of FIGS. 14A, 14B, and 14C, the tile 1304 can communicate with the DPE 304 using a variety of different data paths. In one example, the tile 1304 can communicate with the DPE 304 using the DMA engine 1412. For example, the tile 1304 can use the DMA engine 1412 to communicate with the DMA engine (e.g., DMA engine 816) of one or more DPEs of the DPE array 102. Communication can flow from the DPE to the tile of the SoC interface block 104 or from the tile of the SoC interface block 104 to the DPE. In another example, the DMA engine 1412 can communicate with the cores of one or more DPEs of the DPE array 102 via the stream switch within each DPE. Communication can flow from the core to the tile of the SoC interface block 104 and / or from the tile of the SoC interface block 104 to the cores of one or more DPEs of the DPE array 102.

[0243] FIG. 15 shows an exemplary implementation of the PL interface 1410. In the example of FIG. 15, the PL interface 1410 includes a plurality of channels that couple the PL310 to the stream switch 1406 and / or the stream multiplexer / demultiplexer 1408, depending on the particular tile architecture being used. The particular number of channels shown within the PL interface 1410 in FIG. 15 is for illustration purposes only and not a limitation. In other embodiments, the PL interface 1410 can include fewer or more channels than those shown in FIG. 15. Further, although the PL interface 1410 is shown as connecting to the PL310, in one or more other embodiments, the PL interface 1410 can be coupled to one or more other subsystems and / or circuit blocks. For example, the PL interface 1410 can also be coupled to an ASIC, an analog / mixed-signal circuit, and / or other circuits or subsystems.

[0244] In one or more embodiments, the PL310 operates at a different reference voltage and a different clock speed than the DPE304. Thus, in the example of FIG. 15, the PL interface 1410 includes a plurality of shift and isolation circuits 1502 and a plurality of asynchronous FIFO memories 1504. Each channel includes a shift isolation circuit 1502 and an asynchronous FIFO memory 1504. A first subset of the channels conveys data from the PL310 (and / or other circuits) to the stream switch 1406 and / or the stream multiplexer / demultiplexer 1408. A second subset of the channels conveys data from the stream switch 1406 and / or the stream multiplexer / demultiplexer 1408 to the PL310 and / or other circuits.

[0245] The shift and isolation circuit 1502 can interface between different voltage regions. In this case, the shift and isolation circuit 1502 can provide an interface that transitions between the operating voltage of the PL310 and / or other circuits and the operating voltage of the DPE304. The asynchronous FIFO memory 1504 can interface between two different clock domains. In this case, the asynchronous FIFO memory 1504 can provide an interface that transitions between the clock rate of the PL310 and / or other circuits and the clock rate of the DPE304.

[0246] In one or more embodiments, the asynchronous FIFO memory 1504 has a 32-bit interface to the DPE array 102. The connections between the asynchronous FIFO memory 1504 and the shift and isolation circuit 1502, as well as the connections between the shift and isolation circuit 1502 and the PL310, can be programmable (e.g., configurable) in width. For example, the connections between the asynchronous FIFO memory 1504 and the shift and isolation circuit 1502, as well as the connections between the shift and isolation circuit 1502 and the PL310, can be configured to have a width of 32 bits, 64 bits, or 128 bits. As discussed, the PL interface 1410 can be configured by writing configuration data to the configuration register 1436 by the memory-mapped switch 1402 to achieve the described bit widths. Using the memory-mapped switch 1402, the asynchronous FIFO memory 1504 side on the PL310 side can be configured to use either 32 bits, 64 bits, or 128 bits. The bit widths provided herein are for illustrative purposes. In other embodiments, other bit widths may be used. In any case, the widths described for the various components can be changed based on the configuration data loaded into the configuration register 1436.

[0247] FIG. 16 shows an exemplary implementation of the NoC stream interface 1414. The DPE array 102 has two general ways of communicating via the NoC 308 using the stream interfaces within the DPE. In one aspect, the DPE can access the DMA engine 1412 using the stream switch 1406. The DMA engine 1412 can convert memory-mapped transactions from the NoC 308 into a data stream for transmission to the DPE, and convert a data stream from the DPE into a memory-mapped transaction for transmission via the NoC 308. In another aspect, the data stream can be directed to the NoC stream interface 1414.

[0248] In the example of FIG. 16, the NoC stream interface 1414 includes a plurality of channels that couple the NoC 308 to the stream switch 1406 and / or the stream multiplexer / demultiplexer. Each channel can include a FIFO memory and either an upsizing circuit or a downsizing circuit. A first subset of the channels conveys data from the NoC 308 to the stream switch 1406 and / or the stream multiplexer / demultiplexer 1408. A second subset of the channels conveys data from the stream switch 1406 and / or the stream multiplexer / demultiplexer 1408 to the NoC 308. The specific number of channels within the NoC stream interface 1414 shown in FIG. 16 is for illustrative purposes only and not limiting. In other embodiments, the NoC stream interface 1414 can include fewer or more channels than shown in FIG. 16.

[0249] In one or more embodiments, each of the upsizing circuits 1608 (abbreviated as "US circuit" in FIG. 16) can receive a data stream and increase the width of the received data stream. For example, each upsizing circuit 1608 can receive a 32-bit data stream and output a 128-bit data stream to the corresponding FIFO memory 1610. Each of the FIFO memories 1610 is coupled to an arbitration and multiplexer circuit 1612. The arbitration and multiplexer circuit 1612 can arbitrate between the received data streams using a specific arbitration scheme or priority (e.g., round-robin or other style) to provide the resulting output data stream to the NoC interface 1426. The arbitration and multiplexer circuit 1612 can process and accept new requests every clock cycle. The clock domain crossing between the DPE 304 and the NoC 308 can be processed within the NoC 308 itself. In one or more other embodiments, the clock domain crossing between the DPE 304 and the NoC 308 can be processed within the SoC interface block 104. For example, the clock domain crossing can be processed within the NoC stream interface 1414.

[0250] The demultiplexer 1602 can receive a data stream from the NoC 308. For example, the demultiplexer 1602 can be coupled to the NoC interface 1426. For the sake of explanation, the data stream from the NoC interface 1426 can have a width of 128 bits. The clock domain crossing between the DPE 304 and the NoC 308 can be processed within the NoC 308 and / or within the NoC stream interface 1414 as described above. The demultiplexer 1602 can transfer the received data stream to one of the FIFO memories 1604. The particular FIFO memory 1604 to which the demultiplexer 1602 provides the data stream can be encoded within the data stream itself. The FIFO memory 1604 is connected to a downsize circuit 1606 (abbreviated as "DS circuit" in FIG. 16). The downsize circuit 1606 can downsize the received stream to a narrower width after buffering using time-division multiplexing. For example, the downsize circuit 1606 can downsize the stream from a width of 128 bits to a width of 32 bits.

[0251] As shown, the downsize circuit 1606 and the upsizing circuit 1608 are coupled to the stream switch 1406 or the stream multiplexer / demultiplexer 1408 depending on the particular architecture of the tile of the SoC interface block 104 being used. FIG. 16 is provided for illustrative purposes and is not intended as a limitation. The order and / or connectivity of the components within the channel (e.g., the upsizing / downsizing circuits and the FIFO memories) can vary.

[0252] In one or more other embodiments, the PL interface 1410 described in connection with FIG. 15 may include an upsize circuit and / or a downsize circuit as described in connection with FIG. 16. For example, the downsize circuit may be included in each channel that transmits data from the PL310 (or other circuit) to the stream switch 1406 and / or the stream multiplexer / demultiplexer 1408. The upsize circuit may be included in each channel that transmits data from the stream switch 1406 and / or the stream multiplexer / demultiplexer 1408 to the PL310 (or other circuit).

[0253] In one or more other embodiments, although shown as independent elements, each downsize circuit 1606 may be combined with the corresponding FIFO memory 1604, for example, as a single block or circuit. Similarly, each upsize circuit 1608 may be combined with the corresponding FIFO memory 1610, for example, as a single block or circuit.

[0254] FIG. 17 shows an exemplary implementation of the DMA engine 1412. In the example of FIG. 17, the DMA engine 1412 includes a DMA controller 1702. The DMA controller 1702 can be divided into two separate modules or interfaces. Each module can operate independently of the other. The DMA controller 1702 can include an interface 1704 from memory mapping to stream (interface) and an interface 1706 from stream to memory mapping (interface). Each of the interface 1704 and the interface 1706 can include two or more separate channels. Thus, the DMA engine 1412 can receive two or more input streams from the stream switch 1406 via the interface 1706 and transmit two or more output streams to the stream switch 1406 via the interface 1704. The DMA controller 1702 can further include a master memory-mapped interface 1714. The master memory-mapped interface 1714 couples the NoC 308 to the interface 1704 and the interface 1706.

[0255] The DMA engine 1412 can also include a hardware synchronization circuit 1710 and a buffer descriptor register file 1708. The hardware synchronization circuit 1710 and the buffer descriptor register file 1708 can be accessed via a multiplexer 1712. Thus, both the hardware synchronization circuit 1710 and the buffer descriptor register file 1708 can be accessed via an external control interface. Examples of such control interfaces include, but are not limited to, a memory-mapped interface or a control stream interface from the DPE. An example of the control stream interface of the DPE is a streaming interface output from the core of the DPE.

[0256] The hardware synchronization circuit 1710 can be used to synchronize a plurality of channels included in the DMA engine 1412 and / or a certain channel within the DMA engine 1412 with a master that polls and drives lock requests. For example, the master can be a device implemented within the PS312 or the PL310. In another example, the master can also receive an interrupt generated by the hardware synchronization circuit 1710 within the DMA engine 1412 when a lock is available.

[0257] The DMA transfer can be defined by buffer descriptors stored in the buffer descriptor register file 1708. The interface 1706 can request a read transfer to the NoC308 based on the information within the buffer descriptor. The outgoing stream from the interface 1704 to the stream switch 1406 can be configured as packet switching or circuit switching based on the configuration register for the stream switch.

[0258] FIG. 18 shows an exemplary architecture for a plurality of DPEs. This exemplary architecture shows the DPE 304 that can be included in the DPE array 102. The exemplary architecture of FIG. 18 can be called a checkerboard architecture. The exemplary architecture of FIG. 18 enables a core of a certain DPE to communicate with up to eight other cores of other DPEs using shared memory (for example, a total of nine cores communicate via shared memory). In the example of FIG. 18, each DPE 304 can be implemented as described in relation to FIGS. 6, 7, and 8. Thus, each core 602 can access four different memory modules 604. Each memory module 604 can be accessed by up to four different cores 602.

[0259] As shown, the DPE array 102 includes row 1, row 2, row 3, row 4, and row 5. Each of row 1 to row 5 includes three DPEs 304. The specific number of DPEs 304 and the number of rows in each row shown in FIG. 18 are for illustrative purposes only and not limitations. Referring to row 1, row 3, and row 5, the cores of each DPE in these rows are located on the left side of the memory module. Referring to row 2 and row 4, the cores of each DPE in these rows are located on the right of the memory module. In fact, the orientation of the DPEs in row 2 and row 4 is reversed or flipped horizontally compared to the orientation of the DPEs in row 1, row 3, and row 5. The orientation of the DPEs is reversed as shown in each alternate row.

[0260] In the example of FIG. 18, the DPEs 304 are aligned in columns. However, the cores and memory modules of adjacent rows are not aligned in columns. The architecture of FIG. 18 is an example of a heterogeneous architecture implemented such that the DPEs are different based on the specific row in which the DPEs are located. Due to the horizontal reversal of the DPEs 304, the cores of adjacent rows are not aligned. The cores of adjacent rows are offset from each other. Similarly, the memory modules of adjacent rows are not aligned. The memory modules of adjacent rows are offset from each other. However, the cores of every other row are aligned such that the memory modules of every other row are aligned. For example, the cores and memory modules of row 1, row 3, and row 5 are aligned vertically (e.g., in columns). Similarly, the cores and memory modules of row 2 and row 4 are aligned vertically (e.g., in columns).

[0261] For illustration, the cores of DPE304-2, 304-4, 304-5, 304-7, 304-8, 304-9, 304-10, 304-11, and 304-14 are considered part of a group and can communicate via shared memory. The arrows show how the exemplary architecture of FIG. 18 supports a core that communicates with up to eight other cores in different DPEs using shared memory. Referring to DPE304-8, for example, core 602-8 can access memory modules 604-11, 604-7, 604-8, and 604-5. Through memory module 604-11, core 602-8 can communicate with cores 602-14, 602-10, and 602-11. Through memory module 604-7, core 602-8 can communicate with cores 602-7, 602-4, and 602-10. Through memory module 604-8, core 602-8 can communicate with cores 602-9, 602-11, and 602-5. Through memory module 604-5, core 602-8 can communicate with cores 602-4, 602-5, and 602-2.

[0262] In the example of FIG. 18, excluding core 602-8, within the group, there are four different cores that can access two different memory modules out of the group's shared memory modules. The remaining four other cores share only one memory module out of the group's shared memory modules. The group's shared memory modules include memory modules 604-5, 604-7, 604-8, and 604-11. For example, each of cores 602-10, 602-11, 602-4, and 602-5 can access two different memory modules. Core 602-10 can access memory modules 604-11 and 604-7. Core 602-11 can access memory modules 604-11 and 604-8. Core 602-4 can access memory modules 604-5 and 604-7. Core 602-5 can access memory modules 604-5 and 604-8.

[0263] In the example of FIG. 18, up to nine cores of a total of nine DPEs can communicate via shared memory without using the DPE interconnect network of DPE array 102. As can be discussed, cores 602-8 can view memory modules 604-11, 604-7, 604-5, and 604-8 as an integrated memory space.

[0264] Cores 602-14, 602-7, 602-9, and 602-2 can access only one of the memory modules of the group of shared memory modules. Core 602-14 can access memory module 604-11. Core 602-7 can access memory module 604-7. Core 602-9 can access memory module 604-8. Core 602-2 can access memory module 604-5.

[0265] As described above, in other embodiments where more than four memory interfaces are provided for each memory module, a core can communicate with more than eight other cores via shared memory using the architecture of FIG. 18.

[0266] In one or more other embodiments, a particular row and / or column of DPEs can be offset relative to other rows. For example, row 2 and row 4 can start from a position that is not aligned with the start of row 1, row 3, and / or row 5. For example, row 2 and row 4 can be shifted to the right relative to the start of row 1, row 3, and / or row 5.

[0267] FIG. 19 is a diagram showing another exemplary architecture for a plurality of DPEs. The exemplary architecture shows DPEs 304 that may be included in DPE array 102. The exemplary architecture shown in FIG. 19 may be referred to as a grid architecture. The exemplary architecture of FIG. 19 enables the core of one DPE to communicate with up to 10 other cores of other DPEs using shared memory (e.g., a total of 11 cores communicate via shared memory). In the example of FIG. 19, each DPE 304 may be implemented as described in connection with FIGS. 6, 7, and 8. Thus, each core 602 can access four different memory modules 604. Each memory module 604 can be accessed by up to four different cores 602.

[0268] As shown, DPE array 102 includes rows 1, 2, 3, 4, and 5. Each of rows 1-5 includes three DPEs 304. The particular number of DPEs 304 in each row shown in FIG. 19 and the number of rows are for purposes of illustration and not limitation. In the example of FIG. 19, the DPEs 304 are vertically aligned in columns. Each of rows 1, 2, 3, 4, and 5 has the same starting point that is aligned with each other row of DPEs. Further, the arrangement of cores 602 and memory modules 604 within each respective DPE 304 is the same. In other words, the cores 602 are vertically aligned. Similarly, the memory modules 604 are vertically aligned.

[0269] For illustration, the cores of DPE304-2, 304-4, 304-5, 304-6, 304-7, 304-8, 304-9, 304-10, 304-11, 304-12, and 304-14 are considered part of a group and can communicate via shared memory. The arrows show how the exemplary architecture of FIG. 19 supports a core that uses shared memory to communicate with up to ten other cores within different DPEs. Referring to DPE304-8, for example, core 602-8 can access memory modules 604-11, 604-8, 604-5, and 604-9. Through memory module 604-11, core 602-8 can communicate with cores 602-14, 602-10, and 602-11. Through memory module 604-8, core 602-8 can communicate with cores 602-7, 602-11, and 602-5. Through memory module 604-5, core 602-8 can communicate with cores 602-4, 602-5, and 602-2. Through memory module 604-9, core 602-8 can communicate with cores 602-12, 602-9, and 602-6.

[0270] In the example of FIG. 19, except for core 602-8, within the group, there are two different cores that can access two of the group's shared memory modules. The group's shared memory modules include memory modules 604-5, 604-8, 604-9, and 604-11. The remaining eight cores of the group share only one memory module. For example, each of cores 602-11 and 602-5 can access two different memory modules. Core 602-11 can access memory modules 604-11 and 604-8. Since memory module 604-14 is not accessible by core 604-8, memory module 604-14 is not considered part of the group of shared memory. Core 602-5 can access memory modules 604-5 and 604-8. Since memory module 604-2 is not accessible by core 602-8, memory module 604-2 is not considered part of the group of shared memory modules.

[0271] Cores 602-14, 602-10, 602-12, 602-7, 602-9, 602-4, 602-6, and 602-2 can access only one of the group's shared memory modules. Core 602-14 can access memory module 604-11. Core 602-10 can access memory module 604-11. Core 602-12 can access memory module 604-9. Core 602-7 can access memory module 604-8. Core 602-9 can access memory module 604-9. Core 602-4 can access memory module 604-5. Core 602-6 can access memory module 604-9. Core 602-2 can access memory module 604-5.

[0272] In the example of FIG. 19, up to 11 cores of up to 11 DPEs can communicate via shared memory without using the DPE interconnect network of the DPE array 102. As can be discussed, cores 602-8 can view memory modules 604-11, 604-9, 604-5, and 604-8 as an integrated memory space.

[0273] As described above, in other embodiments where more than four memory interfaces are provided for each memory module, a core can communicate with more than 10 other cores via shared memory using the architecture of FIG. 19.

[0274] FIG. 20 shows an exemplary method 2000 for constructing a DPE array. Method 2000 is provided for illustrative purposes and is not intended to limit the configuration of the invention described within this disclosure.

[0275] In block 2002, configuration data regarding the DPE array is loaded into the device. The configuration data can be provided from any of a variety of different sources, regardless of whether it is a computer system (e.g., a host), off-chip memory, or other suitable source.

[0276] In block 2004, the configuration data is provided to the SoC interface block. In certain embodiments, the configuration data is provided via the NoC. The tiles of the SoC interface block are capable of receiving the configuration data and converting the configuration data into memory-mapped data, and the memory-mapped data can be provided to the memory-mapped switches included within the tile.

[0277] In block 2006, the configuration data propagates between tiles of the SoC interface block to a particular tile that operates as or provides an interface to the target DPE. The target DPE is the DPE to which the configuration data is addressed. For example, the configuration data includes an address that specifies a particular DPE to which different portions of the configuration data should be directed. Memory-mapped switches within the tiles of the SoC interface block can propagate different portions of the configuration data to a particular tile that operates as an interface for the target DPE (e.g., a subset of DPEs that includes the target DPE).

[0278] In block 2008, a tile of the SoC interface block that operates as an interface to the target DPE can direct the portion of the configuration data for the target DPE to the target DPE. For example, a tile that provides an interface to one or more target DPEs can direct a portion of the configuration data to a subset of the DPEs for which the tile provides an interface. As described above, the subset of DPEs includes one or more target DPEs. Each tile, upon receiving the configuration data, can determine whether any portion of the configuration data is addressed to other DPEs within the same subset of DPEs for which the tile provides an interface. The tile directs any configuration data addressed to DPEs within the subset of DPEs to such DPEs.

[0279] In block 2010, the configuration data is loaded into the target DPE and programs the elements of the DPE contained therein. For example, the configuration data is loaded into configuration registers to program elements of the target DPE such as a stream interface, a core (e.g., a stream interface, a cascade interface, a core interface), a memory module (e.g., a DMA engine, a memory interface, an arbiter, etc.), a broadcast event switch, and / or broadcast logic. The configuration data may also include executable program code that can be loaded into the program memory of the core and / or data that can be loaded into the memory banks of the memory module.

[0280] It should be understood that the received configuration data may also include portions that are addressed to one or more or all of the tiles of the SoC interface block 104. In that case, the memory-mapped switches within each tile can transmit the configuration data to the appropriate (e.g., target) tile, extract such data, and write such data to the appropriate configuration registers within each tile.

[0281] FIG. 21 shows an exemplary method 2100 of operation of a DPE array. Method 2100 is provided for illustrative purposes and is not intended to limit the configuration of the invention described within the present disclosure. Method 2100 begins with the configuration data loaded into the DPE and / or the SoC interface block. For illustration, refer to FIG. 3.

[0282] In block 2102, the core 602-15 (e.g., "the first core") of DPE304-15 (e.g., "the first DPE") generates data. The generated data may be application data. For example, the core 602-15 may process data stored in a memory module accessible by the core. The memory module may be within DPE304-15 or within a different DPE, as described herein. The data may be received, for example, using the SoC interface block 104, from another DPE and / or another subsystem of the device.

[0283] In block 2104, the core 602-15 stores data in the memory module 604-15 of DPE304-15. In block 2106, one or more cores within a neighboring DPE (e.g., DPE304-25, 304-16, and / or 304-5) read data from the memory module 604-15 of DPE304-15. Cores within the neighboring DPE can utilize the data read from the memory module 604-15 for further calculations.

[0284] In block 2108, DPE304-15 optionally sends data to one or more other DPEs via a stream interface. The DPE to which the data is sent can be a non-neighboring DPE. For example, DPE304-15 can send data from memory module 604-15 to one or more other DPEs such as DPE304-35, 304-36. As will be discussed, in one or more embodiments, DPE304-15 can broadcast and / or multicast application data via a stream interface within the DPE interconnect network of DPE array 102. In another example, the data sent to different DPEs can be different portions of the data, and each different portion of the data targets a different destination DPE. Although not shown in FIG. 21, core 602-15 can also send data to another core and / or DPE of DPE array 102 directly from the core using a cascade interface and / or using a stream switch.

[0285] In block 2110, core 602-15 optionally sends data to a neighboring core via a cascade interface and / or receives data from a neighboring core. The data can be application data. For example, core 602-15 can directly receive data from core 602-14 of DPE304-14 via a cascade interface and / or directly send data to core 602-16 of DPE304-16.

[0286] In block 2112, DPE304-15 optionally transmits data to one or more subsystems via an SoC interface block and / or receives data from one or more subsystems. The data may be application data. For example, DPE304-15 can transmit data to PS312 via NoC308, to a circuit implemented in PL310, to a selected hardwired circuit block via NoC308, to a selected hardwired circuit block via PL310, and / or to another external subsystem such as an external memory. Similarly, DPE304-15 can receive application data from such other subsystems via an SoC interface block.

[0287] FIG. 22 shows another exemplary method 2200 of operation of a DPE array. Method 2200 is provided for illustrative purposes and is not intended to limit the configuration of the invention described within this disclosure. Method 2200 begins with the configuration data loaded into the DPE array.

[0288] In block 2202, a first core, e.g., a core within a first DPE, requests a lock on a target memory region from a hardware synchronization circuit. The first core can request a lock from the hardware synchronization circuit for a target memory region within a memory module located, for example, within the same DPE as the first DPE, e.g., the first core, or within a memory module located within a DPE different from the first core. The first core can request a lock from a specific hardware synchronization circuit located in the same DPE as the target memory region to be accessed.

[0289] In block 2204, the first core acquires the requested lock. The hardware synchronization circuit grants, for example, the requested lock on the target memory region to the first core.

[0290] In block 2206, in response to acquiring the lock, the first core writes data to the target memory region. For example, if the target memory region is within the first DPE, the first core can write data to the target memory region via a memory interface located within a memory module in the first DPE. In another example, if the target memory region is located in a DPE different from the first core, the first core can write data to the target memory region using any of the various techniques described herein. For example, the first core can write data to the target memory region via any of the mechanisms described in connection with FIG. 10.

[0291] In block 2208, the first core releases the lock on the target memory region. In block 2210, the second core requests a lock on the target memory region that contains the data written by the first core. The second core may be located in the same DPE as the target memory region or in a DPE different from the target memory region. The second core requests the lock from the same hardware synchronization circuit that granted the lock to the first core. In block 2212, the second core acquires the lock from the hardware synchronization circuit. The hardware synchronization circuit grants the lock to the second core. In block 2214, the second core can access the data from the target memory region and utilize that data for processing. In block 2216, the second core releases the lock on the target memory region, for example, if access to the target memory region is no longer required.

[0292] The example of FIG. 22 is described in relation to access to a memory region. In certain embodiments, the first core can write data directly to a target memory region. In other embodiments, the first core can move data from a source memory region (e.g., in a first DPE) to a target memory region (e.g., located within a second or different DPE). In this case, the first core acquires locks on the source memory region and the target memory region to perform the data transfer.

[0293] In other embodiments, the first core can acquire a lock on the second core to stall the operation of the second core and then release the lock to resume the operation of the second core. For example, the first core may acquire a lock on the second core in addition to a lock on the target memory region to stall the operation of the second core while data is being written to the target memory region for use by the second core. When the first core finishes writing the data, the first core may release the lock on the target memory region and the lock on the second core, whereupon the second core can process the data when it acquires the lock on the target memory region.

[0294] In yet other embodiments, as shown in FIG. 10C, the first core can initiate a direct data transfer to another core from a memory module within the same DPE, e.g., via a DMA engine within the memory module.

[0295] FIG. 23 shows another exemplary method of operation of the DPE array. Method 2300 is provided for illustrative purposes and is not intended to limit the configuration of the invention described within this disclosure. Method 2300 begins with the configuration data being loaded into the DPE array.

[0296] In block 2302, the first core places data in the accumulation register included therein. For example, the first core may execute a calculation regarding whether a part of the calculation, whether an intermediate result or a final result, will be directly provided to another core. In this case, the first core can load the data to be sent to the second core into the accumulation register included therein.

[0297] In block 2304, the first core transmits data from the accumulation register included therein to the second core from the cascade interface output of the first core. In block 2306, the second core receives the data from the first core on the cascade interface input of the second core. Then, the second core can process the data or store the data in memory.

[0298] In one or more embodiments, the use of the cascade interface by the core can be controlled by loading configuration data. For example, the cascade interface may be enabled or disabled between successive core pairs as needed for a particular application based on the configuration data. In a particular embodiment, when the cascade interface is enabled, the use of the cascade interface can be controlled based on the program code loaded into the program memory of the core. In other cases, the use of the cascade interface may be controlled by dedicated circuitry and configuration registers included in the core.

[0299] FIG. 24 shows another exemplary method of operation of the DPE array. Method 2400 is provided for illustrative purposes and is not intended to limit the configuration of the invention described within the present disclosure. Method 2400 begins with the state where configuration data is loaded into the DPE array.

[0300] In block 2402, the event logic in the first DPE locally detects one or more events within the first DPE. The events can be detected from the cores, from the memory modules, or from both the cores and the memory modules. In block 2404, the event broadcast circuit in the first DPE broadcasts events based on the configuration data loaded into the first DPE. The broadcast circuit can broadcast selected events among those generated in block 2402. The event broadcast circuit can also broadcast selected events that can be received from one or more other DPEs within the DPE array 102.

[0301] In block 2406, events from the DPE are propagated to tiles within the SoC interface block. For example, the events can be propagated in each of the four basic cardinal directions through the DPE in a pattern and / or route determined by the configuration data. The broadcast circuit within a particular DPE can be configured to propagate events to tiles within the SoC interface block.

[0302] In block 2408, the event logic in the tile of the SoC interface block optionally generates events. In block 2410, the tile of the SoC interface block optionally broadcasts events to other tiles within the SoC interface block. The broadcast circuit in the tile of the SoC interface block can broadcast selected events among those events generated by the tile itself and / or received from other sources (such as other tiles of the SoC interface block or DPEs).

[0303] In block 2412, the tile of the SoC interface block optionally generates one or more interrupts. The interrupts can be generated, for example, by interrupt handler 1434. The interrupt handler can generate one or more interrupts in response to receiving a particular event, combination of events, and / or sequence of events over time. The interrupt handler can send the generated interrupts to other circuits such as PS312 and / or circuits implemented within PL310.

[0304] In block 2414, the tile of the SoC interface block optionally sends an event to one or more other circuits. For example, the CDT circuit 1420 can packetize the event and send the event from the tile of the SoC interface block to PS312, a circuit within PL310, an external memory, or another destination having the SoC.

[0305] In one or more embodiments, PS312 can respond to an interrupt generated by a tile of the SoC interface block 104. For example, PS312 can reset the DPE array 102 in response to receiving a particular interrupt. In another example, PS312 can reconfigure (e.g., perform a partial reconfiguration) the DPE array 102 or a portion of the DPE array 102 in response to a particular interrupt. In another example, PS312 can take other actions such as loading new data into different memory modules of the DPE for use by cores within the DPE.

[0306] In the example of FIG. 24, PS312 operates in response to an interrupt. In other embodiments, PS312 may operate as a global controller for DPE array 102. PS312 can control application parameters stored in a memory module and used by one or more DPEs (e.g., cores) of DPE array 102 during runtime. As an illustrative and non-limiting example, one or more DPEs may operate as kernels implementing a filter. In this case, PS312 can execute program code that enables PS312 to calculate and / or modify filter coefficients, for example, dynamically during runtime of DPE array 102. PS312 can calculate and / or update coefficients in response to specific conditions and / or signals detected within the SoC. For example, PS312 can calculate new coefficients for a filter and / or write such coefficients to application memory (e.g., to one or more memory modules) in response to some detected situation. Examples of conditions that may cause PS312 to write data such as coefficients to a memory module can include receiving specific data from DPE array 102, receiving an interrupt from the SoC interface block, receiving event data from DPE array 102, receiving a signal from a source external to the SoC, receiving another signal from within the SoC, and / or receiving new and / or updated coefficients from a source within the SoC or external to the SoC, but are not limited thereto. PS312 can calculate new coefficients and / or write new coefficients to application data, e.g., to a memory module utilized by a core.

[0307] In another example, the PS312 can execute a debugger application that can perform actions such as starting, stopping, and / or single-stepping the DPE. The PS312 can control the start, stop, and / or single-stepping of the DPE via the NoC308. In other examples, the circuitry implemented in the PL310 may be able to control the operation of the DPE using debug operations.

[0308] For purposes of explanation, specific names are set forth in order to provide a thorough understanding of the various concepts of the invention disclosed herein. However, the terminology used herein is for the purpose of describing particular embodiments of the invention only and is not intended to be limiting.

[0309] As defined herein, the singular forms "a", "an", and "the" are intended to include the plural forms as well, unless the context clearly indicates otherwise.

[0310] As defined herein, the terms "at least one", "one or more", and "and / or" are open-ended expressions that are both conjunctive and disjunctive in operation, unless otherwise expressly stated. For example, each of the expressions "at least one of A, B, and C", "at least one of A, B, or C", "one or more of A, B, and C", "one or more of A, B, or C", and "A, B, and / or C" means either only A, only B, only C, A and B together, A and C together, B and C together, or A, B, and C together.

[0311] As defined herein, the term "automatically" means without human intervention.

[0312] As defined herein, the term "if" means, depending on the context, "when" or "upon" or "in response to" or "responsive to". Thus, the expressions "if it is determined" or "if [a stated condition or event] is detected" can be interpreted, depending on the context, as "upon determining" or "in response to determining" or "upon detecting [the stated condition or event]" or "in response to detecting [the stated condition or event]" or "responsive to detecting [the stated condition or event]".

[0313] As defined herein, the phrase "in response to" and similar phrases as described above, e.g., "if", "when", or "upon", mean to readily respond or react to an action or event. The response or reaction is performed automatically. Thus, when a second action is performed "in response to" a first action, there is a causal relationship between the occurrence of the first action and the occurrence of the second action. The expression "in response to" indicates a causal relationship.

[0314] As defined herein, the phrases "one embodiment," "an embodiment," "one or more embodiments," "a particular embodiment," or similar language mean that a particular feature, structure, or characteristic described in connection with the embodiment is included in at least one embodiment within the disclosure. Thus, the appearances of the phrases "in one embodiment," "in an embodiment," "in one or more embodiments," "in a particular embodiment," and similar language throughout this disclosure are not necessarily all referring to the same embodiment. The terms "embodiment" and "arrangement" are used interchangeably within the disclosure.

[0315] As defined herein, the term "substantially" means that the recited characteristic, parameter, or value need not be achieved exactly, but that deviations or variations, including, for example, tolerances, measurement error, measurement accuracy limitations, and other factors known to those of skill in the art, may occur in amounts that do not preclude the effect intended to be provided by the characteristic.

[0316] The terms first, second, etc. may be used herein to describe various elements. These elements should not be limited by these terms, as these terms are only used to distinguish one element from another, unless specifically stated otherwise or unless it is made clear otherwise by the context.

[0317] Flowcharts and block diagrams of the drawings illustrate the architecture, functionality, and operation of possible realizations of systems, devices, and / or methods according to various aspects of the present invention. In some alternative realizations, the operations recited in the blocks may be performed out of the order shown in the figures. For example, two blocks shown in succession may be executed substantially simultaneously or, depending on the functionality involved, may sometimes be executed in the reverse order. In other instances, the blocks may generally be executed in ascending numerical order, but in still other instances, one or more of the blocks may be executed in various orders and the results thereof may be stored and utilized in subsequent blocks or other blocks that do not immediately follow.

[0318] All means or step plus function elements corresponding structural, material, acts, and equivalents found in the claims are intended to include any structural, material, or act for performing the function in combination with other claimed elements specifically claimed.

[0319] In one or more embodiments, a device may include a plurality of DPEs. Each DPE may include a core and a memory module. Each core may be configured to access a memory module within the same DPE and a memory module within at least one other DPE of the plurality of DPEs.

[0320] In one aspect, each core may be configured to access memory modules of a plurality of neighboring DPEs.

[0321] In another aspect, the cores of the plurality of DPEs may be directly coupled. In another aspect, each of the plurality of DPEs is a hardwired and programmable circuit block.

[0322] In another aspect, each DPE may include an interconnect circuit including a stream switch configured to communicate with one or more DPEs selected from a plurality of DPEs. The stream switch may be programmable to communicate with one or more selected DPEs, e.g., other DPEs.

[0323] The device may also include a subsystem and a SoC interface block configured to couple the plurality of DPEs to the subsystem of the device. In one aspect, the subsystem includes programmable logic. In another aspect, the subsystem includes a processor configured to execute program code. In yet another aspect, the subsystem includes an application specific integrated circuit and / or an analog / mixed signal circuit.

[0324] In another aspect, the stream switch is coupled to the SoC interface block and configured to communicate with the subsystem of the device.

[0325] In another aspect, the interconnect circuit of each DPE may include a memory-mapped switch coupled to the SoC interface block, and the memory-mapped switch is configured to communicate configuration data for programming the DPE from the SoC interface block. The memory-mapped switch may be configured to communicate at least one of control data or debug data with the SoC interface block.

[0326] In another aspect, the plurality of DPEs may be interconnected by an event broadcast network.

[0327] In another aspect, the SoC interface block may be configured to exchange events between the subsystem and an event broadcast network of the plurality of DPEs.

[0328] In one or more embodiments, a method may include a first core of a first data processing engine generating data, the first core writing the data to a first memory module within the first data processing engine, and a second core of a second data processing engine reading the data from the first memory module.

[0329] In one aspect, the method may include the first DPE and the second DPE being neighboring DPEs.

[0330] In another aspect, the method may further include the first core being able to directly provide additional application data to the second core via a cascade interface.

[0331] In another aspect, the method may further include the first core being able to provide application data to a third DPE via a stream switch.

[0332] In another aspect, the method may include programming the first DPE to communicate with one or more other selected DPEs including the second DPE.

[0333] In one or more embodiments, a device may include a plurality of data processing engines, a subsystem, and a SoC interface block coupled to the plurality of data processing engines and the subsystem. The SoC interface block may be configured to exchange data between the subsystem and the plurality of data processing engines.

[0334] In one aspect, the subsystem includes programmable logic. In another aspect, the subsystem includes a processor configured to execute program code. In another aspect, the subsystem includes an application specific integrated circuit and / or an analog / mixed signal circuit.

[0335] In another aspect, the SoC interface block includes a plurality of tiles, and each tile is configured to communicate with a subset of the plurality of DPEs.

[0336] In another aspect, each tile may include a memory-mapped switch configured to provide a first portion of configuration data to at least one neighboring tile and a second portion of configuration data to at least one of the subsets of the plurality of DPEs.

[0337] In another aspect, each tile can include a stream switch configured to provide first data to at least one neighboring tile and second data to at least one of the plurality of DPEs.

[0338] In another aspect, each tile may include an event broadcast circuit configured to receive events generated within the tile and events from circuits external to the tile, and the event broadcast circuit is programmable to provide selected ones of the events to selected destinations.

[0339] In another aspect, the SoC interface block may include control, debug, and trace circuitry configured to packetize selected events and provide the packetized selected events to a subsystem.

[0340] In another aspect, the SoC interface block can include an interface coupling the event broadcast circuit to a subsystem.

[0341] In one or more embodiments, a tile for a SoC interface block can include a memory-mapped switch configured to provide a first portion of configuration data to neighboring tiles and a second portion of configuration data to a data processing engine among a plurality of data processing engines. The tile can include a stream switch configured to provide first data to at least one neighboring tile and second data to a data processing engine among a plurality of data processing engines. The tile can include an event broadcast circuit configured to receive events generated within the tile and events from circuitry external to the tile, and the event broadcast circuit is programmable to provide selected ones of the events to selected destinations. The tile can include an interface circuit coupling the memory-mapped switch, the stream switch, and the event broadcast circuit to a subsystem of a device including the tile.

[0342] In one aspect, the subsystem includes programmable logic. In another aspect, the subsystem includes a processor configured to execute program code. In another aspect, the subsystem includes an application specific integrated circuit and / or an analog / mixed signal circuit.

[0343] In another aspect, the event broadcast circuit is programmable to provide an event generated within the tile or an event received from at least one of a plurality of DPEs to the subsystem.

[0344] In another aspect, the event broadcast circuit is programmable to provide an event generated within the subsystem to at least one neighboring tile or at least one of a plurality of DPEs.

[0345] In another aspect, the tile may include an interrupt handler configured to selectively generate an interrupt to the device's processor based on an event received from the event broadcast circuit.

[0346] In another aspect, the tile can include a clock generation circuit configured to generate a clock signal distributed to a plurality of DPEs.

[0347] In another aspect, the interface circuit may include a stream multiplexer / demultiplexer, a programmable logic interface, a direct memory access engine, and a NoC stream interface. The stream multiplexer / demultiplexer can couple a stream switch to the programmable logic interface, the direct memory access engine, and the network-on-chip stream interface. The stream multiplexer / demultiplexer is programmable to route data between the stream switch, the programmable logic interface, the direct memory access engine, and the NoC stream interface.

[0348] In another aspect, the tile can include a switch coupled to a DMA engine and a NoC stream interface, the switch selectively coupling the DMA engine or the NoC stream interface to the NoC. The tile may also include a bridge circuit coupling the NoC to a memory-mapped switch. The bridge circuit is configured to convert data from the NoC into a format usable by the memory-mapped switch.

[0349] In one or more embodiments, a device may include a plurality of data processing engines. Each of the data processing engines may include a core and a memory module. The plurality of data processing engines may be arranged in a plurality of columns. Each core may be configured to communicate with other neighboring data processing engines among the plurality of data processing engines by shared access to the memory modules of the other neighboring data processing engines.

[0350] In one aspect, the memory module of each DPE includes a memory and a plurality of memory interfaces to the memory. The first memory interface among the plurality of memory interfaces may be coupled to a core within the same DPE, and each of the other memory interfaces among the plurality of memory interfaces may be coupled to a core of a different DPE among the plurality of DPEs.

[0351] In another aspect, the plurality of DPEs may be further arranged in a plurality of columns, the cores of the plurality of DPEs within a column are aligned, and the memory modules of the plurality of DPEs within a column are aligned.

[0352] In another aspect, the memory module of a selected DPE can include a first memory interface coupled to a core of the DPE immediately above the selected DPE, a second memory interface coupled to a core within the selected DPE, a third memory interface coupled to a core of the DPE immediately adjacent to the selected DPE, and a fourth memory interface coupled to a core of the DPE immediately below the selected DPE.

[0353] In another aspect, a selected DPE is configured to communicate with a group consisting of at least 10 DPEs among the plurality of DPEs via shared access to a memory module.

[0354] In another aspect, at least two DPEs among the group are configured to access two or more memory modules among the group consisting of at least 10 DPEs among the plurality of DPEs.

[0355] In another aspect, the plurality of rows of DPEs can include a first row including a first subset of the plurality of DPEs and a second row including a second subset of the plurality of DPEs, and the orientation of each DPE in the second row is horizontally inverted with respect to the orientation of each DPE in the first row.

[0356] In another aspect, the memory module of a selected DPE can include a first memory interface coupled to the core of the DPE immediately above the selected DPE, a second memory interface coupled to the core within the selected DPE, a third memory interface coupled to the core of the DPE immediately adjacent to the selected DPE, and a fourth memory interface coupled to the core of the DPE immediately below the selected DPE.

[0357] In another aspect, a selected DPE can be configured to communicate with a group of at least eight DPEs among the plurality of DPEs via shared access to a memory module.

[0358] In another aspect, at least four DPEs in the group are configured to access two or more memory modules in a group of at least eight DPEs among the plurality of DPEs.

[0359] In one or more embodiments, a device can include a plurality of data processing engines. Each data processing engine can include a memory pool having a plurality of memory banks, a plurality of cores each coupled to the memory pool and configured to access the plurality of memory banks, a memory-mapped switch coupled to the memory pool and a memory-mapped switch of at least one neighboring data processing engine, and a stream switch coupled to each of the plurality of cores and a stream switch of at least one neighboring data processing engine.

[0360] In one aspect, the memory pool may include a crossbar coupled to each of a plurality of memory banks, and an interface coupled to each of a plurality of cores and the crossbar.

[0361] In another aspect, each DPE can include a direct memory access engine coupled to a memory pool and a stream switch, the direct memory access engine being configured to provide data from the memory pool to the stream switch and write data from the stream switch to the memory pool.

[0362] In another aspect, the memory pool may include a further interface coupled to the crossbar and the direct memory access engine.

[0363] In another aspect, each of the plurality of cores has shared access to the plurality of memory banks. In another aspect, within each DPE, a memory-mapped switch may be configured to receive configuration data for programming the DPE.

[0364] In another aspect, the stream switch is programmable to establish connections with different ones of the plurality of DPEs based on the configuration data.

[0365] In another aspect, the plurality of cores within each tile may be directly coupled. In another aspect, within each DPE, the first of the plurality of cores may be directly coupled to a core within a first neighboring DPE, and the last of the plurality of cores may be directly coupled to a core within a second neighboring DPE.

[0366] In another aspect, each of the plurality of cores may be programmable to be deactivated. The description of the configuration of the present invention provided in this specification is for illustrative purposes and is not intended to be exhaustive or limited to the disclosed forms and examples. The technical terms used in this specification are selected to explain the principles of the configuration of the present invention, the practical applications to the technologies seen in the market or the technical improvements, and / or to enable those skilled in the art to understand the configuration of the present invention disclosed in this specification. Modifications and variations will be apparent to those skilled in the art without departing from the scope and spirit of the described configuration of the invention. Therefore, rather than the foregoing disclosure, reference should be made to the claims as indicating the scope of such features and examples of implementation.

Claims

1. A device comprising: a plurality of data processing engines, each data processing engine including a core and a memory module, each core being configured to access the memory module in the same data processing engine and the memory module in at least one other data processing engine among the plurality of data processing engines, the cores of the plurality of data processing engines being further serially connected by a cascade connection, each cascade connection directly connecting a pair of the cores, a first core of the pair connected by the cascade connection being configured to write data to the memory module in the same data processing engine or the memory module in the at least one other data processing engine and / or write the data to a second core via the cascade connection.

2. The device according to claim 1, wherein each cascade connection connects a register in the first core of the pair to the second core of the pair.

3. The device according to claim 1, wherein each cascade connection is implemented using a cascade interface of the first core of the pair and a cascade interface of the second core of the pair, and each cascade interface is independently activated based on configuration data loaded into a corresponding configuration register of the corresponding core.

4. The device according to any one of claims 1 to 3, wherein each of the plurality of data processing engines is a hardwired and programmable circuit block.

5. Each data processing engine includes an interconnect circuit including a stream switch configured to communicate with one or more data processing engines selected from the plurality of data processing engines, the stream switch being programmable to communicate with the one or more selected data processing engines. The device according to any one of claims 1 to 4.

6. a subsystem, and a system-on-chip (SoC) interface block configured to couple the plurality of data processing engines to the subsystem of the device.

7. The device according to claim 6, wherein the subsystem includes programmable logic.

8. The device according to claim 6 or claim 7, wherein the subsystem includes a processor configured to execute program code. **Claim 9** The device according to any one of claims 6 to 8, wherein the subsystem includes at least one of an application-specific integrated circuit or an analog / mixed-signal circuit. **Claim 10** The device according to any one of claims 6 to 9, wherein the stream switch is coupled to the SoC interface block and is configured to communicate with the subsystem of the device. **Claim 11** A device, comprising a plurality of data processing engines, each data processing engine including a core and a memory module, each core being configured to access the memory module in the same data processing engine and the memory module in at least one other data processing engine of the plurality of data processing engines, each data processing engine including an interconnect circuit including a stream switch configured to communicate with one or more data processing engines selected from the plurality of data processing engines, the device further comprising a subsystem and a system-on-chip (SoC) interface block configured to couple the plurality of data processing engines to the subsystem of the device, the interconnect circuit of each data processing engine further including a memory-mapped switch coupled to the SoC interface block, the memory-mapped switch being configured to communicate configuration data to program the data processing engine from the SoC interface block. **Claim 12** The device according to claim 11, wherein the memory-mapped switch is further configured to communicate at least one of control data or debug data with the SoC interface block. **Claim 13** The device according to claim 6, wherein the plurality of data processing engines are interconnected by an event broadcast network. **Claim 14** The device according to claim 13, wherein the SoC interface block is configured to exchange events between the subsystem and the event broadcast network of the plurality of data processing engines.

Citation Information

Patent Citations

  • Multiprocessor computer memory architecture

    JP1989500306A

  • Array type processor

    JP2001312481A

  • Programmable single-chip devices and related development environments

    JP2003534596A

  • Integrated circuits with programmable circuits and embedded processor systems

    JP2014515843A

  • Automated modification of configuration settings of an integrated circuit

    US9652410B1