Data processing engine arrangement in a device
By designing multiple data processing engines in integrated circuit equipment, and sharing and processing across engines through memory-mapping switches and stream switches, the problems of insufficient resource utilization and inefficiency in processing caused by the arrangement of data processing engines in the prior art are solved, and efficient data processing and low power consumption design are achieved.
Patent Information
- Application Number
- CN202510251193.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2018-04-03
- Filing Date
- 2019-04-02
- Publication Date
- 2025-06-06
AI Technical Summary
In the prior art, it is difficult to implement efficient data sharing and processing in the arrangement of data processing engines, resulting in insufficient resource utilization and low processing efficiency.
An integrated circuit device including multiple data processing engines is designed, each data processing engine is composed of core and memory modules, and data sharing and processing across the engine is realized through memory-mapping switches and stream switches.
Through this arrangement, efficient data sharing and processing between multiple data processing engines is realized, resource utilization and processing efficiency are improved, and power consumption and die area are reduced.
Smart Images

Figure CN120104553A_ABST
Abstract
Description
[0001] Divisional Application Instructions
[0002] This application is a divisional application of the Chinese national phase of the PCT patent application with an international application date of April 2, 2019, an international application number of PCT / US2019 / 025414, and an invention name of “Data Processing Engine Arrangement in a Device” and an application number of 201980024322.0. Technical Field
[0003] The present disclosure relates to integrated circuit devices and, more particularly, to devices including data processing engines and / or arrays of data processing engines. Background Art
[0004] A programmable integrated circuit (IC) refers to an IC that includes a programmable circuit device. An example of a programmable IC is a field programmable gate array (FPGA). An FPGA is characterized by including programmable circuit blocks. Examples of programmable circuit blocks include, but are not limited to, input / output blocks (IOBs), configurable logic blocks (CLBs), dedicated random access memory blocks (BRAMs), digital signal processing blocks (DSPs), processors, clock managers, and delay locked loops (DLLs).
[0005] By loading configuration data, sometimes referred to as a configuration bitstream, into the device, a circuit design can be physically implemented within the programmable circuitry of a programmable IC. The configuration data can be loaded into internal configuration memory cells of the device. The collective state of the various configuration memory cells determines the functionality of the programmable IC. For example, once the configuration data is loaded, the specific operations performed by the various programmable circuit blocks and the connectivity between the programmable circuit blocks of the programmable IC are defined by the collective state of the configuration memory cells. Summary of the invention
[0006] In one or more embodiments, the device may include multiple data processing engines. Each data processing engine may include a core and a memory module. Each core may be configured to access the memory module in the same data processing engine and the memory module in at least one other data engine in the multiple data processing engines.
[0007] In one or more embodiments, the method may include a first core of a first data processing engine generating data, the first core writing the data to a first memory module within the first data processing engine, and a second core of a second data processing engine reading the data from the first memory module.
[0008] In one or more embodiments, the device may include multiple data processing engines, subsystems, and a system-on-chip (SoC) interface block, and the system-on-chip (SoC) interface block is coupled to the multiple data processing engines and the subsystems. The SoC interface block can be configured to exchange data between the subsystem and the multiple data processing engines.
[0009] In one or more embodiments, a tile for a SoC interface block may include a memory-mapped switch configured to provide a first portion of configuration data to an adjacent tile and to provide a second portion of configuration data to a data processing engine in a plurality of data processing engines. The tile may include a stream switch configured to provide first data to at least one adjacent tile and to provide second data to a data processing engine in a plurality of data processing engines. The tile may include an event broadcast circuit device configured to receive events generated within the tile and events from circuit devices external to the tile, wherein the event broadcast circuit device is programmable to provide selected events in the event to a selected destination. The tile may include an interface circuit device that couples the memory-mapped switch, the stream switch, and the event broadcast circuit device to a subsystem of a device including the tile.
[0010] In one or more embodiments, the device may include multiple data processing engines. Each of the data processing engines may include a core and a memory module. The multiple data processing engines may be organized in multiple rows. Each core may be configured to communicate with an adjacent data processing engine through shared access to the memory modules of other adjacent data processing engines in the multiple data processing engines.
[0011] In one or more embodiments, the device may include a plurality of data processing engines. Each of the data processing engines may include: a memory pool having a plurality of memory banks; a plurality of cores, each coupled to the memory pool and configured to access the plurality of memory banks; a memory-mapped switch, the memory-mapped switch coupled to the memory pool and the memory-mapped switch of at least one adjacent data processing engine; and a stream switch, the stream switch coupled to each of the plurality of cores and coupled to the stream switch of at least one adjacent data processing engine.
[0012] This summary is provided only to introduce certain concepts, rather than to identify any key features or essential features of the claimed subject matter. Other features of the inventive arrangement will become apparent from the accompanying drawings and the following detailed description. BRIEF DESCRIPTION OF THE DRAWINGS
[0013] The arrangement of the present invention is shown by way of example in the accompanying drawings. However, the accompanying drawings should not be interpreted as limiting the arrangement of the present invention to only the specific implementation shown. After reading the following detailed description and referring to the accompanying drawings, various aspects and advantages will become apparent.
[0014] Figure 1 An example of a device including an array of data processing engines (DPEs) is shown.
[0015] Figure 2A , 2B , 2C, and 2D show example architectures for a device having one or more DPE arrays.
[0016] Figure 3 Another example architecture for a device including a DPE array is shown.
[0017] Figure 4A and Figure 4B An example of a multi-die implementation of a device with one or more DPE arrays is shown.
[0018] Figure 5A , 5B , 5C, 5D, 5E, 5F, and 5G illustrate example multi-die implementations of devices having DPE arrays.
[0019] Figure 6 An example architecture of a DPE for a DPE array is shown.
[0020] Figure 7 Example connectivity between multiple DPEs is shown.
[0021] Figure 8 Shows Figure 6 Further aspects of the example DPE architecture.
[0022] Fig. 9 Example connectivity of the cascade interfaces of the core of the DPE is shown.
[0023] Fig. 10A , 10B , 10C, 10D and 10E show examples of connectivity between DPEs.
[0024] Fig.11 An example of event handling circuitry within a DPE is shown.
[0025] Fig.12 Another example architecture for a DPE is shown.
[0026] Fig.13 An example architecture for a DPE array is shown.
[0027] Fig.14A ,14B 14C show a block example architecture for implementing a system-on-chip (SoC) interface block.
[0028] Fig.15 An example implementation of a programmable logic interface for a block diagram of a SoC interface block is shown.
[0029] Fig.16 An example implementation of a network-on-chip (NoC) streaming interface for a block diagram of a SoC interface block is shown.
[0030] Fig.17 An example implementation of a direct memory access (DMA) engine of a block diagram of a SoC interface block is shown.
[0031] Fig.18 An example architecture for multiple DPEs is shown.
[0032] Fig.19 Another example architecture for multiple DPEs is shown.
[0033] Fig. 20 An example method of configuring a DPE array is shown.
[0034] Fig.21 An example method of operation of a DPE array is shown.
[0035] Fig. 22 Another example method of operation of a DPE array is shown.
[0036] Fig.23 Another example method of operation of a DPE array is shown.
[0037] Fig.24 Another example method of operation of a DPE array is shown. DETAILED DESCRIPTION
[0038] Although the present disclosure is concluded with claims defining novel features, it is believed that the various features described in the present disclosure will be better understood when the specification is considered in conjunction with the accompanying drawings. The processes, machines, manufactures, and any variations thereof described herein are provided for illustrative purposes. The specific structural and functional details described in the present disclosure are not to be construed as limiting, but merely as a basis for the claims and as a representative basis for teaching those skilled in the art to employ the described features in various aspects with virtually any appropriate detailed structure. In addition, the terms and phrases used in the present disclosure are not intended to be limiting, but rather to provide an understandable description of the described features.
[0039] The present disclosure relates to an integrated circuit device (device) including one or more data processing engines (DPEs) and / or DPE arrays. A DPE array refers to a plurality of hard-wired circuit blocks. The plurality of circuit blocks may be programmable. A DPE array may include a plurality of DPEs and a system-on-chip (SoC) interface block. Typically, a DPE includes a core capable of providing data processing capabilities. The DPE also includes a memory module that can be accessed by a core or multiple cores in the DPE. In a particular embodiment, the memory module of the DPE may also be accessed by one or more cores in different DPEs of the DPE array.
[0040] The DPE may also include a DPE interconnect. A DPE interconnect refers to a circuit device that enables communication with other DPEs of a DPE array and / or communication with different subsystems of a device that includes the DPE array. The DPE interconnect may also support configuration of the DPE. In certain embodiments, the DPE interconnect may be capable of transmitting control data and / or debug data.
[0041] The DPE array may be organized using any of a variety of different architectures. In one or more embodiments, the DPE array may be organized in one or more rows and in one or more columns. In some cases, the columns and / or rows of the DPEs are aligned. In some embodiments, each DPE may include a single core coupled to a memory module. In other embodiments, one or more DPEs or each DPE in the DPE array may be implemented to include two or more cores coupled to a memory module.
[0042] In one or more embodiments, the DPE array is implemented as a homogeneous structure, where each DPE is identical to each other. In other embodiments, the DPE array is implemented as a heterogeneous structure, where the DPE array includes two or more different types of DPEs. For example, the DPE array may include a DPE with a single core, a DPE with multiple cores, a DPE with different types of cores, and / or a DPE with different physical architectures.
[0043] The DPE array may be implemented in various sizes. For example, the DPE array may be implemented to span the entire width and / or length of the die of the device. In another example, the DPE array may be implemented to span a portion of the entire width and / or length of such a die. In further embodiments, more than one DPE array may be implemented within the die, wherein different DPE arrays in the DPE array are distributed to different areas on the die, have different sizes, have different shapes, and / or have different architectures as described herein (e.g., aligned rows and / or columns, homogeneous and / or heterogeneous). In addition, the DPE array may include different numbers of rows of DPEs and / or different numbers of columns of DPEs.
[0044] The DPE array can be used with and coupled to any of a variety of subsystems within a device. Such subsystems may include, but are not limited to, processors and / or processor systems, programmable logic, and / or networks on chip (NoCs). In particular embodiments, the NoC may be programmable. Additional examples of subsystems that may be included in a device and coupled to a DPE array may include, but are not limited to, application specific integrated circuits (ASICs), hardwired circuit blocks, analog and / or mixed signal circuit devices, graphics processing units (GPUs), and / or general purpose processors (e.g., central processing units or CPUs). An example of a CPU is a processor having an x86-type architecture. In this specification, the term "ASIC" may refer to an IC, a die, and / or a portion of a die that includes a dedicated circuit combined with another type or types of circuits; and / or to an IC and / or die formed entirely of dedicated circuits.
[0045] In certain embodiments, a device including one or more DPE arrays may be implemented using a single die architecture. In that case, the DPE array and any other subsystems used with the DPE array are implemented on the same die of the device. In other embodiments, a device including one or more DPE arrays may be implemented as a multi-die device including two or more dies. In some multi-die devices, a DPE array or multiple DPE arrays may be implemented on one die, while one or more other subsystems are implemented in one or more other dies. In other multi-die devices, a DPE array or multiple DPE arrays may be implemented in one or more dies in combination with one or more other subsystems of the multi-die device (e.g., where the DPE array is implemented in the same die as at least one subsystem).
[0046] The DPE array as described in the present disclosure is capable of implementing an optimized digital signal processing (DSP) architecture. The DSP architecture is capable of efficiently performing any of a variety of different operations. Examples of the types of operations that can be performed by the architecture include, but are not limited to, operations related to radio, decision feedback equalization (DFE), 5G / baseband, wireless backhaul, machine learning, automotive driver assistance, embedded vision, cable access, and / or radar. Compared to other solutions utilizing conventional programmable (e.g., FPGA-type) circuits, the DPE array described herein is capable of performing such operations while consuming less power. In addition, compared to other solutions utilizing conventional programmable circuit devices, a DPE array-based solution can be implemented using less die area. The DPE array is also capable of performing the operations described herein while meeting predictable and guaranteed data throughput and latency metrics.
[0047] Other aspects of the arrangement of the present invention are described in more detail below with reference to the accompanying drawings. For the purpose of simplifying the description, the elements shown in the accompanying drawings need not be drawn to scale. For example, for the sake of clarity, the size of some elements may be enlarged relative to other elements. Further, where appropriate, reference numerals are repeated in the accompanying drawings to indicate corresponding, similar or similar features.
[0048] Figure 1 An example of a device 100 including a DPE array 102 is shown. Figure 1 In the example of , the DPE array 102 includes a SoC interface block 104. The device 100 also includes one or more subsystems 106-1 to 106-N. In one or more embodiments, the device 100 is implemented as a system-on-chip (SoC) type device. Generally, a SoC refers to an IC that includes two or more subsystems that can interact with each other. As an example, the SoC may include a processor that executes program code and one or more other circuits. The other circuits may be implemented as hard-wired circuit devices, programmable circuit devices, other subsystems, and / or any combination thereof. The circuits may operate in cooperation with each other and / or with the processor.
[0049] The DPE array 102 is formed by a plurality of interconnected DPEs. Each DPE in the DPE is a hardwired circuit block. Each DPE may be programmable. The SoC interface block 104 may include one or more tiles. Each tile in the tiles of the SoC interface block 104 may be hardwired. Each tile of the SoC interface block 104 may be programmable. The SoC interface block 104 provides an interface between the DPE array 102 (e.g., DPE) and other parts of the SoC (such as the subsystem 106 of the device 100). For example, the subsystems 106-1 to 106-N may represent: one or more processors and / or processor systems (e.g., CPUs, general-purpose processors, and / or GPUs), programmable logic, NoCs, ASICs, analog and / or mixed-signal circuit devices, and / or hardwired circuit blocks, or any combination thereof.
[0050] In one or more embodiments, the device 100 is implemented using a single-die architecture. In that case, the DPE array 102 and at least one subsystem 106 may be included or implemented in a single die. In one or more other embodiments, the device 100 is implemented using a multi-die architecture. In that case, the DPE array 102 and the subsystem 106 may be implemented across two or more dies. For example, the DPE array 102 may be implemented in one die and the subsystem 106 may be implemented in one or more other dies. In another example, the SoC interface block 104 may be implemented in a different die than the DPE of the DPE array 102. In yet another example, the DPE array 102 and at least one subsystem 106 may be implemented in the same die, while other subsystems and / or other DPE arrays may be implemented in other dies. The following is a description of the various embodiments of the present invention in conjunction with FIG. 2, Figure 3 4 and 5 describe additional examples of single-die and multi-die architectures in more detail.
[0051] Figure 2A , Figure 2B , Figure 2C and Figure 2D (collectively referred to as "FIG. 2") illustrates an example architecture for a device including one or more DPE arrays 102. More particularly, FIG2 illustrates an example of a single die architecture for device 100. For purposes of illustration, SoC interface block 104 is not shown in FIG2.
[0052] Figure 2A An example architecture for a device 100 including a single DPE array is shown. Figure 2A In the example of FIG. 1 , DPE array 102 is implemented in device 100 with subsystem 106-1. DPE array 102 and subsystem 106-1 are implemented in the same die. DPE array 102 can extend across the entire width of the die of device 100 or partially across the die of device 100. As shown, DPE array 102 is implemented in the top area of device 100. However, it should be understood that DPE array 102 can be implemented in another area of device 100. In this way, Figure 2A The placement and / or size of DPE array 102 in is not intended to be limiting. DPE array 102 may be coupled to subsystem 106-1 via SoC interface block 104 (not shown).
[0053] Figure 2B An example architecture for a device 100 including multiple DPE arrays is shown. Figure 2B In the example of FIG. 1 , multiple DPE arrays are implemented and are depicted as DPE array 102 - 1 and DPE array 102 - 2 . Figure 2BMultiple DPE arrays are shown that may be implemented in the same die of device 100 along with subsystem 106-1. DPE array 102-1 and / or DPE array 102-2 may extend across the entire width of the die of device 100 or only partially across the die of device 100. As shown, DPE array 102-1 is implemented in the top region of device 100 and DPE array 102-2 is implemented in the bottom region of device 100. As noted, Figure 2B The placement and / or size of DPE arrays 102-1 and 102-2 in FIG. 1 are not intended to be limiting.
[0054] In one or more embodiments, DPE array 102-1 and DPE array 102-2 may be substantially similar or identical. For example, DPE array 102-1 may be identical to DPE array 102-2 in terms of size, shape, number of DPEs, and whether the DPEs in each respective DPE array are homogeneous or of the same type and order. In one or more other embodiments, DPE array 102-1 may be different from DPE array 102-2. For example, DPE array 102-1 may be different from DPE array 102-2 in terms of size, shape, number of DPEs, type of core, and whether the DPEs in each respective DPE array are homogeneous or of different types and / or order.
[0055] In one or more embodiments, each of DPE arrays 102-1 and 102-2 is coupled to subsystem 106-1 via its own SoC interface block (not shown). For example, a first SoC interface block may be included and used to couple DPE array 102-1 to subsystem 106-1, while a second SoC interface block may be included and used to couple DPE array 102-2 to subsystem 106-1. In another embodiment, a single SoC interface block may be used to couple both DPE array 102-1 and DPE array 102-2 to subsystem 106-1. In the latter case, for example, one of the DPE arrays may not include a SoC interface block. The DPEs in the array may be coupled to subsystem 106-1 using the SoC interface block of another DPE array.
[0056] Figure 2C An example architecture for a device 100 including multiple DPE arrays and multiple subsystems is shown. Figure 2C In the example of FIG. 1 , multiple DPE arrays are implemented and are depicted as DPE array 102 - 1 and DPE array 102 - 2 . Figure 2CIt is shown that multiple DPE arrays may be implemented in the same die of device 100 and that the placement or location of DPE array 102 may vary. Additionally, DPE arrays 102-1 and 102-2 are implemented in the same die as subsystems 106-1 and 106-2.
[0057] exist Figure 2C In the example of , DPE array 102-1 and DPE array 102-2 do not extend across the entire width of the die of device 100. Instead, each of DPE arrays 102-1 and 102-2 extends partially across the die of device 100 and is also implemented in an area that is a portion of the width of the die of device 100. Figure 2B As in the example, Figure 2C DPE array 102-1 and DPE array 102-2 may be substantially similar or identical, or may be different.
[0058] In one or more embodiments, each of DPE array 102-1 and DPE array 102-2 is coupled to subsystem 106-1 and / or subsystem 106-2 via its own SoC interface block (not shown). In an illustrative and non-limiting example, a first SoC interface block may be included and used to couple DPE array 102-1 to subsystem 106-1, while a second SoC interface block may be included and used to couple DPE array 102-2 to subsystem 106-2. In that case, each DPE array communicates with a subset of the available subsystems of device 100. In another example, a first SoC interface block may be included and used to couple DPE array 102-1 to subsystems 106-1 and 106-2, while a second SoC interface block may be included and used to couple DPE array 102-2 and subsystems 106-1 and 106-2. In yet another example, a single SoC interface block may be used to couple both DPE array 102-1 and DPE array 102-2 to subsystem 106-1 and / or subsystem 106-2. Figure 2C The placement and / or size of DPE arrays 102-1 and 102-2 in FIG. 1 are not intended to be limiting.
[0059] Figure 2D Another example architecture for a device 100 including multiple DPE arrays and multiple subsystems is shown. Figure 2D In the example of FIG. 1 , multiple DPE arrays are implemented and are depicted as DPE array 102 - 1 and DPE array 102 - 2 . Figure 2D It is also shown that multiple DPE arrays can be implemented in the same die of device 100, and the placement and / or location of DPE array 102 can vary. Figure 2DIn the example of , DPE array 102-1 and DPE array 102-2 do not extend across the entire width of the die of device 100. Instead, each of DPE arrays 102-1 and 102-2 is implemented in an area that is a portion of the width of the die of device 100. In addition, Figure 2D The device 100 includes subsystems 106-1, 106-2, 106-3, and 106-4 within the same die as the DPE arrays 102-1 and 102-2. Figure 2B As in the example, Figure 2D DPE array 102-1 and DPE array 102-2 may be substantially similar or identical, or may be different.
[0060] exist Figure 2D In the example of , the connectivity between the DPE array and the subsystems can vary. In some cases, the DPE array can be coupled to only a subset of the available subsystems in the device 100. In other cases, the DPE array can be coupled to more than one subsystem or every subsystem in the device 100.
[0061] The example of FIG. 2 is provided for purposes of illustration and not limitation. A device having a single die may include one or more different DPE arrays located in different regions of the die. The number, placement, and / or size of the DPE arrays may vary. Additionally, the DPE arrays may be the same or different. One or more DPE arrays may be implemented in conjunction with one or more of the different types of subsystems described within the present disclosure and / or any combination thereof.
[0062] In one or more embodiments, two or more DPE arrays may be configured to communicate directly with each other. For example, DPE array 102-1 may be able to communicate directly with DPE array 102-2 and / or with additional DPE arrays. In certain embodiments, DPE arrays may communicate with DPE array 102-2 and / or with other DPE arrays via one or more SoC interface blocks.
[0063] Figure 3 Another example architecture for device 100 is shown. Figure 3 In the example of FIG. 1 , the DPE array 102 is implemented as a two-dimensional array of DPEs 304 including the SoC interface block 104. The DPE array 102 may be implemented using any of a variety of different architectures described in more detail below. For purposes of illustration and not limitation, Figure 3 Shown as a combination Fig.19The DPEs 304 are described in more detail as being arranged in aligned rows and aligned columns. However, in other embodiments, the DPEs 304 may be arranged in which the DPEs in selected rows and / or columns are horizontally inverted or flipped relative to the DPEs in adjacent rows and / or columns. Fig.18 An example of horizontal inversion of DPEs is described. In one or more other embodiments, rows and / or columns of DPEs may be offset relative to adjacent rows and / or columns. One or more or all DPEs 304 may be implemented to include, for example, a combination of Figure 6 and Figure 8 A single core as generally described may be implemented as including a combination of Fig.12 Two or more cores of overall description.
[0064] The SoC interface block 104 can couple the DPE 304 to one or more other subsystems of the device 100. In one or more embodiments, the SoC interface block 104 is coupled to adjacent DPEs 304. For example, the SoC interface block 104 can be directly coupled to each DPE 304 in the bottom row of DPEs in the DPE array 102. In the illustration, the SoC interface block 104 can be directly connected to the DPEs 304-1, 304-2, 304-3, 304-4, 304-5, 304-6, 304-7, 304-8, 304-9, and 304-10.
[0065] Figure 3 104 is provided for illustration purposes. In other embodiments, the SoC interface block 104 may be located at the top of the DPE array 102, on the left side of the DPE array 102 (e.g., as a column), on the right side of the DPE array 102 (e.g., as a column), or at multiple locations in and around the DPE array 102 (e.g., as one or more intermediate rows and / or columns within the DPE array 102). Depending on the layout and location of the SoC interface block 104, the specific DPEs coupled to the SoC interface block 104 may vary.
[0066] For purposes of illustration and not limitation, if SoC interface block 104 is located on the left side of DPE 304, SoC interface block 104 may be directly coupled to the left column of DPEs including DPE 304-1, DPE 304-11, DPE 304-21, and DPE 304-31. If SoC interface block 104 is located on the right side of DPE 304, SoC interface block 104 may be directly coupled to the right column of DPEs including DPE 304-10, DPE 304-20, DPE 304-30, and DPE 304-40. If SoC interface block 104 is located at the top of DPE 304, SoC interface block 104 may be coupled to a top row of DPEs including DPE 304-31, DPE 304-32, DPE 304-33, DPE 304-34, DPE 304-35, DPE 304-36, DPE 304-37, DPE 304-38, DPE 304-39, and DPE 304-40. If SoC interface block 104 is located at multiple locations, the specific DPEs directly connected to SoC interface block 104 may vary. For example, if SoC interface blocks are implemented as rows and / or columns within DPE array 102, the DPEs directly coupled to SoC interface block 104 may be those DPEs adjacent to SoC interface block 104 on one or more or each side of SoC interface block 104.
[0067] The DPEs 304 are interconnected via DPE interconnects (not shown) which, when considered collectively, form a DPE interconnect network. Thus, the SoC interface block 104 is able to communicate with any one of the DPEs 304 of the DPE array 102 by communicating with one or more selected DPEs 304 of the DPE array 102 directly connected to the SoC interface block 104 and utilizing the DPE interconnect network formed by the DPE interconnects implemented within each respective DPE 304.
[0068] SoC interface block 104 can couple each DPE 304 within DPE array 102 with one or more other subsystems of device 100. For purposes of illustration, device 100 includes a subsystem (e.g., subsystem 106) such as: NoC 308, programmable logic (PL) 310, processor system (PS) 312, and / or any of hardwired circuit blocks 314, 316, 318, 320, and / or 322. For example, SoC interface block 104 can establish a connection between a selected DPE 304 and PL 310. SoC interface block 104 can also establish a connection between a selected DPE 304 and NoC 308. Through NoC 308, a selected DPE 304 can communicate with PS 312 and / or hardwired circuit blocks 320 and 322. The selected DPE 304 is capable of communicating with hardwired circuit blocks 314-318 via the SoC interface block 104 and the PL 310. In certain embodiments, the SoC interface block 104 may be directly coupled to one or more subsystems of the device 100. For example, the SoC interface block 104 may be directly coupled to the PS 312 and / or other hardwired circuit blocks. In certain embodiments, the hardwired circuit blocks 314-322 may be considered examples of ASICs.
[0069] In one or more embodiments, the DPE array 102 includes a single clock domain. Other subsystems (such as NoC 308, PL 310, PS 312, and various hardwired circuit blocks) may be in one or more separate or different clock domains. Moreover, the DPE array 102 may include additional clocks that may be used to interface with other subsystems in the subsystem. In a particular embodiment, the SoC interface block 104 includes a clock signal generator that is capable of generating one or more clock signals that may be provided or distributed to the DPEs 304 of the DPE array 102.
[0070] The DPE array 102 can be programmed by loading configuration data into internal configuration memory cells (also referred to herein as "configuration registers") that define the connectivity between the DPE 304 and the SoC interface block 104 and how the DPE 304 and the SoC interface block 104 operate. For example, the DPE 304 and the SoC interface block 104 are programmed for a specific DPE 304 or a group of DPEs 304 to communicate with a subsystem. Similarly, the DPE is programmed for one or more specific DPEs 304 to communicate with one or more other DPEs 304. The DPE 304 and the SoC interface block 104 can be programmed by loading configuration data into configuration registers within the DPE 304 and the SoC interface block 104, respectively. In another example, a clock signal generator that is part of the SoC interface block 104 can be programmed using configuration data to change the clock frequency provided to the DPE array 102.
[0071] NoC 308 provides connectivity to PL 310, PS 312, and to selected ones of the hardwired circuit blocks (e.g., circuit blocks 320 and 322). Figure 3 In the example of , NoC 308 is programmable. In the case of a programmable NoC used with other programmable circuit devices, the network to be routed through NoC 308 is unknown until a user circuit design is created for implementation within device 100. NoC 308 can be programmed by loading configuration data into internal configuration registers that define how elements within NoC 308 (such as switches and interfaces) are configured and how to operate to pass data between switches and between NoC interfaces.
[0072] The NoC 308 is manufactured as part of the device 100 and, although not physically modifiable, can be programmed to establish connectivity between different master circuits and different slave circuits of a user circuit design. In this regard, the NoC 308 is capable of accommodating different circuit designs, each of which has a different combination of master circuits and slave circuits implemented at different locations in the device 100 that can be coupled via the NoC 308. The NoC 308 can be programmed to route data, such as application data and / or configuration data, between the master circuits and slave circuits of the user circuit design. For example, the NoC 308 can be programmed to couple a user-specified circuit device implemented within the PL 310 to the PS 312, to different DPEs in the DPE 304 via the SoC interface block 104, to different hardwired circuit blocks, and / or to different circuits and / or systems external to the device 100.
[0073] PL 310 is a circuit device that can be programmed to perform specified functions. As an example, PL 310 can be implemented as a field programmable gate array (FPGA) circuit. PL 310 may include an array of programmable circuit blocks. Examples of programmable circuit blocks within PL 310 include, but are not limited to: input / output blocks (IOBs), configurable logic blocks (CLBs), dedicated random access memory blocks (BRAMs), digital signal processing blocks (DSPs), clock managers, and / or delay locked loops (DLLs).
[0074] Each programmable circuit block within PL 310 typically includes both programmable interconnect circuitry and programmable logic circuitry. The programmable interconnect circuitry typically includes a number of interconnect wires of varying lengths interconnected by programmable interconnect points (PIPs). Typically, the interconnect wires are configured to provide connectivity on a bit-by-bit basis (e.g., each wire transmits a single bit of information) (e.g., on a per-wire basis). The programmable logic circuitry implements user-designed logic using programmable elements, for example, the programmable elements may include: lookup tables, registers, arithmetic logic, etc. The programmable interconnect circuitry and programmable logic circuitry may be programmed by loading configuration data into internal configuration memory cells, which define how the programmable elements are configured and operate.
[0075] exist Figure 3 In the example of , PL 310 is shown in two separate parts. In another example, PL 310 may be implemented as a unified region of a programmable circuit device. In yet another example, PL 310 may be implemented as more than two different regions of a programmable circuit device. The specific organization of PL 310 is not intended to be limiting.
[0076] exist Figure 3 In the example of, PS 312 is implemented as a hard-wired circuit device manufactured as a part of equipment 100. PS 312 can be implemented as or include any one of a variety of different processor types. For example, PS 312 can be implemented as a separate processor, such as a single core capable of executing program code. In another example, PS 312 can be implemented as a multi-core processor. In yet another example, PS 312 can include one or more cores, modules, coprocessors, interfaces and / or other resources. Any one of the various different types of architectures can be used to implement PS 312. The example architecture that can be used to implement PS 312 can include, but is not limited to: ARM processor architecture, x86 processor architecture, GPU architecture, mobile processor architecture, DSP architecture or other suitable architectures capable of executing computer-readable instructions or program code.
[0077] Circuit blocks 314-322 can be implemented as any of a variety of hardwired circuit blocks. Hardwired circuit blocks 314-322 can be customized to perform dedicated functions. Examples of circuit blocks 314-322 include, but are not limited to, input / output blocks (IOBs), transceivers, or other dedicated circuit blocks. As noted, circuit blocks 314-322 can be considered examples of ASICs.
[0078] Figure 3 The example of FIG. 1 illustrates an architecture that may be implemented in a device that includes a single die. Although DPE array 102 is shown as occupying the entire width of device 100, in other embodiments, DPE array 102 may occupy less than the entire width of device 100 and / or be located in different areas of device 100. In addition, the number of DPEs 304 included may vary. Thus, the specific number of columns and / or rows of DPEs 304 may vary. Figure 3 Different than shown.
[0079] In one or more embodiments, a device such as device 100 may include two or more DPE arrays 102 located in different areas of device 100. For example, additional DPE arrays may be located below circuit blocks 320 and 322.
[0080] As noted, Figure 2- Figure 3 An example architecture for a device including a single die is shown. In one or more other embodiments, device 100 may be implemented as a multi-die device including one or more DPE arrays 102.
[0081] Figure 4A and Figure 4B (collectively referred to as "FIG. 4") shows an example of a multi-die implementation of device 100. A multi-die device is a device or IC that includes two or more dies within a single package.
[0082] Figure 4A A topographic map of the device 100 is shown. Figure 4A In an example of the present invention, device 100 is implemented as a "stacked die" type device formed by stacking multiple dies. Device 100 includes interposer 402, die 404, die 406, and substrate 408. Each of dies 404 and 406 is attached to a surface (e.g., a top surface) of interposer 402. In one aspect, dies 404 and 406 are attached to interposer 402 using flip chip technology. Interposer 402 is attached to the top surface of substrate 408.
[0083] exist Figure 4AIn the example of FIG. 4 , interposer 402 is a die having a flat surface, and dies 404 and 406 are stacked horizontally on the flat surface. As shown, dies 404 and 406 are located side by side on the flat surface of interposer 402. Figure 4A The number of dies shown on interposer 402 in FIG. 4 is for purposes of illustration and not limitation. In other embodiments, more than two dies may be mounted on interposer 402 .
[0084] Interposer 402 provides a common mounting surface and electrical coupling for each of dies 404 and 406. The manufacture of interposer 402 may include one or more process steps that allow the deposition of one or more conductive layers, which are patterned to form wires. These conductive layers may be formed of aluminum, gold, copper, nickel, various silicides, and / or other suitable materials. Interposer 402 may be manufactured using one or more additional process steps that allow the deposition of one or more dielectric layers or insulating layers (such as silicon dioxide). Interposer 402 may also include through-holes and through-holes (TVs). TVs may be through-silicon vias (TSVs), through-glass vias (TGVs), or other through-hole structures based on the specific materials used to implement interposer 402 and its substrate. If interposer 402 is implemented as a passive die, interposer 402 may only have various types of solder bumps, through-holes, wires, TVs, and under-bump metallization (UBM). If implemented as an active die, interposer 402 may include additional process layers that form one or more active devices with reference to electrical devices (such as transistors including PN junctions, diodes, etc.).
[0085] Each of die 404 and 406 may be implemented as a passive die or an active die including one or more active devices. For example, when implemented as an active die, one or more DPE arrays may be implemented in one or both of die 404 and / or die 406. In one or more embodiments, die 404 may include one or more DPE arrays, while die 406 implements any of the different subsystems described herein. The examples provided herein are for illustrative purposes and are not intended to be limiting. For example, device 100 may include more than two dies, where the dies have different types and / or functions.
[0086] Figure 4B yes Figure 4A 1 is a cross-sectional side view of the device 100. Figure 4B The image taken along the cutting line 4B-4B is shown. Figure 4A4. A view of device 100 of FIG. 4. Each of dies 404 and 406 is electrically and mechanically coupled to the first planar surface of interposer 402 via solder bumps 410. In one example, solder bumps 410 are implemented as micro bumps. Additionally, any of a variety of other techniques may be used to attach dies 404 and 406 to interposer 402. For example, bonding wires or edge wires may be used to mechanically and electrically attach dies 404 and 406 to interposer 402. In another example, an adhesive material may be used to mechanically attach dies 404 and 406 to interposer 402. Figure 4B As shown, dies 404 and 406 are attached to interposer 402 using solder bumps 410 for purposes of illustration and not limitation.
[0087] The interposer 402 includes one or more conductive layers 412 shown in dashed or dotted lines in the interposer 402. The conductive layer 412 is implemented using any of the various metal layers described above. The conductive layer 412 is processed to form a patterned metal layer that implements the wires 414 of the interposer 402. The wires implemented within the interposer 402 that couple at least two different dies (e.g., dies 404 and 406) are referred to as inter-die wires. Figure 4B Wires 414 are shown, which are considered inter-die wires for illustration purposes. Wires 414 transmit inter-die signals between die 404 and die 406. For example, each of wires 414 couples a solder bump 410 under die 404 with a solder bump 410 under die 406, thereby allowing the exchange of inter-die signals between die 404 and die 406. Wires 414 can be data lines or power lines. The power lines can be wires that carry a voltage potential or wires that have a ground or reference voltage potential.
[0088] Different ones of the conductive layers 412 may be coupled together using vias 416. Typically, via structures are used to implement vertical conductive paths (e.g., conductive paths that are perpendicular to the processing layers of the device). In this regard, the vertical portion of the wire 414 that contacts the solder bump 410 is implemented as a via 416. Using multiple conductive layers to implement interconnections within the interposer 402 allows a greater number of signals to be routed within the interposer 402 and enables more complex signal routing.
[0089] Solder bumps 418 may be used to mechanically and electrically couple the second planar surface of interposer 402 to substrate 408. In a particular embodiment, solder bumps 418 may be implemented as controlled collapse chip connection (C4) balls. Substrate 408 includes conductive paths (not shown) that couple different solder bumps 418 to one or more nodes below substrate 408. Thus, one or more of solder bumps 418 couple circuitry within interposer 402 to nodes external to device 100 through circuitry or wires within substrate 408.
[0090] TVs 420 are vias that form vertical lateral (e.g., extending through a majority of, if not all of, the interposer 402) electrical connections. TVs 420 (e.g., wires and vias) may be formed of any of a variety of different conductive materials, including, but not limited to, copper, aluminum, gold, nickel, various silicides, and / or other suitable materials. As shown, each of TVs 420 extends from the bottom surface of the interposer 402 to the conductive layer 412 of the interposer 402. TVs 420 may also be coupled to solder bumps 410 through one or more conductive layers 412 in conjunction with one or more vias 416.
[0091] Figure 5A , Figure 5B , Figure 5C , Figure 5D , Figure 5E , Fig. 5F and Figure 5G (collectively referred to as "FIG. 5") shows an example multi-die implementation of device 100. The example of FIG. 5 may be implemented as described in conjunction with the example of FIG.
[0092] refer to Figure 5A , die 404 includes one or more DPE arrays 102 , while die 406 implements PS 312 .
[0093] refer to Figure 5B , die 404 includes one or more DPE arrays 102, while die 406 implements ASIC 504. ASIC 504 may be implemented as any of a variety of different custom circuits suitable for performing specific or dedicated operations.
[0094] refer to Figure 5C , die 404 includes one or more DPE arrays 102 , while die 406 implements PL 310 .
[0095] refer to Figure 5D, die 404 includes one or more DPE arrays 102, and die 406 implements analog and / or mixed (analog / mixed) signal circuitry 508. Analog / mixed signal circuitry 508 may include one or more wireless receivers, wireless transmitters, amplifiers, analog-to-digital converters, digital-to-analog converters, or other analog and / or digital circuitry.
[0096] Figure 5E , Fig. 5F and Figure 5G An example of a device 100 having three dies 404, 406, and 510 is shown. Figure 5E , device 100 includes die 404, 406, and 510. Die 404 includes one or more DPE arrays 102. Die 406 includes PL 310. Die 510 includes ASIC 504.
[0097] refer to Fig. 5F , die 404 includes one or more DPE arrays 102. Die 406 includes PL 310. Die 510 includes analog / mixed signal circuitry 508.
[0098] refer to Figure 5G , die 404 includes one or more DPE arrays 102. Die 406 includes ASIC 504. Die 510 includes analog / mixed signal circuitry 508. In one or more embodiments, a PS (eg, PS 312) is an example of an ASIC.
[0099] 5, each of the dies 406 and / or 510 is depicted as including a particular type of subsystem. In other embodiments, the dies 404, 406, and / or 510 may include one or more subsystems in combination with one or more DPE arrays 102. Additionally, the dies 404, 406, and / or 510 may include two or more different types of subsystems. Thus, any one or more of the dies 404, 406, and / or 510 may include one or more DPE arrays 102 in combination with one or more subsystems in any combination.
[0100] In one or more embodiments, the interposer 402 and the die 404, 406, and / or 510 may be implemented using the same IC manufacturing technology (e.g., feature size). In one or more other embodiments, the interposer 402 may be implemented using a particular IC manufacturing technology, while the die 404, 406, and / or 510 may be implemented using a different IC manufacturing technology. In yet other embodiments, the die 404, 406, and / or 510 may be implemented using a different IC manufacturing technology that is the same as or different from the IC manufacturing technology used to implement the interposer 402. By using different IC manufacturing technologies for different die and / or interposers, a lower cost and / or more reliable IC manufacturing technology may be used for certain die, while other IC manufacturing technologies capable of producing smaller feature sizes may be used for other die. For example, the interposer 402 may be implemented using a more mature manufacturing technology, while other technologies capable of forming smaller feature sizes may be used to implement active die and / or die including the DPE array 102.
[0101] 5 shows a multi-die implementation of device 100 including two or more dies mounted on an interposer. The number of dies shown is for illustration purposes and not limitation. In other embodiments, device 100 may include more than three dies mounted on interposer 402.
[0102] In one or more other embodiments, a multi-die version of the device 100 may be implemented using an architecture other than the stacked die architecture of FIG. 4. For example, the device 100 may be implemented as a multi-chip module (MCM). An MCM implementation of the device 100 may be implemented using one or more pre-packaged ICs mounted on a circuit board, wherein the circuit board has a form factor and / or package that is intended to mimic an existing chip package. In another example, an MCM implementation of the device 100 may be implemented by integrating two or more dies on a high-density interconnect substrate. In yet another example, an MCM implementation of the device 100 may be implemented as a "chip stack" package.
[0103] The use of the DPE arrays described herein in conjunction with one or more other subsystems implemented in a single die device or a multi-die device can increase the processing power of the device while maintaining low area usage and power consumption. For example, one or more DPE arrays can be used to hardware accelerate specific operations and / or perform functions offloaded from one or more subsystems of the device described herein. For example, when used with a PS, the DPE array can be used as a hardware accelerator. The PS can offload operations that would be performed by a DPE array or a portion of a DPE array. In other examples, the DPE array can be used to perform computationally resource intensive operations, such as generating digital predistortion to be provided to an analog / mixed signal circuit device.
[0104] It should be understood that this article combines Figure 1 Any of the various combinations of DPE arrays and / or other subsystems described in , 2, 3, 4 and / or 5 may be implemented in a single die type device or a multi-die type device.
[0105] In various examples described herein, the SoC interface block is implemented within the DPE array. In one or more other embodiments, the SoC interface block can be implemented outside the DPE array. For example, the SoC interface block can be implemented as a circuit block (e.g., an independent circuit block) that is separate from the circuit blocks that implement the multiple DPEs.
[0106] Figure 6 An example architecture for a DPE 304 of a DPE array 102 is shown. Figure 6 In the example of , DPE 304 includes core 602 , memory module 604 , and DPE interconnect 606 .
[0107] Core 602 provides the data processing capabilities of DPE 304. Core 602 may be implemented as any of a variety of different processing circuits. Figure 6 In the example of the embodiment of the present invention, the core 602 includes an optional program memory 608. In one or more embodiments, the core 602 is implemented as a processor capable of executing program code (e.g., computer readable instructions). In that case, the program memory 608 is included and can store instructions executed by the core 602. For example, the core 602 can be implemented as a CPU, a DSP, a vector processor, or other types of processors capable of executing instructions. The core can be implemented using any of the various CPU and / or processor architectures described herein. In another example, the core 602 is implemented as a very long instruction word (VLIW) vector processor or a DSP.
[0108] In a particular embodiment, program memory 608 is implemented as a dedicated program memory dedicated to core 602. Program memory 608 can be used only by cores of the same DPE 304. Therefore, program memory 608 can be accessed only by core 602 and is not shared with any other DPE or components of another DPE. Program memory 608 may include a single port for read and write operations. Program memory 608 may support program compression and may be addressed using a memory-mapped network portion of DPE interconnect 606 described in more detail below. For example, via the memory-mapped network of DPE interconnect 606, program memory 608 may be loaded with program code that can be executed by core 602.
[0109] In one or more embodiments, program memory 608 can support one or more error detection and / or error correction mechanisms. For example, program memory 608 can be implemented to support parity by adding parity bits. In another example, program memory 608 can be an error correction code (ECC) memory that can detect and correct various types of data damage. In another example, program memory 608 can support ECC and parity. Providing different types of error detection and / or error correction described herein is for illustrative purposes, and is not intended to limit the described embodiments. Except those listed, other error detection and / or error correction techniques can be used together with program memory 608.
[0110] In one or more embodiments, core 602 may have a customized architecture to support a dedicated instruction set. For example, core 602 may be customized for wireless applications and configured to execute wireless-specific instructions. In another example, core 602 may be customized for machine learning and configured to execute machine learning-specific instructions.
[0111] In one or more other embodiments, core 602 is implemented as a hardwired circuit device dedicated to performing one specific operation or multiple specific operations (such as an enhanced intellectual property (IP) core). In that case, core 602 may not execute program code. In embodiments where core 602 does not execute program code, program memory 608 may be omitted. As an illustrative and non-limiting example, core 602 may be implemented as an enhanced forward error correction (FEC) engine or other circuit block.
[0112] Core 602 may include configuration registers 624. Configuration registers 624 may be loaded with configuration data to control the operation of core 602. In one or more embodiments, core 602 may be activated and / or deactivated based on the configuration data loaded into configuration registers 624. Figure 6 In the example of , configuration registers 624 are addressable (eg, can be read from and / or written to) via a memory-mapped network of DPE interconnect 606 described in greater detail below.
[0113] In one or more embodiments, the memory module 604 can store data used by the core 602 or generated by the core 602. For example, the memory module 604 can store application data. The memory module 604 can include read / write memory (such as random access memory). Therefore, the memory module 604 can store data that can be read and consumed by the core 602. The memory module 604 can also store data written by the core 602 (e.g., results).
[0114] In one or more other embodiments, the memory module 604 can store data (e.g., application data) that can be used and / or generated by one or more other cores of other DPEs within the DPE array. One or more other cores of the DPE can also read from and / or write to the memory module 604. In a particular embodiment, the other cores that can read from or write to the memory module 604 can be cores of one or more adjacent DPEs. Another DPE that shares an edge or border (e.g., adjacent) with the DPE 304 is referred to as a "neighboring" DPE relative to the DPE 304. By allowing the core 602 and one or more other cores from the adjacent DPE to read and / or write to the memory module 604, the memory module 604 implements a shared memory that supports communication between different DPEs and / or cores that can access the memory module 604.
[0115] refer to Figure 3 , for example, DPEs 304-14, 304-16, 304-5, and 304-25 are considered to be adjacent DPEs to DPE 304-15. In one example, a core within each of DPEs 304-16, 304-5, and 304-25 is able to read and write to a memory module within DPE 304-15. In certain embodiments, only those adjacent DPEs that are adjacent to a memory module may access the memory module of DPE 304-15. For example, although DPE 304-14 is adjacent to DPE 304-15, DPE 304-14 may not be adjacent to a memory module of DPE 304-15 because a core of DPE 304-15 may be located between a core of DPE 304-14 and a memory module of DPE 304-15. Likewise, in certain embodiments, the core of DPE 304-14 may not have access to the memory modules of DPE 304-15.
[0116] In a particular embodiment, whether a DPE core can access a memory module of another DPE depends on the number of memory interfaces included in the memory module and whether such a core is connected to an available memory interface in the memory interface of the memory module. In the above example, the memory module of DPE 304-15 includes four memory interfaces, wherein the core of each DPE in DPE 304-16, 304-5 and 304-15 is connected to such a memory interface. The core 602 of DPE 304-15 itself is connected to the fourth memory interface. Each memory interface may include one or more read and / or write channels. In a particular embodiment, each memory interface includes multiple read channels and multiple write channels, so that the specific core attached thereto can read and / or write multiple memory banks in the memory module 604 at the same time.
[0117] In other examples, more than four memory interfaces may be used. Such other memory interfaces may be used to allow DPEs on a diagonal line of DPE 304-15 to access memory modules of DPE 304-15. For example, if cores in DPEs (such as DPEs 304-14, 304-24, 304-26, 304-4, and / or 304-6) are also coupled to available memory interfaces of memory modules in DPE 304-15, such other DPEs may also be able to access memory modules of DPE 304-15.
[0118] The memory module 604 may include configuration registers 636. The configuration registers 636 may be loaded with configuration data to control the operation of the memory module 604. Figure 6 In the example of , configuration registers 636 (and 624) are addressable (eg, can be read and / or written) via a memory-mapped network of DPE interconnect 606 described in greater detail below.
[0119] exist Figure 6 In the example of FIG. 4 , DPE interconnect 606 is dedicated to DPE 304. DPE interconnect 606 facilitates various operations including communication between DPE 304 and one or more other DPEs of DPE array 102 and / or with other subsystems of device 100. DPE interconnect 606 also enables configuration, control, and debugging of DPE 304.
[0120] In a particular embodiment, the DPE interconnect 606 is implemented as an on-chip interconnect. An example of an on-chip interconnect is an Advanced Microcontroller Bus Architecture (AMBA) Extensible Interface (AXI) bus (e.g., or a switch). The AMBA AXI bus is an embedded microcontroller bus interface for establishing on-chip connections between circuit blocks and / or systems. An AXI bus is provided herein as an example of an interconnect circuit device, which can be used with the inventive arrangements described in the present disclosure, and is not intended to be limiting as such. Other examples of interconnect circuit devices may include other types of buses, crossbar switches, and / or other types of switches.
[0121] In one or more embodiments, the DPE interconnect 606 includes two different networks. The first network can exchange data with other DPEs of the DPE array 102 and / or other subsystems of the device 100. For example, the first network can exchange application data. The second network can exchange data, such as configuration data, control data, and / or debug data for the DPE.
[0122] exist Figure 6 In the example of FIG. 6 , the first network of DPE interconnect 606 is formed by a stream switch 626 and one or more stream interfaces. As shown, the stream switch 626 includes multiple stream interfaces (in Figure 6 In one or more embodiments, each stream interface may include one or more master interfaces (e.g., master interfaces or outputs) and / or one or more slave interfaces (e.g., slave interfaces or inputs). Each master interface may be an independent output with a specific bit width. For example, each master interface included in the stream interface may be an independent AXI master interface. Each slave interface may be an independent input with a specific bit width. For example, each slave interface included in the stream interface may be an independent AXI slave interface.
[0123] Stream interfaces 610-616 are used to communicate with other DPEs in DPE array 102 and / or SoC interface block 104. For example, each of stream interfaces 610, 612, 614, and 616 is capable of communicating in a different basic direction. Figure 6 In the example of , stream interface 610 communicates with the DPE on the left (west). Stream interface 612 communicates with the DPE above (north). Stream interface 614 communicates with the DPE on the right (east). Stream interface 616 communicates with the DPE or SoC interface block 104 below (south).
[0124] The stream interface 628 is used to communicate with the core 602. For example, the core 602 includes a stream interface 638 connected to the stream interface 628, thereby allowing the core 602 to communicate directly with other DPEs 304 via the DPE interconnect 606. For example, the core 602 may include instructions or hard-wired circuit devices that enable the core 602 to directly send and / or receive data via the stream interface 638. The stream interface 638 can be blocking or non-blocking. In one or more embodiments, the core 602 may stall in the event that the core 602 attempts to read from an empty stream or write to a full stream. In other embodiments, attempting to read from an empty stream or write to a full stream does not cause the core 602 to stall. Instead, the core 602 can continue to execute or operate.
[0125] The stream interface 630 is used to communicate with the memory module 604. For example, the memory module 604 includes a stream interface 640 connected to the stream interface 630, thereby allowing other DPEs to communicate with the memory module 604 via the DPE interconnect 606. The stream switch 626 can allow non-adjacent DPEs and / or DPEs that are not coupled to the memory interface of the memory module 604 to communicate with the core 602 and / or the memory module 604 via a DPE interconnect network formed by the DPE interconnects of the corresponding DPEs 304 of the DPE array 102.
[0126] Reference again Figure 3 And using DPE 304-15 as a reference point, stream interface 610 is coupled to and can communicate with another stream interface located in the DPE interconnect of DPE 304-14. Stream interface 612 is coupled to and can communicate with another stream interface located in the DPE interconnect of DPE 304-25. Stream interface 614 is coupled to and can communicate with another stream interface located in the DPE interconnect of DPE 304-16. Stream interface 616 is coupled to and can communicate with another stream interface located in the DPE interconnect of DPE 304-5. In this way, core 602 and / or memory module 604 can also communicate with any of the DPEs in DPE array 102 via the DPE interconnects in the DPEs.
[0127] Stream switch 626 is also used to interface with subsystems such as PL 310 and / or NoC 308. In general, stream switch 626 can be programmed to function as a circuit-switched stream interconnect or a packet-switched stream interconnect. Circuit-switched stream interconnects enable point-to-point dedicated streams suitable for high-bandwidth communications between DPEs. Packet-switched stream interconnects allow for shared streams to time-division multiplex multiple logical streams onto one physical stream for medium-bandwidth communications.
[0128] Stream switch 626 may include configuration registers (in Figure 6Configuration data may be written to configuration registers 634 via a memory-mapped network of DPE interconnect 606. The configuration data loaded into configuration registers 634 indicates which other DPEs and / or subsystems (e.g., NoC 308, PL 310, and / or PS 312) DPE 304 will communicate with, and whether such communications are to be established as circuit-switched point-to-point connections or packet-switched connections.
[0129] It should be understood that Figure 6 The number of stream interfaces shown is for the purpose of illustration and not limitation. In other embodiments, stream switch 626 may include fewer stream interfaces. In a specific embodiment, stream switch 626 may include more stream interfaces that facilitate other components and / or subsystems in the device. For example, additional stream interfaces may be coupled to other non-adjacent DPEs (such as, DPE 304-24, 304-26, 304-4 and / or 304-6). In one or more other embodiments, stream interfaces may be included to couple DPEs (such as DPE 304-15) to other DPEs outside one or more DPEs. For example, one or more stream interfaces may be included to allow DPE 304-15 to be directly coupled to a stream interface in DPE 304-13, DPE 304-16 or other non-adjacent DPEs.
[0130] The second network of DPE interconnects 606 is formed by a memory-mapped switch 632. The memory-mapped switch 632 includes a plurality of memory-mapped interfaces (in Figure 6 In one or more embodiments, each memory-mapped switch may include one or more master interfaces (e.g., master interface interfaces or outputs) and / or one or more slave interfaces (e.g., slave interface interfaces or inputs). Each master interface may be an independent output having a specific bit width. For example, each master interface included in the memory-mapped interface may be an independent AXI master interface. Each slave interface may be an independent input having a specific bit width. For example, each slave interface included in the memory-mapped interface may be an independent AXI slave interface.
[0131] exist Figure 6In the example of , the memory mapped switch 632 includes memory mapped interfaces 620, 622, 642, 644, and 646. It should be understood that the memory mapped switch 632 may include additional or fewer memory mapped interfaces. For example, for each component of the DPE that can be read and / or written using the memory mapped switch 632, the memory mapped switch 632 may include a memory mapped interface coupled to such component. In addition, the components themselves may include memory mapped interfaces that are coupled to corresponding memory mapped interfaces in the memory mapped switch 632 to facilitate reading and / or writing of memory addresses.
[0132] Memory mapped interfaces 620 and 622 may be used to exchange configuration data, control data, and debug data for DPE 304. Figure 6 304, the memory mapped interface 620 can receive configuration data that is used to configure the DPE 304. The memory mapped interface 620 can receive the configuration data from a DPE located below the DPE 304 or from the SoC interface block 104. The memory mapped interface 622 can forward the configuration data received by the memory mapped interface 620 to one or more other DPEs above the DPE 304, the core 602 (e.g., program memory 608 and / or configuration registers 624), the memory module 604 (e.g., memory and / or configuration registers 636 within the memory module 604), and / or the configuration registers 634 within the stream switch 626.
[0133] In certain embodiments, the memory mapped interface 620 communicates with a DPE or tile of the SoC interface block 104 described below. The memory mapped interface 622 communicates with the upper DPE. Figure 3 And using DPE 304-15 as a reference point, memory mapped interface 620 is coupled to and can communicate with another memory mapped interface located in the DPE interconnect of DPE 304-5. Memory mapped interface 622 is coupled to and can communicate with another memory mapped interface located in the DPE interconnect of DPE 304-25. In one or more embodiments, memory mapped switch 632 transfers control and / or debug data from south to north. In other embodiments, memory mapped switch 632 can also transfer data from north to south.
[0134] Memory mapped interface 646 may be coupled to a memory mapped interface (not shown) in memory module 604 to facilitate reading and / or writing of configuration registers 636 and / or memory within memory module 604. Memory mapped interface 644 may be coupled to a memory mapped interface (not shown) in core 602 to facilitate reading and / or writing of program memory 608 and / or configuration registers 624. Memory mapped interface 642 may be coupled to configuration registers 634 to read and / or write to configuration registers 634.
[0135] exist Figure 6 In the example of , memory mapped switch 632 is capable of communicating with circuit devices above (e.g., to the north) and below (e.g., to the south). In one or more other embodiments, memory mapped switch 632 includes additional memory mapped interfaces that are coupled to memory mapped interfaces of memory mapped switches of left and / or right DPEs. Using DPE 304-15 as a reference point, such additional memory mapped switches can be connected to memory mapped switches located in DPE 304-14 and / or 304-16, thereby facilitating communication of configuration data, control data, and debug data between DPEs in the horizontal direction as well as the vertical direction.
[0136] In other embodiments, memory mapped switch 632 may include additional memory mapped interfaces that are connected to memory mapped switches in DPEs that are diagonal to DPE 304. For example, using DPE 304-15 as a reference point, such additional memory mapped interfaces may be coupled to memory mapped switches located in DPEs 304-24, 304-26, 304-4, and / or 304-6, thereby facilitating communication of configuration information, control information, and debug information between the DPE diagonals.
[0137] Depending on the location of the DPE 304, the DPE interconnect 606 is coupled to the DPE interconnect of each adjacent DPE and / or SoC interface block 104. The DPE interconnects of the DPE 304 together form a DPE interconnect network (the DPE interconnect network may include a stream network and / or a memory-mapped network). The configuration registers of the stream switch of each DPE may be programmed by loading configuration data of the memory-mapped switch. Through configuration, the stream switch and / or stream interface is programmed to establish a packet-switched or circuit-switched connection with one or more other DPEs 304 and / or other endpoints in the SoC interface block 104.
[0138] In one or more embodiments, DPE array 102 is mapped into the address space of a processor system, such as PS 312. Thus, any configuration registers and / or memory within DPE 304 may be accessed via a memory mapped interface. For example, memory in memory module 604, program memory 608, configuration registers 624 in core 602, configuration registers 636 in memory module 604, and / or configuration registers 634 may be read and / or written via memory mapped switch 632.
[0139] exist Figure 6 In the example of , the memory mapped interface can receive configuration data for DPE 304. The configuration data can include: program code loaded into program memory 608 (if included), configuration data for loading into configuration registers 624, 634 and / or 636, and / or data loaded into memory (e.g., memory bank) of memory module 604. Figure 6 In the example of FIG. 6 , configuration registers 624 , 634 , and 636 are shown as being located within a particular circuit structure, the configuration registers being intended to control, for example, core 602 , stream switch 626 , and memory module 604 . Figure 6 The examples are for illustration purposes only and illustrate that elements within core 602, memory module 604, and / or stream switch 626 may be programmed by loading configuration data into corresponding configuration registers. In other embodiments, configuration registers may be consolidated within specific areas of DPE 304, although controlling the operation of components distributed throughout DPE 304.
[0140] Thus, the stream switch 626 may be programmed by loading configuration data into the configuration registers 634. The configuration data programs the stream switch 626 and / or the stream interfaces 610-616 and / or the stream interfaces 628-630 to function as a circuit-switched stream interface between two different DPEs and / or other subsystems, or as a packet-switched stream interface coupled to a selected DPE and / or other subsystem. Thus, the connections established by the stream switch 626 with other stream interfaces are programmed by loading appropriate configuration data into the configuration registers 634 to establish actual connections or application data paths within the DPE 304, with other DPEs, and / or with other subsystems of the device 100.
[0141] Figure 7 Example connectivity between multiple DPEs 304 is shown. Figure 7 In the example, Figure 6 The architecture shown in is used to implement each of the DPEs 304-14, 304-15, 304-24, and 304-25. Figure 7An embodiment is shown where stream interfaces are interconnected between adjacent DPEs (on each side and above and below) and where memory mapped interfaces are connected to the above and below DPEs. For illustration purposes, stream switches and memory mapped switches are not shown.
[0142] As noted, in other embodiments, additional memory mapped interfaces may be included to couple the DPE in both the vertical and horizontal directions as shown. Additionally, the memory mapped interface may support bidirectional communication in the vertical and / or horizontal directions.
[0143] The memory mapped interfaces 620 and 622 enable a shared transaction exchange network in which transactions are propagated from memory mapped switch to memory mapped switch. For example, each of the memory mapped switches can dynamically route transactions based on addresses. Transactions can be stalled at any given memory mapped switch. The memory mapped interfaces 620 and 622 allow other subsystems of the device 100 to access resources (e.g., components) of the DPE 304.
[0144] In certain embodiments, a subsystem of device 100 is able to read the internal state of any register and / or memory element of the DPE via memory mapped interfaces 620 and / or 622. Through memory mapped interfaces 620 and / or 622, a subsystem of device 100 is able to read and / or write program memory 608 and read and / or write any configuration register within DPE 304.
[0145] Stream interfaces 610-616 (e.g., stream switch 626) can provide a determined throughput with a guaranteed and fixed delay from source to destination. In one or more embodiments, stream interfaces 610 and 614 can receive four 32-bit streams and output four 32-bit streams. In one or more embodiments, stream interface 614 can receive four 32-bit streams and output six 32-bit streams. In a specific embodiment, stream interface 616 can receive four 32-bit streams and output four 32-bit streams. The number of streams of each stream interface and the size of the stream are provided for illustration and not for limitation purposes.
[0146] Figure 8 Shows Figure 6 Additional aspects of the example architecture. Figure 8 In the example, details related to the interconnection with the DPE are not shown. Figure 8 The connectivity of core 602 to other DPEs through shared memory is shown. Figure 8 Additional aspects of the memory module 604 are also shown. For illustration purposes, Figure 8 Refer to DPE304-15.
[0147] As shown, the memory module 604 includes a plurality of memory interfaces 802, 804, 806, and 808. Figure 8 , memory interfaces 802 and 808 are abbreviated as "MI". The memory module 604 also includes a plurality of memory banks 812-1 to 812-N. In a particular embodiment, the memory module 604 includes eight memory banks. In other embodiments, the memory module 604 may include fewer or more memory banks 812. In one or more embodiments, each memory bank 812 is single-ported, thereby allowing up to one access to each memory bank per clock cycle. In the case where the memory module 604 includes eight memory banks 812, this configuration supports eight parallel accesses per clock cycle. In other embodiments, each memory bank 812 is dual-ported or multi-ported, thereby allowing a large number of parallel accesses per clock cycle.
[0148] In one or more embodiments, memory module 604 can support one or more error detection and / or error correction mechanisms. For example, memory bank 812 can be implemented to support parity by additional parity bits. In another example, memory bank 812 can be an ECC memory capable of detecting and correcting various types of data corruption. In another example, memory bank 812 can support both ECC and parity. Providing different types of error detection and / or error correction described herein is for illustrative purposes, and is not intended to limit the described embodiments. In addition to those listed technologies, other error detection and / or error correction techniques can be used together with memory module 604.
[0149] In one or more other embodiments, error detection and / or error correction mechanisms may be implemented based on each memory bank 812. For example, one or more memory banks 812 may include parity checks, while one or more other memory banks in the memory banks 812 may be implemented as ECC memory. In addition, other memory banks in the memory banks 812 may support both ECC and parity checks. In this way, different combinations of error detection and / or error correction may be supported by different memory banks 812 and / or combinations of memory banks 812.
[0150] exist Figure 8 In the example of FIG. 1 , each of the memory banks 812-1 to 812-N has a corresponding arbiter 814-1 to 814-N. Each of the arbiters 814 can generate a stall signal in response to a detected conflict. Each arbiter 814 can include arbitration logic. In addition, each arbiter 814 can include a crossbar switch. Therefore, any master interface can write to any specific one or more memory banks in the memory banks 812. As shown in FIG. Figure 6As noted, the memory module 604 may include a memory mapped interface (not shown) that communicates with the memory mapped interface 646 of the memory mapped switch 632. The memory mapped interface in the memory module 604 may be connected to communication lines in the memory module 604 that couple the DMA engine 816, the memory interfaces 802, 804, 806, and 808, and the arbiter 814 to read and / or write to the memory bank 812.
[0151] The memory module 604 also includes a direct memory access (DMA) engine 816. In one or more embodiments, the DMA engine 816 includes at least two interfaces. For example, one or more interfaces can receive input data streams from the DPE interconnect 606 and write the received data to the memory bank 812. One or more other interfaces can read data from the memory bank 812 and send data via the stream interface of the DPE interconnect 606. For example, the DMA engine 816 can include Figure 6 The stream interface 640.
[0152] The memory module 604 can be used as a shared memory that can be accessed by multiple different DPEs. Figure 8 , memory interface 802 is coupled to core 602 via core interface 828 included in core 602. Memory interface 802 provides core 602 with access to memory bank 812 through arbiter 814. Memory interface 804 is coupled to the core of DPE 304-25. Memory interface 804 provides the core of DPE 304-25 with access to memory bank 812. Memory interface 806 is coupled to the core of DPE 304-16. Memory interface 806 provides the core of DPE 304-16 with access to memory bank 812. Memory interface 808 is coupled to the core of DPE 304-5. Memory interface 808 provides the core of DPE 304-5 with access to memory bank 812. Thus, in Figure 8 In the example of , each DPE having a shared boundary with the memory module 604 of the DPE 304-15 can read and write to the memory bank 812. Figure 8 In the example of , the core of DPE 304 - 14 does not have direct access to the memory module 604 of DPE 304 - 15 .
[0153] The memory mapped switch 632 is capable of writing data to the memory bank 812. For example, the memory mapped switch 632 may be coupled to a memory mapped interface (not shown) located in the memory module 604, which in turn is coupled to the arbiter 814. In this way, specific data stored in the memory module 604 may be controlled, e.g., written, as part of a configuration process, a control process, and / or a debug process.
[0154] Core 602 can access memory modules of other adjacent DPEs via core interfaces 830, 832, and 834. Figure 8 , core interface 834 is coupled to a memory interface of DPE 304-25. Thus, core 602 is able to access the memory modules of DPE 304-25 via core interface 834 and the memory interfaces contained within the memory modules of DPE 304-25. Core interface 832 is coupled to a memory interface of DPE 304-14. Thus, core 602 is able to access the memory modules of DPE 304-14 via core interface 832 and the memory interfaces contained within the memory modules of DPE 304-14. Core interface 830 is coupled to a memory interface within DPE 304-5. Thus, core 602 is able to access the memory modules of DPE 304-5 via core interface 830 and the memory interfaces contained within the memory modules of DPE 304-5. As discussed, core 602 is able to access memory modules 604 within DPE 304-15 via core interface 828 and memory interface 802.
[0155] exist Figure 8 In the example of , the core 602 can read and write any of the memory modules of the DPEs that share a boundary with the core 602 in the DPE 304-15 (e.g., DPEs 304-25, 304-14, and 304-5). In one or more embodiments, the core 602 can treat the memory modules within the DPEs 304-25, 304-15, 304-14, and 304-5 as a single continuous memory. The core 602 can generate addresses for reading and writing assuming this continuous memory model. The core 602 can direct the read and / or write requests to the appropriate core interface 828, 830, 832, and / or 834 based on the generated addresses.
[0156] In one or more other embodiments, memory module 604 includes additional memory interfaces that can be coupled to other DPEs. For example, memory module 604 can include a memory interface that is coupled to the core of DPE 304-24, 304-26, 304-4, and / or 304-5. In one or more other embodiments, memory module 604 can include one or more memory interfaces that are used to connect to the core of a DPE that is not an adjacent DPE. For example, such additional memory interfaces can be connected to the core of a DPE that is separated from DPE 304-15 by one or more other DPEs in the same row, the same column, or a diagonal direction. Thus, as Figure 8 The number of memory interfaces in the memory module 604 and the specific DPEs connected to the memory interfaces are shown for purposes of illustration and not limitation.
[0157] As noted, core 602 can map read and / or write operations in the correct direction based on the addresses of such operations through core interfaces 828, 830, 832, and / or 834. When core 602 generates an address for a memory access, core 602 can decode the address to determine the direction (e.g., the particular DPE to be accessed) and forward the memory operation to the correct core interface in the determined direction.
[0158] Thus, core 602 can communicate with the core of DPE 304-25 via shared memory, which can be a memory module within DPE 304-25 and / or a memory module 604 of DPE 304-15. Core 602 can communicate with the core of DPE 304-14 via shared memory, which is a memory module within DPE 304-14. Core 602 can communicate with the core of DPE 304-5 via shared memory, which is a memory module within DPE 304-5 and / or a memory module 604 of DPE 304-15. Additionally, core 602 can communicate with the core of DPE 304-16 via shared memory, which is a memory module 604 within DPE 304-15.
[0159] As discussed, the DMA engine 816 may include one or more stream-to-memory interfaces (e.g., stream interface 640). Through the DMA engine 816, application data may be received from other sources within the device 100 and stored in the memory module 604. For example, data may be received from other DPEs that share a boundary with the DPE 304-15 and / or do not share a boundary with the DPE 304-15 via the stream switch 626. Data may also be received from other subsystems of the device 100 (e.g., NoC 308, hardwired circuit blocks, PL 310, and / or PS 312) via the stream switch of the DPE via the SoC interface block 104. The DMA engine 816 is capable of receiving such data from the stream switch and writing the data to the appropriate memory bank or banks 812 within the memory module 604.
[0160] The DMA engine 816 may include one or more memory-to-stream interfaces (e.g., stream interface 630). Through the DMA engine 816, data may be read from a memory bank or banks 812 of the memory module 604 and sent to other destinations via a stream interface. For example, the DMA engine 816 may read data from the memory module 604 via a stream switch and send such data to other DPEs that share a boundary with the DPE 304-15 and / or do not share a boundary with the DPE 304-15 by means of a stream switch. The DMA engine 816 may also send such data to other subsystems (e.g., NoC 308, hardwired circuit blocks, PL 310, and / or PS 312) via a stream switch and the SoC interface block 104.
[0161] In one or more embodiments, the DMA engine 816 may be programmed by a memory mapped switch 632 within a DPE 304-15. For example, the DMA engine 816 may be controlled by a configuration register 636. The configuration register 636 may be written using the memory mapped switch 632 of the DPE interconnect 606. In a particular embodiment, the DMA engine 816 may be controlled by a stream switch 626 within a DPE 304-15. For example, the DMA engine 816 may include control registers that may be written by a stream switch 626 connected thereto (e.g., via a stream interface 640). Depending on the configuration data loaded into the configuration registers 624, 634, and / or 636, a stream received via the stream switch 626 within the DPE interconnect 606 may be connected to the DMA engine 816 in the memory module 604 and / or directly to the core 602. Depending on the configuration data loaded into the configuration registers 624 , 634 , and / or 636 , streams may be sent from the DMA engine 816 (eg, the memory module 604 ) and / or the core 602 .
[0162] The memory module 604 may also include a hardware synchronization circuit device 820 (in Figure 8 Typically, the hardware synchronization circuit device 820 is capable of synchronizing different cores (e.g., cores of adjacent DPEs), Figure 8 The hardware synchronization circuit device 820 can synchronize the operation of the core 602, the DMA engine 816, and other external master interfaces (e.g., PS 312) that can communicate via the DPE interconnect 606. As an illustrative and non-limiting example, the hardware synchronization circuit device 820 can synchronize access to two different cores in different DPEs, such as access to a shared buffer in the memory module 604.
[0163] In one or more embodiments, the hardware synchronization circuit device 820 may include multiple different locks. The specific number of locks included in the hardware synchronization circuit device 820 may depend on the number of entities that can access the memory module, but is not intended to be limiting. In a specific embodiment, each different hardware lock may have an arbiter that can handle simultaneous requests. In addition, each hardware lock can handle a new request every clock cycle. The hardware synchronization circuit device 820 may have multiple requesters, such as: core 602, cores from each DPE in DPE304-25, 304-16 and 304-5, DMA engine 816, and / or a master interface that communicates via DPE interconnect 606. For example, a requester obtains a lock on a specific portion of memory in a memory module from a local hardware synchronization circuit device before accessing the specific portion of memory. The requester can release the lock so that another requester can obtain the lock before accessing the same portion of memory.
[0164] In one or more embodiments, the hardware synchronization circuit device 820 can synchronize accesses by multiple cores to the memory module 604, and in particular to the memory bank 812. For example, the hardware synchronization circuit device 820 can synchronize Figure 8 The core 602 shown in FIG. 1 , the core of DPE 304-25, the core of DPE 304-16, and the core of Figure 8 In certain embodiments, hardware synchronization circuitry 820 can synchronize access to memory bank 812 for any core that can directly access memory module 604 via memory interfaces 802, 804, 806, and / or 808. For example, each core (e.g., Figure 8604 and the cores of one or more adjacent DPEs) can access hardware synchronization circuitry 820 to request and acquire a lock before accessing a particular portion of memory in memory module 604, and subsequently release the lock to allow another core to access that portion of memory once the core acquires the lock. In a similar manner, core 602 can access hardware synchronization circuitry 820, hardware synchronization circuitry within DPE 304-14, synchronization circuitry within DPE 304-25, and hardware synchronization circuitry within DPE 304-5 to request and acquire a lock to access portions of memory in the memory modules of each respective DPE and subsequently release the lock. Hardware synchronization circuitry 820 efficiently manages the operation of shared memory between DPEs by regulating and synchronizing access to the memory modules of the DPEs.
[0165] The hardware synchronization circuit device 820 can also be accessed via the memory mapped switch 632 of the DPE interconnect 606. In one or more embodiments, the lock transaction is implemented as an atomic acquire (e.g., test whether to unlock or set the lock) and release (e.g., unset the lock) operation on the resource. The lock of the hardware synchronization circuit device 820 provides a way to efficiently transfer the ownership of the resource between two participants. The resource can be any of a variety of circuit components, such as a buffer in a local memory (e.g., a buffer in the memory module 604).
[0166] While the hardware synchronization circuitry 820 is capable of synchronizing access to memory to support communication through shared memory, the hardware synchronization circuitry 820 is also capable of synchronizing any of a variety of other resources and / or agents including other DPEs and / or other cores. For example, because the hardware synchronization circuitry 820 provides a shared pool of locks, a lock may be used by a DPE (e.g., a core of a DPE) to start and / or stop the operation of another DPE or core. For example, the locks of the hardware synchronization circuitry 820 may be assigned for different purposes based on configuration data, such as depending on the specific application implemented by the DPE array 102, different agents and / or resources may need to be synchronized.
[0167] In certain embodiments, DPE access and DMA access to the lock of the hardware synchronization circuit device 820 are blocking. In the case where the lock cannot be acquired immediately, such access can cause the requesting core or DMA engine to stall. Once the hardware lock is available, the core or DMA engine acquires the lock and automatically logs out.
[0168] In an embodiment, memory mapped accesses may be non-blocking, enabling the memory mapped master interface to poll the status of a lock of the hardware synchronization circuit device 820. For example, a memory mapped switch may send a lock "get" request to the hardware synchronization circuit device 820 as a normal memory read operation. The read address may encode an identifier of the lock and other request data. For example, the read data in response to the read request may signal the success of the get request operation. The "get" sent as a memory read may be sent in a loop until successful. In another example, the hardware synchronization circuit device 820 may issue an event so that the memory mapped master interface receives an interrupt when the status of the requested lock changes.
[0169] Thus, when two adjacent DPEs share a data buffer via a memory module 604, hardware synchronization circuitry 820 within the particular memory module 604 that includes the buffer synchronizes the access. Typically, but not necessarily, memory blocks may be double buffered to increase throughput.
[0170] In the case where the two DPEs are not adjacent, the two DPEs cannot access a common memory module. In that case, application data can be transferred via a data stream (the terms "data stream" and "stream" may be used interchangeably from time to time in this disclosure). In this way, the local DMA engine can convert a transfer from a local memory transfer-based transfer to a stream-based transfer. In that case, the core 602 and the DMA engine 816 can be synchronized using the hardware synchronization circuit device 820.
[0171] Core 602 can also access hardware synchronization circuitry (e.g., locks of hardware synchronization circuitry) of adjacent DPEs to facilitate communication of shared memory. Thus, hardware synchronization circuitry in such other or adjacent DPEs can synchronize access to resources (e.g., memory) between cores of adjacent DPEs.
[0172] PS 312 can communicate with core 602 via memory mapped switch 632. For example, PS 312 can access memory module 604 and hardware synchronization circuit device 820 by initiating memory reads and writes. In another embodiment, hardware synchronization circuit device 820 can also send interrupts to PS 312 when the state of the lock changes to avoid polling of PS 312 by hardware synchronization circuit device 820. PS 312 can also communicate with DPE 304-15 via a stream interface.
[0173] The examples provided herein with respect to entities sending memory mapping requests and / or transmissions are for purposes of illustration and not limitation. In certain embodiments, any entity external to the DPE array 102 can send memory mapping requests and / or transmissions. For example, a circuit block implemented in a PL 310, an ASIC, or other circuitry external to the DPE array 102 as described herein can send memory mapping requests and / or transmissions to the DPE 304 and can access hardware synchronization circuitry of a memory module within such a DPE.
[0174] In addition to communicating with adjacent DPEs through a shared memory module and communicating with adjacent and / or non-adjacent DPEs via DPE interconnect 606, core 602 may also include a cascade interface. Figure 8 In the example of FIG. 6 , core 602 includes cascade interfaces 822 and 824 (in Figure 8 Cascade interfaces 822 and 824 can provide direct communication with other cores. As shown, cascade interface 822 of core 602 directly receives input data streams from the core of DPE 304-14. The data stream received via cascade interface 822 can be provided to data processing circuit devices within core 602. Cascade interface 824 of core 602 can send output data directly to the core of DPE 304-16.
[0175] exist Figure 8 In the example of , each of cascade interface 822 and cascade interface 824 may include a first-in-first-out (FIFO) interface for buffering. In a particular embodiment, cascade interfaces 822 and 824 are capable of transmitting data streams with a width of hundreds of bits. The specific bit widths of cascade interfaces 822 and 824 are not intended to be limiting. Figure 8 In the example of FIG. 8 , the cascade interface 824 is coupled to the accumulator register 836 within the core 602 (in Figure 8 Cascade interface 824 can output the contents of accumulator register 836 every clock cycle. Accumulator register 836 can store data generated and / or operated on by data processing circuit devices within core 602.
[0176] exist Figure 8 In the example of , cascade interfaces 822 and 824 can be programmed based on configuration data loaded into configuration register 624. For example, based on configuration register 624, cascade interface 822 can be activated or deactivated. Similarly, based on configuration register 624, cascade interface 824 can be activated or deactivated. Cascade interface 822 can be activated and / or deactivated independently of cascade interface 824.
[0177] In one or more other embodiments, the cascade interfaces 822 and 824 are controlled by the core 602. For example, the core 602 may include instructions to read / write to the cascade interfaces 822 and / or 824. In another example, the core 602 may include hardwired circuitry that is capable of reading and / or writing to the cascade interfaces 822 and / or 824. In a particular embodiment, the cascade interfaces 822 and 824 may be controlled by an entity external to the core 602.
[0178] In the embodiments described in the present disclosure, DPE 304 does not include cache memory. By omitting cache memory, DPE array 102 is able to achieve predictable (e.g., deterministic) performance. In addition, since there is no need to maintain consistency between caches located in different DPEs, considerable processing overhead is avoided.
[0179] According to one or more embodiments, the core 602 of the DPE 304 does not have input interrupts. Therefore, the core 602 of the DPE 304 can operate uninterruptedly. Omitting input interrupts to the core 602 of the DPE 304 also allows the DPE array 102 to achieve predictable (eg, deterministic) performance.
[0180] In cases where one or more DPEs 304 communicate with an external agent implemented in PS 312, PL 310, a hardwired circuit block, and / or another subsystem (e.g., an ASIC) of device 100 through a shared buffer in an external read-write (e.g., DDR) memory, a coherency mechanism may be implemented using a coherent interconnect in PS 312. In these cases, application data transfers between DPE array 102 and the external agent may traverse both NoC 308 and / or PL 310.
[0181] In one or more embodiments, the DPE array 102 may be functionally isolated into multiple groups of one or more DPEs. For example, specific memory interfaces may be enabled and / or disabled via configuration data to create one or more groups of DPEs, where each group includes one or more DPEs (e.g., a subset) of the DPEs of the DPE array 102. In another example, a stream interface may be independently configured for each group to communicate with other cores of the DPEs in the group and / or with specified input sources and / or output destinations.
[0182] In one or more embodiments, the core 602 can support debug functionality via a memory mapped interface. As discussed, the program memory 608, memory module 604, core 602, DMA engine 816, stream switch 626, and other components of the DPE are memory mapped. The memory mapped registers can be read and / or written by any source capable of generating a memory mapping request (such as, PS 312, PL 310, and / or a platform management controller within the IC). The request can reach the intended DPE or target DPE within the DPE array 102 through the SoC interface module 104.
[0183] Functions such as pausing the core, resuming the core, single-stepping the core, and / or resetting the core can be performed via the memory-mapped switch within the DPE. In addition, such operations can be initiated for multiple different DPEs. For example, other example debug operations that can be performed include: reading status and / or setting the status of the hardware synchronization circuit device 820 and / or the DMA engine 816 via the memory-mapped interface described herein.
[0184] In one or more embodiments, the stream interface of the DPE can generate trace information that can be output from the DPE array 102. For example, the stream interface can be configured to extract trace information from the DPE array 102. The trace information can be generated as a packet-switched stream that contains timestamp data marking event occurrences and / or limited branch traces of execution flow. In one aspect, the trace generated by the DPE can be pushed to a local trace buffer implemented in the PL 310 or pushed to an external RAM using the SoC interface block 104 and the NoC 308. In another aspect, the trace generated by the DPE can be sent to a debug subsystem implemented on-chip.
[0185] In a particular embodiment, each core 602 and memory module 604 of each DPE may include an additional stream interface that is capable of outputting trace data directly to the stream switch 626. The stream interfaces for trace data may be in addition to those interfaces already discussed. The stream switch 626 may be configured to direct trace data onto a packet-switched stream so that trace information from multiple cores and memory modules of different DPEs may travel on a single data stream. As noted, the stream portion of the DPE interconnect network may be configured to send trace data to an on-chip debug system via the PL 310, to external memory via the SoC interface block 104, or directly to a gigabit transceiver via the NoC 308. Examples of different types of trace streams that may be generated include: a program counter (PC) trace stream that produces a PC value that is the opposite of each change in the PC at a branch instruction; and an application data trace stream, including intermediate results within the DPE (e.g., from the core and / or memory module via a corresponding trace data stream).
[0186] Fig. 9 An example connectivity of the cascaded interfaces of the cores in multiple DPEs is shown. Fig. 9 In the example of FIG. 6 , only the core 602 of the DPE is shown. For illustration purposes, other parts of the DPE (such as DPE interconnects and memory modules) are omitted.
[0187] As shown in the figure, the core is combined with Figure 8The cascade interfaces described are connected serially. Core 602-1 is coupled to core 602-2, core 602-2 is coupled to core 602-3, and core 602-3 is coupled to core 602-4. Therefore, application data can be directly propagated from core 602-1 to core 602-2, to core 602-3, to core 602-4. Core 602-4 is coupled up to core 602-8 in the next row. Core 602-8 is coupled to core 602-7, core 602-7 is coupled to core 602-6, and core 602-6 is coupled to core 602-5. Therefore, application data can be directly propagated from core 602-4 to core 608-8, to core 602-7, to core 602-6, to core 602-5. Core 602-5 is coupled up to core 602-9 in the next row. Core 602-9 is coupled to core 602-10, which is coupled to core 602-11, which is coupled to core 602-12. Thus, application data can propagate directly from core 602-5 to core 602-9, to core 602-10, to core 602-11, to core 602-12. Core 602-12 is coupled up in the next row to core 602-16. Core 602-16 is coupled to core 602-15, which is coupled to core 602-14, which is coupled to core 602-13. Thus, application data can propagate directly from core 602-12 to core 602-16, to core 602-15, to core 602-14, to core 602-13.
[0188] Fig. 9 It is intended to illustrate how the cascade interface of the core of a DPE is coupled from one row of DPEs to another row of DPEs within a DPE array.The specific number of columns and / or rows of cores (eg, DPEs) shown is not intended to be limiting. Fig. 9 It is shown that connections between cores using the cascade interface can be made in an "S" shape or zigzag shape at alternating ends of the row of DPEs.
[0189] In embodiments where DPE array 102 implements two or more different clusters of DPEs 304, a first cluster of DPEs is not coupled to a second cluster of DPEs via a cascade and / or stream interface. For example, if a first two rows of DPEs form a first cluster and a second two rows of DPEs form a second cluster, the cascade interface of core 602-5 may be programmed to be disabled so as not to pass data to the cascade input of core 602-9.
[0190] In combination Figure 8 and Fig. 9In the described examples, each core is shown as having a cascade interface used as an input and a cascade interface used as an output. In one or more other embodiments, the cascade interface can be implemented as a bidirectional interface. In a specific embodiment, the core can include an additional cascade interface so that the core can communicate directly with other cores above, below, on the left and / or on the right via the cascade interface. As noted, such an interface can be unidirectional or bidirectional.
[0191] Fig. 10A , Fig. 10B , Fig. 10C , Fig. 10D and Fig.10E An example of connectivity between DPEs is shown. Fig. 10A Example connectivity between DPEs using shared memory is shown. Fig. 10A In the example of , a function or kernel implemented in core 602-15 (e.g., a user circuit design implemented in a DPE and / or DPE array) uses a core interface and a memory interface in DPE 304-15 to operate and place data 1005 (e.g., application data) in memory module 604-15. DPE 304-15 and DPE 304-16 are adjacent DPEs. Therefore, core 602-16 is able to access data 1005 from memory module 604-15 based on a lock obtained from a hardware synchronization circuit device (not shown) in memory module 604-15 for a buffer including data 1005. Shared access to memory module 604-15 by cores 602-15 and 602-16 facilitates high-speed transaction processing because data does not need to be physically transferred from one memory to another in order for core 602-16 to operate on application data.
[0192] Fig. 10B Example connectivity between DPEs using stream switches is shown. Fig. 10BIn the example of FIG. 1 , DPEs 304-15 and 304-17 are non-adjacent DPEs and are likewise separated by one or more intermediate DPEs. Functions or kernels implemented in core 602-15 operate on data 1005 and place data 1005 in memory module 604-15. DMA engine 816-15 of memory module 604-15 retrieves data 1005 based on acquisition of a lock on a buffer for storing data 1005 within memory module 604-15. DMA engine 816-15 sends data 1005 to DPE 304-17 via a stream switch of the DPE interconnect. The DMA engine 816-17 within the memory module 604-17 can retrieve the data 1005 from the stream switch within the DPE 304-17 and store the data 1005 in a buffer of the memory module 604-17 after acquiring a lock for the buffer within the memory module 604-17 from the hardware synchronization circuitry in the memory module 604-17. Fig. 10B , to configure the respective stream switches and DMA engines 816-15 and 816-17 within DPEs 304-15 and 304-17 to operate as described.
[0193] Fig. 10C Another example of inter-DPE connectivity using a stream switch is shown. Fig. 10C In the example of FIG. 3 , DPE 304 - 15 and DPE 304 - 17 are non-adjacent DPEs and are separated by one or more intermediate DPEs. Fig. 10C It is shown that data 1005 can be provided directly from DMA 816-15 to the core of another DPE via a stream switch. As shown, DMA 816-15 places data 1005 on the stream switch of DPE 304-15. Core 602-17 can receive data 1005 directly from the stream switch in DPE 304-17 using the stream interface included therein, and data 1005 does not traverse to memory module 604-17. Configuration data can be loaded by Fig. 10C 816 - 15 to configure the DPEs 304 - 15 and 304 - 17 and the corresponding stream switches of the DMA 816 - 15 to operate as described.
[0194] generally, Fig. 10CAn example of DMA to core data transfer is shown. It should be understood that core to DMA data transfer can also be implemented. For example, the core 602-17 can send data to the DPE 304-15 via the stream interface included therein and the stream switch of the DPE 304-17. The DMA engine 816-15 can extract data from the stream switch included in the DPE 304-15 and store the data in the memory module 604-15.
[0195] Fig. 10D Another example of connectivity between DPEs using a stream switch is shown. Fig. 10D , the cores 602-15, 602-17, and 602-19 of each different, non-adjacent DPE can communicate directly with each other via the stream interface of each corresponding DPE. Fig. 10D In the example of , core 602-15 is capable of broadcasting the same data stream to core 602-17 and core 602-19. The broadcast functionality of the stream interface within each corresponding DPE including cores 602-15, 602-17, and 602-19 can be programmed by loading configuration data to configure the corresponding stream switches and / or stream interfaces. In one or more other embodiments, core 602-15 is capable of multicasting data to cores of other DPEs.
[0196] Fig.10E An example of inter-DPE connectivity using a stream switch and cascade interface is shown. Fig.10E , DPE 304 - 15 and DPE 304 - 16 are adjacent DPEs. In some cases, a kernel may be split to run on multiple cores. In that case, the intermediate accumulation results of one sub-kernel can be transferred to the sub-kernel in the next core via the cascade interface.
[0197] exist Fig.10E In the example of , the core 602-15 receives data 1005 via a stream switch and operates on the data 1005. The core 602-15 generates intermediate result data 1010 and outputs the intermediate result data 1010 directly from the accumulation register therein to the core 602-16. In a specific embodiment, the cascade interface of the core 602-15 is capable of transmitting accumulator values in each clock cycle of the DPE 304-15. The data 1005 received by the core 602-15 is also propagated to the core 602-16 via the stream switch in the DPE interconnect, thereby allowing the core 602-16 to operate on both the data 1005 (e.g., original data) and the intermediate result data 1010 generated by the core 602-15.
[0198] In the example of FIG. 10 , the sending of a data stream, the broadcasting of a data stream, and / or the multicasting of a data stream are shown in the horizontal direction. It should be understood that a data stream can be sent, broadcasted, and / or multicasted from a DPE to any other DPE in the array of DPEs. Thus, based on the configuration data loaded into each such DPE, a data stream can be sent, broadcasted, or multicasted to the left, right, upward, downward, and / or to the DPE in the diagonal required to reach the intended destination DPE.
[0199] Fig.11 An example of event processing circuitry within a DPE is shown. A DPE may include event processing circuitry that is interconnected with event processing circuitry of other DPEs. Fig.11 In the example of , event processing circuitry is implemented in core 602 and within memory module 604. Core 602 may include event broadcasting circuitry 1102 and event logic 1104. Memory module 604 may include separate event processing circuitry including event broadcasting circuitry 1106 and event logic 1108.
[0200] Event broadcast circuitry 1102 may be connected to Fig.11 The event broadcast circuitry 1102 may also be connected to the event broadcast circuitry within each core of the adjacent DPEs above and below the example DPE shown. Fig.11 100 is a schematic diagram of an example DPE in which an event broadcast circuit is provided in a memory module of an adjacent DPE to the left of the example DPE shown. As shown, event broadcast circuit 1102 is connected to event broadcast circuit 1106. Event broadcast circuit 1106 may be connected to Fig.11 The event broadcast circuitry 1106 may also be connected to the memory modules of the adjacent DPEs above and below the example DPE shown. Fig.11 The example DPE shown has event broadcast circuitry within the core of an adjacent DPE to the right.
[0201] In this way, the event processing circuit device of the DPE can form an independent event broadcast network within the DPE array. The event broadcast network within the DPE array can exist independently of the DPE interconnection network. In addition, the event broadcast network can be configured separately by loading appropriate configuration data into the configuration register 624 and / or 536.
[0202] exist Fig.11In the example of , event broadcast circuitry 1102 and event logic 1104 may be configured by configuration registers 624. Event broadcast circuitry 1106 and event logic 1108 may be configured by configuration registers 636. Configuration registers 624 and 636 may be written to via a memory region mapping switch of DPE interconnect 606. Fig.11 In the example of , configuration register 624 programs event logic 1104 to detect specific types of events occurring within core 602. For example, the configuration data loaded into configuration register 624 determines which of a plurality of different types of predetermined events are detected by event logic 1104. Examples of events may include, but are not limited to, the beginning and / or end of a read operation of core 602, the beginning and / or end of a write operation of core 602, and the occurrence of other operations performed by core 602. Similarly, configuration register 636 programs event logic 1108 to detect specific types of events occurring within memory module 604. Examples of events may include, but are not limited to, the beginning and / or end of a read operation of DMA engine 816, the beginning and / or end of a write operation of DMA engine 816, a stall, and the occurrence of other operations performed by memory module 604. For example, the configuration data loaded into configuration register 636 determines which of a plurality of different types of predetermined events are detected by event logic 1108. It should be understood that event logic 1104 and / or event logic 1108 can detect events originating from and / or relative to DMA engine 816, memory mapped switch 632, stream switch 626, memory interface of memory module 604, core interface of core 602, cascade interface of core 602 and / or other components located in the DPE.
[0203] Configuration registers 624 can also program event broadcast circuitry 1102, while configuration registers 636 can program event broadcast circuitry 1106. For example, configuration data loaded into configuration registers 624 can determine which of the events received by event broadcast circuitry 1102 from other event broadcast circuitry are propagated to other event broadcast circuitry and / or SoC interface block 104. Configuration data can also specify which events generated internally by event logic 1104 are propagated to other event broadcast circuitry and / or SoC interface block 104.
[0204] Similarly, the configuration data loaded into the configuration registers 636 may determine which of the events received by the event broadcast circuitry 1106 from other event broadcast circuitry are propagated to the other event broadcast circuitry and / or the SoC interface block 104. The configuration data may also specify which events generated internally by the event logic 1108 are propagated to the other event broadcast circuitry and / or the SoC interface block 104.
[0205] Thus, events generated by event logic 1104 may be provided to event broadcast circuitry 1102 and may be broadcast to other DPEs. Fig.11 In the example of , the event broadcasting circuit device 1102 can broadcast events, whether generated internally or received from other DPEs, to the DPE above, to the DPE to the left, to the DPE below, or to the SoC interface block 104. The event broadcasting circuit device 1102 can also broadcast events to the event broadcasting circuit device 1106 within the memory module 604.
[0206] Events generated by event logic 1108 may be provided to event broadcast circuitry 1106 and may be broadcast to other DPEs. Fig.11 In the example of , event broadcast circuit device 1106 can broadcast events, whether generated internally or received from other DPEs, to the DPE above, to the DPE to the right, to the DPE below, or to the SoC interface block 104. Event broadcast circuit device 1106 can also broadcast events to event broadcast circuit device 1102 within core 602.
[0207] exist Fig.11 In the example of , an event broadcast circuit device located in a core vertically communicates with an event broadcast circuit device in a core of an adjacent DPE located above and / or below. In a case where a DPE is immediately (or adjacent) above the SoC interface block 104, the event broadcast circuit device in the core of the DPE is able to communicate with the SoC interface block 104. Similarly, an event broadcast circuit device located in a memory module vertically communicates with an event broadcast circuit device in a memory module of an adjacent DPE located above and / or below. In a case where a DPE is immediately (e.g., adjacent) above the SoC interface block 104, the event broadcast circuit device in the memory module of the DPE is able to communicate with the SoC interface block 104. The event broadcast circuit device is also able to communicate with an event broadcast circuit device immediately to the left and / or right of the event broadcast circuit device, regardless of whether such event broadcast circuit device is located in another DPE and / or within a core or memory module.
[0208] Once configuration registers 624 and 636 are written, event logic 1104 and 1108 can run in the background. In a particular embodiment, event logic 1104 generates events only in response to specific conditions detected within core 602; and event logic 1108 generates events only in response to specific conditions detected within memory module 604.
[0209] Fig.12 Another example architecture for DPE 304 is shown. Fig.12In the example of , DPE 304 includes multiple different cores and may be referred to as a "cluster" type of DPE architecture. Fig.12 In FIG. 1 , DPE 304 includes cores 1202, 1204, 1206, and 1208. Each of cores 1202-1208 is connected to a core interface 1210, 1212, 1214, and 1216 (in Fig.12 Each of the core interfaces 1210-1216 is coupled to a plurality of memory banks 1222-1 to 1222-N via a crossbar switch 1224. Through the crossbar switch 1224, any one of the cores 1202-1208 can access any one of the memory banks 1222-1 to 1222-N. Fig.12 In the example architecture of , cores 1202 - 1208 are able to communicate with each other via a shared memory bank 1222 of a memory pool 1220 .
[0210] In one or more embodiments, the memory pool 1220 may include 32 memory banks. The number of memory banks included in the memory pool 1220 is provided for purposes of illustration and not limitation. In other embodiments, the number of memory banks included in the memory pool 1220 may be more than 32 or less than 32.
[0211] exist Fig.12 In the example of , DPE 304 includes a memory-mapped switch 1226. Memory-mapped switch 1226 includes multiple memory-mapped interfaces (not shown) that are capable of coupling to memory-mapped switches within adjacent DPEs and to memory pool 1220 in each cardinal direction (e.g., north, south, west, and east). Each memory-mapped interface may include one or more master interfaces and one or more slave interfaces. For example, memory-mapped switch 1226 is coupled to crossbar switch 1224 via a memory-mapped interface. Memory-mapped switch 1226 is capable of transmitting configuration data, control data, and debug data as described in conjunction with other example DPEs within the present disclosure. In this way, memory-mapped switch 1226 is capable of loading configuration registers (not shown) in DPE 304. In Fig.12 In the example of , DPE 304 may include configuration registers for controlling the operation of stream switch 1232 , cores 1202 - 1208 , and DMA engine 1234 .
[0212] exist Fig.12In the example of , the memory mapped switch 1226 is capable of communicating in each of the four cardinal directions. In other embodiments, the memory mapped switch 1226 is capable of communicating only in the north and south directions. In other embodiments, the memory mapped switch 1226 may include additional memory mapped interfaces that allow the memory mapped switch 1226 to communicate with more than four other entities, thereby allowing communication with other DPEs in diagonal directions and / or other non-adjacent DPEs.
[0213] DPE 304 also includes a stream switch 1232. Stream switch 1232 includes multiple stream interfaces (not shown) that can be coupled to stream switches in adjacent DPEs in each cardinal direction (e.g., north, south, west, and east) and coupled to cores 1202-1208. Each stream switch can include one or more master interfaces and one or more slave interfaces. Stream switch 1232 also includes a stream interface coupled to DMA engine 1234.
[0214] DMA engine 1234 is coupled to crossbar switch 1224 via interface (IF) 1218. DMA engine 1234 may include two interfaces. For example, DMA engine 1234 may include a memory-to-stream interface that can read data from one or more of memory banks 1222 and send the data on stream switch 1232. DMA engine 1234 may also include a stream-to-memory interface that receives data via stream switch 1232 and stores the data in one or more of memory banks 1222. Each of the interfaces, whether memory-to-stream interface or stream-to-memory interface, may support one input / output stream or multiple concurrent input / output streams.
[0215] Fig.12 The example architecture supports inter-DPE communication via both memory-mapped switch 1226 and stream switch 1232. As shown, memory-mapped switch 1226 can communicate with memory-mapped switches of adjacent DPEs above, below, to the left, and to the right. Similarly, stream switch 1232 can communicate with stream switches of adjacent DPEs above, below, to the left, and to the right.
[0216] In one or more embodiments, both the memory mapped switch 1226 and the stream switch 1232 can support data transfer between cores of other (adjacent and non-adjacent) DPEs to share application data. The memory mapped switch 1226 can also support the transfer of configuration data, control data, and debug data for the purpose of configuring the DPE 304. In a specific embodiment, the stream switch 1232 supports the transfer of application data, while the memory mapped switch 1226 only supports the transfer of configuration data, control data, and debug data.
[0217] exist Fig.12 In the example of FIG. 1 , cores 1202-1208 are serially connected via a cascade interface as described above. In addition, core 1202 is coupled to Fig.12 The cascade interface (e.g., output) of the rightmost core in the adjacent DPE to the left of the DPE, and core 1208 is coupled to Fig.12 The cascade interface (e.g., input) of the leftmost core in the adjacent DPE to the right of the DPE. Fig. 9 As shown, the cascade interfaces of the DPEs using the cluster architecture can be connected row by row. In one or more other embodiments, instead of and / or in addition to the horizontal cascade connections, one or more cores 1202-1208 can be connected to the cores in the adjacent DPEs above and / or below via the cascade interfaces.
[0218] Fig.12 The example architecture of can be used to implement DPE and form a DPE array as described herein. Compared with other example DPE architectures described in this disclosure, Fig.12 The example architecture of increases the amount of memory available to the core. Therefore, for applications where the core needs to access a larger amount of memory, the Fig.12 The architecture, Fig.12 The architecture clusters multiple cores in a single DPE. For illustration purposes, Fig.12 Therefore, one or more (e.g., less than all of the cores 1202-1208 of the DPE 304) may access the memory pool 1220 and may access more than the cores 1202-1208 of the DPE 304 based on the configuration of the DPE 304. Fig.12 In the example of a configuration register (not shown) a larger number of memory devices are used for configuration data.
[0219] Fig.13 Shown for Figure 1 An example architecture of a DPE array 102 is shown in FIG. Fig.13In the example of , the SoC interface block 104 provides an interface between the DPE 304 and other subsystems of the device 100. The SoC interface block 104 integrates the DPE into the device. The SoC interface block 104 can transfer configuration data to the DPE 304, transfer events from the DPE 304 to other subsystems, transfer events from other subsystems to the DPE 304, generate interrupts and transfer interrupts to entities external to the DPE array 102, transfer application data between other subsystems and the DPE 304, and / or transfer trace data and / or debug data between other subsystems and the DPE 304.
[0220] exist Fig.13 In the example of , SoC interface block 104 includes a plurality of interconnected tiles. For example, SoC interface block 104 includes tiles 1302, 1304, 1306, 1308, 1310, 1312, 1314, 1316, 1318, and 1320. Fig.13 In the example of , tiles 1302-1320 are arranged in a row. In other embodiments, the tiles may be arranged in columns, a grid, or another layout. For example, the SoC interface block 104 may be implemented as a column of tiles on the left side of the DPE 304, a column of tiles on the right side of the DPE 304, a column of tiles between the columns of the DPE 304, and the like. In another embodiment, the SoC interface block 104 may be located above the DPE array 102. The SoC interface block 104 may be implemented such that the tiles are located in any combination below the DPE array 102, to the left of the DPE array 102, to the right of the DPE array 102, and / or above the DPE array 102. In this regard, Fig.13 It is provided for purposes of illustration and not limitation.
[0221] In one or more embodiments, blocks 1302-1320 have the same architecture. In one or more other embodiments, blocks 1302-1320 may be implemented using two or more different architectures. In a particular embodiment, blocks within SoC interface block 104 may be implemented using different architectures, where each different block architecture supports communication with a different type of subsystem or combination of subsystems of device 100.
[0222] exist Fig.13In the example of , tiles 1302-1320 are coupled so that data can be propagated from one tile to another. For example, data can propagate from tile 1302 through tiles 1304, 1306, and down the line of tiles to tile 1320. Similarly, data can propagate from tile 1320 to tile 1302 in the reverse direction. In one or more embodiments, each of tiles 1302-1320 can be used as an interface for multiple DPEs. For example, each of tiles 1302-1320 can be used as an interface for a subset of DPEs 304 of the DPE array 102. The subset of DPEs to which each tile provides an interface can be mutually exclusive so that more than one tile of the SoC interface block 104 does not provide an interface to a DPE.
[0223] In one example, each of tiles 1302-1320 provides an interface for a column of DPEs 304. For purposes of illustration, tile 1302 provides an interface to the DPEs of column A. Tile 1304 provides an interface to the DPEs of column B, and so on. In each case, the tile includes a direct connection to the adjacent DPE in the column of DPEs, in this example the bottom DPE. For example, referring to column A, tile 1302 is directly connected to DPE 304-1. Other DPEs within column A may communicate with tile 1302, but may communicate through the DPE interconnects of the intermediate DPEs in the same column.
[0224] For example, block 1302 can receive data from another source (such as PS 312, PL 310) and / or another hardwired circuit block (e.g., an ASIC block). Block 1302 can provide those portions of the data addressed to a DPE in rank A to such DPE, while sending data addressed to DPEs in other ranks (e.g., DPEs that are not interfaces to block 1302) to block 1304. Block 1304 can perform the same or similar processing, wherein data received from block 1302 addressed to a DPE in rank B is provided to such DPE, while data addressed to DPEs in other ranks is sent to block 1306, and so on.
[0225] In this manner, data may propagate between tiles of the SoC interface block 104 until reaching a tile that serves as an interface for the DPE to which the data is addressed (e.g., a "target DPE"). The tile that serves as an interface for the target DPE can direct the data to the target DPE using the DPE's memory mapped switches and / or the DPE's stream switches.
[0226] As previously mentioned, the use of columns is an example implementation. In other embodiments, each tile of the SoC interface block 104 can provide an interface to a row of DPEs of the DPE array 102. In the case where the SoC interface block 104 is implemented as a column of tiles, this configuration can be used whether on the left side, right side, or between columns of the DPE 304. In other embodiments, the subset of DPEs to which each tile provides an interface can be any combination of less than all DPEs of the DPE array 102. For example, DPE 304 can be assigned to the tiles of the SoC interface block 104. The specific physical layout of such DPEs can vary based on the connectivity of the DPEs established by the DPE interconnection. For example, tile 1302 can provide interfaces to DPE 304-1, 304-2, 304-11, and 304-12. Another tile of the SoC interface block 104 can provide interfaces to four other DPEs, and so on.
[0227] Fig.14A , 14B 14C show an example architecture of a block for implementing the SoC interface block 104 . Fig.14A An example implementation of block 1304 is shown. Fig.14A The architecture shown in can also be used to implement any other blocks included in the SoC interface block 104.
[0228] Tile 1304 includes a memory mapped switch 1402. The memory mapped switch 1402 may include multiple memory mapped interfaces for communicating in each of multiple different directions. As an illustrative and non-limiting example, the memory mapped switch 1402 may include one or more memory mapped interfaces, wherein the memory mapped interface has a master interface that is vertically connected to the memory mapped interface of the DPE immediately above. In this way, the memory mapped switch 1402 is capable of serving as a master interface for the memory mapped interfaces of one or more of the DPEs. In a specific example, the memory mapped switch 1402 may serve as a master interface for a subset of the DPEs. For example, the memory mapped switch 1402 may serve as a master interface for a column of DPEs above tile 1304 (e.g., Fig.13 It should be understood that the memory-mapped switch 1402 may include additional memory-mapped interfaces to connect to multiple different circuits (e.g., DPEs) within the DPE array 102. The memory-mapped interfaces of the memory-mapped switch 1402 may also include one or more slave interfaces capable of communicating with circuits located above the tile 1304 (e.g., one or more DPEs).
[0229] exist Fig.14AIn the example of, the memory mapped switch 1402 may include one or more memory mapped interfaces that facilitate communication with memory mapped switches in adjacent tiles (e.g., tiles 1302 and 1306) in a horizontal direction. For illustrative purposes, the memory mapped switch 1402 may be connected to adjacent tiles in a horizontal direction via memory mapped interfaces, wherein each such memory mapped interface includes one or more master interfaces and / or one or more slave interfaces. Thus, the memory mapped switch 1402 is capable of moving data (e.g., configuration data, control data, and / or debug data) from one tile to another to reach the correct DPE and / or subset of multiple DPEs and directing the data to the target DPE, whether such DPE is in a column above tile 1304 or in another subset of another tile of the SoC interface block 104 used as an interface. For example, if a memory mapped transaction is received from the NoC 308, the memory mapped switch 1402 is capable of distributing the transaction horizontally to other tiles within (e.g.) the SoC interface module 104.
[0230] The memory mapped switch 1402 may also include a memory mapped interface having one or more master and / or slave interfaces coupled to configuration registers 1436 within the tile 1304. Through the memory mapped switch 1402, configuration data may be loaded into the configuration registers 1436 to control various functions and operations performed by components within the tile 1304. Fig.14A , 14B 14C show connections between configuration register 1436 and one or more elements of block 1304. However, it should be understood that configuration register 1436 may control other elements of block 1304 and thus have connections to such other elements, even though such connections are not shown. Fig.14A , 14B and / or as shown in 14C.
[0231] The memory mapped switch 1402 may include a memory mapped interface coupled to the NoC interface 1426 via a bridge 1418. The memory mapped interface may include one or more master interfaces and / or slave interfaces. The bridge 1418 is capable of converting memory mapped data (e.g., configuration data, control data, and / or debug data) transmitted from the NoC 308 into memory mapped data that can be received by the memory mapped switch 1402.
[0232] Block 1304 may also include event processing circuitry. For example, block 1304 includes event logic 1432. Event logic 1432 may be configured by configuration registers 1436. Fig.14AIn the example of FIG. 14 , event logic 1432 is coupled to control, debug, and trace (CDT) circuitry 1420. Configuration data loaded into configuration registers 1436 defines specific events that can be detected locally within tile 1304. Event logic 1432 is capable of detecting a variety of different events for each configuration register 1436, the events originating from and / or involving DMA engine 1412, memory mapped switch 1402, stream switch 1406, first-in-first-out (FIFO) memory located at PL interface 1410, and / or NoC stream interface 1414. Examples of events may include, but are not limited to, DMA complete transfer, lock release, lock acquisition, PL transfer end, or other events related to the start or end of data flow through tile 1304. Event logic 1432 may provide such events to event broadcast circuitry 1404 and / or CDT circuitry 1420. For example, in another embodiment, the event logic 1432 may not have a direct connection to the CDT circuit 1420 , but rather a connection to the CDT circuit 1420 via the event broadcasting circuitry 1404 .
[0233] Tile 1304 includes event broadcast circuitry 1404 and event broadcast circuitry 1430. Each of event broadcast circuitry 1404 and event broadcast circuitry 1430 provides an interface between the event broadcast network of DPE array 102, other tiles of SoC interface block 104, and PL 310 of device 100. Event broadcast circuitry 1404 is coupled to event broadcast circuitry of adjacent or neighboring tile 1302 and is coupled to event broadcast circuitry 1430. Event broadcast circuitry 1430 is coupled to event broadcast circuitry of adjacent or neighboring tile 1306. In one or more other embodiments, the tiles of SoC interface block 104 are arranged in a grid or array, and event broadcast circuitry 1404 and / or event broadcast circuitry 1430 may be connected to event broadcast circuitry in other tiles located above and / or below tile 1304.
[0234] exist Fig.14A In the example of , event broadcast circuitry 1404 is coupled to an event broadcast circuitry in a core that is proximate to a DPE of tile 1304 (e.g., in column B, DPE 304-2 is immediately above tile 1304). Event broadcast circuitry 1404 is also coupled to PL interface 1410. Event broadcast circuitry 1430 is coupled to an event broadcast circuitry in a memory module that is proximate to a DPE of tile 1304 (e.g., in column B, DPE 304-2 is immediately above tile 1304). Although not shown, in another embodiment, event broadcast circuitry 1430 may also be coupled to PL interface 1410.
[0235] Event broadcast circuitry 1404 and event broadcast circuitry 1430 are capable of sending events generated internally by event logic 1432, events received from other tiles of SoC interface block 104, and / or events received from DPEs in column B (or other DPEs of DPE array 102) to other tiles. Event broadcast circuitry 1404 is also capable of sending such events to PL 310 via PL interface 1410. In another example, events may be sent from event broadcast circuitry 1404 to other blocks and / or subsystems in device 100 (such as, ASICs and / or PL circuit blocks external to DPE array 102) using PL interface block 1410. Further, PL interface 1410 may receive events from PL 310 and provide such events to event broadcast switch 1404 and / or stream switch 1406. In one aspect, the event broadcasting circuitry 1404 is capable of sending any events received from the PL 310 to other tiles of the SoC interface block 104 and / or to DPEs in Column B and / or other DPEs of the DPE array 102 via the PL interface 1410. In another example, events received from the PL 310 may be sent from the event broadcasting circuitry 1404 to other blocks and / or subsystems (such as, ASICs) in the device 100. Because events may be broadcast between tiles in the SoC interface block 104, the event may reach any DPE in the DPE array 102 by traversing the tiles in the SoC interface block 104 and the event broadcasting circuitry to the target (e.g., intended) DPE. For example, the event broadcasting circuitry in a tile of the SoC interface block 104 below the column (or subset) of DPEs managed by the tile including the target DPE may propagate the event to the target DPE.
[0236] exist Fig.14A 1406. In the example of FIG. 1404, event broadcast circuitry 1404 and event logic 1432 are coupled to CDT circuitry 1420. Event broadcast circuitry 1404 and event logic 1432 are capable of sending events to CDT circuitry 1420. CDT circuitry 1420 is capable of grouping received events and sending events from event broadcast circuitry 1404 and / or event logic 1432 to stream switch 1406. In a particular embodiment, event broadcast circuitry 1430 may be connected to stream switch 1406 and / or also connected to CDT circuitry 1420.
[0237] In one or more embodiments, event broadcast circuitry 1404 and event broadcast circuitry 1430 are capable of Fig.14A One or more or all directions shown (e.g., via Fig.14A) to collect broadcast events. In a particular embodiment, event broadcast circuit device 1404 and / or event broadcast circuit device 1430 is capable of performing a logical "OR" of signals and forwarding the results in one or more or all directions (e.g., including to CDT circuit 1420). Each output from event broadcast circuit device 1404 and event broadcast circuit device 1430 can include a bit mask, which can be configured by configuration data loaded into configuration register 1436. The bit mask determines the events broadcast in each direction respectively. For example, such a bit mask can eliminate unnecessary or duplicate event propagation.
[0238] The interrupt handler 1434 is coupled to the event broadcasting circuitry 1404 and is capable of receiving events broadcasted from the event broadcasting circuitry 1404. In one or more embodiments, the interrupt handler 1434 may be configured by configuration data loaded into the configuration register 1436 to generate interrupts in response to events (e.g., DPE generated events, events generated within the tile 1304, and / or events generated by the PL 310) and / or combinations of events selected from the event broadcasting circuitry 1404. The interrupt handler 1434 is capable of generating interrupts based on to the PS 312 and / or to other device-level management blocks within the device 100. In this way, the interrupt handler 1434 is capable of notifying the PS 312 and / or such other device-level management blocks of events occurring in the DPE array 102, events occurring in the tile of the SoC interface block 104, and / or events occurring in the PL 310 based on the interrupts generated by the interrupt handler 1434.
[0239] In certain embodiments, the interrupt handler 1434 may be coupled to an interrupt handler or interrupt port of the PS 312 and / or other device-level management blocks via a direct connection. In one or more other embodiments, the interrupt handler 1434 may be coupled to the PS 312 and / or other device-level management blocks via another interface.
[0240] PL interface 1410 couples to and provides an interface to PL 310 of device 100. In one or more embodiments, PL interface 410 provides asynchronous clock domain crossing between the DPE array clock and the PL clock. PL interface 1410 may also provide level shifters and / or isolation cells to integrate with PL power rails. In certain embodiments, PL interface 1410 may be configured to provide a 32-bit, 64-bit, and / or 128-bit interface with FIFO support to handle back pressure. The specific width of PL interface 1410 may be controlled by configuration data loaded into configuration register 1436. In Fig.14AIn the example of , PL interface 1410 is directly coupled to one or more PL interconnect blocks 1422. In a particular embodiment, PL interconnect blocks 1422 are implemented as hardwired circuit blocks coupled to interconnect circuit devices located in PL 310.
[0241] In one or more other embodiments, PL interface 1410 is coupled to other types of circuit blocks and / or subsystems. For example, PL interface 1410 may be coupled to an ASIC, analog / mixed signal circuitry, and / or other subsystems. Thus, PL interface 1410 is capable of transferring data between block 1304 and such other subsystems and / or blocks.
[0242] exist Fig.14A In the example of , block 1304 includes a stream switch 1406. Stream switch 1406 is coupled to a stream switch of an adjacent or neighboring block 1302 and is coupled to a stream switch of an adjacent or neighboring block 1306 via one or more stream interfaces. Each stream interface may include one or more master interfaces and / or one or more slave interfaces. In a particular embodiment, each pair of adjacent stream switches is capable of exchanging data via one or more streams in each direction. Stream switch 1406 is also coupled to a stream switch in a DPE (e.g., DPE304-2) in column B that is immediately above block 1304 via one or more stream interfaces. As discussed, the stream interface may include one or more stream slave interfaces and / or stream master interfaces. Stream switch 1406 also communicates with each other via a stream multiplexer / demultiplexer 1408 (in Fig.14A The stream switch 1406 may include a stream multiplexer / demultiplexer 1408 (hereinafter referred to as a stream multiplexer / demultiplexer) coupled to the PL interface 1410, the DMA engine 1412, and / or the NoC stream interface 1414. For example, the stream switch 1406 may include one or more stream interfaces, and the stream interface is used to communicate with each of the PL interface 1410, the DMA engine 1412, and / or the NoC stream interface 1414 through the stream multiplexer / demultiplexer 1408.
[0243] In one or more other embodiments, stream switch 1406 may be coupled to other circuit blocks in other directions and / or diagonally, depending on the data and / or tile of the stream interface included and / or the arrangement of DPEs and / or other circuit blocks around tile 1304 .
[0244] In one or more embodiments, the stream switch 1406 is configured by configuration data loaded into the configuration register 1436. For example, the stream switch 1406 can be configured to support packet switching and / or circuit switching operations based on the configuration data. In addition, the configuration data defines the specific DPE and / or DPEs within the DPE array 102 with which the stream switch 1406 communicates. In one or more embodiments, the configuration data defines the specific DPE and / or a subset of the DPEs of the DPE array 102 (e.g., DPEs within column B) with which the stream switch 1406 communicates.
[0245] The stream multiplexer / demultiplexer 1408 can direct data received from the PL interface 1410, the DMA engine 1412, and / or the NoC stream interface 1414 to the stream switch 1406. Similarly, the stream multiplexer / demultiplexer 1408 can direct data received from the stream switch 1406 to the PL interface 1410, the DMA engine 1412, and / or to the NoC stream interface 1414. For example, the stream multiplexer / demultiplexer 1408 can be programmed by configuration data stored in the configuration registers 1436 to route selected data to the PL interface 1410, to route selected data to the DMA engine 1412, where such data is sent over the NoC 308 as a memory mapped transaction, and / or to route selected data to the NoC stream interface 1414, where the data is sent over the NoC 308 as a data stream or multiple data streams.
[0246] DMA engine 1412 can act as a master interface to direct data into NoC 308 and onto NoC interface 1426 through selector block 1416. DMA engine 1412 can receive data from the DPE and provide such data to NoC 308 as memory mapped data transactions. In one or more embodiments, DMA engine 1412 includes hardware synchronization circuitry that can be used to synchronize multiple channels included in DMA engine 1412 and / or channels within DMA engine 1412 with a master interface that polls and drives lock requests. For example, the master interface can be a device implemented within PS 312 or PL 310. The master interface can also receive interrupts generated by hardware synchronization circuitry within DMA engine 1412.
[0247] In one or more embodiments, the DMA engine 1412 can access external memory. For example, the DMA engine 1412 can receive a data stream from the DPE and send the data stream to the external memory through the NoC 308 to a memory controller located within the SoC. The memory controller then directs the data received in the data stream to the external memory (e.g., starts reading and / or writing to the external memory as requested by the DMA engine 1412). Similarly, the DMA engine 1412 can receive data from the external memory, where the data can be distributed to other tiles of the SoC interface block 104 and / or distributed upward to the target DPE.
[0248] In certain embodiments, the DMA engine 1412 includes security bits that can be set using the DPE global control settings register (DPE GCS register) 1438. External memory can be divided into different areas or partitions, where the DPE array 102 is only allowed to access specific areas of the external memory. The security bits within the DMA engine 1412 can be set so that the DPE array 102, with the aid of the DMA engine 1412, can only access specific areas of external memory as allowed by each security bit. For example, an application implemented by the DPE array 102 can be restricted to only access specific areas of external memory, restricted to only read from specific areas of external memory, and / or restricted to completely write to external memory using this mechanism.
[0249] Security bits within the DMA engine 1412 that control access to external memory may be implemented as a whole to control the DPE array 102 or may be implemented in a more granular manner where access to external memory may be specified and / or controlled on a per-DPE basis (e.g., core by core or on a core group basis) where a core group is configured to operate in a coordinated manner, such as to implement a kernel and / or other application.
[0250] The NoC stream interface 1414 can receive data from the NoC 308 via the NoC interface 1426 and forward the data to the stream to the multiplexer / demultiplexer 1408. The NoC stream interface 1414 can also receive data from the stream multiplexer / demultiplexer 1408 and forward the data to the NoC interface 1426 through a selector block 1416. The selector block 1416 can be configured to pass data from the DMA engine 1412 or from the NoC stream interface 1414 to the NoC interface 1426.
[0251] The CDT circuit 1420 can perform control operations, debug operations, and trace operations within the tile 1304. With respect to debugging, each of the registers located in the tile 1304 is mapped to a memory map accessible via the memory-mapped switch 1402. For example, the CDT circuit 1420 can include circuit devices such as: trace hardware, trace buffers, performance counters, and / or pause logic. The trace hardware of the CDT circuit 1420 can collect trace data. The trace buffer of the CDT circuit 1420 can buffer the trace data. The CDT circuit 1420 can also output the trace data to the stream switch 1406.
[0252] In one or more embodiments, the CDT circuit 1420 can collect data (e.g., trace data and / or debug data), package such data, and then output the packaged data through the stream switch 1406. For example, the CDT circuit 1420 can output the packaged data and provide such data to the stream switch 1406. Additionally, the configuration registers 1436 or other registers can be read or written during debugging via memory-mapped transactions through the memory-mapped switch 1402 of the corresponding tile. Similarly, performance counters within the CDT circuit 1420 can be read or written during profiling via memory-mapped transactions through the memory-mapped switch 1402 of the corresponding tile.
[0253] In one or more embodiments, the CDT circuit 1420 is capable of receiving any event propagated by the event broadcast circuit device 1404 (or the event broadcast circuit device 1430) or the selected event of each bit mask utilized by the interface of the event broadcast circuit device 1404 coupled to the CDT circuit 1420. The CDT circuit 1420 is also capable of receiving events generated by the event logic 1432. For example, the CDT circuit 1420 is capable of receiving broadcast events from other blocks of the PL 310, the DPE 304, the tile 1304 (e.g., the event logic 1432 and / or the event broadcast switch 1404) and / or the SoC interface block 104. The CDT circuit 1420 is capable of packaging (e.g., grouping) multiple such events into data packets and associating the grouped events with timestamps. The CDT circuit 1420 is also capable of sending the grouped events to a destination outside the tile 1304 over the stream switch 1406. Events may be sent through the PL interface 1410 , the DMA engine 1412 , and / or the NoC stream interface 1414 by way of the stream switch 1406 and the stream multiplexer / demultiplexer 1408 .
[0254] The DPE GCS register 1438 may store a DPE global control setting / bit (also referred to herein as a "security bit") used to enable or disable secure access to and / or from the DPE array 102. Fig. 14C The SoC security / initialization interface described in more detail is used to program the DPE GCS register 1438. The security bits received from the SoC security / initialization interface can be passed to Fig.14A The bus shown in FIG. 1 propagates from one tile to the next tile of the SoC interface block 104 .
[0255] In one or more embodiments, external memory mapped data transfers to the DPE array 102 (e.g., using the NoC 308) are not secure or trusted. Without setting the secure bit within the DPE GCS register 1438, any entity in the device 100 that is capable of communicating via memory mapped data transfers (e.g., over the NoC 308) is able to communicate with the DPE array 102. By setting the secure bit within the DPE GCS register 1438, specific entities that are allowed to communicate with the DPE array 102 can be limited, such that only designated entities that are capable of generating secure traffic can communicate with the DPE array 102.
[0256] For example, the memory mapped interface of the memory mapped switch 1402 is capable of communicating with the NoC 308. The memory mapped data transfer may include additional sideband signals, such as a bit that specifies whether the transaction is secure or unsecure. When the secure bit within the DPE GCS register 1438 is set, then the memory mapped transaction entering the SoC interface block 104 must have a sideband signal that is set to indicate that the memory mapped transaction arriving from the NoC 308 to the SoC interface block 104 is secure. When the memory mapped transaction arriving at the SoC interface block 104 does not have the sideband bit set and the secure bit is set within the DPE GCS register 1438, then the SoC interface block 104 does not allow the transaction to enter or pass to the DPE 304.
[0257] In one or more embodiments, the SoC includes a security agent (e.g., circuitry) that serves as a root of trust. When the security bit of the DPE GCS register 1438 is set, the security agent is able to configure permissions for different entities (e.g., circuitry) within the SoC to set the sideband bit within a memory-mapped transaction in order to access the DPE array 102. When configuring the SoC, the security agent grants permissions to different master interfaces that may be implemented in the PL 310 or PS 312, thereby enabling such master interfaces to have the ability to issue secure transactions to the DPE array 102 over (or without) the NoC 308.
[0258] Fig. 14B Another example implementation of block 1304 is shown. Fig. 14B The example architecture shown in can also be used to implement any other blocks included in the SoC interface block 104. Fig. 14B The example shows Fig.14A A simplified version of the architecture shown in . Fig. 14B The block architecture provides connectivity between the DPE and other subsystems and / or blocks within the device 100. For example, Fig. 14B Block 1304 may provide an interface between the DPE and PL 310, analog / mixed signal circuit blocks, ASICs, or other subsystems described herein. Fig. 14B The tile architecture of does not provide connectivity to the NoC 308. Therefore, the DMA engine 1412, NoC interface 1414, selector block 1416, bridge 1418, and stream multiplexer / demultiplexer 1408 are omitted. In this way, a smaller area of the SoC can be used to implement Fig. 14B 1304. Additionally, as shown, stream switch 1406 is directly coupled to PL interface 1410.
[0259] To configure the DPE from the NoC 308, Fig. 14B The example architecture of FIG. 1404 is unable to receive memory mapped data (e.g., configuration data) for configuring the DPEs from the NoC 308. Such configuration data may be received from an adjacent tile via the memory mapped switch 1402 and directed to a subset of the DPEs managed by the tile 1304 (e.g., up into the Fig. 14B 1304).
[0260] Fig. 14C Another example implementation of block 1304 is shown. In certain embodiments, Fig. 14C The architecture shown in can be used to implement only one block within the SoC interface block 104. For example, Fig. 14C The architecture shown in can be used to implement block 1302 within SoC interface block 104. Fig. 14C The architecture shown in Fig. 14B The architecture shown in Fig. 14C In, additional components, such as: SoC security / initialization interface 1440, clock signal generator 1442 and global timer 1444, are included.
[0261] exist Fig. 14CIn the example of , the SoC security / initialization interface 1440 provides an additional interface for the SoC interface block 104. In one or more embodiments, the SoC security / initialization interface 1440 is implemented as a NoC peripheral interconnect. The SoC security / initialization interface 1440 can provide access to a global reset register for the DPE array 102 (not shown) and to the DPE GCS registers 1438. In a specific embodiment, the DPE GCS registers 1438 include configuration registers for the clock signal generator 1442. As shown, the SoC security / initialization interface 1440 can provide security bits to the DPE GCS registers 1438 and propagate the security bits to other DPE GCS registers 1438 in other tiles of the SoC interface block 104. In a specific embodiment, the SoC security / initialization interface 1440 implements a single slave endpoint for the SoC interface block 104.
[0262] exist Fig. 14C In the example of , the clock signal generator 1442 can generate one or more clock signals 1446 and / or one or more reset signals 1450. The clock signal 1446 and / or the reset signal 1450 can be distributed to each DPE in the DPE 304 and / or distributed to other blocks of the SoC interface block 104 of the DPE array 102. In one or more embodiments, the clock signal generator 1442 can include one or more phase-locked loop circuits (PLL). As shown, the clock signal generator 1442 can receive a reference clock signal, which is generated by another circuit external to the DPE array 102 and located on the SoC. The clock signal generator 1442 can generate a clock signal 1446 based on the received reference clock signal.
[0263] exist Fig. 14C In the example of FIG. 14 , the clock signal generator 1442 is configured through the SoC security / initialization interface 1440. For example, the clock signal generator 1442 may be configured by loading data into the DPE GCS register 1438. Thus, the generation of a clock frequency or multiple clock frequencies and a reset signal 1450 for the DPE array 102 may be set by writing appropriate configuration data to the DPE GCS register 1438 through the SoC security / initialization interface 1440. For testing purposes, the clock signal 1446 and / or the reset signal 1450 may also be routed directly to the PL 310.
[0264] The SoC security / initialization interface 1440 can be coupled to a SoC control / debug (circuitry) block (e.g., a control and / or debug subsystem of the device 100, not shown). In one or more embodiments, the SoC security / initialization interface 1440 can provide a status signal to the SoC control / debug block. As an illustrative and non-limiting example, the SoC security / initialization interface 1440 can provide a "PLL locked" signal generated from within the clock signal generator 1440 to the SoC control / debug block. The PLL lock signal can indicate when the PLL acquires lock on the reference clock signal.
[0265] The SoC security / initialization interface 1440 can receive instructions and / or data via the interface 1448. The data can include the security bits described herein, clock signal generator configuration data, and / or other data that can be written to the DPE GCS register 1438.
[0266] The global timer 1444 can interface with the CDT circuit 1420. For example, the global timer 1444 can be coupled to the CDT circuit 1420. The global timer 1444 can provide a signal used by the CDT circuit 1420 to track the timestamp events used. In one or more embodiments, the global timer 1444 can be coupled to the CDT circuit 1420 in other tiles of the tiles of the SoC interface circuit device 104. For example, the global timer 1444 can be coupled to Fig.14A , 14B and / or the CDT circuit 1420 in the example block of 14C. The global timer 1444 may also be coupled to the SoC control / debug block.
[0267] Common Reference Fig.14A , 14B 14C architecture, the tile 1304 can communicate with the DPE 304 using a variety of different data paths. In one example, the tile 1304 can communicate with the DPE 304 using the DMA engine 1412. For example, the tile 1304 can communicate with the DMA engine (e.g., DMA engine 816) of one or more DPEs of the DPE array 102 using the DMA engine 1412. Communications can flow from the DPE to the tile of the SoC interface block 104, or from the tile of the SoC interface block 104 to the DPE. In another example, the DMA engine 1412 can communicate with the core of one or more DPEs of the DPE array 102 with the help of a stream switch within the corresponding DPE. Communications can flow from the core to the tile of the SoC interface block 104 and / or from the tile of the SoC interface block 104 to the core of one or more DPEs of the DPE array 102.
[0268] Fig.15An example implementation of PL interface 1410 is shown. Fig.15 In the example of FIG. 1 , PL interface 1410 includes a plurality of channels that couple PL 310 to stream switch 1406 and / or stream multiplexer / demultiplexer 1408 depending on the particular tile architecture used. Fig.15 The specific number of channels shown in FIG. 1 is for purposes of illustration and not limitation. In other embodiments, PL interface 1410 may include more than Fig.15 14. In addition, although PL interface 1410 is shown as being connected to PL 310, in one or more other embodiments, PL interface 1410 can be coupled to one or more other subsystems and / or circuit blocks. For example, PL interface 1410 can also be coupled to an ASIC, analog / mixed signal circuitry, and / or other circuits or subsystems.
[0269] In one or more embodiments, PL 310 operates at a different reference voltage and a different clock speed than DPE 304. Fig.15 In the example of , PL interface 1410 includes a plurality of shift and isolation circuits 1502 and a plurality of asynchronous FIFO memories 1504. Each of the channels includes a shift isolation circuit 1502 and an asynchronous FIFO memory 1504. A first subset of channels transmits data from PL 310 (and / or other circuitry) to stream switch 1406 and / or stream multiplexer / demultiplexer 1408. A second subset of channels transmits data from stream switch 1406 and / or stream multiplexer / demultiplexer 1408 to PL 310 and / or other circuitry.
[0270] Shift and isolation circuit 1502 can interface between domains of different voltages. In this case, shift and isolation circuit 1502 can provide an interface that converts between the operating voltage of PL 310 and / or other circuit devices and the operating voltage of DPE 304. Asynchronous FIFO memory 1504 can interface between two different clock domains. In this case, asynchronous FIFO memory 1504 can provide an interface that converts between the clock rate of PL 310 and / or other circuit devices and the clock rate of DPE 304.
[0271] In one or more embodiments, the asynchronous FIFO memory 1504 has a 32-bit interface to the DPE array 102. The connection between the asynchronous FIFO memory 1504 and the shift and isolation circuit 1502 and the connection between the shift and isolation circuit 1502 and the PL 310 is programmable (e.g., configurable) in width. For example, the connection between the asynchronous FIFO memory 1504 and the shift and isolation circuit 1502 and the connection between the shift and isolation circuit 1502 and the PL 310 can be configured to be 32 bits, 64 bits, or 128 bits in width. As discussed, the PL interface 1410 can be configured by writing configuration data to the configuration register 1436 with the aid of the memory mapped switch 1402 to achieve the described bit width. Using the memory mapped switch 1402, one side of the asynchronous FIFO memory 1504 on one side of the PL 310 can be configured to use 32 bits, 64 bits, or 128 bits. The bit widths provided herein are for illustrative purposes. In other embodiments, other bit widths can be used. In any case, the widths used for the various component descriptions may be altered based on the configuration data loaded into configuration registers 1436 .
[0272] Fig.16 An example implementation of a NoC stream interface 1414 is shown. The DPE array 102 has two general ways of communicating via the NoC 308 using the stream interface in the DPE. In one aspect, the DPE can access a DMA engine 1412 using the stream switch 1406. The DMA engine 1412 can convert memory-mapped transactions from the NoC 308 into data streams to send to the DPE, and convert data streams from the DPE into memory-mapped transactions to send over the NoC 308. In another aspect, the data stream can be directed to the NoC stream interface 1414.
[0273] exist Fig.16 In the example of , the NoC flow interface 1414 includes a plurality of channels that couple the NoC 308 to the flow switch 1406 and / or the flow multiplexer / demultiplexer. Each channel may include a FIFO memory and an amplification circuit or a reduction circuit. A first subset of the channels transmits data from the NoC 308 to the flow switch 1406 and / or the flow multiplexer / demultiplexer 1408. A second subset of the channels transmits data from the flow switch 1406 and / or the flow multiplexer / demultiplexer 1408 to the NoC 308. The NoC flow interface 1414 includes a plurality of channels that couple the NoC 308 to the flow switch 1406 and / or the flow multiplexer / demultiplexer. Each channel may include a FIFO memory and an amplification circuit or a reduction circuit. Fig.16 The specific number of channels shown in is for purposes of illustration and not limitation. In other embodiments, the NoC flow interface 1414 may include more than Fig.16 Fewer or more channels are shown.
[0274] In one or more embodiments, the amplifier circuit 1608 (in Fig.16 Each amplifier circuit in the FIFO memory 1608 (hereinafter referred to as "US circuit") can receive a data stream and increase the width of the received data stream. For example, each amplifier circuit 1608 can receive a 32-bit data stream and output a 128-bit data stream to a corresponding FIFO memory 1610. Each FIFO memory in the FIFO memory 1610 is coupled to an arbitration and multiplexer circuit 1612. The arbitration and multiplexer circuit 1612 can arbitrate between the received data streams using a specific arbitration scheme or priority (e.g., round-robin or other manner) to provide the resulting output data stream to the NoC interface 1426. The arbitration and multiplexer circuit 1612 can process and accept new requests at each clock cycle. The clock domain crossing between the DPE 304 and the NoC 308 can be handled within the NoC 308 itself. In one or more other embodiments, the clock domain crossing between the DPE 304 and the NoC 308 can be handled within the SoC interface block 104. For example, the clock domain crossing can be handled in the NoC stream interface 1414.
[0275] The demultiplexer 1602 can receive a data stream from the NoC 308. For example, the demultiplexer 1602 can be coupled to the NoC interface 1426. For purposes of illustration, the data stream from the NoC interface 1426 can be 128 bits in width. As previously described, clock domain crossings between the DPE 304 and the NoC 308 can be handled within the NoC 308 and / or within the NoC stream interface 1414. The demultiplexer 1602 can forward the received data stream to one of the FIFO memories 1604. The particular FIFO memory 1604 to which the demultiplexer 1602 provides the data stream can be encoded within the data stream itself. The FIFO memory 1604 is coupled to the scale-down circuit 1606 (in Fig.16 After buffering using time division multiplexing, the reduction circuit 1606 can reduce the received stream to a smaller width. For example, the reduction circuit 1606 can reduce the stream from a 128-bit width to a 32-bit width.
[0276] As shown, the down-scaling circuit 1606 and the up-scaling circuit 1608 are coupled to the stream switch 1406 or the stream multiplexer / demultiplexer 1408, depending on the particular architecture of the SoC interface block 104 block being used. Fig.16 This is for purposes of illustration and not limitation. The order and / or connectivity of components in a channel (eg, up / down circuits and FIFO memories may vary).
[0277] In one or more other embodiments, such as in combination with Fig.15As described above, the PL interface 1410 may include a combination of Fig.16 The amplification circuit and / or reduction circuit described above. For example, a reduction circuit may be included in each channel that transmits data from the PL 310 (or other circuit) to the stream switch 1406 and / or to the stream multiplexer / demultiplexer 1408. Amplification circuit may be included in each channel that transmits data from the stream switch 1406 and / or the stream multiplexer / demultiplexer 1408 to the PL 310 (or other circuit).
[0278] In one or more other embodiments, although shown as separate elements, each scale-down circuit 1606 can be combined with the corresponding FIFO memory 1604, for example, as a single block or circuit. Similarly, each scale-up circuit 1608 can be combined with the corresponding FIFO memory 1610, for example, as a single block or circuit.
[0279] Fig.17 An example implementation of a DMA engine 1412 is shown. Fig.17 In the example of , DMA engine 1412 includes DMA controller 1702. DMA controller 1702 can be divided into two separate modules or interfaces. Each module can operate independently of the other modules. DMA controller 1702 can include a memory mapped to stream interface (interface) 1704 and a stream to memory mapped interface (interface) 1706. Each of interface 1704 and interface 1706 can include two or more separate channels. Therefore, DMA engine 1412 can receive two or more input streams from stream switch 1406 via interface 1706, and send two or more output streams to stream switch 1406 via interface 1704. DMA controller 1702 can also include main memory mapped interface 1714. Main memory mapped interface 1714 couples NoC 308 to interface 1704 and interface 1706.
[0280] DMA engine 1412 may also include hardware synchronization circuitry 1710 and buffer descriptor register file 1708. Hardware synchronization circuitry 1710 and buffer descriptor register file 1708 may be accessed via multiplexer 1712. Thus, both hardware synchronization circuitry 1710 and buffer descriptor register file 1708 may be externally accessed via a control interface. Examples of such control interfaces include, but are not limited to, a memory mapped interface or a control flow interface from a DPE. An example of a control flow interface of a DPE is a stream interface output from the core of the DPE.
[0281] Hardware synchronization circuitry 1710 may be used to synchronize multiple channels included in DMA engine 1412 and / or channels within DMA engine 1412 with a master interface that polls and drives a lock request. For example, the master interface may be PS 312 or a device implemented within PL 310. In another example, the master interface may also receive an interrupt generated by hardware synchronization circuitry 1710 within DMA engine 1412 when a lock is available.
[0282] DMA transfers may be defined by buffer descriptors stored in buffer descriptor register file 1708. Interface 1706 can request read transfers to NoC 308 based on the information in the buffer descriptors. Registers may configure the output flow from interface 1704 to stream switch 1406 as packet switched or circuit switched based on the configuration for the stream switch.
[0283] Fig.18 An example architecture for multiple DPEs is shown. The example architecture shows a DPE 304 that may be included in a DPE array 102. Fig.18 The example architecture can be referred to as a chessboard architecture. Fig.18 The example architecture allows a core of a DPE to communicate with up to eight other cores of other DPEs using shared memory (e.g., a total of nine cores communicating via shared memory). Fig.18 In the example of Figure 6 , 7 8. Thus, each core 602 is able to access four different memory modules 604. Each memory 604 can be accessed by up to four different cores 602.
[0284] As shown, DPE array 102 includes rows 1, 2, 3, 4, and 5. Each of rows 1-5 includes three DPEs 304. Fig.18 The specific number of DPEs 304 in each row and the number of rows shown in FIG. 3 are for purposes of illustration and not limitation. Referring to rows 1, 3, and 5, the core of each DPE in these rows is located on the left side of the memory module. Referring to rows 2 and 4, the core of each DPE in these rows is located on the right side of the memory module. In practice, the orientation of the DPEs in rows 2 and 4 is horizontally inverted or flipped compared to the orientation of the DPEs in rows 1, 3, and 5. As shown in each alternating row, the orientation of the DPEs is reversed.
[0285] exist Fig.18 In the example of FIG. 3 , DPEs 304 are aligned in columns. However, cores and memory modules in adjacent rows are not aligned in columns. Fig.18The architecture of is an example of a heterogeneous architecture in which the DPEs are implemented differently based on the specific rows in which they are located. Due to the horizontal inversion of DPE 304, the cores in adjacent rows are not aligned. The cores in adjacent rows are offset from each other. Similarly, the memory modules in adjacent rows are not aligned. The memory modules in adjacent rows are offset from each other. However, the cores in alternating rows are aligned with the memory modules in alternating rows. For example, the cores and memory modules (e.g., in columns) of rows 1, 3, and 5 are vertically aligned. Similarly, the cores and memory modules (e.g., in columns) of rows 2 and 4 are vertically aligned.
[0286] For purposes of illustration, the cores of DPEs 304-2, 304-4, 304-5, 304-7, 304-8, 304-9, 304-10, 304-11, and 304-14 are considered part of a group and are able to communicate via shared memory. Fig.18 4-10. Example architecture of how to use shared memory to support cores communicating with up to eight other cores in different DPEs. For example, with reference to DPE 304-8, core 602-8 is able to access memory modules 604-11, 604-7, 604-8, and 604-5. Through memory module 604-11, core 602-8 is able to communicate with cores 602-14, 602-10, and 602-11. Through memory module 604-7, core 602-8 is able to communicate with cores 602-7, 602-4, and 602-10. Through memory module 604-8, core 602-8 is able to communicate with cores 602-9, 602-11, and 602-5. Through memory module 604-5, core 602-8 is able to communicate with cores 602-4, 602-5, and 602-2.
[0287] exist Fig.18 In the example of , in addition to core 602-8, there are four different cores in the group that can access two different memory modules of the shared memory module of the group. The remaining four cores only share one memory module of the shared memory module of the group. The shared memory module of the group includes memory modules 604-5, 604-7, 604-8, and 604-11. For example, each core in cores 602-10, 602-11, 602-4, and 602-5 can access two different memory modules. Core 602-10 can access memory modules 604-11 and 604-7. Core 602-11 can access memory modules 604-11 and 604-8. Core 602-4 can access memory modules 604-5 and 604-7. Core 602-5 can access memory modules 604-5 and 604-8.
[0288] exist Fig.18In the example of , up to nine cores of a total of nine DPEs can communicate through shared memory without utilizing the DPE interconnect network of DPE array 102. As discussed, core 602-8 can view memory modules 604-11, 604-7, 604-5, and 604-8 as a unified memory space.
[0289] Cores 602-14, 602-7, 602-9, and 602-2 can access only one memory module of the group's shared memory. Core 602-14 can access memory module 604-11. Core 602-7 can access memory module 604-7. Core 602-9 can access memory module 604-8. Core 602-2 can access memory module 604-5.
[0290] As previously mentioned, in other embodiments, where more than four memory interfaces are provided for each memory module, the core may use Fig.18 The architecture communicates with more than eight other cores via shared memory.
[0291] In one or more other embodiments, certain rows and / or columns of the DPE may be offset relative to other rows. For example, rows 2 and 4 may start at positions that are not aligned with the start of rows 1, 3, and / or 5. For example, rows 2 and 4 may be shifted to the right relative to the start of rows 1, 3, and / or 5.
[0292] Fig.19 Another example architecture for multiple DPEs is shown. The example architecture shows a DPE 304 that may be included in a DPE array 102. Fig.19 The example architecture shown in can be referred to as a grid architecture. Fig.19 The example architecture allows a core of a DPE to communicate with up to ten other cores of other DPEs using shared memory (e.g., a total of 11 cores communicating via shared memory). Fig.19 In the example of Figure 6 , 7 8. Thus, each core 602 is able to access four different memory modules 604. Each memory module 604 can be accessed by up to four different cores 602.
[0293] As shown, DPE array 102 includes rows 1, 2, 3, 4, and 5. Each of rows 1-5 includes three DPEs 304. Fig.19 The specific number of DPEs 304 in each row and the number of rows shown in FIG. 3 are for purposes of illustration and not limitation. Fig.19, the DPEs 304 are aligned vertically in columns. Each of rows 1, 2, 3, 4, and 5 has the same starting point aligned with each other row of DPEs. In addition, the arrangement of the cores 602 and memory modules 604 within each respective DPE 304 is the same. In other words, the cores 602 are aligned vertically. Similarly, the memory modules 604 are aligned vertically.
[0294] For purposes of illustration, the cores of DPEs 304-2, 304-4, 304-5, 304-6, 304-7, 304-8, 304-9, 304-10, 304-11, 304-12, and 304-14 are considered part of a group and are able to communicate via shared memory. Fig.19 4-10 and 602-2. Through memory module 604-9, core 602-8 can communicate with cores 602-14, 602-10 and 602-11. Through memory module 604-8, core 602-8 can communicate with cores 602-7, 602-11 and 602-5. Through memory module 604-5, core 602-8 can communicate with cores 602-4, 602-10 and 602-2. Through memory module 604-9, core 602-8 can communicate with cores 602-12, 602-9 and 602-6.
[0295] exist Fig.19 In the example of , in addition to core 602-8, two different cores in the group cores can access two memory modules of the shared memory module of the group. The shared memory modules of the group include memory modules 604-5, 604-8, 604-9 and 604-11. The remaining eight cores of the group share only one memory module. For example, each of cores 602-11 and 602-5 can access two different memory modules. Core 602-11 can access memory modules 604-11 and 604-8. Memory module 604-14 is not considered to be part of the shared memory group because memory module 604-14 cannot be accessed by core 604-8. Core 602-5 can access memory modules 604-5 and 604-8. Memory module 604-2 is not considered to be part of the shared memory module group because memory module 604-2 cannot be accessed by core 602-8.
[0296] Cores 602-14, 602-10, 602-12, 602-7, 602-9, 602-4, 602-6, and 602-2 can access only one memory module of the group's shared memory module. Core 602-14 can access memory module 604-11. Core 602-10 can access memory module 604-11. Core 602-12 can access memory module 604-9. Core 602-7 can access memory module 604-8. Core 602-9 can access memory module 604-9. Core 602-4 can access memory module 604-5. Core 602-6 can access memory module 604-9. Core 602-2 can access memory module 604-5.
[0297] exist Fig.19 In the example of , up to 11 cores of 11 DPEs can communicate through shared memory without utilizing the DPE interconnect network of DPE array 102. As discussed, core 602-8 can view memory modules 604-11, 604-9, 604-5, and 604-8 as a unified memory space.
[0298] As previously mentioned, in other embodiments, where more than four memory interfaces are provided for each memory module, the core may use Fig.19 The architecture communicates with more than 10 other cores via shared memory.
[0299] Fig. 20 An example method 2000 of configuring a DPE array is shown. The method 2000 is provided for purposes of illustration and is not intended to limit the inventive arrangements described within the present disclosure.
[0300] Configuration data for the DPE array is loaded into the device in block 2002. The configuration data may be provided from any of a variety of different sources, whether a computer system (eg, a host), off-chip memory, or other suitable source.
[0301] In block 2004, configuration data is provided to the SoC interface block. In a particular embodiment, the configuration data is provided via the NoC. The tile of the SoC interface block can receive the configuration data and convert the configuration data into memory-mapped data, which can be provided to a memory-mapped switch contained within the tile.
[0302] In block 2006, configuration data is propagated between tiles of the SoC interface block to a particular tile that serves as an interface to or provides an interface to a target DPE. The target DPE is the DPE to which the configuration data is addressed. For example, the configuration data includes an address that specifies a particular DPE to which different portions of the configuration data should be directed. Memory-mapped switches within the tiles of the SoC interface block are capable of propagating different portions of the configuration data to a particular tile that serves as an interface to a target DPE (e.g., a subset of the DPEs that include the target DPE).
[0303] In block 2008, a tile of a SoC interface block that serves as an interface to a target DPE is capable of directing portions of configuration data for a target DPE to the target DPE. For example, a tile that provides an interface to one or more target DPEs is capable of directing portions of configuration data to a subset of DPEs to which the tile provides an interface. As described above, the subset of DPEs includes one or more target DPEs. As each tile receives configuration data, the tile is capable of determining whether to address any portion of the configuration data to other DPEs in the same subset of DPEs to which the tile provides an interface. The tile directs any configuration data addressed to a DPE in the subset of DPEs to such DPE.
[0304] In block 2010, configuration data is loaded into the target DPE to program elements included therein of the DPE. For example, the configuration data is loaded into configuration registers to program elements of the target DPE, such as: stream interfaces, cores (e.g., stream interfaces, cascade interfaces, core interfaces), memory modules (e.g., DMA engines, memory interfaces, arbiters, etc.), broadcast event switches, and / or broadcast logic. The configuration data may also include executable program code, which may be loaded into a program memory of a core and / or loaded into a memory bank of a memory module.
[0305] It should be understood that the received configuration data may also include portions of one or more or all of the tiles addressed to the SoC interface block 104. In that case, a memory-mapped switch within the corresponding tile can transfer the configuration data to the appropriate (e.g., target) tile, extract such data, and write such data to the appropriate configuration registers within the corresponding tile.
[0306] Fig.21 An example method 2100 of the operation of a DPE array is shown. The method 2100 is provided for illustrative purposes and is not intended to limit the inventive arrangements described within the present disclosure. The method 2100 begins in a state where the DPE and / or SoC interface block has been loaded with configuration data. For illustrative purposes, reference is made to Figure 3 .
[0307] In block 2102, a core 602-15 (e.g., a "first core") of a DPE 304-15 (e.g., a "first DPE") generates data. The generated data may be application data. For example, the core 602-15 may operate on data stored in a memory module accessible by the core. The memory module may be in the DPE 304-15 or in a different DPE as described herein. For example, the data may have been received from another DPE and / or another subsystem of the device using the SoC interface block 104.
[0308] In block 2104, core 602-15 stores data in memory module 604-15 of DPE 304-15. In block 2106, one or more cores in adjacent DPEs (e.g., DPE 304-25, 304-16, and / or 304-5) read data from memory module 604-15 of DPE 304-15. Cores in adjacent DPEs may utilize the data read from memory module 604-15 in further computations.
[0309] In block 2108, DPE 304-15 optionally sends data to one or more other DPEs via a stream interface. The DPE to which the data is sent may be a non-adjacent DPE. For example, DPE 304-15 may be able to send data from memory module 604-15 to one or more other DPEs, such as: DPE 304-35, 304-36, etc. As discussed, in one or more embodiments, DPE 304-15 may be able to broadcast and / or multicast application data via a stream interface in a DPE interconnect network of DPE array 102. In another example, the data sent to different DPEs may be different portions of the data, where each different portion of the data is intended for a different target DPE. Although in Fig.21 Not shown, but core 602-15 can also send data directly from the core to another core and / or DPE of DPE array 102 using a cascade interface and / or using a stream switch.
[0310] In block 2110, core 602-15 optionally sends data to and / or receives data from an adjacent core via a cascade interface. The data may be application data. For example, core 602-15 may be able to receive data directly from core 602-14 of DPE 304-14 and / or send data directly to core 602-16 of DPE 304-16 via a cascade interface.
[0311] In block 2112, DPE 304-15 optionally sends data to and / or receives data from one or more subsystems via the SoC interface block. The data may be application data. For example, DPE 304-15 can send data to PS 312 via NoC 308, to circuits implemented in PL 310, to selected hardwired circuit blocks via PL 310, and / or to other external subsystems (such as external memory). Similarly, DPE 304-15 can receive application data from such other subsystems via the SoC interface block.
[0312] Fig. 22 Another example method 2200 of operation of a DPE array is shown. The method 2200 is provided for illustration purposes and is not intended to limit the inventive arrangements described within the present invention. The method 2200 begins in a state where the DPE array has been loaded with configuration data.
[0313] In block 2202, a first core (e.g., a core within a first DPE) requests a lock for a target region of memory from a hardware synchronization circuit device. For example, the first core can request a lock from the hardware synchronization circuit device for a target region of memory within a memory module located in a first DPE (e.g., the same DPE as the first core), or request a lock for a target region of memory within a memory module located in a different DPE than the first core. The first core can request a lock from a specific hardware synchronization circuit device located in the same DPE as the target region of memory to be accessed.
[0314] In block 2204, the first core acquires the requested lock. For example, the hardware synchronization circuitry grants the first core the requested lock for the target region of memory.
[0315] In block 2206, in response to the acquired lock, the first core writes data to the target area in the memory. For example, if the target area of the memory is in the first DPE, the first core can write the data to the target area in the memory via a memory interface within a memory module located within the first DPE. In another example, where the target area of the memory is located in a different DPE than the first core, the first core can write the data to the target area of the memory using various techniques described herein. For example, the first core can write the data to the target area of the memory via any of the mechanisms described in conjunction with FIG. 10 .
[0316] In block 2208, the first core releases a lock on a target region of memory. In block 2210, a second core requests a lock on a target region of memory containing data written by the first core. The second core may be located in the same DPE as the target region of memory, or in a different DPE than the target region of memory. The second core requests a lock from the same hardware synchronization circuit device that granted the lock to the first core. In block 2212, the second core acquires the lock from the hardware synchronization circuit device. The hardware synchronization circuit device grants the lock to the second core. In block 2214, the second core is able to access the data from the target region of memory and utilize the data for processing. In block 2216, the second core releases the lock on the target region of memory, for example, when access to the target region of memory is no longer required.
[0317] Describes the area of memory access Fig. 22 In certain embodiments, the first core can write data directly to a target area of memory. In other embodiments, the first core can move data from a source area of memory (e.g., in a first DPE) to a target area of memory (e.g., located in a second or different DPE). In that case, the first core acquires locks on the source area of memory and the target area of memory in order to perform the data transfer.
[0318] In other embodiments, the first core can acquire a lock for the second core in order to quiesce the operation of the second core, and then release the lock to allow the operation of the second core to continue. For example, in addition to acquiring a lock on a target area of memory, the first core can also acquire a lock on the second core in order to quiesce the operation of the second core while data is written to the target area of memory used by the second core. Once the first core has completed writing the data, the first core can release the lock on the target area of memory and the lock on the second core, thereby allowing the second core to operate on the data once the second core has acquired the lock on the target area of memory.
[0319] In yet another embodiment, Fig. 10C As shown, a first core can initiate a transfer of data from a memory module in the same DPE directly to another core, for example, via a DMA engine in the memory module.
[0320] Fig.23 Another example method of operation of a DPE array is shown. The method 2300 is provided for illustrative purposes and is not intended to limit the inventive arrangements described within the present disclosure. The method 2300 begins in a state where the DPE array has been loaded with configuration data.
[0321] In block 2302, the first core places data into the accumulation register contained therein. For example, the first core may be performing a computation in which some portion of the computation (whether intermediate or final results) is provided directly to another core. In that case, the first core can load the data to be sent to the second core into the accumulation register contained therein.
[0322] In block 2304, the first core sends data from the accumulation register contained therein from the cascade interface output of the first core to the second core. In block 2306, the second core receives the data from the first core on the cascade interface input of the second core. The second core may then process the data or store the data in memory.
[0323] In one or more embodiments, the utilization of the cascade interface by the core can be controlled by the loading of configuration data. For example, the cascade interface can be enabled or disabled between consecutive pairs of cores as needed for a specific application based on the configuration data. In a specific embodiment, in the case where the cascade interface is enabled, the use of the cascade interface can be controlled based on program code loaded into the program memory of the core. In other cases, the use of the cascade interface can be controlled by dedicated circuit devices and configuration registers contained in the core.
[0324] Fig.24 Another example method of operation of a DPE array is shown. Method 2400 is provided for illustration purposes and is not intended to limit the inventive arrangements described within the present disclosure. Method 2400 begins in a state where the DPE array has been loaded with configuration data.
[0325] In block 2402, event logic within a first DPE locally detects one or more events within the first DPE. The events may be detected from a core, from a memory module, or from both a core and a memory module. In block 2404, event broadcasting circuitry within the first DPE broadcasts the events based on configuration data loaded into the first DPE. The broadcasting circuitry is capable of broadcasting selected ones of the events generated in block 2402. The event broadcasting circuitry is also capable of broadcasting selected events that may be received from one or more other DPEs within the DPE array 102.
[0326] In block 2406, events from the DPE are propagated to tiles within the SoC interface block. For example, events may be propagated through the DPE in each of four cardinal directions in a pattern and / or route determined by the configuration data. Broadcast circuitry within a particular DPE may be configured to propagate events down to tiles in the SoC interface block.
[0327] In block 2408, event logic within the tile of the SoC interface block optionally generates an event. In block 2410, the tile of the SoC interface block optionally broadcasts the event to other tiles within the SoC interface block. Broadcast circuitry within the tile of the SoC interface block can broadcast selected events from events generated by the tile itself and / or events received from other sources (e.g., whether other tiles of the SoC interface block or the DPE) to other tiles of the SoC interface block.
[0328] In block 2412, the block of the SoC interface block optionally generates one or more interrupts. For example, the interrupt may be generated by the interrupt processor 1434. The interrupt processor can generate one or more interrupts in response to receiving a particular event, combination of events, and / or sequence of events that vary over time. The interrupt processor may send the generated interrupt to other circuits (such as, PS 312) and / or to circuits implemented within PL 310.
[0329] The tile of the SoC interface block optionally sends the events to one or more other circuits in block 2414. For example, CDT circuit 1420 can group the events and send the events from the tile of the SoC interface block to PS 312, to circuits within PL 310, to external memory, or to another destination with the SoC.
[0330] In one or more embodiments, the PS 312 can respond to interrupts generated by tiles of the SoC interface block 104. For example, the PS 312 can reset the DPE array 102 in response to receiving a particular interrupt. In another example, the PS 312 can reconfigure the DPE array 102 or a portion of the DPE array 102 (e.g., perform a partial reconfiguration) in response to a particular interrupt. In another example, the PS 312 can take other actions, such as loading new data into different memory modules of the DPE for use by cores within such DPE.
[0331] exist Fig.24In the example of , PS 312 performs operations in response to an interrupt. In other embodiments, PS 312 can be used as a global controller for DPE array 102. PS 312 can control application parameters, which are stored in memory modules and used by one or more DPEs (e.g., cores) of DPE array 102 during operation. As an illustrative and non-limiting example, one or more DPEs can be used as a kernel to implement a filter. In that case, PS 312 can execute program code that allows PS 312 to calculate and / or modify coefficients of the filter during operation of DPE array 102 (e.g., dynamically at runtime). PS 312 can calculate and / or update coefficients in response to specific conditions and / or signals detected within the SoC. For example, PS 312 can calculate new coefficients for the filter in response to some detected conditions and / or write such coefficients to application memory (e.g., to one or more memory modules). Examples of conditions that may cause PS 312 to write data (such as coefficients) to a memory module include, but are not limited to: receiving specific data from DPE array 102, receiving an interrupt from the SoC interface block, receiving event data from DPE array 102, receiving a signal from a source external to the SoC, receiving another signal from within the SoC, and / or receiving new and / or updated coefficients from a source internal to the SoC or from a source external to the SoC. PS 312 is capable of calculating new coefficients and / or writing new coefficients to application data, such as a memory module utilized by one or more cores.
[0332] In another embodiment, PS 312 can execute a debugger application that can perform actions such as starting, stopping, and / or stepping of the DPE. PS 312 can control the starting, stopping, and / or stepping of the DPE via NoC 308. In other examples, circuitry implemented in PL 310 can also control the operation of the DPE using debug operations.
[0333] For purposes of explanation, specific terminology is set forth to provide a thorough understanding of the various inventive concepts disclosed herein. However, the terminology used herein is for the purpose of describing particular aspects of the present arrangements only and is not intended to be limiting.
[0334] As defined herein, the singular forms "a", "an" and "the" are intended to include the plural forms as well, unless the context clearly indicates otherwise.
[0335] As defined herein, unless the context clearly indicates otherwise, the terms "at least one", "one or more", and "and / or" are open-ended expressions that are both conjunctive and disjunctive in operation. For example, the expressions "at least one of A, B, and C", "at least one of A, B, or C", "one or more of A, B, and C", "one or more of A, B, or C", and "A, B, and / or C" mean A alone, B alone, C alone, A and B together, A and C together, B and C together, or A, B, and C together.
[0336] As defined herein, the term "automatically" means without human intervention.
[0337] As defined herein, the term "if" means "when" or "upon" or "in response to" or "in response to," depending on the context. Thus, the phrase "if it is determined" or "if [a stated condition or event] is detected" may be interpreted to mean "upon the determination" or "in response to the determination" or "upon detecting [a stated condition or event]" or "in response to detecting [a stated condition or event]" or "in response to detecting [a stated condition or event]," depending on the context.
[0338] As defined herein, the term "in response to" and similar language as described above, such as "if," "when," or "upon," refers to the tendency to respond or react to an action or event. A response or reaction is automatically performed. Thus, if a second action is performed "in response to" a first action, there is a causal relationship between the occurrence of the first action and the occurrence of the second action. The term "in response to" indicates a causal relationship.
[0339] As defined herein, the terms "one embodiment," "an embodiment," "one or more embodiments," "a specific embodiment," or similar language refer to a particular feature, structure, or characteristic in conjunction with an embodiment that is included in at least one embodiment described within the present disclosure. Thus, appearances of the phrases "in one embodiment," "in an embodiment," "in one or more embodiments," "in a specific embodiment," and similar language throughout the present disclosure may, but do not necessarily, all refer to the same embodiment. Within the present disclosure, the terms "an embodiment" and "an arrangement" may be used interchangeably.
[0340] As defined herein, the term "substantially" means that the recited characteristics, parameters or values need not be achieved precisely, but deviations or changes (including, for example, tolerances, measurement errors, measurement precision limitations and other factors known to those skilled in the art) may occur by amounts that do not eliminate the effect that the characteristic is intended to provide.
[0341] The terms first, second, etc. may be used herein to describe various elements. These elements should not be limited by these terms, because unless otherwise stated or the context clearly states otherwise, these terms are only used to distinguish one element from another.
[0342] Flowchart and block diagram in the accompanying drawings show the possible implementation of the system, equipment and / or method of each aspect of arrangement according to the present invention, architecture, function and operation.In some alternative implementations, the operation pointed out in the frame may not occur in the order pointed out in the figure.For example, depending on the function involved, two frames shown in succession can be performed substantially simultaneously or can sometimes be performed in reverse order.In other examples, frame can be performed usually in the numerical order of increase, and in other examples, one or more frames can be performed in the order of variation, and in other frames subsequently or not immediately followed, the result is stored and utilized.
[0343] The corresponding structures, materials, acts, and equivalents of all means or step plus function elements found in the claims below are intended to include any structure, material, or act for performing the function in combination with other claimed elements as specifically claimed.
[0344] In one or more embodiments, the device may include multiple DPEs. Each DPE may include a core and a memory module. Each core may be configured to access the memory module in the same DPE and the memory module in at least one other DPE in the multiple DPEs.
[0345] In one aspect, each core may be configured to access memory modules of multiple adjacent DPEs.
[0346] In another aspect, cores of multiple DPEs may be directly coupled.
[0347] In another aspect, each DPE of the plurality of DPEs is a hardwired and programmable block of circuitry.
[0348] In another aspect, each DPE may include interconnect circuitry including a stream switch configured to communicate with one or more DPEs selected from the plurality of DPEs. The stream switch may be programmable to communicate with one or more selected DPEs (eg, other DPEs).
[0349] The device may also include a subsystem and a SoC interface block configured to couple the plurality of DPEs to the subsystem of the device. In one aspect, the subsystem includes programmable logic. In another aspect, the subsystem includes a processor configured to execute program code. In yet another aspect, the subsystem includes an application specific integrated circuit and / or analog / mixed signal circuit device.
[0350] In another aspect, a stream switch is coupled to the SoC interface block and is configured to communicate with a subsystem of the device.
[0351] In another aspect, the interconnect circuitry of each DPE may include a memory-mapped switch coupled to the SoC interface block, wherein the memory-mapped switch is configured to communicate configuration data for programming the DPE from the SoC interface block. The memory-mapped switch may be configured to communicate at least one of control data or debug data with the SoC interface block.
[0352] In another aspect, multiple DPEs may be interconnected by an event broadcast network.
[0353] In another aspect, the SoC interface block can be configured to exchange events between the subsystem and an event broadcast network of the plurality of DPEs.
[0354] In one or more embodiments, a method may include a first core of a first data processing engine generating data, the first core writing the data to a first memory module within the first data processing engine, and a second core of a second data processing engine reading the data from the first memory module.
[0355] In one aspect, a method may include a first DPE and a second DPE, the first DPE and the second DPE being adjacent DPEs.
[0356] In another aspect, a method may include a first core capable of providing further application data to a second core via a cascade interface.
[0357] In another aspect, a method may include the first core being capable of providing further application data to a third DPE via a stream switch.
[0358] In another aspect, a method may include programming a first DPE to communicate with selected other DPEs including a second DPE.
[0359] In one or more embodiments, the device may include multiple data processing engines, subsystems, and a SoC interface block coupled to the multiple data processing engines and the subsystems. The SoC interface block may be configured to exchange data between the subsystems and the multiple data processing engines.
[0360] In one aspect, the subsystem includes programmable logic. In another aspect, the subsystem includes a processor configured to execute program code. In another aspect, the subsystem includes an application specific integrated circuit and / or analog / mixed signal circuitry.
[0361] In another aspect, the SoC interface block includes a plurality of tiles, wherein each tile is configured to communicate with a subset of the plurality of DPEs.
[0362] In another aspect, each tile may include a memory mapped switch configured to provide a first portion of the configuration data to at least one adjacent tile and to provide a second portion of the configuration data to at least one subset of the plurality of DPEs.
[0363] In another aspect, each tile may include a stream switch configured to provide the first data to at least one adjacent tile and configured to provide the second data to at least one DPE of the plurality of DPEs.
[0364] In another aspect, each tile may include an event broadcasting circuit device configured to receive events generated within the tile and events from circuit devices external to the tile, wherein the event broadcasting circuit device is programmable to provide selected ones of the events to a selected destination.
[0365] In another aspect, the SoC interface block may include control, debug, and trace circuitry configured to group selected events and provide the grouped selected events to the subsystem.
[0366] In another aspect, the SoC interface block may include an interface to couple the event broadcasting circuitry to the subsystem.
[0367] In one or more embodiments, a tile for a SoC interface block may include a memory-mapped switch configured to provide a first portion of configuration data to an adjacent tile and configured to provide a second portion of configuration data to a data processing engine in a plurality of data processing engines. The tile may include a stream switch configured to provide first data to at least one adjacent tile and configured to provide second data to a data processing engine in a plurality of data processing engines. The tile may include an event broadcast circuit device configured to receive events generated within the tile and events from circuits external to the tile, wherein the event broadcast circuit device is programmable to provide selected events in the event to a selected destination. The tile may include an interface circuit device that couples the memory-mapped switch, the stream switch, and the event broadcast circuit device to a subsystem of a device including the tile.
[0368] In one aspect, the subsystem includes programmable logic. In another aspect, the subsystem includes a processor configured to execute program code. In another aspect, the subsystem includes an application specific integrated circuit and / or analog / mixed signal circuitry.
[0369] In another aspect, the event broadcasting circuitry is programmable to provide events generated within the tile or received from at least one of the plurality of DPEs to the subsystem.
[0370] In another aspect, the event broadcasting circuitry is programmable to provide an event generated within the subsystem to at least one adjacent tile or to at least one DPE of the plurality of DPEs.
[0371] In another aspect, the tile may include an interrupt handler configured to selectively generate an interrupt to a processor of the device based on an event received from the event broadcasting circuitry.
[0372] In another aspect, a tile may include a clock generation circuit configured to generate a clock signal that is distributed to a plurality of DPEs.
[0373] In another aspect, the interface circuit device may include a stream multiplexer / demultiplexer, a programmable logic interface, a direct memory access engine, and a NoC stream interface. The stream multiplexer / demultiplexer may couple the stream switch to the programmable logic interface, the direct memory access engine, and the on-chip network stream interface. The stream multiplexer / demultiplexer may be programmable to route data between the stream switch, the programmable logic interface, the direct memory access engine, and the NoC stream interface.
[0374] In another aspect, the tile may include a switch coupled to the DMA engine and the NoC stream interface, wherein the switch selectively couples the DMA engine or the NoC stream interface to the NoC. The tile may also include a bridge circuit coupling the NoC with the memory-mapped switch. The bridge circuit is configured to convert data from the NoC into a format usable by the memory-mapped switch.
[0375] In one or more embodiments, the device may include multiple data processing engines. Each of the data processing engines may include a core and a memory module. Multiple data processing engines may be organized in multiple rows. Each core may be configured to communicate with adjacent data processing engines through shared access to the memory modules of other adjacent data processing engines in the multiple data processing engines.
[0376] In one aspect, the memory module of each DPE includes a memory and a plurality of memory interfaces to the memory. A first memory interface of the plurality of memory interfaces may be coupled to a core within the same DPE, and each other memory interface of the plurality of memory interfaces may be coupled to a core of a different DPE in the plurality of DPEs.
[0377] In another aspect, the plurality of DPEs may also be organized in a plurality of columns, wherein the cores of the plurality of DPEs in a column are aligned and the memory modules of the plurality of DPEs in a column are aligned.
[0378] In another aspect, the memory module of the selected DPE may include: a first memory interface coupled to a core of a DPE immediately above the selected DPE; a second memory interface coupled to a core within the selected DPE; a third memory interface coupled to a core of a DPE immediately adjacent to the selected DPE; and a fourth memory interface coupled to a core of a DPE immediately below the selected DPE.
[0379] In another aspect, the selected DPE is configured to communicate with a group of at least ten DPEs of the plurality of DPEs via shared access to the memory module.
[0380] In another aspect, at least two DPEs in the group are configured to access more than one memory module of the group of at least ten DPEs in the plurality of DPEs.
[0381] In another aspect, the plurality of rows of DPEs may include a first row including a first subset of the plurality of DPEs and a second row including a second subset of the DPEs, wherein an orientation of each DPE in the second row is horizontally reversed relative to an orientation of each DPE in the first row.
[0382] In another aspect, the memory module of the selected DPE may include: a first memory interface coupled to a core of a DPE immediately above the selected DPE; a second memory interface coupled to a core within the selected DPE; a third memory interface coupled to a core of a DPE immediately adjacent to the selected DPE; and a fourth memory interface coupled to a core of a DPE immediately below the selected DPE.
[0383] In another aspect, the selected DPE may be configured to communicate with a group of at least eight DPEs of the plurality of DPEs via shared access to the memory module.
[0384] In another aspect, wherein the group of at least four DPEs are configured to access more than one memory module of the group of at least eight DPEs of the plurality of DPEs.
[0385] In one or more embodiments, the device may include a plurality of data processing engines. Each of the data processing engines may include: a memory pool having a plurality of memory banks; a plurality of cores, each coupled to the memory pool and configured to access the plurality of memory banks; a memory-mapped switch coupled to the memory pool and the memory-mapped switch of at least one adjacent data processing engine; and a stream switch coupled to each of the plurality of cores and coupled to the stream switch of at least one adjacent data processing engine.
[0386] In one aspect, the memory pool may include a crossbar switch coupled to each of the plurality of memory banks and an interface coupled to each of the plurality of cores and to the crossbar switch.
[0387] In another aspect, each DPE may include a direct memory access engine coupled to the memory pool and the stream switch, wherein the direct memory access engine is configured to provide data from the memory pool to the stream switch and write data from the stream switch to the memory pool.
[0388] In another aspect, the memory pool may include additional interfaces coupled to the crossbar switch and the direct memory access engine.
[0389] In another aspect, each core of the plurality of cores has shared access to the plurality of memory banks.
[0390] In another aspect, within each DPE, a memory mapped switch may be configured to receive configuration data for programming the DPE.
[0391] In another aspect, the stream switch is programmable to establish connections with different ones of the plurality of DPEs based on the configuration data.
[0392] In another aspect, multiple cores within each tile may be directly coupled.
[0393] In another aspect, within each DPE, a first core of the plurality of cores may be directly coupled to a core in a first adjacent DPE, and a last core of the plurality of cores is directly coupled to a core in a second adjacent DPE.
[0394] In another aspect, each core of the plurality of cores may be programmable to be disabled.
[0395] The description of the inventive arrangement provided herein is for illustrative purposes and is not intended to be exhaustive or limited to the disclosed forms and examples. The terms used herein are selected to explain the principles of the inventive arrangement, the practical application of the technology found on the market or the improvement of the technology, and / or to enable other persons of ordinary skill in the art to understand the inventive arrangement disclosed herein. Modifications and variations will be apparent to persons of ordinary skill in the art without departing from the scope and spirit of the described inventive arrangement. Therefore, reference should be made to the appended claims, rather than to the preceding disclosure, to indicate the scope of such features and implementations.
Claims
1. A device, include: Multiple data processing engines; Each data processing engine includes: A memory pool having a plurality of memory banks; a plurality of cores, each core being coupled to the memory pool and configured to access the plurality of memory banks; a memory mapped switch coupled to the memory pool and to a memory mapped switch of at least one adjacent data processing engine; and a stream switch coupled to each of the plurality of cores and to the at least one adjacent data processing engine, The plurality of cores communicate with each other via a shared memory bank of the memory pool.
2. The device of claim 1, wherein the memory pool include: a crossbar switch coupled to each memory bank of the plurality of memory banks; as well as An interface is coupled to each of the plurality of cores and to the crossbar switch.
3. The apparatus according to claim 2, wherein each data processing engine further include: A direct memory access engine is coupled to the memory pool and to the stream switch, wherein the direct memory access engine is configured to provide data from the memory pool to the stream switch and to write data from the stream switch to the memory pool.
4. The apparatus of claim 3, wherein the memory pool further comprises: include: An additional interface is coupled to the crossbar switch and to the direct memory access engine. 5 . The apparatus of claim 1 , wherein each core of the plurality of cores has shared access to the plurality of memory banks.
6. The device according to claim 1, in, Within each data processing engine, the memory mapped switch is configured to receive configuration data for programming the data processing engine. 7 . The apparatus of claim 6 , wherein the stream switch is programmable to establish connections with different ones of the plurality of data processing engines based on the configuration data. The apparatus of claim 1 , wherein the plurality of cores within each tile are directly coupled.
9. The device according to claim 8, in, Within each data processing engine, a first core of the plurality of cores is directly coupled to a core of a first adjacent data processing engine, and a last core of the plurality of cores is directly coupled to a core of a second adjacent data processing engine.
10. The apparatus of claim 1, wherein each core of the plurality of cores is programmable to be disabled.
11. A device, include: Multiple data processing engines; Each data processing engine includes a core and a memory module; wherein the plurality of data processing engines are organized in a plurality of rows; wherein each core is configured to communicate with an adjacent data processing engine in the plurality of data processing engines by sharing access to the memory modules in other adjacent data processing engines, and The cores in the plurality of data processing engines lack input interrupts to provide deterministic and uninterrupted operation.
12. The apparatus of claim 11 , wherein the memory module in each data processing engine comprises a random access memory and a plurality of memory interfaces to the random access memory, wherein a first memory interface of the plurality of memory interfaces is directly connected to the core within the same data processing engine, and each other memory interface of the plurality of memory interfaces is directly connected to the core of a different data processing engine of the plurality of data processing engines.
13. The apparatus of claim 11 , wherein the plurality of data processing engines are further organized in a plurality of columns, wherein the cores of the plurality of data processing engines in the columns are aligned, and the memory modules of the plurality of data processing engines in the columns are aligned, and wherein the memory modules of selected ones of the plurality of data processing engines are aligned. include: a first memory interface coupled to a core of a data processing engine of the plurality of data processing engines immediately north of the selected data processing engine; a second memory interface coupled to a core within the selected data processing engine; a third memory interface coupled to a core of a data processing engine of the plurality of data processing engines that is adjacent to the east or west of the selected data processing engine; as well as A fourth memory interface is coupled to a core of a data processing engine immediately south of the selected data processing engine among the plurality of data processing engines.
14. The apparatus of claim 11, wherein the plurality of data processing engines are further organized in a plurality of columns, wherein the cores of the plurality of data processing engines in the columns are aligned, and the memory modules of the plurality of data processing engines in the columns are aligned, and Wherein a selected data processing engine of the plurality of data processing engines is configured to communicate with a group of at least ten data processing engines of the plurality of data processing engines through shared access to a memory module.
15. The apparatus of claim 11, wherein the plurality of rows of the data processing engine include: a first row having a first subset of the plurality of data processing engines, wherein each data processing engine in the first subset has the memory module located on a first side of the core; as well as A second row having a second subset of the plurality of data processing engines, wherein each data processing engine in the second subset has the memory module located on a second side of the core, wherein the first side and the second side are opposite.
16. The apparatus of claim 15, wherein the memory modules of a first selected one of the plurality of data processing engines in the first row and a second selected one of the plurality of data processing engines in the second row are each include: a first memory interface coupled to a core of a data processing engine of the plurality of data processing engines immediately north of the selected data processing engine; a second memory interface coupled to a core within the selected data processing engine; a third memory interface coupled to a core of a data processing engine of the plurality of data processing engines that is adjacent to the east or west of the selected data processing engine; as well as A fourth memory interface is coupled to a core of a data processing engine immediately south of the selected data processing engine among the plurality of data processing engines.
17. The apparatus of claim 15, wherein a selected one of the plurality of data processing engines is configured to communicate with a group of at least eight of the plurality of data processing engines via shared access to a memory module.
18. A device having multiple dies, the device include: a first die including one or more electronic subsystems; as well as A second die is coupled to the first die, wherein the second die comprises a data processing engine (DPE) array having a plurality of data processing engines (DPEs), and wherein each of the plurality of DPEs is configurable to share data with one or more other DPEs of the plurality of DPEs using one or more of a plurality of data sharing techniques, the data sharing techniques comprising: The core of the selected DPE accesses the memory module of the adjacent DPE via a memory interface of the selected DPE connected to the memory module of the adjacent DPE; as well as The selected DPE uses a direct memory access DMA engine and a stream switch of the selected DPE to access a memory module of a non-adjacent DPE.
19. The apparatus of claim 18, wherein each DPE comprises a core and a memory module, and wherein in at least one DPE in the array of DPEs, the core is implemented as a graphics processing unit (GPU).
20. A device, include: A first data processing engine DPE array has a plurality of data processing engines DPE; as well as an electronics subsystem coupled to the DPE array; Each of the plurality of DPEs can be configured to share data with one or more other DPEs of the plurality of DPEs using one or more of a plurality of data sharing technologies, wherein the data sharing technologies include: The core of the selected DPE accesses the memory module of the adjacent DPE via a memory interface of the selected DPE connected to the memory module of the adjacent DPE; as well as The selected DPE uses a direct memory access DMA engine and a stream switch of the first DPE to access a memory module of a non-adjacent DPE.
21. The device of claim 20, wherein the electronic subsystem is at least one of a graphics processing unit (GPU), an application specific integrated circuit, or programmable logic.