Hardware-Software Design Flow for Heterogeneous and Programmable Devices
Patent Information
- Application Number
- KR1020217040001
- Authority / Receiving Office
- KR · KR
- Patent Type
- Patents
- Current Assignee / Owner
- Priority Date
- 2019-05-23
- Filing Date
- 2020-05-11
- Publication Date
- 2026-09-09
- Estimated Expiration
- 2040-05-11
Smart Images

Figure 112021141175905-PCT00020_ABST
Abstract
Description
Technology Field Holding rights to copyrighted materials
[0001] Parts of the disclosures of this patent document include materials protected by copyright. The copyright holder does not object to anyone facsimile reproducing this patent document or the patent disclosures as they appear in the patent files or records of the Patent and Trademark Office, but all other materials are protected by copyright.
[0002] The present disclosure relates to integrated circuits (ICs), and more specifically to implementing applications including hardware and software parts within heterogeneous and programmable ICs. Background Technology
[0003] A programmable integrated circuit (IC) refers to a type of IC that includes programmable logic. An example of a programmable IC is a field programmable gate array (FPGA). An FPGA is characterized by including programmable circuit blocks. Examples of programmable circuit blocks include, but are not limited to, input / output blocks (IOBs), configurable logic blocks (CLBs), dedicated random access memory blocks (BRAMs), digital signal processing blocks (DSPs), processors, clock managers, and delay lock loops (DLLs).
[0004] Modern programmable ICs have evolved to include programmable logic along with one or more other subsystems. For example, some programmable ICs have evolved into System-on-Chips (SoCs) that include both programmable logic and a hardwired processor system. Other variations of programmable ICs include additional and / or different subsystems. As the heterogeneity of the subsystems included in programmable ICs increases, there are difficulties in implementing applications within these devices. Conventional design flows for ICs containing both hardware and software-based subsystems (e.g., programmable logic circuits and processors) have relied on hardware designers first creating a monolithic hardware design for the IC. The hardware design serves as a platform where the software design is subsequently created, compiled, and executed. This approach is often overly restrictive.
[0005] In other cases, software and hardware design processes can be decoupled. However, decoupling hardware and software design processes fails to provide any indication of the placement of interfaces between various subsystems within the IC or software requirements. Consequently, the hardware and software design processes may not converge into an operable implementation of the application within the IC.
[0006] In one aspect, the method may include the step of using a processor to generate a first interface solution and a logical architecture for an application that specifies a mapping of logical resources to the hardware of an interface circuit block between a DPE array and programmable logic, for an application that specifies a software part for implementation within a data processing engine (DPE) array of the device and a hardware part for implementation within programmable logic (PL) of the device. The method may include the step of constructing a block diagram of the hardware part based on the logical architecture and the first interface solution and the step of performing an implementation flow on the block diagram using a processor. The method may include the step of compiling the software part of the application for implementation in one or more DPEs of the DPE array using a processor.
[0007] In another aspect, the system includes a processor configured to initiate operations. The operations may include a first interface solution that specifies the mapping of logical resources to the hardware of an interface circuit block between a DPE array and a PL for an application that specifies a software part for implementation within a DPE array of the device and a hardware part for implementation within a PL of the device, and an operation to generate a logical architecture for the application. The operations may include an operation to construct a block diagram of the hardware part based on the logical architecture and the first interface solution, an operation to perform an implementation flow on the block diagram, and an operation to compile the software part of the application for implementation in one or more DPEs of the array of DPEs.
[0008] In another aspect, a computer program product comprises a computer-readable storage medium storing program code. The program code is executable by computer hardware to initiate operations. The operations may include a first interface solution that specifies a mapping of logical resources to the hardware of an interface circuit block between a DPE array and a logical architecture for an application that specifies a software part for implementation within a DPE array of the device and a hardware part for implementation within a PL of the device, for an application that specifies a software part for implementation within a DPE array of the device and a hardware part for implementation within a PL of the device, and an operation to generate a logical architecture for the application. The operations may include an operation to construct a block diagram of the hardware part based on the logical architecture and the first interface solution, an operation to perform an implementation flow on the block diagram, and an operation to compile the software part of the application for implementation in one or more DPEs of the array of DPEs.
[0009] In another aspect, the method may include the step of performing an implementation flow for a hardware part based on an interface block solution that maps logical resources used by the software part to the hardware of an interface block coupling the DPE array to the PL, using a processor running a hardware compiler for an application having a software part for implementation in a DPE array of the device and a hardware part for implementation in a PL of the device. The method may include the step of providing interface block constraints to a DPE compiler using a processor running a hardware compiler in response to failure to satisfy design metrics during the implementation flow. The method may also include the step of generating an updated interface block solution using a processor running a DPE compiler in response to receiving interface block constraints, and providing the updated interface block solution from the DPE compiler to the hardware compiler.
[0010] In another aspect, the system includes a processor configured to initiate operations. The operations may include, for an application having a software part for implementation in a DPE array of the device and a hardware part for implementation in a PL of the device, an operation to perform an implementation flow for the hardware part based on an interface block solution that maps logical resources used by the software part to the hardware of the interface block coupling the DPE array to the PL using a hardware compiler. The operations may include an operation to provide interface block constraints to the DPE compiler using the hardware compiler in response to failure to satisfy design metrics during the implementation flow. The operations may further include an operation to generate an updated interface block solution using the DPE compiler in response to receiving interface block constraints, and an operation to provide the updated interface block solution from the DPE compiler to the hardware compiler.
[0011] In another aspect, a computer program product comprises a computer-readable storage medium storing program code. The program code is executable by computer hardware to initiate operations. For an application having a software part for implementation in a DPE array of a device and a hardware part for implementation in a PL of a device, operations may include performing an implementation flow for the hardware part based on an interface block solution that maps logical resources used by the software part to the hardware of an interface block coupling the DPE array to the PL using a hardware compiler. Operations may include providing interface block constraints to a DPE compiler using a hardware compiler in response to failure to satisfy design metrics during the implementation flow. Operations may further include generating an updated interface block solution using a DPE compiler in response to receiving interface block constraints, and providing the updated interface block solution from the DPE compiler to the hardware compiler.
[0012] In another aspect, the method may include the step of using a processor to generate a first interface solution for an application that specifies a software part for implementation within a DPE array of a device and a hardware part having HLS kernels for implementation within a PL of the device, which maps logical resources used by the software part to hardware resources of an interface block coupling the DPE array and the PL. The method may include the step of using a processor to generate a connectivity graph that specifies connectivity between nodes of the software part to be implemented in the DPE array and HLS kernels, and using a processor to generate a block diagram based on the connectivity graph and HLS kernels, wherein the block diagram is synthesizable. The method may further include the step of using a processor to perform an implementation flow on the block diagram based on the first interface solution and using a processor to compile the software part of the application for implementation in one or more DPEs of the DPE array.
[0013] In another aspect, the system includes a processor configured to initiate operations. The operations may include an operation to generate a first interface solution that maps logical resources used by the software part to the hardware resources of an interface block coupling the DPE array and the PL, for an application that specifies a software part for implementation within the device's DPE array and a hardware part having HLS kernels for implementation within the device's PL. The operations may include an operation to generate a connection graph that specifies the connectivity between the nodes of the software part to be implemented in the DPE array and the HLS kernels, and an operation to generate a block diagram based on the connection graph and the HLS kernels, wherein the block diagram is synthesizable. The operations may further include an operation to perform an implementation flow on the block diagram based on the first interface solution and an operation to compile the software part of the application for implementation in one or more DPEs of the DPE array.
[0014] In another aspect, a computer program product comprises a computer-readable storage medium storing program code. The program code is executable by computer hardware to initiate operations. The operations may include generating a first interface solution that maps logical resources used by the software part to the hardware resources of an interface block coupling the DPE array and the PL, for an application that specifies a software part for implementation within the DPE array of the device and a hardware part having HLS kernels for implementation within the PL of the device. The operations may include generating a connection graph that specifies the connectivity between the nodes of the software part to be implemented in the DPE array and the HLS kernels, and generating a block diagram based on the connection graph and the HLS kernels, wherein the block diagram is synthesizable. The operations may further include performing an implementation flow on the block diagram based on the first interface solution and compiling the software part of the application for implementation in one or more DPEs of the DPE array.
[0015] This summary paragraph is provided solely to introduce specific concepts, not to identify any important or essential features of the claimed subject matter. Other features of the arrangements of the invention will become apparent from the accompanying drawings and the following detailed description. Brief explanation of the drawing
[0016] Arrangements of the present invention are illustrated by way of example in the accompanying drawings. However, the drawings should not be construed as limiting the arrangements of the present invention only to the specific embodiments depicted. Various aspects and advantages will become apparent upon review of the following detailed description and reference to the drawings.
[0017] FIG. 1 illustrates an example of a computing node for use with one or more embodiments described herein.
[0018] Figure 2 illustrates an exemplary architecture for a System-on-Chip (SoC) type integrated circuit (IC).
[0019] Figure 3 illustrates an exemplary architecture for a DPE of the data processing engine (DPE) array of Figure 2.
[0020] FIG. 4 illustrates additional aspects of the exemplary architecture of FIG. 3.
[0021] Figure 5 illustrates another exemplary architecture for a DPE array.
[0022] Figure 6 illustrates an exemplary architecture for the tiles of the SoC interface block of the DPE array.
[0023] FIG. 7 illustrates an exemplary implementation of the Network-on-Chip (NoC) of FIG. 1.
[0024] Figure 8 is a block diagram illustrating the connections between the endpoint circuits of the SoC of Figure 1 through the NoC.
[0025] Figure 9 is a block diagram illustrating NoC according to another example.
[0026] Figure 10 illustrates an exemplary method for programming a NoC.
[0027] Figure 11 illustrates another exemplary method of programming a NoC.
[0028] Figure 12 illustrates an exemplary data path through the NoC between endpoint circuits.
[0029] FIG. 13 illustrates an exemplary method for processing read / write requests and responses related to NoC.
[0030] FIG. 14 illustrates an exemplary implementation of a NoC master unit.
[0031] FIG. 15 illustrates an exemplary implementation of a NoC slave unit.
[0032] FIG. 16 illustrates an exemplary software architecture executable by the system described in relation to FIG. 1.
[0033] FIGS. 17a and FIGS. 17b illustrate examples of applications mapped to an SoC using the system described in relation to FIG. 1.
[0034] Figure 18 illustrates an exemplary implementation of another application mapped to the SoC.
[0035] FIG. 19 illustrates another exemplary software architecture executable by the system described in relation to FIG. 1.
[0036] Figure 20 illustrates an exemplary method for performing a design flow to implement an application on an SoC.
[0037] Figure 21 illustrates another exemplary method of performing a design flow to implement an application on an SoC.
[0038] Figure 22 illustrates an exemplary method of communication between a hardware compiler and a DPE compiler.
[0039] FIG. 23 illustrates an exemplary method for handling SoC interface block solutions.
[0040] Figure 24 illustrates another example of an application for implementation in an SoC.
[0041] Figure 25 illustrates an example of an SoC interface block solution generated by a DPE compiler.
[0042] Figure 26 illustrates an example of routable SoC interface block constraints received by the DPE compiler.
[0043] Figure 27 illustrates examples of non-routable SoC interface block constraints.
[0044] FIG. 28 illustrates an example where the DPE compiler ignores the soft type SoC interface block constraints from FIG. 27.
[0045] FIG. 29 illustrates another example of non-routable SoC interface block constraints.
[0046] FIG. 30 illustrates an exemplary mapping of the DPE nodes of FIG. 29.
[0047] Figure 31 illustrates another example of non-routable SoC interface block constraints.
[0048] FIG. 32 illustrates an exemplary mapping of the DPE nodes of FIG. 31.
[0049] FIG. 33 illustrates another exemplary software architecture executable by the system of FIG. 1.
[0050] Figure 34 illustrates another exemplary method of performing a design flow to implement an application on an SoC.
[0051] Figure 35 illustrates another exemplary method of performing a design flow to implement an application on an SoC. Specific details for implementing the invention
[0052] Although the present disclosure is concluded by claims defining novel features, it is thought that the various features described in the present disclosure will be better understood when considered in conjunction with the detailed description in the drawings. The process(s), machine(s), manufacture(s), and any variations thereof described herein are provided for illustrative purposes. Specific structural and functional details described in the present disclosure are not to be interpreted as limitations, but solely as a basis for the claims and as a representative basis to instruct those skilled in the art to make various uses of the features described in any appropriately detailed structure. Additionally, terms and phrases used in the present disclosure are not intended to be limiting, but rather to provide an understandable description of the described features.
[0053] The present disclosure relates to integrated circuits (ICs), and more specifically to implementing applications comprising hardware and software parts within heterogeneous and programmable ICs. An example of a heterogeneous and programmable IC is a device, such as an integrated circuit, comprising a programmable circuit portion referred herein as “programmable logic” or “PL” and a plurality of hardwired and programmable data processing engines (DPEs). The plurality of DPEs may be arranged in an array linked to communicate with the PL of the IC through a System-on-Chip (SoC) interface block. As defined in the present disclosure, a DPE is a hardwired and programmable circuit block comprising a core capable of executing program code and a memory module coupled to the core. The DPEs may communicate with each other as described in more detail in the present disclosure.
[0054] As described, an application intended for implementation on a device comprises a hardware portion implemented using the device’s PL and a software portion implemented on and executed by the device’s DPE array. The device may also include a hardwired processor system or “PS” capable of executing additional program code, such as other software portions of the application. For example, the PS includes a central processing unit or “CPU,” or other hardwired processors capable of executing program code. Thus, the application may also include additional software portions intended to be executed by the CPU of the PS.
[0055] According to the arrangements of the invention described in this disclosure, design flows that can be performed by a data processing system are provided. The design flows can implement both hardware and software parts of an application within a heterogeneous and programmable IC comprising a PL, a DPE array and / or a PS. The IC may also include a programmable Network-on-Chip (NoC).
[0056] In some implementations, the application is specified as a data flow graph containing multiple interconnected nodes. The nodes of the data flow graph are specified for implementation within a DPE array or within a PL. Nodes implemented in a DPE are, for example, ultimately mapped to a specific DPE in the DPE array. Object code executed by each DPE in the array used by the application is generated to implement the node(s). Nodes implemented in a PL may be implemented, for example, by being synthesized in the PL, or by using a pre-built core (e.g., a register transfer level or "RTL" core).
[0057] The arrangements of the present invention provide exemplary design flows capable of coordinating the construction and integration of different parts of an application for implementation in different heterogeneous subsystems of an IC. Different stages within the exemplary design flows target specific subsystems. For example, one or more stages of the design flows aim to implement the hardware part of the application in a PL, while one or more other stages of the design flows aim to implement the software part of the application in a DPE array. Furthermore, one or more other stages of the design flows aim to implement other software parts of the application in a PS. Still other stages of the design flows aim to implement routes and data transfers between different subsystems and / or circuit blocks through a NoC.
[0058] Different stages of exemplary design flows corresponding to different subsystems may be executed by different compilers specific to the subsystem. For example, software parts may be implemented using a DPE compiler and / or a PS compiler. Hardware parts to be implemented in the PL may be implemented by a hardware compiler. Roots for the NoC may be implemented by a NoC compiler. Various compilers may communicate and interact with each other while implementing the individual subsystems specified by the application so that the application converges to a solution that is executable on the IC. For example, compilers may exchange design data during operation to converge to a solution that satisfies the design metrics specified for the application. Furthermore, the achieved solution (e.g., the implementation of the application on the device) is a solution in which various parts of the application are mapped to the individual subsystems of the device and the interfaces between the different subsystems are consistent and mutually compatible.
[0059] By using the exemplary design flows described in this disclosure, a system can implement an application within heterogeneous and programmable ICs in a short time (e.g., with a short runtime), in contrast to cases where all parts of the application are implemented jointly on a device. Furthermore, the exemplary design flows described in this disclosure achieve feasibility and quality (e.g., closure of design metrics such as timing, area, and power) for the resulting implementation of the application within heterogeneous and programmable ICs that are often superior to results obtained using other prior art in which each part of the application is mapped completely independently and subsequently stitched or combined together. The exemplary design flows achieve these results through loosely-coupled joint convergence techniques described herein, which rely, at least in part, on shared interface constraints between different subsystems.
[0060] Further aspects of arrangements of the present invention are described in more detail below with reference to the drawings. For the purposes of simplification and clarification of the example, the elements depicted in the drawings are not necessarily drawn to actual scale. For instance, the dimensions of some elements may be exaggerated relative to others for clarification. Additionally, where deemed appropriate, reference numbers are repeated between the drawings to indicate corresponding, similar, or identical features.
[0061] FIG. 1 illustrates an example of a computing node (100). The computing node (100) may include a host data processing system (host system) (102) and a hardware acceleration board (104). The computing node (100) is just one exemplary implementation of a computing environment that can be used with the hardware acceleration board. In this regard, the computing node (100) may be used as a standalone capacity as a bare metal server, as part of a computing cluster, or as a cloud computing node within a cloud computing environment. FIG. 1 is not intended to imply any limitation on the use or scope of function of the examples described herein. The computing node (100) is an example of a system and / or computer hardware capable of performing various operations described in this disclosure in relation to implementing applications within an SoC (200). For example, the computing node (100) may be used to implement an Electronic Design Automation (EDA) system.
[0062] The host system (102) operates with a number of other general-purpose or special-purpose computing system environments or configurations. Examples of computing systems, environments and / or configurations that may be suitable for use with the host system (102) include, but are not limited to, personal computer systems, server computer systems, thin clients, thick clients, hand-held or laptop devices, multiprocessor systems, microprocessor-based systems, set-top boxes, programmable consumer electronics, network PCs, minicomputer systems, mainframe computer systems, and distributed cloud computing environments including any of the foregoing systems or devices.
[0063] As illustrated, the host system (102) is illustrated in the form of a computing device, such as a computer or a server. The host system (102) may be implemented as a standalone device in a distributed cloud computing environment or cluster where tasks are performed by remote processing devices linked via a communication network. In a distributed cloud computing environment, program modules may be located on both local and remote computer system storage media, including memory storage devices. Components of the host system (102) may include, but are not limited to, one or more processors (106) (e.g., a central processing unit), memory (108), and a bus (110) coupling various system components, including memory (108), to the processor (106). The processor(s) (106) may include any of the various processors capable of executing program code. Exemplary processor types include, but are not limited to, processors having an x86 type architecture (IA-32, IA-64, etc.), Power Architecture, ARM processors, etc.
[0064] The bus (110) represents one or more of any various types of communication bus structures, including a memory bus or memory controller, a peripheral bus, an accelerated graphics port, and a processor or local bus using any of the various available bus architectures. As examples, but not limited to, these architectures include the Industry Standard Architecture (ISA) bus, Micro Channel Architecture (MCA) bus, Enhanced ISA (EISA) bus, Video Electronics Standards AsSoCiation (VESA) local bus, Peripheral Component Interconnect (PCI) bus, and PCIe (PCI Express) bus.
[0065] The host system (102) typically includes various computer-readable media. These media may be any available media accessible by the host system (102) and may include any combination of volatile media, non-volatile media, removable media, and / or non-removable media.
[0066] Memory (108) may include computer-readable media in the form of volatile memory, such as cache memory (114) and / or random-access memory (RAM) (112). The host system (102) may also include other removable / non-removable, volatile / non-volatile computer system storage media. For example, a storage system (116) may be provided to read from and write to non-removable, non-volatile magnetic media (not shown, typically referred to as a "hard drive"). Although not shown, a magnetic disk drive for reading from and writing to a removable non-volatile magnetic disk (e.g., a "floppy disk") and an optical disk drive for reading from and writing to a removable, non-volatile optical disk, such as a CD-ROM, DVD-ROM, or other optical media may be provided. In such cases, each may be connected to the bus (110) by one or more data media interfaces. As further illustrated and described below, the memory (108) may include at least one computer program product having a set (e.g., at least one) of program modules (e.g., program code) configured to perform the functions and / or operations described in the present disclosure.
[0067] A program / utility (118) having a set (at least one) of program modules (120), as well as an operating system, one or more application programs, other program modules, and program data may be stored in memory (108) as an example, but not limited to. The program modules (120) generally perform the functions and / or methodologies of the embodiments of the invention described herein. For example, the program modules (120) may include a driver or daemon and one or more applications for communicating with a hardware acceleration board (104) and / or an SoC (200).
[0068] A program / utility (118) is executable by a processor (106). Any data items used, created, and / or operated by the program / utility (118) and the processor (106) are functional data structures that are given functionality when used by the processor (106). As defined in this disclosure, a “data structure” is a physical implementation of the organization of a data model of data within physical memory. Thus, a data structure is formed by specific electrical or magnetic structure elements of memory. A data structure introduces a physical organization to data stored in memory as used by an application program executed using the processor.
[0069] The host system (102) may include one or more input / output (I / O) interfaces (128) that are communicably linked to the bus (110). The I / O interface(s) (128) may enable the host system (102) to communicate with external devices, to be coupled with external devices that enable user(s) to interact with the host system (102), and to be coupled with external devices that enable the host system (102) to communicate with other computing devices, etc. For example, the host system (102) may be communicably linked to a display (130) and a hardware acceleration board (104) through the I / O interface(s) (128). The host system (102) may be coupled to other external devices, such as a keyboard (not shown), through the I / O interface(s) (128). Examples of I / O interfaces (128) may include, but are not limited to, network cards, modems, network adapters, hardware controllers, etc.
[0070] In an exemplary implementation, the I / O interface (128) that enables the host system (102) to communicate with the hardware acceleration board (104) is a PCIe adapter. The hardware acceleration board (104) may be implemented as a circuit board, such as a card, coupled to the host system (102). The hardware acceleration board (104) may be inserted, for example, into a card slot, such as an available bus and / or PCIe slot of the host system (102).
[0071] The hardware acceleration board (104) includes an SoC (200). The SoC (200) is a heterogeneous and programmable IC and thus has multiple heterogeneous subsystems. An exemplary architecture for the SoC (200) is described in more detail in relation to FIG. 2. The hardware acceleration board (104) also includes volatile memory (134) coupled to the SoC (200) and non-volatile memory (136) also coupled to the SoC (200). The volatile memory (134) may be implemented as RAM and is considered as "local memory" of the SoC (200), whereas memory (108) within the host system (102) is considered to be local to the host system (102) rather than local to the SoC (200). In some implementations, the volatile memory (134) may include several gigabytes of RAM, e.g., 64 GB of RAM. Examples of non-volatile memory (136) include flash memory.
[0072] In the example of FIG. 1, the computing node (100) can operate on an application for the SoC (200) and can implement the application within the SoC (200). The application may include hardware and software parts corresponding to different heterogeneous subsystems available on the SoC (200). Generally, the computing node (100) can map the application to the SoC (200) for execution by the SoC (200).
[0073] FIG. 2 illustrates an exemplary architecture for an SoC (200). The SoC (200) is an example of a programmable IC and an integrated programmable device platform. In the example of FIG. 2, various different subsystems or regions of the illustrated SoC (200) may be implemented on a single die provided within a single integrated package. In other examples, different subsystems may be implemented on multiple interconnected dies provided as a single integrated package.
[0074] In the example, the SoC (200) includes a plurality of regions having circuit portions having different functionalities. In the example, the SoC (200) optionally includes a data processing engine (DPE) array (202). The SoC (200) includes programmable logic (PL) regions (214) (hereinafter, PL region(s) or PL), a processing system (PS) (212), a network-on-chip (NoC) (208), and one or more hardwired circuit blocks (210). The DPE array (202) is implemented as a plurality of interconnected, hardwired, and programmable processors having interfaces to other regions of the SoC (200).
[0075] PL (214) is a circuit that can be programmed to perform specified functions. For example, PL (214) can be implemented as a circuit of the field programmable gate array (FPGA) type. PL (214) may include an array of programmable circuit blocks. Examples of programmable circuit blocks within PL (214) include, but are not limited to, configurable logic blocks (CLB), dedicated random access memory blocks (BRAM and / or UltraRAM or URAM), digital signal processing blocks (DSP), clock managers, and / or delay lock loops (DLL).
[0076] Each programmable circuit block within the PL (214) typically includes both programmable interconnect circuits and programmable logic circuits. The programmable interconnect circuits typically include a large number of interconnect wires of varying lengths interconnected by programmable interconnect points (PIPs). Typically, the interconnect wires are configured to provide connections on a bit-by-bit basis (e.g., wire-by-wire basis) (e.g., each wire carries a single bit of information). The programmable logic circuits implement user-designed logic using programmable elements, which may include, for example, lookup tables, registers, arithmetic logic, etc. The programmable interconnects and programmable logic circuits can be programmed by loading configuration data that defines how the programmable elements are configured and operate into internal configuration memory cells.
[0077] The PS (212) is implemented as a hardwired circuit that is manufactured as part of the SoC (200). The PS (212) may be implemented as or include any of the various different processor types capable of executing program code. For example, the PS (212) may be implemented as an individual processor, such as a single core capable of executing program code. In another example, the PS (212) may be implemented as a multi-core processor. In yet another example, the PS (212) may include one or more cores, modules, co-processors, interfaces, and / or other resources. The PS (212) may be implemented using any of the various different types of architectures. Exemplary architectures that may be used to implement PS (212) may include, but are not limited to, ARM processor architectures, x86 processor architectures, GPU architectures, mobile processor architectures, DSP architectures, other suitable architectures capable of executing computer-readable instructions or program code, and / or different processors and / or combinations of processor architectures.
[0078] The NoC (208) includes an interconnect network for sharing data between endpoint circuits of the SoC (200). The endpoint circuits may be placed in a DPE array (202), PL regions (214), PS (212), and / or hardwired circuit blocks (210). The NoC (208) may include high-speed data paths using dedicated switching. In an example, the NoC (208) includes horizontal paths, vertical paths, or both horizontal and vertical paths. The arrangement and number of regions shown in FIG. 1 are merely examples. The NoC (208) is an example of a common infrastructure available within the SoC (200) to connect selected components and / or subsystems.
[0079] The NoC (208) provides connectivity to selected circuit blocks among the PL (214), PS (212), and hardwired circuit blocks (210). The NoC (208) is programmable. In the case of a programmable NoC used with other programmable circuits, the nets and / or data transfers to be routed through the NoC (208) are not known until a user circuit design is created for implementation within the SoC (200). The NoC (208) can be programmed by loading configuration data into internal configuration registers that define how elements within the NoC (208), such as switches and interfaces, are configured and operated to transfer data from switch to switch and between NoC interfaces.
[0080] NoC (208) is manufactured as part of SoC (200) and, although not physically modifiable, can be programmed to establish connectivity between different master circuits and different slave circuits of a user circuit design. NoC (208) may include a plurality of programmable switches capable of establishing a packet switched network connecting, for example, user-specified master circuits and slave circuits. In this regard, NoC (208) may be adapted to different circuit designs, each having different combinations of master circuits and slave circuits implemented at different locations within the SoC (200) that can be coupled by NoC (208). NoC (208) can be programmed to route data, for example, application data and / or configuration data, between the master and slave circuits of the user circuit design. For example, the NoC (208) can be programmed to couple different user-defined circuits implemented within the PL (214) with the PS (212) and / or DPE array (202), different hardwired circuit blocks, and / or different circuits and / or systems outside the SoC (200).
[0081] Hardwired circuit blocks (210) may include input / output (I / O) blocks and / or transceivers for transmitting and receiving signals to and from circuits and / or systems, memory controllers, etc., outside the SoC (200). Examples of different I / O blocks may include single-ended and pseudo-differential I / Os and high-speed differential clock transceivers. Furthermore, hardwired circuit blocks (210) may be implemented to perform specific functions. Additional examples of hardwired circuit blocks (210) include, but are not limited to, cryptographic engines, digital-to-analog converters, analog-to-digital converters, etc. Hardwired circuit blocks (210) within the SoC (200) may sometimes be referred to herein as application-specific blocks.
[0082] In the example of FIG. 2, PL (214) is depicted as two distinct regions. In another example, PL (214) may be implemented as an integrated region of the programmable circuitry. In yet another example, PL (214) may be implemented as more than two different regions of the programmable circuitry. The specific organization of PL (214) is not intended as a limitation. In this regard, SoC (200) includes one or more PL regions (214), PS (212), and NoC (208).
[0083] In other exemplary implementations, the SoC (200) may include two or more DPE arrays (202) located in different regions of the IC. In other examples, the SoC (200) may be implemented as a multi-die IC. In this case, each subsystem may be implemented on a different die. The different dies may be communicably linked using any of the various available multi-die IC technologies of stacking dies side by side on an interposer, such as a stacked-die architecture in which the IC is implemented as a Multi-Chip Module (MCM). In the multi-die IC example, it should be recognized that each die may include a single subsystem, two or more subsystems, a subsystem and other partial subsystems, or any combination thereof.
[0084] The DPE array (202) is implemented as a two-dimensional array of DPEs (204) including an SoC interface block (206). The DPE array (202) may be implemented using any of the various different architectures described in more detail herein below. For the purposes of illustration rather than limitation, FIG. 2 illustrates DPEs (204) arranged in aligned rows and aligned columns. However, in other embodiments, DPEs (204) may be arranged such that DPEs in selected rows and / or columns are horizontally inverted or flipped relative to DPEs in adjacent rows and / or columns. In one or more other embodiments, the rows and / or columns of DPEs may be offset relative to adjacent rows and / or columns. One or more or all DPEs (204) may be implemented to include one or more cores capable of executing program code. The number of DPEs (204), the specific arrangement of DPEs (204), and / or the orientation of DPEs (204) are not intended to be limited.
[0085] The SoC interface block (206) can couple DPEs (204) to one or more other subsystems of the SoC (200). In one or more embodiments, the SoC interface block (206) is coupled to adjacent DPEs (204). For example, the SoC interface block (206) can be directly coupled to each DPE (204) in the bottom row of DPEs within the DPE array (202). In an example, the SoC interface block (206) can be directly connected to DPEs (204-1, 204-2, 204-3, 204-4, 204-5, 204-6, 204-7, 204-8, 204-9, and 204-10).
[0086] FIG. 2 is provided for illustrative purposes. In other embodiments, the SoC interface block (206) may be located at the top of the DPE array (202), to the left of the DPE array (202) (e.g., as a column), to the right of the DPE array (202) (e.g., as a column), or at a number of locations within and around the DPE array (202) (e.g., as one or more interposed rows and / or columns within the DPE array (202). Depending on the layout and location of the SoC interface block (206), the specific DPEs coupled to the SoC interface block (206) may vary.
[0087] For the purposes of example, if the SoC interface block (206) is located to the left of the DPEs (204), the SoC interface block (206) can be directly coupled to the left column of DPEs including DPE (204-1), DPE (204-11), DPE (204-21), and DPE (204-31). If the SoC interface block (206) is located to the right of the DPEs (204), the SoC interface block (206) can be directly coupled to the right column of DPEs including DPE (204-10), DPE (204-20), DPE (204-30), and DPE (204-40). If the SoC interface block (206) is located at the top of the DPEs (204), the SoC interface block (206) may be coupled to the top row of DPEs including DPE (204-31), DPE (204-32), DPE (204-33), DPE (204-34), DPE (204-35), DPE (204-36), DPE (204-37), DPE (204-38), DPE (204-39), and DPE (204-40). If the SoC interface block (206) is located at multiple locations, the specific DPEs directly connected to the SoC interface block (206) may vary. For example, if the SoC interface block is implemented as a row and / or column within the DPE array (202), the DPEs directly coupled to the SoC interface block (206) may be one or more sides of the SoC interface block (206) or DPEs adjacent to the SoC interface block (206) on each side.
[0088] The DPEs (204) are interconnected by DPE interconnects (not shown) that, when taken collectively, form a DPE interconnect network. Accordingly, the SoC interface block (206) can communicate with one or more selected DPEs (204) of the DPE array (202) directly connected to the SoC interface block (206) and communicate with any DPE (204) of the DPE array (202) by utilizing a DPE interconnect network formed by DPE interconnects implemented within each individual DPE (204).
[0089] The SoC interface block (206) can couple each DPE (204) within the DPE array (202) to one or more other subsystems of the SoC (200). For example, the SoC interface block (206) can couple the DPE array (202) to the NoC (208) and PL (214). Thus, the DPE array (202) can communicate with circuit blocks implemented in any of the PL (214), PS (212), and / or hardwired circuit blocks (210). For example, the SoC interface block (206) can establish connections between selected DPEs (204) and PL (214). The SoC interface block (206) can also establish connections between selected DPEs (204) and NoC (208). Through the NoC (208), selected DPEs (204) can communicate with the PS (212) and / or hardwired circuit blocks (210). The selected DPEs (204) can communicate with the hardwired circuit blocks (210) through the SoC interface block (206) and PL (214). In certain embodiments, the SoC interface block (206) can be directly coupled to one or more subsystems of the SoC (200). For example, the SoC interface block (206) can be directly coupled to the PS (212) and / or hardwired circuit blocks (210).
[0090] In one or more embodiments, the DPE array (202) includes a single clock domain. Other subsystems, such as the NoC (208), PL (214), PS (212), and various hardwired circuit blocks (210), may be in one or more separate or different clock domain(s). Furthermore, the DPE array (202) may include additional clocks that can be used to interface with other subsystems among the subsystems. In specific embodiments, the SoC interface block (206) includes a clock signal generator capable of generating one or more clock signals that can be provided to or distributed to the DPEs (204) of the DPE array (202).
[0091] The DPE array (202) can be programmed by loading configuration data into internal configuration memory cells (also referred to herein as “configuration registers”) that define the connectivity between the DPEs (204) and the SoC interface block (206), and how the DPEs (204) and the SoC interface block (206) operate. For example, for a specific DPE (204) or a group of DPEs (204) to communicate with a subsystem, the DPE(s) (204) and the SoC interface block (206) are programmed to do so. Similarly, for one or more specific DPEs (204) to communicate with one or more other DPEs (204), the DPEs are programmed to do so. The DPE(s) (204) and the SoC interface block (206) can be programmed by loading configuration data into the configuration registers within the DPE(s) (204) and the SoC interface block (206), respectively. In another example, a clock signal generator that is part of the SoC interface block (206) may be programmable using configuration data to change the clock frequencies provided to the DPE array (202).
[0092] FIG. 3 illustrates an exemplary architecture for a DPE (204) of the DPE array (202) of FIG. 2. In the example of FIG. 3, the DPE (204) includes a core (302), a memory module (304), and a DPE interconnect (306). Each DPE (204) is implemented as a hardwired and programmable circuit block on an SoC (200).
[0093] The core (302) provides the data processing capabilities of the DPE (204). The core (302) may be implemented as any of the various different processing circuits. In the example of FIG. 3, the core (302) includes an optional program memory (308). In an exemplary implementation, the core (302) is implemented as a processor capable of executing program code, such as computer-readable instructions. In that case, the program memory (308) is included, and the program memory (308) may store instructions executed by the core (302). For example, the core (302) may be implemented as a CPU, GPU, DSP, vector processor, or other type of processor capable of executing instructions. The core (302) may be implemented using any of the various CPU and / or processor architectures described herein. In another example, the core (302) is implemented as a very long instruction word (VLIW) vector processor or DSP.
[0094] In certain implementations, the program memory (308) is implemented as a dedicated program memory that is exclusive to the core (302) (e.g., accessed only by the core (302)). The program memory (308) can be used only by the core of the same DPE (204). Thus, the program memory (308) can be accessed only by the core (302) and is not shared with any other DPE or components of another DPE. The program memory (308) may include a single port for read and write operations. The program memory (308) may support program compression and is addressable using the memory-mapped network portion of the DPE interconnect (306), which is described in more detail below. For example, through the memory-mapped network of the DPE interconnect (306), program code that can be executed by the core (302) may be loaded into the program memory (308).
[0095] The core (302) may include configuration registers (324). Configuration data may be loaded into the configuration registers (324) to control the operation of the core (302). In one or more embodiments, the core (302) may be enabled and / or disabled based on the configuration data loaded into the configuration registers (324). In the example of FIG. 3, the configuration registers (324) are addressable (e.g., readable and / or written) through a memory-mapped network of the DPE interconnect (306) described in more detail below.
[0096] In one or more embodiments, the memory module (304) may store data used and / or generated by the core (302). For example, the memory module (304) may store application data. The memory module (304) may include a read / write memory such as random-access memory (RAM). Thus, the memory module (304) may store data that can be read and consumed by the core (302). The memory module (304) may also store data (e.g., results) that is written by the core (302).
[0097] In one or more other embodiments, the memory module (304) may store data, such as application data, that may be used and / or generated by one or more other cores of other DPEs within the DPE array. One or more other cores of the DPEs may also read from and / or write to the memory module (304). In certain embodiments, the other cores that may read from and / or write to the memory module (304) may be cores of one or more neighboring DPEs. Another DPE that shares an edge or boundary with (e.g., adjacent) DPE (204) is referred to as the “neighboring” DPE with respect to DPE (204). By enabling one or more other cores from the core (302) and neighboring DPEs to read the memory module (304) and / or write to the memory module (304), the memory module (304) implements a shared memory that supports communication between different DPEs and / or cores that can access the memory module (304).
[0098] Referring to FIG. 2, for example, DPEs (204-14, 204-16, 204-5, and 204-25) are considered neighboring DPEs of DPE (204-15). In one example, a core within each of DPEs (204-16, 204-5, and 204-25) can read and write to a memory module within DPE (204-15). In certain embodiments, only these neighboring DPEs adjacent to the memory module can access the memory module of DPE (204-15). For example, DPE (204-14) may be adjacent to DPE (204-15) but may not be adjacent to the memory module of DPE (204-15) because the core of DPE (204-15) may be located between the core of DPE (204-14) and the memory module of DPE (204-15). Therefore, in certain embodiments, the core of DPE (204-14) cannot access the memory module of DPE (204-15).
[0099] In specific embodiments, whether a core of a DPE can access a memory module of another DPE depends on the number of memory interfaces included in the memory module and whether such cores are connected to an available memory interface among the memory interfaces of the memory module. In the preceding example, the memory module of DPE (204-15) includes four memory interfaces, wherein the core of each DPE of DPEs (204-16, 204-5, and 204-25) is connected to such memory interfaces. The core (302) within DPE (204-15) itself is connected to the fourth memory interface. Each memory interface may include one or more read and / or write channels. In specific embodiments, each memory interface includes multiple read channels and multiple write channels that allow a specific core attached thereto to simultaneously read and / or write to multiple banks within the memory module (304).
[0100] In other examples, more than four memory interfaces may be available. Such other memory interfaces may be used to enable DPEs diagonally opposite to DPE (204-15) to access the memory module of DPE (204-15). For example, if cores within DPEs such as DPEs (204-14, 204-24, 204-26, 204-4, and / or 204-6) are also coupled to available memory interfaces of the memory module within DPE (204-15), such other DPEs will also be able to access the memory module of DPE (204-15).
[0101] The memory module (304) may include configuration registers (336). Configuration data may be loaded into the configuration registers (336) to control the operation of the memory module (304). In the example of FIG. 3, the configuration registers (336 (and 324)) are addressable (e.g., readable and / or written) through the memory-mapped network of the DPE interconnect (306), which is described in more detail below.
[0102] In the example of FIG. 3, the DPE interconnect (306) is specific to the DPE (204). The DPE interconnect (306) enables various operations including communication between the DPE (204) and one or more other DPEs of the DPE array (202) and / or communication with other subsystems of the SoC (200). The DPE interconnect (306) further enables configuration, control, and debugging of the DPE (204).
[0103] In certain embodiments, the DPE interconnect (306) is implemented as an on-chip interconnect. An example of an on-chip interconnect is an AMBA AXI (Advanced Microcontroller Bus Architecture eXtensible Interface) bus (e.g., or switch). The AMBA AXI bus is an embedded microcontroller bus interface for use in establishing on-chip connections between circuit blocks and / or systems. The AXI bus is provided herein as an example of an interconnect circuit that can be used in the arrangements of the invention described herein, and is therefore not intended as a limitation. Other examples of interconnect circuits may include other types of buses, crossbars, and / or other types of switches.
[0104] In one or more embodiments, the DPE interconnect (306) includes two different networks. The first network can exchange data with other DPEs of the DPE array (202) and / or other subsystems of the SoC (200). For example, the first network can exchange application data. The second network can exchange data such as configuration, control, and / or debugging data for the DPE(s).
[0105] In the example of FIG. 3, the first network of the DPE interconnection unit (306) is formed by a stream switch (326) and one or more stream interfaces (not shown). For example, the stream switch (326) includes a core (302), a memory module (304), a memory-mapped switch (332), and stream interfaces for connecting to each of the upper DPE, the left DPE, the right DPE, and the lower DPE. Each stream interface may include one or more masters and one or more slaves.
[0106] The stream switch (326) can enable DPEs that are not coupled to the memory interface of the memory module (304) and / or non-neighboring DPEs to communicate with the core (302) and / or memory module (304) through a DPE interconnection network formed by the DPE interconnection portions of the individual DPEs (204) of the DPE array (202).
[0107] Referring again to FIG. 2 and using DPE (204-15) as a reference point, the stream switch (326) can be coupled to and communicate with another stream switch located at the DPE interconnect of DPE (204-14). The stream switch (326) can be coupled to and communicate with another stream switch located at the DPE interconnect of DPE (204-25). The stream switch (326) can be coupled to and communicate with another stream switch located at the DPE interconnect of DPE (204-16). The stream switch (326) can be coupled to and communicate with another stream switch located at the DPE interconnect of DPE (204-5). Thus, the core (302) and / or memory module (304) can also communicate with any DPE among the DPEs in the DPE array (202) through the DPE interconnects within the DPEs.
[0108] The stream switch (326) may also be used to interface with subsystems such as PL (214) and / or NoC (208). Generally, the stream switch (326) may be programmed to operate as a circuit-switching stream interconnect or a packet-switched stream interconnect. A circuit-switching stream interconnect can implement point-to-point dedicated streams suitable for high-bandwidth communication between DPEs. A packet-switched stream interconnect allows streams to be shared to time-multiplex multiple logical streams into a single physical stream for medium-bandwidth communication.
[0109] The stream switch (326) may include configuration registers (abbreviated as “CR” in FIG. 3) (334). Configuration data may be written to the configuration registers (334) by the memory-mapped network of the DPE interconnect (306). Configuration data loaded into the configuration registers (334) indicates which other DPEs (204) will communicate with other DPEs and / or subsystems (e.g., NoC (208), PL (214), and / or PS (212)), and whether such communications will be established as circuit-switching point-to-point connections or as packet-switching connections.
[0110] The second network of the DPE interconnection section (306) is formed by a memory-mapped switch (332). The memory-mapped switch (332) includes a plurality of memory-mapped interfaces (not shown). Each memory-mapped interface may include one or more masters and one or more slaves. For example, the memory-mapped switch (332) includes a memory-mapped interface for connecting to each of the core (302), the memory module (304), the memory-mapped switch in the DPE above the DPE (204), and the memory-mapped switch in the DPE below the DPE (204).
[0111] The memory-mapped switch (332) is used to transmit configuration, control, and debugging data for the DPE (204). In the example of FIG. 3, the memory-mapped switch (332) can receive configuration data used to configure the DPE (204). The memory-mapped switch (332) can receive configuration data from a DPE located below the DPE (204) and / or from the SoC interface block (206). The memory-mapped switch (332) can forward the received configuration data to one or more other DPEs above the DPE (204), to the core (302) (e.g., to the program memory (308) and / or configuration registers (324)), to the memory module (304) (e.g., to the memory within the memory module (304) and / or configuration registers (336)), and / or to the configuration registers (334) within the stream switch (326).
[0112] The DPE interconnects (306) are coupled to the DPE interconnects of each adjacent DPE and / or SoC interface block (206) according to the location of the DPE (204). When taken collectively, the DPE interconnects of the DPEs (204) form a DPE interconnect network (which may include a stream network and / or a memory-mapped network). The configuration registers of the stream switches of each DPE can be programmed by loading configuration data through the memory-mapped switches. Through configuration, the stream switches and / or stream interfaces are programmed to establish connections with other endpoints, whether to one or more other DPEs (204) and / or the SoC interface block (206), either via packet-switching or circuit-switching.
[0113] In one or more embodiments, the DPE array (202) is mapped to the address space of a processor system such as a PS (212). Accordingly, any configuration registers and / or memories within the DPE (204) can be accessed through a memory-mapped interface. For example, memory within the memory module (304), program memory (308), configuration registers (324) within the core (302), configuration registers (336) within the memory module (304), and / or configuration registers (334) can be read and / or written through a memory-mapped switch (332).
[0114] In the example of FIG. 3, the memory-mapped switch (332) may receive configuration data for the DPE (204). The configuration data may include program code to be loaded into the program memory (308) (if included), configuration data to be loaded into configuration registers (324, 334, and / or 336), and / or data to be loaded into memory (e.g., memory banks) of the memory module (304). In the example of FIG. 3, the configuration registers (324, 334, and 336) are depicted as being located within specific circuit structures, such as the core (302), stream switch (326), and memory module (304), which the configuration registers are intended to control. The example of FIG. 3 is for illustrative purposes only and illustrates that elements within the core (302), memory module (304), and / or stream switch (326) can be programmed by loading configuration data into the corresponding configuration registers. In other embodiments, the configuration registers may be integrated within a specific area of the DPE (204), even though they control the operation of components distributed throughout the DPE (204).
[0115] Accordingly, the stream switch (326) can be programmed by loading configuration data into the configuration registers (334). The configuration data programs the stream switch (326) to operate in a circuit-switching mode between two different DPEs and / or other subsystems or in a packet-switching mode between selected DPEs and / or other subsystems. Accordingly, the connections established for other stream interfaces and / or switches by the stream switch (326) are programmed by loading appropriate configuration data into the configuration registers (334) to establish actual connections or application data paths with other DPEs and / or other subsystems of the IC (300) within the DPE (204).
[0116] FIG. 4 illustrates additional aspects of the exemplary architecture of FIG. 3. In the example of FIG. 4, details regarding the DPE interconnection (306) are not illustrated. FIG. 4 illustrates the connectivity of the core (302) with other DPEs through shared memory. FIG. 4 also illustrates additional aspects of the memory module (304). For the purpose of illustration, FIG. 4 refers to the DPE (204-15).
[0117] As illustrated, the memory module (304) includes a plurality of memory interfaces (402, 404, 406, and 408). In FIG. 4, the memory interfaces (402 and 408) are abbreviated as “MI”. The memory module (304) further includes a plurality of memory banks (412-1 to 412-N). In certain embodiments, the memory module (304) includes eight memory banks. In other embodiments, the memory module (304) may include fewer or more memory banks (412). In one or more embodiments, each memory bank (412) is single-ported to allow up to one access to each memory bank per clock cycle. When the memory module (304) includes eight memory banks (412), such a configuration supports eight parallel accesses per clock cycle. In other embodiments, each memory bank (412) is dual-ported or multi-ported to allow a greater number of parallel accesses per clock cycle.
[0118] In the example of FIG. 4, each of the memory banks (412-1 to 412-N) has individual arbitrators (414-1 to 414-N). Each arbitrator (414) may generate a stall signal in response to detecting collisions. Each arbitrator (414) may include arbitration logic. Furthermore, each arbitrator (414) may include a crossbar. Thus, any master may write to any one or more specific memory banks among the memory banks (412). As mentioned in relation to FIG. 3, the memory module (304) is connected to a memory-mapped switch (332) to enable reading and writing of data to the memory bank (412). Thus, specific data stored in the memory module (304) may be controlled, e.g., written, through the memory-mapped switch (332) as part of a configuration, control, and / or debugging process.
[0119] The memory module (304) further includes a direct memory access (DMA) engine (416). In one or more embodiments, the DMA engine (416) includes at least two interfaces. For example, one or more interfaces may receive input data streams from the DPE interconnect (306) and write the received data to memory banks (412). One or more other interfaces may read data from the memory banks (412) and transmit the data externally through a stream interface (e.g., a stream switch) of the DPE interconnect (306). For example, the DMA engine (416) may include a stream interface for accessing the stream interface (326) of FIG. 3.
[0120] The memory module (304) can operate as a shared memory that can be accessed by a plurality of different DPEs. In the example of FIG. 4, the memory interface (402) is coupled to the core (302) through a core interface (428) included in the core (302). The memory interface (402) provides the core (302) with access to memory banks (412) through mediators (414). The memory interface (404) is coupled to the core of the DPE (204-25). The memory interface (404) provides the core of the DPE (204-25) with access to memory banks (412). The memory interface (406) is coupled to the core of the DPE (204-16). The memory interface (406) provides the core of the DPE (204-16) with access to memory banks (412). The memory interface (408) is coupled to the core of the DPE (204-5). The memory interface (408) provides access to the memory banks (412) to the core of the DPE (204-5). Thus, in the example of FIG. 4, each DPE having a shared boundary with the memory module (304) of the DPE (204-15) can read and write to the memory banks (412). In the example of FIG. 4, the core of the DPE (204-14) does not have direct access to the memory module (304) of the DPE (204-15).
[0121] The core (302) can access the memory modules of other neighboring DPEs through core interfaces (430, 432, and 434). In the example of FIG. 4, the core interface (434) is coupled to the memory interface of DPE (204-25). Thus, the core (302) can access the memory module of DPE (204-25) through the memory interface and core interface (434) contained within the memory module of DPE (204-25). The core interface (432) is coupled to the memory interface of DPE (204-14). Thus, the core (302) can access the memory module of DPE (204-14) through the memory interface and core interface (432) contained within the memory module of DPE (204-14). The core interface (430) is coupled to the memory interface within DPE (204-5). Accordingly, the core (302) can access the memory module of the DPE (204-5) through the memory interface and core interface (430) included within the memory module of the DPE (204-5). As discussed, the core (302) can access the memory module (304) within the DPE (204-15) through the core interface (428) and the memory interface (402).
[0122] In the example of FIG. 4, the core (302) can read and write to any memory module among the memory modules of DPEs (e.g., DPEs (204-25, 204-14, and 204-5)) that share a boundary with the core (302) of DPE (204-15). In one or more embodiments, the core (302) may regard the memory modules within the DPEs (204-25, 204-15, 204-14, and 204-5) as a single continuous memory (e.g., as a single address space). Thus, the process of the core (302) reading and / or writing to the memory modules of these DPEs is the same as the core (302) reading and / or writing to the memory module (304). The core (302) may generate addresses for reads and writes by assuming this continuous memory model. The core (302) can send read and / or write requests to the appropriate core interface (428, 430, 432, and / or 434) based on the generated addresses.
[0123] As mentioned, the core (302) can map read and / or write operations in the correct direction through the core interface (428, 430, 432, and / or 434) based on the addresses of such operations. When the core (302) generates an address for memory access, the core (302) can decode the address to determine the direction (e.g., a specific DPE to be accessed) and forward the memory operation to the correct core interface in the determined direction.
[0124] Accordingly, the core (302) can communicate with the core of the DPE (204-25) through shared memory, which may be a memory module within the DPE (204-25) and / or a memory module (304) of the DPE (204-15). The core (302) can communicate with the core of the DPE (204-14) through shared memory, which may be a memory module within the DPE (204-14). The core (302) can communicate with the core of the DPE (204-5) through shared memory, which may be a memory module within the DPE (204-5) and / or a memory module (304) of the DPE (204-15). Furthermore, the core (302) can communicate with the core of the DPE (204-16) through shared memory, which may be a memory module (304) within the DPE (204-15).
[0125] As discussed, the DMA engine (416) may include one or more stream-memory interfaces. Through the DMA engine (416), application data may be received from other sources within the SoC (200) and stored in the memory module (304). For example, data may be received by stream switches (326) from other DPEs that share boundaries with and / or do not share boundaries with DPE (204-15). Data may also be received by the SoC interface block (206) through the stream switches of the DPEs from other subsystems of the SoC (e.g., NoC (208), hardwired circuit blocks (210), PL (214), and / or PS (212)). The DMA engine (416) may receive such data from the stream switches and write the data to a suitable memory bank or memory banks (412) within the memory module (304).
[0126] The DMA engine (416) may include one or more memory-to-stream interfaces. Through the DMA engine (416), data may be read from a memory bank or memory banks (412) of the memory module (304) and transmitted to other destinations through the stream interfaces. For example, the DMA engine (416) may read data from the memory module (304) and transmit such data to other DPEs that share a boundary with and / or do not share a boundary with the DPE (204-15) by stream switches. The DMA engine (416) may also transmit such data to other subsystems (e.g., NoC (208), hardwired circuit blocks (210), PL (214), and / or PS (212)) by stream switches and the SoC interface block (206).
[0127] In one or more embodiments, the DMA engine (416) may be programmed by a memory-mapped switch (332) within the DPE (204-15). For example, the DMA engine (416) may be controlled by configuration registers (336). The configuration registers (336) may be written using a memory-mapped switch (332) of the DPE interconnect (306). In specific embodiments, the DMA engine (416) may be controlled by a stream switch (326) within the DPE (204-15). For example, the DMA engine (416) may include control registers that can be written by a stream switch (326) connected to control registers. Streams received through the stream switch (326) in the DPE interconnect (306) can be directly connected to the DMA engine (416) in the memory module (304) and / or the core (302) according to the configuration data loaded in the configuration registers (324, 334, and / or 336). Streams can be transmitted from the DMA engine (416) (e.g., the memory module (304)) and / or the core (302) according to the configuration data loaded in the configuration registers (324, 334, and / or 336).
[0128] The memory module (304) may further include a hardware synchronization circuit (420) (abbreviated as “HSC” in FIG. 4). Generally, the hardware synchronization circuit (420) can synchronize the operation of different cores (e.g., cores of neighboring DPEs) that can communicate through the DPE interconnect (306), the core (302) of FIG. 4, the DMA engine (416), and other external masters (e.g., PS (212)). As an exemplary and non-limiting example, the hardware synchronization circuit (420) can synchronize two different cores, stream switches, memory-mapped interfaces, and / or DMAs of DPEs (204-15) and / or different DPEs accessing the same, e.g., shared buffer of the memory module (304).
[0129] If two DPEs are not neighbors, the two DPEs do not have access to a common memory module. In that case, application data may be transmitted via a data stream (the terms “data stream” and “stream” may be used interchangeably from time to time within this disclosure). Accordingly, the local DMA engine may convert transmission from local memory-based transmission to stream-based transmission. In that case, the core (302) and the DMA engine (416) may be synchronized using a hardware synchronization circuit (420).
[0130] The PS (212) can communicate with the core (302) through a memory-mapped switch (332). The PS (212) can access the memory module (304) and the hardware synchronization circuit (420), for example, by initiating memory reads and writes. In another embodiment, the hardware synchronization circuit (420) can also send an interrupt to the PS (212) when the state of the lock changes in order to avoid polling by the PS (212) of the hardware synchronization circuit (420). The PS (212) can also communicate with the DPE (204-15) through stream interfaces.
[0131] In addition to communicating with neighboring DPEs through shared memory modules and with neighboring and / or non-neighboring DPEs through the DPE interconnect (306), the core (302) may include cascade interfaces. In the example of FIG. 4, the core (302) includes cascade interfaces (422 and 424) (abbreviated as “CI” in FIG. 4). The cascade interfaces (422 and 424) may provide direct communication with other cores. As illustrated, the cascade interface (422) of the core (302) receives a direct input data stream from the core of the DPE (204-14). The data stream received through the cascade interface (422) may be provided to the data processing circuitry within the core (302). The cascade interface (424) of the core (302) can directly transmit an output data stream to the core of the DPE (204-16).
[0132] In the example of FIG. 4, each of the cascade interfaces (422) and (424) may include a first-in-first-out (FIFO) interface for buffering. In certain embodiments, the cascade interfaces (422 and 424) may transmit data streams that may be hundreds of bits wide. The specific bit width of the cascade interfaces (422 and 424) is not intended as a limitation. In the example of FIG. 4, the cascade interface (424) is coupled to an accumulator register (436) (abbreviated as “AC” in FIG. 4) within the core (302). The cascade interface (424) may output the contents of the accumulator register (436) and may do so for each clock cycle. The accumulator register (436) may store data generated and / or operated by the data processing circuitry within the core (302).
[0133] In the example of FIG. 4, the cascade interfaces (422 and 424) can be programmed based on configuration data loaded into the configuration registers (324). For example, based on the configuration registers (324), the cascade interface (422) can be enabled or disabled. Similarly, based on the configuration registers (324), the cascade interface (424) can be enabled or disabled. The cascade interface (422) can be enabled and / or disabled independently of the cascade interface (424).
[0134] In one or more other embodiments, the cascade interfaces (422 and 424) are controlled by the core (302). For example, the core (302) may include commands for reading / writing to the cascade interfaces (422 and / or 424). In another example, the core (302) may include a hardwired circuit capable of reading and / or writing to the cascade interfaces (422 and / or 424). In certain embodiments, the cascade interfaces (422 and 424) may be controlled by an entity outside the core (302).
[0135] In the embodiments described herein, the DPEs (204) do not include cache memories. By omitting cache memories, the DPE array (202) can achieve predictable, for example, deterministic performance. Furthermore, since coherency between cache memories located in different DPEs is not required, significant processing overhead is avoided.
[0136] According to one or more embodiments, the cores (302) of the DPEs (204) do not have input interrupts. Therefore, the cores (302) of the DPEs (204) can operate without being interrupted. Omitting input interrupts for the cores (302) of the DPEs (204) also enables the DPE array (202) to achieve predictable, for example, deterministic performance.
[0137] FIG. 5 illustrates another exemplary architecture for a DPE array. In the example of FIG. 5, the SoC interface block (206) provides an interface between the DPEs (204) and other subsystems of the SoC (200). The SoC interface block (206) integrates the DPEs into the device. The SoC interface block (206) can transmit configuration data to the DPEs (204), transmit events from the DPEs (204) to other subsystems, transmit events from other subsystems to the DPEs (204), generate interrupts and transmit them to entities outside the DPE array (202), transmit application data between the other subsystems and the DPEs (204), and / or transmit trace and / or debug data between the other subsystems and the DPEs (204).
[0138] In the example of FIG. 5, the SoC interface block (206) includes a plurality of interconnected tiles. For example, the SoC interface block (206) includes tiles (502, 504, 506, 508, 510, 512, 514, 516, 518, and 520). In the example of FIG. 5, the tiles (502-520) are arranged in rows. In other embodiments, the tiles may be arranged in columns, a grid, or other layouts. For example, the SoC interface block (206) may be implemented as a column of tiles to the left of the DPEs (204), to the right of the DPEs (204), between the columns of the DPEs (204), etc. In another embodiment, the SoC interface block (206) may be located on the DPE array (202). The SoC interface block (206) may be implemented such that the tiles are positioned below the DPE array (202), to the left of the DPE array (202), to the right of the DPE array (202), and / or above the DPE array (202) in any combination. In this regard, FIG. 5 is provided for illustrative purposes rather than for limitation.
[0139] In one or more embodiments, the tiles (502-520) have the same architecture. In one or more other embodiments, the tiles (502-520) may be implemented using two or more different architectures. In certain embodiments, different architectures may be used to implement the tiles within the SoC interface block (206), where each different tile architecture supports communication with a different type of subsystem or combination of subsystems of the SoC (200).
[0140] In the example of FIG. 5, the tiles (502-520) are coupled so that data can be propagated from one tile to another. For example, data can be propagated from tile (502) to tile (520) along the lines of tiles through tiles (504, 506). Similarly, data can be propagated in the reverse direction from tile (520) to tile (502). In one or more embodiments, each of the tiles (502-520) can function as an interface for a plurality of DPEs. For example, each of the tiles (502-520) can function as an interface for a subset of DPEs (204) of the DPE array (202). The subset of DPEs to which each tile provides an interface may be mutually exclusive so that no DPE is provided with an interface by more than one tile of the SoC interface block (206).
[0141] In one example, each of the tiles (502-520) provides an interface to a column of DPEs (204). For the purpose of illustration, tile (502) provides an interface to the DPEs in column A. Tile (504) provides an interface to the DPEs in column B, and so on. In each case, the tile includes a direct connection to an adjacent DPE in the column of DPEs, which is the lowest DPE in this example. Referring to column A, for example, tile (502) is directly connected to DPE (204-1). Other DPEs in column A may communicate with tile (502), but this may be done through the DPE interconnects of the intervening DPEs in the same column.
[0142] For example, the tile (502) may receive data from other sources, such as PS (212), PL (214), and / or other hardwired circuit blocks (210), such as application-specific circuit blocks. The tile (502) may transmit data addressed to DPEs in other columns (e.g., DPEs that are not interfaces with the tile (502)) to the tile (504), while providing these portions of data addressed to DPEs in column A to those DPEs. The tile (504) may perform the same or similar processing, wherein the data received from the tile (502) addressed to DPEs in column B is provided to those DPEs, and the data addressed to DPEs in other columns is transmitted to the tile (506).
[0143] In this way, data can be propagated from tile to tile of the SoC interface block (206) until it reaches a tile acting as an interface to the DPEs to which the data is addressed (e.g., "target DPE(s)"). The tile acting as an interface to the target DPE(s) can send data to the target DPE(s) using the memory-mapped switches of the DPEs and / or the stream switches of the DPEs.
[0144] As mentioned, the use of columns is an exemplary implementation. In other embodiments, each tile of the SoC interface block (206) may provide an interface to a row of DPEs of the DPE array (202). Such a configuration may be used in cases where the SoC interface block (206) is implemented as a column of tiles, whether to the left or right of the DPEs (204), or between the columns of the DPEs (204). In other embodiments, the subset of DPEs to which each tile provides an interface may be any combination of fewer DPEs than all DPEs of the DPE array (202). For example, DPEs (204) may be distributed among the tiles of the SoC interface block (206). The specific physical layout of such DPEs may vary based on the connectivity of the DPEs, such as that established by the DPE interconnects. For example, the tile (502) may provide an interface to the DPEs (204-1, 204-2, 204-11, and 204-12). Another tile of the SoC interface block (206) may provide an interface to four different DPEs.
[0145] FIG. 6 illustrates an exemplary architecture for tiles of an SoC interface block (206). In the example of FIG. 6, two different types of tiles for the SoC interface block (206) are illustrated. Tile (602) is configured to serve as an interface between DPEs and only PL (214). Tile (610) is configured to serve as an interface between DPEs and NoC (208) and between DPEs and PL (214). The SoC interface block (206) may include a combination of tiles using both architectures as exemplified for tile (602) and tile (610), or in another example, may include only tiles having the architecture as exemplified for tile (610).
[0146] In the example of FIG. 6, the tile (602) includes a stream switch (604) connected to the PL interface (606) and to a DPE such as the DPE (204-1) immediately above. The PL interface (606) is connected to BLI (Boundary Logic Interface) circuits (620) and BLI circuits (622) located in the PL (214), respectively. The tile (610) includes a stream switch (612) connected to the NoC and PL interface (614) and to a DPE such as the DPE (204-5) immediately above. The NoC and PL interface (614) is connected to the BLI circuits (624 and 626) of the PL (214) and also to the NoC master unit (NMU) (630) and NoC slave unit (NSU) (632) of the NoC (208).
[0147] In the example of FIG. 6, each stream interface (604) can output six different 32-bit data streams to the DPE connected thereto and receive four different 32-bit data streams from the DPE connected thereto. Each of the PL interface (606), NoC, and PL interface (614) can provide six different 64-bit data streams to PL (214) through BLI (620) and BLI (624), respectively. Generally, each of the BLIs (620, 622, 624, and 626) provides an interface or connection point within PL (214) to which PL interface (606) and / or NoC and PL interface (614) are connected. Each of the PL interface (606), NoC, and PL interface (614) can receive eight different 64-bit data streams from PL (214) through BLI (622) and BLI (624), respectively.
[0148] The NoC and PL interface (614) is also connected to the NoC (208). In the example of FIG. 6, the NoC and PL interface (614) is connected to one or more NMUs (630) and one or more NSUs (632). In one example, the NoC and PL interface (614) can provide two different 128-bit data streams to the NoC (208), each data stream being provided to a different NMU (630). The NoC and PL interface (614) can receive two different 128-bit data streams from the NoC (208), each data stream being received from a different NSU (632).
[0149] The stream switches (604) of adjacent tiles are connected. In one example, the stream switches (604) of adjacent tiles can communicate through four different 32-bit data streams in the left direction and right direction, respectively (e.g., while the tile is on the right or left depending on the case).
[0150] Each of the tiles (602 and 610) may include one or more memory-mapped switches to transmit configuration data. For the purposes of illustration, memory-mapped switches are not illustrated. Memory-mapped switches may be vertically connected, for example, to a memory-mapped switch of the DPE immediately above, to memory-mapped switches of other adjacent tiles of the SoC interface block (206) in the same or similar manner as the stream switches (604), to configuration registers (not illustrated) of the tiles (602 and 610), and / or, in some cases, to the PL interface (608) or the NoC and PL interface (614).
[0151] The various bit widths and numbers of the data streams described in relation to the various switches included in the tiles (602 and / or 610) and / or DPEs (204) of the SoC interface block (206) are provided for illustrative purposes and are not intended to limit the arrangements of the invention described in this disclosure.
[0152] FIG. 7 illustrates an exemplary implementation of a NoC (208). The NoC (208) includes NMUs (702), NSUs (704), a network (714), a NoC peripheral interconnect (NPI) (710), and registers (712). Each NMU (702) is an inlet circuit that connects an endpoint circuit to the NoC (208). Each NSU (704) is an outlet circuit that connects the NoC (208) to an endpoint circuit. The NMUs (702) are connected to the NSUs (704) via the network (714). In one example, the network (714) includes NoC packet switches (NPSs) (706) and routing (708) between the NPSs (706). Each NPS (706) performs switching of NoC packets. The NPSs (706) are connected to each other and to the NMUs (702) and NSUs (704) via routing (708) to implement multiple physical channels. The NPSs (706) also support multiple virtual channels per physical channel.
[0153] The NPI (710) includes circuitry for programming NMUs (702), NSUs (704), and NPSs (706). For example, the NMUs (702), NSUs (704), and NPSs (706) may include registers (712) that determine their functionality. The NPI (710) includes peripheral interconnects coupled to the registers (712) for its own programming to set functionality. The registers (712) of the NoC (208) support interrupts, quality of service (QoS), error handling and reporting, transaction control, power management, and address mapping control. The registers (712) may be initialized to a usable state before being reprogrammed by writing to the registers (712), for example, using write requests. Configuration data for the NoC (208) can be stored in non-volatile memory (NVM), for example, as part of a programming device image (PDI), and can be provided to the NPI (710) to program the NoC (208) and / or other endpoint circuits.
[0154] NMUs (702) are traffic entry points. NSUs (704) are traffic exit points. Endpoint circuits coupled to NMUs (702) and NSUs (704) may be reinforced circuits (e.g., hardwired circuit blocks (210)) or circuits implemented in PL (214). A given endpoint circuit may be coupled to more than one NMU (702) or more than one NSU (704).
[0155] FIG. 8 is a block diagram illustrating connections between endpoint circuits of an SoC (200) through a NoC (208) according to one example. In this example, endpoint circuits (802) are connected to endpoint circuits (804) through the NoC (208). Endpoint circuits (802) are master circuits coupled to the NMUs (702) of the NoC (208). Endpoint circuits (804) are slave circuits coupled to the NSUs (704) of the NoC (208). Each endpoint circuit (802, 804) may be a circuit of the PS (212), a circuit of the PL area (214), or a circuit of another subsystem (e.g., hardwired circuit blocks (210)).
[0156] The network (714) includes a plurality of physical channels (806). The physical channels (806) are implemented by programming the NoC (208). Each physical channel (806) includes one or more NPSs (706) and associated routing (708). The NMU (702) is connected to the NSU (704) through at least one physical channel (806). The physical channel (806) may also have one or more virtual channels (808).
[0157] Connections through the network (714) use a master-slave arrangement. In one example, the most basic connection through the network (714) is a single master connected to a single slave. However, in other examples, more complex structures may be implemented.
[0158] FIG. 9 is a block diagram illustrating a NoC (208) according to another example. In this example, the NoC (208) includes vertical sections (902) (VNoC) and horizontal sections (904) (HNoC). Each VNoC (902) is positioned between PL regions (214). The HNoC (904) is positioned between the PL regions (214) and I / O banks (910) (e.g., transceivers and / or I / O blocks corresponding to hardwired circuit blocks (210)). The NoC (208) is connected to memory interfaces (908) (e.g., hardwired circuit blocks (210)). A PS (212) is coupled to the HNoC (904).
[0159] In this example, the PS (212) includes a plurality of NMUs (702) coupled to the HNoC (904). The VNoC (902) includes both NMUs (702) and NSUs (704) placed in the PL regions (214). Memory interfaces (908) include NSUs (704) connected to the HNoC (904). Both the HNoC (904) and the VNoC (902) include NPSs (706) connected by routing (708). In the VNoC (902), routing (708) extends vertically. In the HNoC (904), routing extends horizontally. In each VNoC (902), each NMU (702) is coupled to an NPS (706). Likewise, each NSU (704) is coupled to an NPS (706). The NPSs (706) are coupled to each other to form a matrix of switches. Some NPSs (706) of each VNoC (902) are coupled to other NPSs (706) of the HNoC (904).
[0160] Although only a single HNoC (904) is illustrated, in other examples, the NoC (208) may include more than one HNoC (904). Furthermore, although two VNoCs (902) are illustrated, the NoC (208) may include more than two VNoCs (902). Although memory interfaces (908) are illustrated as examples, it should be understood that hardwired circuit blocks (210) or other hardwired circuit blocks (210) may be used instead of or in addition to the memory interfaces (908).
[0161] FIG. 10 illustrates an exemplary method (1000) for programming the NoC (208). Although described independently of other subsystems of the SoC (200), the method (1000) may be included and / or used as part of a larger boot or programming process for the SoC (200).
[0162] In block (1002), a Platform Management Controller (PMC) implemented in the SoC (200) receives NoC programming data at boot time. The NoC programming data may be part of the PDI. The PMC is responsible for managing the SoC (200). The PMC can maintain a secure environment during normal operations, boot the SoC (200), and manage the SoC (200).
[0163] In block (1004), the PMC loads NoC programming data into registers (712) via NPI (710) to create physical channels (806). In one example, the programming data may also include information for configuring routing tables in NPSs (706). In block (1006), the PMC boots the SoC (200). In this way, the NoC (208) includes configuration information for physical channels (806) between NMUs (702) and NSUs (704). The remaining configuration information for the NoC (208) may be received during runtime as described further below. In another example, all or part of the configuration information described below as being received during runtime may be received at boot time.
[0164] FIG. 11 illustrates an exemplary method (1100) for programming a NoC (208). In block (1102), the PMC receives NoC programming data during runtime. In block (1104), the PMC loads the programming data into the NoC registers (712) via the NPI (710). In one example, in block (1106), the PMC configures routing tables in the NPSs (706). In block (1108), the PMC configures QoS paths through the physical channels (806). In block (1110), the PMC configures address space mappings. In block (1112), the PMC configures the inlet / outlet interface protocol, width, and frequency. The QoS paths, address space mappings, routing tables, and inlet / outlet configurations are discussed further below.
[0165] FIG. 12 illustrates an exemplary data path (1200) through a NoC (208) between endpoint circuits. The data path (1200) includes an endpoint circuit (1202), an AXI master circuit (1204), an NMU (1206), NPSs (1208), an NSU (1210), an AXI slave circuit (1212), and an endpoint circuit (1214). The endpoint circuit (1202) is coupled to the AXI master circuit (1204). The AXI master circuit (1204) is coupled to the NMU (1206). In another example, the AXI master circuit (1204) is part of the NMU (1206).
[0166] The NMU (1206) is coupled to the NPS (1208). The NPSs (1208) are coupled to each other to form a chain of NPSs (1208) (e.g., a chain of five NPSs (1208) in this example). Generally, there is at least one NPS (1208) between the NMU (1206) and the NSU (1210). The NSU (1210) is coupled to one of the NPSs (1208). The AXI slave circuit (1212) is coupled to the NSU (1210). In another example, the AXI slave circuit (1212) is part of the NSU (1210). The endpoint circuit (1214) is coupled to the AXI slave circuit (1212).
[0167] The endpoint circuits (1202 and 1214) may each be a circuit configured in an enhanced circuit (e.g., a PS circuit, a hardwired circuit (210), one or more DPEs (204)) or a PL (214). The endpoint circuit (1202) functions as a master circuit and transmits read / write requests to the NMU (1206). In this example, the endpoint circuits (1202 and 1214) communicate with the NoC (208) using the AXI protocol. While AXI is described in this example, it should be understood that the NoC (208) may be configured to receive communications from the endpoint circuits using other types of protocols known in the art. For clarity as an example, the NoC (208) is described herein as supporting the AXI protocol. The NMU (1206) relays the request through a set of NPSs (1208) to reach the destination NSU (1210). The NSU (1210) transmits a request to an attached AXI slave circuit (1212) to process data and distribute it to an endpoint circuit (1214). The AXI slave circuit (1212) can send read / write responses back to the NSU (1210). The NSU (1210) can forward responses to the NMU (1206) via a set of NPSs (1208). The NMU (1206) communicates responses to the AXI master circuit (1204), and the AXI master circuit (1204) distributes the data to the endpoint circuit (1202).
[0168] FIG. 13 illustrates an exemplary method (1300) for processing read / write requests and responses. The method (1300) begins in block (1302), where the endpoint circuit (1202) transmits a request (e.g., a read request or a write request) to the NMU (1206) via the AXI master (1204). In block (1304), the NMU (1206) processes the response. In one example, the NMU (1206) performs asynchronous crossing and rate-matching between the clock domains of the NoC (208) and the endpoint circuit (1202). The NMU (1206) determines the destination address of the NSU (1210) based on the request. The NMU (1206) may perform address remapping if virtualization is used. The NMU (1206) also performs AXI conversion of the request. The NMU (1206) further packetizes the request into a stream of packets.
[0169] In block (1306), the NMU (1206) transmits packets for the request to the NPSs (1208). Each NPS (1208) performs a table lookup for the target output port based on the destination address and routing information. In block (1308), the NSU (1210) processes the packets of the request. In one example, the NSU (1210) depackets the request, performs AXI conversion, and performs asynchronous crossing and rate-matching from the NoC clock domain to the clock domain of the endpoint circuit (1214). In block (1310), the NSU (1210) transmits the request to the endpoint circuit (1214) through the AXI slave circuit (1212). The NSU (1210) can also receive a response from the endpoint circuit (1214) through the AXI slave circuit (1212).
[0170] In block (1312), the NSU (1210) processes the response. In one example, the NSU (1210) performs asynchronous crossing and rate-matching from the clock domain of the endpoint circuit (1214) and the clock domain of the NoC (208). The NSU (1210) further packetizes the response into a stream of packets. In block (1314), the NSU (1210) transmits the packets through the NPSs (1208). Each NPS (1208) performs a table lookup for the target output port based on the destination address and routing information. In block (1316), the NMU (1206) processes the packets. In one example, the NMU (1206) reverse packetizes the response, performs an AXI conversion, and performs asynchronous crossing and rate-matching from the NoC clock domain to the clock domain of the endpoint circuit (1202). In block (1318), the NMU (1206) transmits the response to the endpoint circuit (1202) through the AXI master circuit (1204).
[0171] FIG. 14 illustrates an exemplary implementation of an NMU (702). The NMU (702) includes an AXI master interface (1402), a packetization circuit (1404), an address map (1406), a de-packetization circuit (1408), a QoS circuit (1410), a VC mapping circuit (1412), and a clock management circuit (1414). The AXI master interface (1402) provides the NMU (702) with an AXI interface for an endpoint circuit. In other examples, different protocols may be used, and thus the NMU (702) may have different master interfaces that comply with the selected protocol. The NMU (702) routes inbound traffic to the packetization circuit (1404), and the packetization circuit (1404) generates packets from the inbound data. The packetization circuit (1404) determines the destination ID used to route packets from the address map (1406). The QoS circuit (1410) may provide incoming rate control to control the injection rate of packets to the NoC (208). The VC mapping circuit (1412) manages QoS virtual channels through each physical channel. The NMU (702) may be configured to select which virtual channels packets are mapped to. The clock management circuit (1414) performs rate matching and asynchronous data crossing to provide an interface between the AXI clock domain and the NoC clock domain. The de-packetization circuit (1408) is configured to receive return packets from the NoC (208) and de-packetize the packets for output by the AXI master interface (1402).
[0172] FIG. 15 illustrates an exemplary implementation of an NMU (704). The NSU (704) includes an AXI slave interface (1502), a clock management circuit (1504), a packetization circuit (1508), a de-packetization circuit (1506), and a QoS circuit (1510). The AXI slave interface (1502) provides the NMU (704) with an AXI interface to an endpoint circuit. In other examples, different protocols may be used, and thus the NSU (704) may have different slave interfaces that comply with the selected protocol. The NSU (704) routes inbound traffic from the NoC (208) to the de-packetization circuit (1506), and the de-packetization circuit (1506) generates de-packetized data. The clock management circuit (1504) performs rate matching and asynchronous data crossing to provide an interface between the AXI clock domain and the NoC clock domain. The packetization circuit (1508) is configured to receive return data from the slave interface (1502) and to packetize the return data for transmission through the NoC (208). The QoS circuit (1510) can provide inlet rate control to control the injection rate of packets to the NoC (208).
[0173] FIG. 16 illustrates an exemplary software architecture executable by the system described in relation to FIG. 1. For example, the architecture of FIG. 16 may be implemented as one or more of the program modules (120) of FIG. 1. The software architecture of FIG. 16 includes a DPE compiler (1602), a NoC compiler (1604), and a hardware compiler (1606). FIG. 16 illustrates examples of various types of design data that may be exchanged between the compilers during operation (e.g., while performing a design flow for implementing an application on the SoC (200)).
[0174] A DPE compiler (1602) may generate one or more binaries from an application that can be loaded into one or more DPEs and / or a subset of DPEs (204) of a DPE array (202). Each binary may contain object code executable by the core(s) of the DPE(s), optionally application data, and configuration data for the DPEs. A NoC compiler (1604) may generate a binary containing configuration data that is loaded into a NoC (208) to generate internal data paths for the application. A hardware compiler (1606) may compile the hardware portion of the application to generate a configuration bitstream for implementation in a PL (214).
[0175] FIG. 16 illustrates an example of how a DPE compiler (1602), a NoC compiler (1604), and a hardware compiler (1606) communicate with each other during operation. The individual compilers communicate in a coordinated manner by exchanging design data and converging into a solution. The solution is an implementation of an application within the SoC (200) that satisfies design metrics and constraints and includes common interfaces that allow various heterogeneous subsystems of the SoC (200) to communicate.
[0176] As defined in this disclosure, the term “design metric” defines the purpose or requirement of an application to be implemented in the SoC (200). Examples of design metrics include, but are not limited to, power consumption requirements, data throughput requirements, timing requirements, etc. Design metrics may be provided through user input, files, or other methods defining the application’s higher or system-level requirements. As defined in this disclosure, “design constraint” is a requirement that an EDA tool may or may not follow to achieve a design metric or requirement. Design constraints may be specified as compiler directives and may typically specify lower-level requirements or suggestions that an EDA tool (e.g., a compiler) must follow. Design constraints may be specified through user input(s), files containing one or more design constraints, command-line input, etc.
[0177] In one aspect, the DPE compiler (1602) can generate an SoC interface block solution and a logical architecture for an application. The DPE compiler (1602) can generate a logical architecture based on high-level user-defined metrics for the software portion of the application to be implemented, for example, in the DPE array (202). Examples of metrics may include, but are not limited to, data throughput, latency, resource utilization, and power consumption. Based on the metrics and the application (e.g., specific nodes to be implemented in the DPE array (202)), the DPE compiler (1602) can generate a logical architecture.
[0178] A logical architecture is a file or data structure that can specify hardware resource block information required by various parts of an application. For example, a logical architecture can specify the number of DPEs (204) required to implement the software part of the application, any Intellectual Property (IP) cores required in the PL (214) to communicate with the DPE array (202), any connections that need to be routed through the NoC (208), and port information for the IP cores of the DPE array (202), NoC (208), and PL (214). An IP core is a reusable block of logic, cells, or IC layout design that can be used in circuit design as a reusable block of circuitry capable of performing a specific function or operation. An IP core can be specified in a format that can be integrated into the circuit design for implementation within the PL (214). Although the present disclosure refers to various types of cores, the term "core" without any other modifiers is intended to refer collectively to these different types of cores.
[0179] Example 1 of the disclosure, located at the end of the detailed description, illustrates an exemplary schema that can be used to specify a logical architecture for an application. Example 1 illustrates information about various types included in the logical architecture for an application. In one aspect, a hardware compiler (1606) may implement the hardware portion of an application based on or using a logical architecture and an SoC interface block solution, as opposed to using the application itself.
[0180] Port information for the DPE array (202), and port information for the IP cores of the NoC (208) and PL (214) may include the logical configuration of the ports, such as whether each port is a stream data port, a memory-mapped port, or a parameter port, and whether the ports are masters or slaves. Other examples of port information for the IP cores include the data width and operating frequency of the ports. Connectivity between the IP cores of the DPE array (202), NoC (208), and PL (214) may be specified as logical connections between the ports of individual hardware resource blocks specified in the logical architecture.
[0181] The SoC interface block solution is a data structure or file that specifies the mapping of connections to the inside and outside of the DPE array (202) to physical data paths (e.g., physical resources) of the SoC interface block (206). For example, the SoC interface block solution maps specific logical connections used for data transfer to and from the inside and outside of the DPE array (202) to specific stream channels of the SoC interface block (206), such as specific tiles, stream switches, and / or stream switch interfaces (e.g., ports) of the SoC interface block (206). Example 2, which follows Example 1 at the end of the detailed description, illustrates an exemplary schema for the SoC interface block solution for an application.
[0182] In one aspect, the DPE compiler (1602) can analyze or simulate data traffic through the NoC (208) based on the application and logical architecture. The DPE compiler (1602) can provide the software part of the application, such as data delivery requirements of "NoC traffic," to the NoC compiler (1604). The NoC compiler (1604) can generate routing for data paths through the NoC (208) based on the NoC traffic received from the DPE compiler (1602). The result from the NoC compiler (1604), described as a "NoC solution," can be provided to the DPE compiler (1602).
[0183] In one aspect, the NoC solution may be an initial NoC solution that specifies only the inlet and / or outlet points of the NoC (208) to which the nodes of the application connected to the NoC (208) will be connected. For example, more detailed routing and / or configuration data for data paths within the NoC (208) (e.g., between inlet and outlet points) may be excluded from the NoC solution for the convergence of compilers. Example 3, which follows Example 2 at the end of the detailed description, illustrates an exemplary schema for a NoC solution for an application.
[0184] A hardware compiler (1606) may operate on a logical architecture to implement the hardware portion of an application in PL (214). If the hardware compiler (1606) cannot generate an implementation of the hardware portion of an application (e.g., using a logical architecture) that satisfies established design constraints (e.g., timing, power, data throughput, etc.), the hardware compiler (1606) may generate one or more SoC interface block constraints and / or receive one or more user-defined SoC interface block constraints. The hardware compiler (1606) may provide SoC interface block constraints to the DPE compiler (1602) as requests. The SoC interface block constraints effectively remap one or more portions of the logical architecture to different stream channels of the SoC interface block (206). The SoC interface block constraints provided by the hardware compiler (1606) are more advantageous for the hardware compiler (1606) to generate an implementation of the hardware portion of an application in PL (214) that satisfies design metrics. Example 4, which follows Example 3 at the end of the detailed description, illustrates exemplary constraints for NoC and / or SoC interface blocks for an application.
[0185] In another aspect, the hardware compiler (1606) may also generate NoC traffic based on the application and logical architecture and provide it to the NoC compiler (1604). The hardware compiler (1606) may analyze or simulate the hardware part of the application to determine the data traffic generated by the hardware part of the design to be delivered to the PS (212), DPE array (202), and / or other parts of the SoC (200) through the NoC (208), for example. The NoC compiler (1604) may generate and / or update the NoC solution based on the information received from the hardware compiler (1606). The NoC compiler (1604) may provide the NoC solution or an updated version thereof to the hardware compiler (1606) and also to the DPE compiler (1602). In this regard, the DPE compiler (1602) may update the SoC interface block solution in response to receiving a NoC solution or an updated NoC solution from the NoC compiler (1604) and / or in response to receiving one or more SoC interface block constraints from the hardware compiler (1606), and may provide the updated solution to the hardware compiler (1606). The DPE compiler (1602) generates an updated SoC interface block solution based on the updated NoC solution from the NoC compiler (1604) and / or the SoC interface block constraint(s) received from the hardware compiler (1606).
[0186] It should be recognized that the data flows between compilers illustrated in the example of FIG. 16 are merely for illustrative purposes. In this regard, the exchange of information between compilers may be performed at various stages of the exemplary design flows described in this disclosure. In other aspects, the exchange of design data between compilers may be performed iteratively so that each compiler can continuously refine the implementation of the part of the application handled by that compiler based on information received from other compilers in order to converge to a solution.
[0187] In one particular example, the hardware compiler (1606), after receiving a logical architecture and SoC interface block solution from the DPE compiler (1602) and a NoC solution from the NoC compiler (1604), may determine that it is not possible to generate an implementation of the hardware part of the application that satisfies the set design metrics. The initial SoC interface block solution generated by the DPE compiler (1602) is generated based on the DPE compiler (1602)'s knowledge of the part of the application to be implemented in the DPE array (202). Similarly, the initial NoC solution generated by the NoC compiler (1604) is generated based on the initial NoC traffic provided to the NoC compiler (1604) by the DPE compiler (1602). Example 5, following Example 4 at the end of the detailed description, illustrates an exemplary schema for the NoC traffic for the application. It should be understood that while schemas are used in Examples 1-5, other formatting and / or data structures may be used to specify the illustrated information.
[0188] The hardware compiler (1606) attempts to perform an implementation flow for the hardware part of the application, which includes synthesizing (if necessary), placing, and routing the hardware part. Thus, the initial SoC interface block solution and the initial NoC solution may result in placements and / or routes within the PL (214) that do not satisfy the set timing constraints. In other cases, the SoC interface block solution and the NoC solution may not have a sufficient number of physical resources, such as wires, to accommodate the data to be transmitted, and thus may cause congestion in the PL (214). In such cases, the hardware compiler (1606) may generate one or more different SoC interface block constraints and / or receive one or more user-defined SoC interface block constraints and provide the SoC interface block constraints to the DPE compiler (1602) as a request to generate the SoC interface block solution. Likewise, the hardware compiler (1606) may generate one or more different NoC constraints and / or receive one or more user-defined NoC constraints and provide the NoC constraints to the NoC compiler (1604) as a request to regenerate the NoC solution. In this way, the hardware compiler (1606) invokes the DPE compiler (1602) and / or the NoC compiler (1604).
[0189] The DPE compiler (1602) takes SoC interface block constraints received from the hardware compiler (1606), updates the SoC interface block solution using the received SoC interface block constraints where possible, and can provide the updated SoC interface block solution back to the hardware compiler (1606). Similarly, the NoC compiler (1604) takes NoC constraints received from the hardware compiler (1606), updates the NoC solution using the received NoC constraints where possible, and can provide the updated NoC solution back to the hardware compiler (1606). Then, the hardware compiler (1606) can continue the implementation flow to generate the hardware part of the application for implementation within the PL (214) using the updated SoC interface block solution received from the DPE compiler (1602) and the updated NoC solution received from the NoC compiler (1604).
[0190] In one aspect, a hardware compiler (1606) that invokes a DPE compiler (1602) and / or a NoC compiler (1604) by providing one or more SoC interface block constraints and one or more NoC constraints, respectively, may be part of a verification process. The hardware compiler (1606) searches from the DPE compiler (1602) and / or the NoC compiler (1604) for verification that the NoC constraints and SoC interface block constraints provided by the hardware compiler (1606) can be used or incorporated into a routable SoC interface block solution and / or NoC solution.
[0191] FIG. 17a illustrates an example of an application (1700) mapped to an SoC (200) using the system described in relation to FIG. 1. For the sake of illustration, only a subset of different subsystems of the SoC (200) is illustrated. The application (1700) includes nodes A, B, C, D, E, and F having illustrated connectivity. Example 6 below illustrates exemplary source code that can be used to specify the application (1700). Example 6
[0192] In one aspect, the application (1700) is specified as a data flow graph comprising multiple nodes. Each node represents a computation corresponding to a function contrasting with a single instruction. The nodes are interconnected by edges representing data flows. The hardware implementation of a node may be executed only in response to receiving data from each of the inputs to that node. Nodes are generally executed in a non-blocking manner. The data flow graph specified by the application (1700) represents a parallel specification to be implemented in the SoC (200) in contrast to a sequential program. The system may operate on the application (1700) (e.g., in the form of a graph as exemplified in Example 1) to map various nodes to appropriate subsystems of the SoC (200) for the implementation herein.
[0193] In one example, the application (1700) is specified in a high-level programming language (HLL) such as C and / or C++. As mentioned, although it is specified in an HLL typically used to generate sequential programs, the application (1700), which is a data flow graph, is of parallel specification. The system may provide a class library used to build data flow graphs and the corresponding application (1700). The data flow graph is defined by the user and compiled into the architecture of the SoC (200). The class library may be implemented as a helper library having predefined classes and constructors for graphs, nodes, and edges that can be used to build the application (1700). The application (1700) is effectively executed on the SoC (200) and includes delegated objects that are executed in the PS (212) of the SoC (200). Objects of the application (1700) running in the PS (212) can be used to direct and monitor actual calculations running on the SoC (200), for example in the PL (214), in the DPE array (202), and / or in the hardwired circuit blocks (210).
[0194] According to the arrangements of the invention described in this disclosure, accelerators (e.g., PL nodes) can be represented as objects in a data flow graph (e.g., applications). The system can automatically synthesize PL nodes for implementation in PL (214) and connect the synthesized PL nodes. In contrast, in conventional EDA systems, users specify applications for hardware acceleration that utilize sequential semantics. Hardware acceleration functions are specified through function calls. The interface to the hardware acceleration function (e.g., PL nodes in this example) is defined by the function call and various arguments provided in the function call, in contrast to the connections in the data flow graph.
[0195] As illustrated in the source code of Example 6, nodes A and F are designated for implementation in PL (214), whereas nodes B, C, D, and E are designated for implementation within the DPE array (202). The connectivity of the nodes is specified by data transfer edges in the source code. The source code of Example 6 also specifies a control program and a top-level test bench running in PS (212).
[0196] Referring to FIG. 17a, the application (1700) is mapped onto the SoC (200). As illustrated, nodes A and F are mapped to PL (214). The shaded DPEs (204-13 and 204-14) represent the DPEs (204) to which nodes B, C, D, and E are mapped. For example, nodes B and C are mapped onto DPE (204-13), while nodes D and E are mapped onto DPE (204-4). Nodes A and F are implemented in PL (214) and are connected to DPEs (204-13 and 204-44) through routing via specific tiles and switches of PL (214) and SoC interface block (206), switches within the DPE interconnection of the interposed DPEs (204), and specific memories of selected neighboring DPEs (204).
[0197] The binary generated for DPE (204-13) includes object code necessary for DPE (204-13) to implement calculations corresponding to nodes B and C, and configuration data for creating data paths between DPE (204-13) and DPE (204-14) and between DPE (204-13) and DPE (204-3). The binary generated for DPE (204-4) includes object code necessary for DPE (204-4) to implement calculations corresponding to nodes D and E, and configuration data for establishing data paths between DPE (204-14) and DPE (204-5).
[0198] Other binaries are generated for other DPEs (204), such as DPEs (204-3, 204-5, 204-6, 204-7, 204-8, and 204-9), to connect DPEs (204-13) and DPE (204-4) to the SoC interface block (206). In a recognizable manner, such binaries will contain arbitrary object code where such other DPEs (204) implement different computations (where they have nodes of the application assigned to them).
[0199] In this example, the hardware compiler (1606) cannot generate an implementation of the hardware part that satisfies the timing constraints due to the long route connecting the DPE (204-14) and node F. Within this disclosure, a specific state of the implementation of the hardware part of the application may be referred to as a state of the hardware design, where the hardware design is generated and / or updated throughout the implementation flow. The SoC interface block solution may, for example, assign a signal crossing for node F to a tile of the SoC interface block under the DPE (204-9). In that case, the hardware compiler (1606) may provide the DPE compiler (1602) with a requested SoC interface block constraint requesting that the crossing through the SoC interface block (206) for node F be moved closer to the DPE (204-4). For example, a requested SoC interface block constraint from the hardware compiler (1606) may request that logical connections to the DPE (204-4) be mapped to the tile immediately below the DPE (204-4) within the SoC interface block (206). This remapping would allow the hardware compiler to place node F much closer to the DPE (204-4) to improve timing.
[0200] FIG. 17b illustrates another exemplary mapping of an application (1700) to an SoC (200). FIG. 17b illustrates a more detailed and alternative example than that illustrated in FIG. 17a. FIG. 17b illustrates, for example, the mapping of nodes of an application (1700) to specific DPEs (204) of a DPE array (202), connectivity established between the DPEs (204) to which the nodes of the application (1700) are mapped, memory allocation of memory modules of the DPEs (204) to the nodes of the application (1700), and mapping of data transfers to the memory and core interfaces (e.g., 428, 430, 432, 434, 402, 404, 406, and 408) of the DPEs (204) (indicated by double arrows) and / or stream switches of the DPE interconnect (306), which is performed by the DPE compiler (1602).
[0201] In the example of FIG. 17b, memory modules (1702, 1706, 1710, 1714, and 1718) are illustrated together with cores (1704, 1708, 1712, 1716, and 1720). The cores (1704, 1708, 1712, 1716, and 1720) each include program memories (1722, 1724, 1726, 1728, and 1730). In the upper row, the core (1704) and memory module (1706) form a DPE (204), whereas the core (1708) and memory module (1710) form a different DPE (204). In the lower row, the memory module (1714) and the core (1716) form a DPE (204), whereas the memory (1718) and the core (1720) form a different DPE (204).
[0202] As illustrated, nodes A and F are mapped to PL (214). Node A is connected to the memory banks (e.g., shaded portions of memory banks) of the memory module (1702) through the mediators and stream switches of the memory module (1702). Nodes B and C are mapped to the core (1704). Instructions for implementing nodes B and C are stored in program memory (1722). Nodes D and E are mapped to the core (1716) along with instructions for implementing nodes D and E stored in program memory (1728). Node B is allocated and accesses the shaded portions of the memory banks of the memory module (1702) through core-memory interfaces, whereas Node C is allocated and accesses the shaded portions of the memory banks of the memory module (1706) through core-memory interfaces. Nodes B, C, and E can be allocated and accessed in shaded portions of memory banks of the memory module (1714) through core-memory interfaces. Node D can access shaded portions of memory banks of the memory module (1718) through core-memory interfaces. Node F is connected to the memory module (1718) through mediators and stream switches.
[0203] FIG. 17b illustrates that connectivity between nodes of an application can be implemented using memory and / or core interfaces that share memory between cores and using DPE interconnects (306).
[0204] FIG. 18 illustrates an exemplary implementation of another application mapped to the SoC (200). For the sake of illustration, only a subset of different subsystems of the SoC (200) is illustrated. In this example, connections to nodes A and F, each implemented in PL (214), are routed through NoC (208). NoC (208) includes inlet / outlet points (1802, 1804, 1806, 1808, 1810, 1812, 1814, and 1816) (e.g., NMUs / NSUs). The example in FIG. 18 illustrates a case where node A is placed relatively close to the inlet / outlet point (1802), whereas node F, accessing volatile memory (134), has a long route through PL (214) to reach the inlet / outlet point (1816). If the hardware compiler (1606) is unable to place node F closer to the inlet / outlet point (1816), the hardware compiler (1606) may request an updated NoC solution from the NoC compiler (1604). In that case, the hardware compiler (1606) may invoke the NoC compiler (1604) with a NoC constraint to generate an updated NoC solution that specifies a different inlet / outlet point for node F, such as inlet / outlet point (1812). The different inlet / outlet point for node F allows the hardware compiler (1606) to place node F closer to the newly specified inlet / outlet point specified in the updated NoC solution and to utilize faster data paths available in the NoC (208).
[0205] FIG. 19 illustrates an exemplary software architecture (1900) executable by the system described in relation to FIG. 1. For example, the architecture (1900) may be implemented as one or more of the program modules (120) of FIG. 1. In the example of FIG. 19, the application (1902) is intended for implementation within the SoC (200).
[0206] In the example of FIG. 19, the user can interact with a user interface (1906) provided by the system. When interacting with the user interface (1906), the user can specify or provide an application (1902), performance and partitioning constraints (1904) for the application (1902), and a base platform (1908).
[0207] The application (1902) may include a plurality of different parts, each corresponding to a different subsystem available in the SoC (200). The application (1902) may be specified, for example, as described in connection with Example 6. The application (1902) includes a software part to be implemented in the DPE array (202) and a hardware part to be implemented in the PL (214). The application (1902) may optionally include an additional software part to be implemented in the PS (212) and a part to be implemented in the NoC (208).
[0208] Partitioning constraints (of performance and partitioning constraints (1904)) optionally specify the location or subsystem where various nodes of the application (1902) are to be implemented. For example, partitioning constraints may indicate on a node-by-node basis for the application (1902) whether a node should be implemented in the DPE array (202) or in the PL (214). In other examples, location constraints may provide more specific or detailed information to the DPE compiler (1602) to perform mapping of kernels to DPEs, networks or data flows to stream switches, and buffers to memory modules of DPEs and / or banks of memory modules.
[0209] As an exemplary example, the implementation of an application may require specific mappings. For instance, in an application where multiple copies of the kernel are implemented in a DPE array and each copy of the kernel operates on different data sets simultaneously, it is desirable that the data sets be located at the same relative address (location in memory) for each copy of the kernel running on different DPEs in the DPE array. This can be accomplished using location constraints. If these conditions are not maintained by the DPE compiler (1602), each copy of the kernel must be programmed individually or independently rather than replicating the same programming across multiple different DPEs in the DPE array.
[0210] Another exemplary example is imposing position constraints on applications that utilize cascade interfaces between DPEs. Since cascade interfaces flow in one direction within each row, it may be desirable to ensure that the beginning of a chain of DPEs coupled using cascade interfaces does not start at a DPE with a missing cascade interface (e.g., a corner DPE) or at a location that cannot be easily replicated elsewhere in the DPE array (e.g., the last DPE in the row). Position constraints can cause the beginning of the application's chain of DPEs to start at a specific DPE.
[0211] Performance constraints (of performance and partitioning constraints (1904)) may specify various metrics such as power requirements, latency requirements, timing and / or data throughput to be achieved by the implementation of the node, whether in the DPE array (202) or in the PL (214).
[0212] The base platform (1908) is a description of an infrastructure circuitry to be implemented in the SoC (200), and this infrastructure circuitry interacts with and / or is connected to a circuitry on a circuit board to which the SoC (200) is coupled. The base platform (1908) is synthesizable. The base platform (1908) specifies, for example, a circuitry to be implemented within the SoC (200), which receives signals from outside the SoC (200) (e.g., outside the SoC (200)) and provides signals to systems and / or circuitry outside the SoC (200). For example, the base platform (1908) may specify circuit resources such as a PCIe (Peripheral Component Interconnect Express) node for communicating with the host system (102) and / or computing node (100) of FIG. 1, memory controllers or controllers for accessing volatile memory (134) and / or non-volatile memory (136), and / or other resources such as internal interfaces coupling a DPE array (202) and / or PL (214) to the PCIe node. The circuits specified by the base platform (1908) are available for any application that can be implemented in the SoC (200) when a specific type of circuit board is given. In this regard, the base platform (1908) is specific to the specific circuit board to which the SoC (200) is coupled.
[0213] In one example, the partitioner (1910) can separate different parts of the application (1902) based on the subsystems of the SoC (200) on which each part of the application (1902) is to be implemented. In an exemplary implementation, the partitioner (1910) is implemented as a user-directed tool that provides an input indicating which of the different parts of the application (1902) (e.g., nodes) correspond to each of the different subsystems of the SoC (200). For example, the provided input may be performance and partitioning constraints (1904). For example, the partitioner (1910) partitions the application (1902) into a PS portion (1912) to be executed on the PS (212), a DPE array portion (1914) to be executed on the DPE array (202), a PL portion (1916) to be implemented on the PL (214), and a NoC portion (1936) to be implemented on the NoC (208). In one aspect, the partitioner (1910) may create each of the PS portion (1912), the DPE array portion (1914), the PL portion (1916), and the NoC portion (1936) as separate files or separate data structures.
[0214] As described, each of the different parts corresponding to different subsystems is processed by a different compiler specific to the subsystem. For example, a PS compiler (1918) may compile the PS part (1912) to generate one or more binaries containing object code executable by the PS (212). A DPE compiler (1602) may compile the DPE array part (1914) to generate one or more binaries containing object code, application data, and / or configuration data executable by a different DPE (204). A hardware compiler (1606) may perform an implementation flow on the PL part (1916) to generate a configuration bitstream that can be loaded into the SoC (200) to implement the PL part (1916) in the PL (214). As defined herein, the term “implementation flow” means a process in which placement, routing, and optionally synthesis are performed. The NoC compiler (1604) can generate binary specified configuration data for the NoC (208), which, when loaded into the NoC (208), generates data paths in the NoC (208) that connect various masters and slaves of the application (1902). These different outputs generated by the compilers (1918, 1602, 1604, and / or 1606) are exemplified as binaries and configuration bitstreams (1924).
[0215] In certain implementations, certain compilers among the compilers (1918, 1602, 1604, and / or 1606) may communicate with each other during operation. By communicating at various stages during the design flow running on the application (1902), the compilers (1918, 1602, 1604, and / or 1606) may converge to a solution. In the example of FIG. 19, the DPE compiler (1602) and the hardware compiler (1606) may communicate during operation while compiling parts (1914 and 1916) of the application (1902), respectively. The hardware compiler (1606) and the NoC compiler (1604) may communicate during operation while compiling parts (1916 and 1936) of the application (1902), respectively. The DPE compiler (1602) may also invoke the NoC compiler (1604) to obtain the NoC routing solution and / or the updated NoC routing solution.
[0216] The resulting binaries and configuration bitstreams (1924) may be provided to any of the various targets. For example, the resulting binaries and configuration bitstream(s) (1924) may be provided to a simulation platform (1926), a hardware emulation platform (1928), an RTL simulation platform (1930), and / or a target IC (1932). In the case of the RTL simulation platform (1930), the hardware compiler (1922) may be configured to output RTL for the PL portion (1916) that can be simulated on the RTL simulation platform (1930).
[0217] Results obtained from the implementation of the application (1902) of the simulation platform (1926), emulation platform (1928), RTL simulation platform (1930), and / or target IC (1932) may be provided to a performance profiler and debugger (1934). Results from the performance profiler and debugger (1934) may be provided to a user interface (1906), whereby the user can view the results of running and / or simulating the application (1902).
[0218] FIG. 20 illustrates an exemplary method (2000) for performing a design flow to implement an application on an SoC (200). The method (2000) can be performed by a system described in relation to FIG. 1. The system can execute a software architecture described in relation to FIG. 16 or FIG. 19.
[0219] In block (2002), the system receives an application. The application may specify a software part for implementation within the DPE array (202) of the SoC (200) and a hardware part for implementation within the PL (214) of the SoC (200).
[0220] In block (2004), the system can generate a logical architecture for the application. For example, a DPE compiler (1602) executed by the system can generate a logical architecture based on the software portion of the application to be implemented in the DPE array (202) and any high-level user-defined metrics. The DPE compiler (1602) can also generate an SoC interface block solution that specifies the mapping of connections to and from the DPE array (202) to the physical data paths of the SoC interface block (206).
[0221] In another aspect, when generating a logical architecture and SoC interface block solution, the DPE compiler (1602) can generate an initial mapping of the application nodes to specific DPEs (204) to be implemented in the DPE array (202) (referred to as "DPE nodes"). The DPE compiler (1602) optionally generates an initial mapping and routing of the application's global memory data structures to global memory (e.g., volatile memory (134)) by providing NoC traffic to global memory to the NoC compiler (1604). As discussed, the NoC compiler (1604) can generate a NoC solution from the received NoC traffic. Using the initial mappings and routings, the DPE compiler (1602) can simulate the DPE part to verify the initial implementation of the DPE part. The DPE compiler (1602) can output the data generated by the simulation to a hardware compiler (1606) corresponding to each stream channel used in the SoC interface block solution.
[0222] In one aspect, generating a logical architecture as performed by the DPE compiler (1602) implements the partitioning previously described in relation to FIG. 19. Various exemplary schemas illustrate how the different compilers of FIG. 19 (DPE compiler (1602), hardware compiler (1606), and NoC compiler (1604)) exchange decisions and constraints while compiling parts of the application assigned to each individual compiler. Various exemplary schemas further illustrate how decisions and / or constraints are logically applied across the various subsystems of the SoC (200).
[0223] In block (2006), the system can construct a block diagram of the hardware part. For example, a hardware compiler (1606) executed by the system can generate the block diagram. The block diagram integrates the hardware part of the application with the base platform for the SoC (200) specified by the logical architecture. For example, the hardware compiler (1606) can connect the hardware part and the base platform when generating the block diagram. Furthermore, the hardware compiler (1606) can generate the block diagram to connect IP cores corresponding to the hardware part of the application to the SoC interface block based on the SoC interface block solution.
[0224] For example, each node of the hardware portion of an application specified by a logical architecture may be mapped to a specific RTL core (e.g., a user-provided or user-specified portion of a custom RTL) or an available IP core. Using the mappings of nodes to the cores specified by the user, the hardware compiler (1606) may construct a block diagram to specify various circuit blocks of the base platform, any IP cores of the PL (214) required to interface with the DPE array (202) for each logical architecture, and / or any additional user-specified IP cores and / or RTL cores to be implemented in the PL (214). Examples of additional IP cores and / or RTL cores that may be manually inserted by the user include, but are not limited to, data-width conversion blocks, hardware buffers and / or clock domain logic. In one aspect, each block of the block diagram may correspond to a specific core (e.g., a circuit block) to be implemented in the PL (214). The block diagram specifies the connectivity of the cores to be implemented in the PL, and the connectivity of the cores with the physical resources of the NoC (208) and / or SoC interface block (206), which is determined from the SoC interface block solution and logical architecture.
[0225] In one aspect, the hardware compiler (1606) may also generate NoC traffic for each logical architecture and run the NoC compiler (1604) to obtain a NoC solution, thereby creating logical connections between the cores of the PL (214) and global memory (e.g., volatile memory (134)). In one example, the hardware compiler (1606) may route the logical connections to verify the capacity of the PL (214) to implement the block diagram and logical connections. In another aspect, the hardware compiler (1606) may use SoC interface block traces (e.g., described in more detail below) along with one or more data traffic generators as part of a simulation to verify the functionality of the block diagram with actual data traffic.
[0226] In block (2008), the system performs an implementation flow for the block diagram. For example, a hardware compiler may perform an implementation flow involving synthesis, placement, and routing, if necessary, for the block diagram to generate a configuration bitstream that can be loaded into the SoC (200) to implement the hardware part of the application in PL (214).
[0227] A hardware compiler (1606) can perform an implementation flow for a block diagram using an SoC interface block solution and a NoC solution. For example, because the SoC interface block solution specifies specific stream channels of the SoC interface block (206) that allow specific DPEs (204) to communicate with the PL (214), a placer can place blocks having connections to the DPEs (204) through the SoC interface block (206) close to (e.g., within a specific distance) the specific stream channels of the SoC interface block (206) to which the blocks of the block diagram will be connected. For example, the ports of the blocks may be correlated with the stream channels specified by the SoC interface block solution. The hardware compiler (1606) can also route connections between the ports of the blocks in the block diagram by routing signals input to and / or output from the ports of the blocks of the block diagram connected to the SoC interface block (206) to the BLIs of the PL (214) connected to specific stream channel(s) coupled to the ports as determined by the SoC interface block solution.
[0228] Similarly, since the NoC solution specifies specific inlet / outlet points to which the circuit blocks of PL (214) will be connected, the placer can place blocks having connections to NoC (208) close to (e.g., within a specific distance) the specific inlet / outlet points to which the blocks of the block diagram will be connected. For example, the ports of the blocks may be correlated with, for example, the inlet / outlet points of the NoC solution. The hardware compiler (1606) can also route connections between the ports of the blocks of the block diagram by routing signals input to and / or output from the ports of the blocks of the block diagram connected to the inlet / outlet points of NoC (208) to the inlet / outlet points of NoC (208) that are logically coupled to the ports as determined by the NoC solution. The hardware compiler (1606) can additionally route any signals connecting the ports of the blocks of PL (214) to each other. However, in some applications, NoC (208) may not be used to transfer data between the DPE array (202) and PL (214).
[0229] In block (2010), during the implementation flow, the hardware compiler optionally exchanges design data with the DPE compiler (1602) and / or the NoC compiler (1604). For example, the hardware compiler (1606), the DPE compiler (1602), and the NoC compiler (1604) may exchange design data as described in relation to FIG. 16, once, as needed, or repeatedly. Block (2010) may be performed optionally. The hardware compiler (1606) may exchange design data with the DPE compiler (1602) and / or the NoC compiler (1604), for example, before or during the construction of the block diagram, before and / or during the construction, and / or before and / or during routing.
[0230] In block (2012), the system exports the final hardware design generated by the hardware compiler (1606) as a hardware package. The hardware package includes a configuration bitstream used to program the PL (214). The hardware package is generated according to the hardware part of the application.
[0231] In block (2014), the user configures a new platform using a hardware package. The user initiates the creation of a new platform based on a user-provided configuration. The platform created in the system using the hardware package is used to compile the software portion of an application.
[0232] In block (2016), the system compiles the software portion of an application for implementation in the DPE array (202). For example, the system runs a DPE compiler (1602) to generate one or more binaries that can be loaded into various DPEs (204) of the DPE array (202). The binaries for the DPEs (204) may include object code, application data, and configuration data for the DPEs (204). Once the configuration bitstream and binaries are generated, the system can load the configuration bitstream and binaries into the SoC (200) to implement the application internally.
[0233] In another aspect, the hardware compiler (1606) may provide the hardware implementation to the DPE compiler (1602). The DPE compiler (1602) may extract the final SoC interface block solution that was relied upon by the hardware compiler (1606) when performing the implementation flow. The DPE compiler (1602) performs compilation using the same SoC interface block solution used by the hardware compiler (1606).
[0234] In the example of FIG. 20, each part of the application is resolved by a subsystem-specific compiler. The compilers may communicate design data, such as constraints and / or proposed solutions, to ensure that the interfaces between the various subsystems (e.g., SoC interface blocks) implemented for the application are compliant and consistent. Although not specifically illustrated in FIG. 20, a NoC compiler (1604) may also be invoked to generate a binary for programming the NoC (208) when used in the application.
[0235] FIG. 21 illustrates another exemplary method (2100) for performing a design flow to implement an application in an SoC (200). The method (2100) may be performed by a system described in relation to FIG. 1. The system may execute a software architecture described in relation to FIG. 16 or FIG. 19. The method (2100) may begin at a block (2102) where the system receives an application. The application may be designated as a data flow graph to be implemented in the SoC (200). The application may include a software part for implementation in the DPE array (202), a hardware part for implementation in the PL (214), and data transfers for implementation in the NoC (208) of the SoC (200). The application may also include an additional software part for implementation in the PS (212).
[0236] In block (2104), the DPE compiler (1602) can generate a logical architecture, an SoC interface block solution, and SoC interface block traces from the application. The logical architecture may be based on any IP cores to be implemented in the PL (214) required to interface with the DPEs (204) and the DPEs (204) required to implement the software portion of the application specified to be implemented within the DPE array (202). As mentioned, the DPE compiler (1602) can generate an initial DPE solution in which the DPE compiler (1602) performs an initial mapping of nodes (of the software portion of the application) to the DPE array (202). The DPE compiler (1602) can generate an initial SoC interface block solution that maps logical resources to physical resources (e.g., stream channels) of the SoC interface block (206). In one aspect, the SoC interface block solution may be generated using an initial NoC solution generated by the NoC compiler (1604) from data transfers. The DPE compiler (1602) may further simulate the initial DPE solution with the SoC interface block solution to simulate data transfers through the SoC interface block (206). The DPE compiler (1602) may capture the data transfers through the SoC interface block during the simulation as "SoC interface block traces" for subsequent use during the design flow illustrated in FIG. 21.
[0237] In block (2104), the hardware compiler (1606) generates a block diagram of the hardware portion of the application to be implemented in PL (214). The hardware compiler (1606) generates the block diagram based on the logical architecture and SoC interface block solution, and optionally additional IP cores specified by the user to be included in the block diagram along with the circuit blocks specified by the logical architecture. In one aspect, the user manually inserts these additional IP cores and connects the IP cores to other circuit blocks of the hardware description specified by the logical architecture.
[0238] In block (2106), the hardware compiler (1606) optionally receives one or more user-defined SoC interface block constraints and provides the SoC interface block constraints to the DPE compiler (1602).
[0239] In one aspect, before implementing the hardware portion of the application, the hardware compiler (1606) may evaluate the physical connections defined between the NoC (208), the DPE array (202), and the PL (214) based on the block diagram and logical architecture. The hardware compiler (1606) may perform an architectural simulation of the block diagram to evaluate the connections between the block diagram (e.g., the PL portion of the design) and the DPE array (202) and / or the NoC (208). For example, the hardware compiler (1606) may perform the simulation using SoC interface block traces generated by the DPE compiler (1602). As an exemplary and non-limiting example, the hardware compiler (1606) may perform a SystemC simulation of the block diagram. In the simulation, data traffic is generated for stream channels (e.g., physical connections) between the PL (214) and the DPE array (202) (via the SoC interface block (206)) and / or the NoC (208) and for the block diagram using SoC interface block traces. The simulation generates system performance and / or debugging information provided to the hardware compiler (1606).
[0240] The hardware compiler (1606) can evaluate system performance data. If the hardware compiler (1606) determines from the system performance data that, for example, one or more design metrics for the hardware part of the application are not satisfied, the hardware compiler (1606) can generate one or more SoC interface block constraints at the user's direction. The hardware compiler (1606) provides the SoC interface block constraints to the DPE compiler (1602) as a request.
[0241] The DPE compiler (1602) can perform an updated mapping of the DPE portion of an application to the DPEs (204) of the DPE array (202) utilizing the SoC interface block constraints provided by the hardware compiler (1606). If, for example, an application is implemented in which the hardware portion of the PL (214) is connected directly to the DPE array (202) through the SoC interface block (206) (for example, without passing through the NoC (208)), the DPE compiler (1602) can generate an updated SoC interface block solution for the hardware compiler (1606) without involving the NoC compiler (1604).
[0242] In block (2108), the hardware compiler (1606) optionally receives one or more user-defined NoC constraints and provides the NoC constraints to the NoC compiler for verification. The hardware compiler (1606) may also provide NoC traffic to the NoC compiler (1606). The NoC compiler (1604) may generate an updated NoC solution using the received NoC constraints and / or NoC traffic. If, for example, an application is implemented in which a hardware part of PL (214) is connected to a DPE array (202), PS (212), hardwired circuit blocks (210), or volatile memory (134) via the NoC (208), the hardware compiler (1606) may call the NoC compiler (1604) by providing the NoC constraints and / or NoC traffic to the NoC compiler (1604). The NoC compiler (1604) can update routing information for data paths through the NoC (208) as an updated NoC solution. The updated routing information can specify updated routes and specific entry / exit points for the routes. The hardware compiler (1606) can obtain the updated NoC solution and, in response, generate updated SoC interface block constraints provided to the DPE compiler (1602). The process may be iterative in nature. The DPE compiler (1602) and the NoC compiler (1604) may operate simultaneously as exemplified by the blocks (2106 and 2108).
[0243] In block (2110), the hardware compiler (1606) may perform synthesis on the block diagram. In block (2112), the hardware compiler (1606) performs placement and routing on the block diagram. In block (2114), while performing placement and / or routing, the hardware compiler may determine whether the current state of the implementation of the block diagram, for example, at any of these various stages of the implementation flow, the implementation of the hardware part (e.g., hardware design) satisfies design metrics for the hardware part of the application. For example, the hardware compiler (1606) may determine whether the current implementation satisfies design metrics before placement, during placement, before routing, or during routing. In response to determining that the current implementation of the hardware part of the application does not satisfy design metrics, the method (2100) continues to block (2116). Otherwise, the method (2100) continues to block (2120).
[0244] In block (2116), the hardware compiler may provide one or more user-defined SoC interface block constraints to the DPE compiler (1602). The hardware compiler (1606) may optionally provide one or more NoC constraints to the NoC compiler (1604). As discussed, the DPE compiler (1602) generates an updated SoC interface block solution using the SoC interface block constraint(s) received from the hardware compiler (1606). The NoC compiler (1604) optionally generates an updated NoC solution. For example, if one or more data paths between the DPE array (202) and the PL (214) flow through the NoC (208), the NoC compiler (1604) may be invoked. In block (2118), the hardware compiler (1606) receives the updated SoC interface block solution and optionally the updated NoC solution. After block (2118), the method (2100) continues to block (2112), where the hardware compiler (1606) continues to perform placement and / or routing using the updated SoC interface block solution and optionally the updated NoC solution.
[0245] FIG. 21 illustrates that the exchange of design data between compilers may be performed repeatedly. For example, at any point among a plurality of different points during the placement and / or routing stages, the hardware compiler (1606) may determine whether the current state of the implementation of the hardware part of the application satisfies the set design metrics. If not, the hardware compiler (1606) may initiate the exchange of design data as described to obtain the updated NoC solution and the updated SoC interface block solution that the hardware compiler (1606) uses for the purpose of placement and routing. It should be recognized that in cases where the configuration of the NoC (208) is updated (e.g., where data from the PL (214) is provided to other circuit blocks through the NoC (208) and / or received from these other circuit blocks), the hardware compiler (1606) needs to invoke only the NoC compiler (1604).
[0246] In block (2120), if the hardware part of the application satisfies the design metrics, the hardware compiler (1606) generates a configuration bitstream that specifies the implementation of the hardware part within the PL (214). The hardware compiler (1606) may additionally provide the final SoC interface block solution (e.g., the SoC interface block solution used for placement and routing) to the DPE compiler (1602) and provide the final NoC solution that may have been used for placement and routing to the NoC compiler (1604).
[0247] In block (2122), the DPE compiler (1602) generates binaries for programming the DPE (202) of the DPE array (204). The NoC compiler (1604) generates binaries for programming the NoCs (208). For example, throughout blocks (2106, 2108, and 2116), the DPE compiler (1602) and the NoC compiler (1604) may perform incremental validation functions in which the used NoC solutions and SoC interface block solutions are generated based on verification procedures that can be performed in a shorter runtime than when complete solutions for the SoC interface block and the NoC are determined. In block (2122), the DPE compiler (1602) and the NoC compiler (1604) may each generate final binaries used to program the DPE array (202) and the NoC (208).
[0248] In block (2124), the PS compiler (1918) generates a PS binary. The PS binary contains object code that is executed by the PS (212). The PS binary implements, for example, a control program that is executed by the PS (212) to monitor the operation of the SoC (200) in which an application is implemented internally. The DPE compiler (1602) can also generate a DPE array driver that can be compiled by the PS compiler (1918) and executed by the PS (212) to read and / or write to the DPEs (204) of the DPE array (202).
[0249] In block (2126), the system can deploy configuration bitstreams and binaries in the SoC (200). For example, the system can combine various binaries and configuration bitstreams into a PDI, and the PDI is provided to and loaded into the SoC (200) so that an application can be implemented therein.
[0250] FIG. 22 illustrates an exemplary method of communication (2200) between a hardware compiler (1606) and a DPE compiler (1602). The method (2200) provides an example of how communications between the hardware compiler (1606) and the DPE compiler (1602) can be handled, as described in relation to FIG. 16, FIG. 19, FIG. 20, and FIG. 21. The method (2200) illustrates an exemplary implementation of a verification call (e.g., a verification procedure) performed between the hardware compiler (1606) and the DPE compiler (1602). An example of the method (2200) provides an alternative to performing full placement and routing on the DPE array (202) and / or NoC (208) to generate updated SoC interface block solutions in response to SoC interface block constraints provided by the hardware compiler (1606). The method (2200) exemplifies an incremental method in which re-routing is attempted before mapping and routing of the software part of the application begins.
[0251] The method (2200) may begin at a block (2202) in which a hardware compiler (1606) provides one or more SoC interface block constraints to a DPE compiler (1602). For example, the hardware compiler (1606) may receive one or more user-defined SoC interface block constraints and / or generate one or more SoC interface block constraints during the implementation flow and in response to a decision that design metrics for the hardware part of the application are not or will not be satisfied. The SoC interface block constraints may specify a preferred mapping of logical resource(s) to physical stream channels of the SoC interface block (206) that is expected to result in improved Quality of Service (QoS) for the hardware part of the application.
[0252] The hardware compiler (1606) provides SoC interface block constraints to the DPE compiler (1602). The SoC interface block constraints provided by the hardware compiler (1606) can be divided into two different categories. The first category of SoC interface block constraints is hard constraints. The second category of SoC interface block constraints is soft constraints. Hard constraints are design constraints that must be satisfied to implement an application within the SoC (200). Soft constraints are design constraints that may be violated in the implementation of an application on the SoC (200).
[0253] In one example, hard constraints are user-defined constraints on the hardware portion of the application to be implemented in PL (214). Hard constraints may include any available constraint types, such as position, power, timing, etc., which are user-defined constraints. Soft constraints may include any available constraints generated by the hardware compiler (1606) and / or the DPE compiler (1602) throughout the implementation flow, such as a constraint specifying a particular mapping of logical resource(s) to the stream channels of the SoC interface block (206) as described.
[0254] In block (2204), the DPE compiler (1602) initiates a verification process to incorporate the received SoC interface block constraints when generating an updated SoC interface block solution in response to receiving the SoC interface block constraint(s). In block (2206), the DPE compiler (1602) can distinguish between the hard constraint(s) and the soft constraint(s) received from the hardware compiler (1606) in relation to the hardware part of the application.
[0255] In block (2208), the DPE compiler (1602) routes the software portion of the application while following both the hard constraint(s) and the soft constraint(s) provided by the hardware compiler. For example, the DPE compiler (1602) may route connections between the DPEs (204) of the DPE array (202) and data paths between the DPEs (204) and the SoC interface block (206) to determine which stream channels (e.g., tiles, stream switches, and ports) of the SoC interface block (206) are used for data path crossings between the DPE array (202) and the PL (214) and / or NoC (208). If the DPE compiler (1602) successfully routes the software portion of the application for implementation in the DPE array (202) while following both the hard constraint(s) and the soft constraint(s), the method (2200) continues to block (2218). If the DPE compiler (1602) cannot generate a root for the software part of the application in the DPE array while following both hard constraint(s) and soft constraint(s), for example, if the constraints are not routable, the method (2200) continues to block (2210).
[0256] In block (2210), the DPE compiler (1602) routes the software portion of the application while following only the hard constraint(s). In block (2210), the DPE compiler (1602) ignores the soft constraint(s) for the routing operation. If the DPE compiler (1602) successfully routes the software portion of the application for implementation in the DPE array (202) while following only the hard constraint(s), the method (2200) continues to block (2218). If the DPE compiler (1602) cannot generate a route for the software portion of the application in the DPE array (202) while following only the hard constraint(s), the method (2200) continues to block (2212).
[0257] Blocks (2208 and 2210) exemplify an approach to a verification operation that aims to generate an updated SoC interface block solution in a shorter time than the full mapping (e.g., placement) and routing of DPE nodes is performed, using SoC interface block constraint(s) provided by the hardware compiler (1606). Thus, blocks (2208 and 2210) involve only routing without attempting to map (e.g., remap) or "place" DPE nodes to the DPEs (204) of the DPE array (202).
[0258] Method (2200) continues to block (2212) in cases where routing alone cannot reach an updated SoC interface block solution using SoC interface block constraint(s) from a hardware compiler. In block (2212), the DPE compiler (1602) can map the software portion of the application to the DPEs of the DPE array (202) using both hard constraint(s) and soft constraint(s). The DPE compiler (1602) is also programmed with the architecture (e.g., connectivity) of the SoC (200). The DPE compiler (1602) performs the actual allocation of logical resources for the physical channels (e.g., stream channels) of the SoC interface block (206) and can also model the structural connectivity of the SoC (200).
[0259] For example, consider a DPE node A communicating with a PL node B. Each block in the block diagram may correspond to a specific core (e.g., a circuit block) to be implemented in the PL (214). The PL node B communicates with the DPE node A via a physical channel X in the SoC interface block (206). The physical channel X carries data stream(s) between the DPE node A and the PL node B. The DPE compiler (1602) can map the DPE node A to a specific DPE Y such that the distance between the DPE Y and the physical channel X is minimized.
[0260] In some implementations of the SoC interface block (206), one or more of the tiles included in the SoC interface block (206) are not connected to the PL (214). The unconnected tiles may be the result of the placement of certain hardwired circuit blocks (210) inside and / or around the PL (214). For example, such an architecture having unconnected tiles in the SoC interface block (206) complicates routing between the SoC interface block (206) and the PL (214). Connectivity information regarding the unconnected tiles is modeled in the DPE compiler (1602). The DPE compiler (1602) may select DPE nodes having connections to the PL (214) as part of the mapping process. The DPE compiler (1602) can minimize the number of selected DPE nodes mapped to DPEs (204) in columns of the DPE array (202) located directly above the unconnected tiles of the SoC interface block (206) as part of the mapping operation. The DPE compiler (1602) maps DPE nodes that do not have connections to the PL (214) (e.g., direct connections) (e.g., nodes that are instead connected to other DPEs (204)) to columns of the DPE array (202) positioned above the unconnected tiles of the SoC interface block (206).
[0261] In block (2214), the DPE compiler (1602) routes the remapped software portion of the application while following only the hard constraint(s). If the DPE compiler (1602) successfully routes the remapped software portion of the application for implementation in the DPE array (202) while following only the hard constraint(s), the method (2200) continues to block (2218). If the DPE compiler (1602) cannot generate a route for the software portion of the application in the DPE array (202) while following only the hard constraint(s), the method (2200) continues to block (2216). In block (2216), the DPE compiler (1602) indicates that the verification operation failed. The DPE compiler (1602) may output a notification and provide a notification to the hardware compiler (1606).
[0262] In block (2218), the DPE compiler (1602) generates an updated SoC interface block solution and a score for the updated SoC interface block solution. The DPE compiler (1602) generates an updated SoC interface block solution based on the updated routing or updated mapping and routing determined in block (2208), block (2210), or blocks (2212 and 2214).
[0263] The score generated by the DPE compiler (1602) indicates the quality of the SoC interface block solution based on the mapping and / or routing operations performed. In one exemplary implementation, the DPE compiler (1602) determines the score based on how many soft constraints are not satisfied and the distance between the stream channel requested by the soft constraint and the actual channel allocated in the updated SoC interface block solution. For example, both the number of unsatisfied soft constraints and the distance may be inversely proportional to the score.
[0264] In another exemplary implementation, the DPE compiler (1602) determines a score based on the quality of the updated SoC interface block solution using one or more design cost metrics. These design cost metrics may include the number of data moves supported by the SoC interface block solution, memory crash costs, and route latency. In one aspect, the number of data moves in the DPE array (202) may be quantified by the number of DMA moves used in the DPE array (202) in addition to those required to transfer data through the SoC interface block (206). Memory crash costs may be determined based on the number of concurrent access circuits (e.g., DPE or DMA) for each memory bank. Route latency may be quantified by the minimum number of cycles required to transfer data between the SoC interface block (206) ports and individual source or destination DPEs (204). The DPE compiler (1602) determines a higher score when the design cost metrics are lower (e.g., when the sum of the design cost metrics is lower).
[0265] In another exemplary implementation, the total score of the updated SoC interface block solution is calculated as a fraction (e.g., 80 / 100), where the numerator is reduced from 100 by the sum of the number of additional DMA transfers, the number of concurrent access circuits for each memory bank exceeding 2, and the number of hops required for the roots between the SoC interface block (206) ports and the DPE (204) cores.
[0266] In block (2220), the DPE compiler (1602) provides the updated SoC interface block solution and score to the hardware compiler (1606). The hardware compiler (1606) can evaluate the various SoC interface block solutions received from the DPE compiler (1602) based on the score of each individual SoC interface block solution. In one aspect, for example, the hardware compiler (1606) may retain the previous SoC interface block solutions. The hardware compiler (1606) may compare the score of the previous SoC interface block solution (e.g., the immediate SoC interface block solution) with the score of the updated SoC interface block solution, and if the score of the updated SoC interface block solution exceeds the score of the previous SoC interface block solution, the updated SoC interface block solution may be used.
[0267] In another exemplary implementation, the hardware compiler (1606) receives an SoC interface block solution with a score of 80 / 100 from the DPE compiler (1602). The hardware compiler (1606) cannot reach the implementation of the hardware part of the application within the PL (214) and provides one or more SoC interface block constraints to the DPE compiler (1602). The updated SoC interface block solution received by the hardware compiler (1606) from the DPE compiler (1602) has a score of 20 / 100. In that case, in response to determining that the score of the newly received SoC interface block solution does not exceed the score of the previous SoC interface block solution (e.g., lower than the score of the previous SoC interface block solution), the hardware compiler (1606) relaxes one or more of the SoC interface block constraints (e.g., soft constraints) and provides the SoC interface block constraints containing the relaxed constraint(s) to the DPE compiler (1602). The DPE compiler (1602) attempts to generate other SoC interface block solutions with scores higher than 20 / 100 and / or 80 / 100 by taking into account relaxed design constraint(s).
[0268] In another example, the hardware compiler (1606) may choose to use the previous SoC interface block solution having a higher score or the highest score. The hardware compiler (1606) may return to the previous SoC interface block solution at any point, for example, in response to receiving a SoC interface block solution having a lower score than the previous SoC interface block solution, or in response to receiving a SoC interface block solution having a lower score than the previous SoC interface block solution after one or more of the SoC interface block constraints have been relaxed.
[0269] FIG. 23 illustrates an exemplary method (2300) for handling SoC interface block solutions. The method (2300) may be performed by a hardware compiler (1606) to evaluate the received SoC interface block solution(s) and to select the SoC interface block solution referred to as the best current SoC interface block solution for use in performing the implementation flow for the hardware part of the application.
[0270] In block (2302), the hardware compiler (1606) receives an SoC interface block solution from the DPE compiler (1602). The SoC interface block solution received in block (2302) may be an initial or first SoC interface block solution provided by the DPE compiler (1602). When providing SoC interface block solutions to the hardware compiler (1606), the DPE compiler (1602) additionally provides a score for the SoC interface block solution. At least initially, the hardware compiler (1606) selects the first SoC interface block solution as the best SoC interface block solution at present.
[0271] In block (2304), the hardware compiler (1606) optionally receives one or more hard SoC interface block constraints from the user. In block (2306), the hardware compiler may generate one or more soft SoC interface block constraints for implementing the hardware portion of the application. The hardware compiler generates soft SoC interface block constraints in an effort to satisfy hardware design metrics.
[0272] In block (2308), the hardware compiler (1606) transmits SoC interface block constraints (e.g., both hard and soft) to the DPE compiler (1602) for verification. In response to receiving the SoC interface block constraints, the DPE compiler can generate an updated SoC interface block solution based on the SoC interface block constraints received from the hardware compiler (1606). The DPE compiler (1602) provides the updated SoC interface block solution to the hardware compiler (1606). Thus, in block (2310), the hardware compiler receives the updated SoC interface block solution.
[0273] In block (2312), the hardware compiler (1606) compares the score of the first (e.g., previously received) SoC interface block solution with the score of the updated SoC interface block solution (e.g., most recently received SoC interface block solution).
[0274] In block (2314), the hardware compiler (1606) determines whether the score of the updated (e.g., most recently received) SoC interface block solution exceeds the score of the previously received (e.g., first) SoC interface block solution. In block (2316), the hardware compiler (1606) selects the most recently received (e.g., updated) SoC interface block solution as the current best SoC interface block solution.
[0275] In block (2318), the hardware compiler (1606) determines whether the improvement goal has been achieved or whether the time budget has been exceeded. For example, the hardware compiler (1606) may determine whether the current implementation state of the hardware part of the application meets a greater number of design metrics and / or nearly meets one or more design metrics. The hardware compiler (1606) may also determine whether the time budget has been exceeded based on the amount of processing time consumed for batching and / or routing, and whether that time exceeds the maximum batch time, the maximum routing time, or the maximum amount of time for both batching and routing. In response to the determination that the improvement goal has been reached or the time budget has been exceeded, the method (2300) continues to block (2324). If not, the method (2300) continues to block (2320).
[0276] In block (2324), the hardware compiler (1606) uses the best current SoC interface block solution to implement the hardware part of the application.
[0277] Continuing to block (2320), the hardware compiler (1606) relaxes one or more of the SoC interface block constraints. The hardware compiler (1606) may relax or change one or more of the soft constraints, for example. Examples of relaxing or changing soft SoC interface block constraints include removing (e.g., deleting) the soft SoC interface block constraints. Another example of relaxing or changing soft SoC interface block constraints includes replacing the soft SoC interface block constraints with different SoC interface block constraints. The replacement soft SoC interface block constraints may be less strict than the original constraints being replaced.
[0278] In block (2322), the hardware compiler (1606) may transmit SoC interface block constraint(s) containing relaxed SoC interface block constraint(s) to the DPE compiler (1602). After block (2322), the method (2300) loops back to block (2310) to continue processing as described. For example, the DPE compiler generates an additionally updated SoC interface block solution based on the SoC interface block constraints received from the hardware compiler in block (2322). In block (2310), the hardware compiler receives the additionally updated SoC interface block solution.
[0279] Method (2300) illustrates situations in which SoC interface block constraint(s) may be relaxed and an exemplary process of selecting from the DPE compiler (1602) an SoC interface block solution to be used to perform the implementation flow. It should be recognized that the hardware compiler (1606) may provide SoC interface block constraints to the DPE compiler (1602) at any point among various points during the implementation flow to obtain an updated SoC interface block solution as part of the adjustment and / or verification process. For example, at any point in time when the hardware compiler (1606) determines (e.g., based on timing, power, or other checks or analyses) that the implementation of the hardware part of the application does not meet or will not meet the design metrics of the application in the current state of that part, the hardware compiler (1606) may request an updated SoC interface block solution by providing the updated SoC interface block constraint(s) to the DPE compiler (1602).
[0280] FIG. 24 illustrates another example of an application (2400) for implementation in an SoC (200). The application (2400) is designated as a directed flow graph. Nodes are differently shaded and shaped to distinguish between PL nodes, DPE nodes, and I / O nodes. In the illustrated example, I / O nodes may be mapped to SoC interface blocks (206). PL nodes are implemented in PL. DPE nodes are mapped to specific DPEs. The application (2400) includes, although not in its entirety, 36 kernels (e.g., nodes) to be mapped to DPEs (204), 72 PLs to be mapped to DPE array data streams, and 36 DPE arrays to be mapped to PL data streams.
[0281] FIG. 25 is an example of an SoC interface block solution generated by a DPE compiler (1602). The SoC interface block solution of FIG. 25 can be generated by the DPE compiler (1602) and provided to a hardware compiler (1606). The example of FIG. 25 illustrates a scenario in which the DPE compiler (1602) generates an initial mapping of DPE nodes to DPEs (204) of a DPE array (202). Furthermore, the DPE compiler (1602) successfully routes the initial mapping of DPE nodes. In the example of FIG. 25, only the columns (6-17) of the DPE array (202) are shown. Furthermore, each column contains four DPEs (204).
[0282] FIG. 25 illustrates the mapping of DPE nodes to DPEs (204) of a DPE array (202) and the routing of data streams to SoC interface block (206) hardware. The mapping of DPE nodes (0-35) of an application (2400) to DPEs (204), determined by the DPE compiler (1602), is illustrated with reference to the DPE array (202). The routing of data streams between specific tiles of the SoC interface block (206) and DPEs is illustrated as a set of arrows. For the sake of illustration when describing FIG. 25 through 30, the key shown in FIG. 25 is used to distinguish between data streams controlled by soft constraints, data streams controlled by hard constraints, and data streams that do not have applicable constraints.
[0283] Referring to FIGS. 25 through 30, soft constraints correspond to routes determined by the DPE compiler (1602) and / or the hardware compiler (1606), whereas hard constraints may include user-defined SoC interface block constraints. All constraints illustrated in FIG. 25 are soft constraints. An example in FIG. 25 illustrates a case where the DPE compiler (1602) has successfully determined an initial SoC interface block solution. In one aspect, the DPE compiler (1602) may be configured to at least initially attempt to use vertical routes for the SoC interface block solution as illustrated before attempting to use other routes traversing along the row DPEs (204) from one column to another (e.g., from left to right).
[0284] FIG. 26 illustrates an example of routable SoC interface block constraints received by the DPE compiler (1602). The DPE compiler (1602) can generate an updated SoC interface block solution that specifies updated routing in the form of updated SoC interface block constraints. In the example of FIG. 26, a larger number of SoC interface block constraints are hard constraints. In this example, the DPE compiler (1602) successfully routes data streams of the DPE array (202) while complying with each of the illustrated types of constraints.
[0285] FIG. 27 illustrates examples of non-routable SoC interface block constraints that the DPE compiler (1602) must comply with. The DPE compiler (1602) cannot generate an SoC interface block solution that complies with the constraints exemplified in FIG. 27.
[0286] FIG. 28 illustrates an example in which the DPE compiler (1602) ignores the soft-type SoC interface block constraints of FIG. 27. In the example of FIG. 28, the DPE compiler (1602) successfully routes the software portion of the application for implementation in the DPE array (202) using only hard constraints. These data streams not controlled by constraints can be routed in any way that the DPE compiler (1602) deems appropriate or capable of doing so.
[0287] FIG. 29 illustrates another example of non-routable SoC interface block constraints. The example in FIG. 29 has only hard constraints. Therefore, a DPE compiler (1602) that cannot ignore hard constraints initiates a mapping (or re-mapping) operation.
[0288] FIG. 30 illustrates an exemplary mapping of the DPE nodes of FIG. 29. In this example, following the remapping, the DPE compiler (1602) can successfully route the DPE nodes to generate an updated SoC interface block solution.
[0289] FIG. 31 illustrates another example of non-routable SoC interface block constraints. The example in FIG. 31 has only hard constraints. Therefore, a DPE compiler (1602) that cannot ignore hard constraints initiates a mapping operation. For the purpose of illustration, the DPE array (202) contains only three rows of DPEs (e.g., three DPEs in each column).
[0290] FIG. 32 illustrates an exemplary mapping of the DPE nodes of FIG. 31. FIG. 32 illustrates the result obtained from the re-mapping operation initiated as described in relation to FIG. 31. In this example, following the re-mapping, the DPE compiler (1602) can successfully route the software solution of the application to generate an updated SoC interface block solution.
[0291] In one aspect, the system can perform the mapping exemplified in FIGS. 25 through 32 by generating an Integer Linear Programming (ILP) formulation of the mapping problem. The ILP formulation may include multiple different variables and constraints defining the mapping problem. The system can also solve the ILP formulation while minimizing cost(s). Costs may be determined at least partially based on the number of DMA engines used. In this way, the system can map the DFG onto the DPE array.
[0292] In another aspect, the system can order the nodes of the DFG in descending order of priority. The system can determine priority based on one or more factors. Examples of factors may include, but are not limited to, the height of the node in the DFG graph, the total node degree (e.g., the sum of all edges entering and leaving the node), and / or the types of edges connected to the node, such as memory, streams, and cascades. The system can place nodes in the best available DPE based on affinity and validity. The system can determine validity based on whether all resource requirements of such nodes can be met for a given DPE (e.g., computing resources, memory buffers, stream resources). The system can determine affinity based on one or more other factors. Examples of similarity factors may include neighbors of these nodes deploying adjacent DPEs or identical DPEs that have been pre-placed to minimize DMA communication, structural constraints such as whether these nodes are part of a cascade chain, and / or finding a DPE with maximum free resources. When a node is placed with all constraints satisfied, the system may increase the priority of neighbor nodes of the placed node so that these nodes are handled next. If available placements are not valid for the current node, the system may attempt to unplace some other nodes from their best candidate DPE(s) to make space for this node. The system may place the unplaced nodes back into a priority queue for re-placement. The system may limit the total effort spent on finding a good solution by tracking the total number of placements and unplacements performed. However, it should be recognized that other mapping techniques may be used and that the examples provided herein are not intended to be limiting.
[0293] FIG. 33 illustrates another exemplary software architecture (3300) executable by the system described in relation to FIG. 1. For example, the architecture (3300) of FIG. 33 may be implemented as one or more of the program modules (120) of FIG. 1. The exemplary software architecture (3300) of FIG. 33 may be used in cases where an application, such as a data flow graph, specifies one or more High-Level Synthesis (HLS) kernels for implementation in PL (214). For example, PL nodes of an application refer to HLS kernels that require HLS processing. In one aspect, the HLS kernels are specified in a high-level language (HLL) such as C and / or C++.
[0294] In the example of FIG. 33, the software architecture (3300) includes a DPE compiler (1602), a hardware compiler (1606), an HLS compiler (3302), and a system linker (3304). The NoC compiler (1604) may be included and used together with the DPE compiler (1602) to perform validation (3306) as previously described in the present disclosure.
[0295] As illustrated, the DPE compiler (1602) receives an application (3312), an SoC architecture description (3310), and optionally a test bench (3314). As discussed, the application (3312) may be specified as a data flow graph containing parallel execution semantics. The application (3312) may include interconnected PL nodes and DPE nodes and may specify runtime parameters. In this example, the PL nodes refer to HLS kernels. The SoC architecture description (3310) may be a data structure or file specifying information such as the size and dimensions of the DPE array (202), the size of the PL (214) and various programmable circuit blocks available within it, the type of the PS (212) such as the types of processors and other devices included in the PS (212), and other physical characteristics of the circuit portion of the SoC (200) where the application (3312) is to be implemented. The SoC architecture description (3310) can also specify connectivity (e.g., interfaces) between subsystems included therein.
[0296] The DPE compiler (1602) can output HLS kernels to the HLS compiler (3302). The HLS compiler (3302) converts the HLS kernels specified in HLL into HLS IPs that can be synthesized by a hardware compiler. For example, the HLS IPs can be specified as RTL (register transfer level) blocks. For example, the HLS compiler (3302) generates an RTL block for each HLS kernel. As illustrated, the HLS compiler (3302) outputs the HLS IPs to the system linker (3304).
[0297] The DPE compiler (1602) generates additional outputs, such as an initial SoC interface block solution and a connectivity graph. The DPE compiler (1602) outputs the connectivity graph to the system linker (3304) and outputs the SoC interface block solution to the hardware compiler (1606). The connectivity graph specifies the connectivity between nodes corresponding to the HLS kernels (currently converted into HLS IPs) to be implemented in the PL (214) and nodes to be implemented in the DPE array (202).
[0298] As illustrated, the system linker (3304) receives the SoC architecture description (3310). The system linker (3304) may also receive one or more HLS and / or RTL blocks directly from the application (3312) that are not processed by the DPE compiler (1602). The system linker (3304) may automatically generate a block diagram corresponding to the hardware portion of the application using a connection graph that specifies the connectivity between the received HLS and / or RTL blocks, HLS IPs, and IP kernels, and the connectivity between the IP kernels and DPE nodes. In one aspect, the system linker (3304) may integrate the block diagram with a base platform (not illustrated) for the SoC (200). For example, the system linker (3304) may create an integrated block diagram by linking the block diagram to the base platform. The block diagram and the linked base platform may be referred to as a synthesizable block diagram.
[0299] In another aspect, HLS IPs and RTL IPs referenced by kernels within an SDF graph (e.g., application (3312)) can be compiled into IPs outside of the DPE compiler (1602). The compiled IPs can be provided directly to the system linker (3304). The system linker (3304) can use the provided IPs to automatically generate a block diagram corresponding to the hardware part of the application.
[0300] In one aspect, the system linker (3304) may include additional hardware-specific details derived from the original SDF (e.g., application (3312)) and the generated link graph within the block diagram. For example, because the application (3312) includes software models, which are actual HLS models that can be converted into IPs from a database of these IPs or correlated (e.g., matched) to IPs using some mechanism (e.g., by name or other matching / correlation techniques), the system linker (3304) may automatically generate the block diagram (e.g., without user intervention). In this example, custom IPs may not be used. When automatically generating the block diagram, the system linker (3304) may automatically insert one or more additional circuit blocks, such as data-width conversion blocks, hardware buffers, and / or clock domain crossing logic that was manually inserted and linked by the user in other cases described herein. For example, the system linker (3304) can analyze data types and software models to determine that one or more additional circuit blocks are needed to create the connections specified by the connection graph as described.
[0301] The system linker (3304) outputs a block diagram to the hardware compiler (1606). The hardware compiler (1606) receives the initial SoC interface block solution and block diagram generated by the DPE compiler (1602). The hardware compiler (1606) may initiate a verification check (3306) with the DPE compiler (1602) and optionally the NoC compiler (1604) as previously described in relation to the block (2010) of FIG. 20 and the blocks (2106, 2108, 2112, 2114, 2116, and 2118) of FIG. 20. Verification may be an iterative process in which a hardware compiler provides design data, such as various types of constraints (which may include relaxed / modified constraints in an iterative approach), to a DPE compiler (1602) and optionally to a NoC compiler (1604), and in turn receives an updated SoC interface block solution from the DPE compiler (1602) and optionally an updated NoC solution from the NoC compiler (1604).
[0302] A hardware compiler (1606) can generate a hardware package containing a configuration bitstream that implements the hardware portion of the application (3312) in PL (214). The hardware compiler (1606) can output the hardware package to a DPE compiler (1602). The DPE compiler (1602) can generate DPE array configuration data (e.g., one or more binaries) that internally programs the software portion of the application (3312) intended for implementation in the DPE array (202).
[0303] FIG. 34 illustrates another exemplary method (3400) for performing a design flow to implement an application on an SoC (200). The method (3400) can be performed by a system as described in relation to FIG. 1. The system can execute a software architecture as described in relation to FIG. 33. In the example of FIG. 34, the application being processed includes nodes that specify HLS kernels for implementation in a PL (214).
[0304] In block (3402), the DPE compiler (1602) receives an application, a description of the SoC architecture of the SoC (200), and optionally a test bench. In block (3404), the DPE compiler (1602) can generate a connection graph and provide the connection graph to the system linker. In block (3406), the DPE compiler (1602) generates an initial SoC interface block solution and provides the initial SoC interface block solution to the hardware compiler (1606). The initial SoC interface block solution can specify the initial mapping of the application's DPE nodes to the DPEs (204) of the DPE array (202), and the mapping of connections to and from the DPE array (202) to the physical data paths of the SoC interface block (206).
[0305] In block (3408), the HLS compiler (3302) can perform HLS on HLS kernels to generate synthesizable IP cores. For example, the DPE compiler (1602) provides the HLS kernels specified by the application nodes to the HLS compiler (3302). The HLS compiler (3302) generates an HLS IP for each of the received HLS kernels. The HLS compiler (3302) outputs the HLS IPs to the system linker.
[0306] In block (3410), the system linker can automatically generate a block diagram corresponding to the hardware part of the application using a link graph, an SoC architecture description, and HLS IPs. In block (3412), the system linker can integrate the base platform and the block diagram for the SoC (200). For example, a hardware compiler (1606) can generate an integrated block diagram by linking the block diagram to the base platform. In one aspect, the block diagram and the linked base platform are referred to as a synthesizable block diagram.
[0307] In block (3414), the hardware compiler (1606) may perform an implementation flow on the integrated block diagram. During the implementation flow, the hardware compiler (1606) may perform verification as described herein in cooperation with the DPE compiler (1602) and optionally the NoC compiler (1604) to converge on the implementation of the hardware part of the application for implementation in PL. For example, as discussed, the hardware compiler (1606) may invoke the DPE compiler (1602) and optionally the NoC compiler (1604) in response to determining that the current implementation state of the hardware part of the application does not satisfy one or more design metrics. The hardware compiler (1606) may invoke the DPE compiler (1602) and optionally the NoC compiler (1604) before deployment, during deployment, before routing, and / or during routing.
[0308] In block (3416), the hardware compiler (1606) exports the hardware implementation to the DPE compiler (1602). In one aspect, the hardware implementation may be output as a device support archive (DSA) file. The DSA file may include platform metadata, emulation data, and one or more configuration bitstreams generated by the hardware compiler (1606) from the implementation flow. The hardware implementation may also include a final SoC interface block solution and optionally a final NoC solution used by the hardware compiler (1606) to generate an implementation of the hardware part of the application.
[0309] In block (3418), the DPE compiler (1602) completes the software generation for the DPE array. For example, the DPE compiler (1602) generates binaries used to program the DPEs used in the application. When generating the binaries, the DPE compiler (1602) may use the final SoC interface block solution and optionally the final NoC solution used by the hardware compiler (1606) to perform the implementation flow. In one aspect, the DPE compiler may determine the SoC interface block solution used by the hardware compiler by examining the configuration bitstream and / or metadata contained in the DSA.
[0310] In block (3420), the NoC compiler (1604) generates a binary or binaries for programming the NoC (208). In block (3422), the PS compiler (1918) generates a PS binary. In block (3424), the system can deploy the configuration bitstreams and binaries from the SoC (200).
[0311] FIG. 35 illustrates another exemplary method (3500) for performing a design flow to implement an application in an SoC (200). The method (3500) may be performed by a system as described in relation to FIG. 1. The application may be designated as a data flow graph as described herein and may include a software part for implementation within a DPE array (202) and a hardware part for implementation within a PL (214).
[0312] In block (3502), the system can generate a first interface solution that maps logical resources used by the software portion to the hardware resources of the interface block coupling the DPE array (202) and PL (214). For example, the DPE compiler (1602) can generate an initial or first SoC interface block solution.
[0313] In block (3504), the system can generate a connectivity graph that specifies the connectivity between the nodes of the software part to be implemented in the DPE array and the HLS kernels. In one aspect, the DPE compiler (1602) can generate the connectivity graph.
[0314] In block (3506), the system can generate a block diagram based on the connection graph and HLS kernels. The block diagram is synthesizable. For example, the system linker can generate a synthesizable block diagram.
[0315] In block (3508), the system can perform an implementation flow for the block diagram using the first interface solution. As discussed, the hardware compiler (1606) can exchange design data with the DPE compiler (1602) and optionally the NoC compiler (1604) during the implementation flow. The hardware compiler (1606) and the DPE compiler (1602) may repeatedly exchange data when the DPE compiler (1602) provides updated SoC interface block solutions to the hardware compiler (1606) in response to being invoked by the hardware compiler (1606). The hardware compiler (1606) may invoke the DPE compiler by providing one or more constraints for the SoC interface block to the DPE compiler. The hardware compiler (1606) and the NoC compiler (1604) may repeatedly exchange data when the NoC compiler (1604) provides updated NoC solutions to the hardware compiler (1606) in response to being invoked by the hardware compiler (1606). The hardware compiler (1606) may invoke the NoC compiler (1604) by providing one or more constraints on the NoC (208) to the NoC compiler (1604).
[0316] In block (3510), the system can compile the software portion of the application using a DPE compiler (1602) for implementation in one or more DPEs (204) of the DPE array (202). The DPE compiler (1602) can receive the results of the implementation flow (e.g., the same SoC interface block solution used during the implementation flow by the hardware compiler (1606)) to use a consistent interface between the DPE array (202) and the PL (214).
[0317] For the purpose of explanation, specific nomenclature is used to provide a complete understanding of the various concepts of the invention disclosed herein. However, the terms used herein are intended only to describe specific aspects of the arrangements of the invention and are not intended to be limiting.
[0318] As defined herein, singular forms are intended to include plural forms as well, unless otherwise clearly indicated in the context.
[0319] As defined herein, the terms “at least one,” “one or more,” and “and / or” are open-ended expressions that are both linking and separable in action, unless otherwise explicitly stated. For example, the expressions “at least one of A, B, and C,” “at least one of A, B, or C,” “one or more of A, B, and C,” “one or more of A, B, or C,” and “A, B, and / or C” each mean A alone, B alone, C alone, both A and B, both A and C, both B and C, or A, B, and C.
[0320] As defined herein, the term “automatically” means the absence of user intervention. As defined herein, the term “user” means a person.
[0321] As defined herein, the term “computer-readable storage medium” means a storage medium containing or storing program code for use by or associated with an instruction execution system, device, or device. As defined herein, “computer-readable storage medium” is not the transient radio signal itself. A computer-readable storage medium may be, but is not limited to, an electronic storage device, a magnetic storage device, an optical storage device, an electromagnetic storage device, a semiconductor storage device, or any suitable combination of the foregoing. As described herein, various forms of memory are examples of computer-readable storage media. A more specific, non-limiting list of computer-readable storage media may include portable computer diskettes, hard disks, RAM, ROM (read-only memory), EPROM (erasable programmable read-only memory or flash memory), EEPROM (electronically erasable programmable read-only memory), SRAM (static random access memory), portable CD-ROM (compact disc read-only memory), DVD (digital versatile disk), memory sticks, floppy disks, etc.
[0322] As defined herein, the term “in the case of” means, depending on the context, “at a time,” “at the time of,” “in response to,” or “in response to.” Accordingly, the phrase “in the case of being determined” or “in the case where [the mentioned condition or event] is detected” may be interpreted, depending on the context, to mean “when determining,” “in response to determining,” “when [the mentioned condition or event] is detected,” “in response to detecting [the mentioned condition or event],” or “in response to detecting [the mentioned condition or event].”
[0323] As defined herein, the term “high-level language” or “HLL” refers to a programming language or set of instructions used to program a data processing system where the instructions have strong abstract concepts from the details of the data processing system, for example, machine language. For example, an HLL can automate or hide aspects of the operation of a data processing system, such as memory management. Although referred to as HLLs, these languages are typically classified as “efficiency-level languages.” HLLs directly expose hardware-assisted programming models. Examples of HLLs include, but are not limited to, C, C++, and other appropriate languages.
[0324] HLL can be contrasted with hardware description languages (HDL), such as Verilog, System Verilog, and VHDL, which are used to describe digital circuits. HDLs allow designers to generate definitions of digital circuit designs that can typically be compiled into a register transfer level (RTL) netlist, which is independent of the technology.
[0325] As defined herein, the terms “in response to” and similar language as described above, such as “in the case of,” “when,” or “at the time of,” mean to readily respond to or react to an action or event. The response or reaction is performed automatically. Thus, where a second action is performed “in response to” a first action, a causal relationship exists between the occurrence of the first action and the occurrence of the second action. The term “in response to” indicates a causal relationship.
[0326] As defined herein, the terms “one embodiment,” “an embodiment,” “one or more embodiments,” “specific embodiments,” or similar language mean that a specific feature, structure, or characteristic described in relation to an embodiment is included in at least one embodiment described within this disclosure. Accordingly, throughout this disclosure, appearances of the phrases “in one embodiment,” “in an embodiment,” “in one or more embodiments,” “in specific embodiments,” and similar language may all refer to the same embodiment, but are not necessarily so. The terms “an embodiment” and “arrangement” are used interchangeably within this disclosure.
[0327] As defined herein, the term “output” means storing in physical memory elements, e.g., devices, writing to a display or other peripheral output device, transmitting or sending to another system, exporting, etc.
[0328] As defined herein, the term “substantially” means that the mentioned characteristic, parameter, or value does not need to be achieved exactly, but that deviations or variations, including, for example, tolerances, measurement errors, limitations on measurement accuracy, and other factors known to those skilled in the art, may occur in amounts that do not exclude the effect the characteristic was intended to provide.
[0329] Terms such as first, second, etc., may be used herein to describe various elements. Unless otherwise noted or otherwise clearly indicated in the context, these terms are used solely to distinguish one element from another, and thus these elements should not be limited by these terms.
[0330] A computer program product may include a computer-readable storage medium (or media) having computer-readable program instructions for a processor to perform aspects of the arrangements of the invention described herein. Within this disclosure, the term “program code” is used interchangeably with the term “computer-readable program instructions.” The computer-readable program instructions described herein may be downloaded from a computer-readable storage medium to individual computing / processing devices, or downloaded to an external computer or external storage device via a network, such as the Internet, LAN, WAN, and / or wireless network. A network may include edge devices, including copper transmission cables, transmission optical fibers, wireless transmissions, routers, firewalls, switches, gateway computers, and / or edge servers. A network adapter card or network interface within each computing / processing device receives computer-readable program instructions from the network and forwards the computer-readable program instructions for storage in a computer-readable storage medium within the individual computing / processing device.
[0331] Computer-readable program instructions for performing operations for the arrangements of the invention described herein may be code of source code or object code written in any combination of one or more programming languages, including assembler instructions, ISA (instruction-set-architecture) instructions, machine instructions, machine-dependent instructions, microcode, firmware instructions, or object-oriented programming languages and / or procedural programming languages. Computer-readable program instructions may include state-setting data. Computer-readable program instructions may be executed wholly on a user's computer, partially on a user's computer, as a standalone software package, partially on a user's computer and partially on a remote computer, or wholly on a remote computer or server. In the latter scenario, the remote computer may be connected to the user's computer through any type of network including a LAN or WAN, or the connection may be made to an external computer (e.g., over the Internet using an Internet service provider). In some cases, an electronic circuit comprising, for example, a programmable logic circuit, an FPGA, or a PLA may execute computer-readable program instructions by utilizing state information of computer-readable program instructions to personalize the electronic circuit in order to perform aspects of the arrangements of the invention described herein.
[0332] Specific aspects of the arrangements of the present invention are described herein with reference to flowcharts and / or block diagrams of methods, apparatuses (systems), and computer program products. It will be understood that each block of the flowcharts and / or block diagrams, and combinations of blocks within the flowcharts and / or block diagrams, may be implemented by computer-readable program instructions, such as program code.
[0333] These computer-readable program instructions may be provided to a processor of a general-purpose computer, a special-purpose computer, or another programmable data processing unit to create a machine, and instructions executed through the processor of the computer or other programmable data processing unit create means for implementing specific functions / operations in flowcharts and / or block diagram blocks or blocks. These computer-readable program instructions, which can instruct a computer, a programmable data processing unit, and / or other devices to function in a specific manner, may also be stored in a computer-readable storage medium, and the computer-readable storage medium in which the instructions are stored comprises a manufactured article containing instructions that implement aspects of specific operations in flowcharts and / or block diagram blocks or blocks.
[0334] Computer-readable program instructions can also be loaded onto a computer, another programmable data processing device, or another device to cause a series of operations to be performed on a computer, another programmable device, or another device to create a process that is implemented by a computer, so that the instructions executed on the computer, another programmable device, or other device implement functions / operations specified in blocks or blocks of a flowchart and / or block diagram.
[0335] The flowcharts and block diagrams in the drawings illustrate the architecture, function, and operation of possible implementations of systems, methods, and computer program products according to various aspects of the arrangements of the present invention. In this regard, each block of the flowcharts or block diagrams may represent a module, segment, or part of instructions comprising one or more executable instructions for implementing specific operations.
[0336] In some alternative implementations, the operations mentioned in the blocks may occur out of the order mentioned in the drawings. For example, depending on the accompanying function, two consecutively drawn blocks may be executed substantially simultaneously, or the blocks may often be executed in reverse order. In other examples, blocks may generally be executed in increasing numerical order, whereas in other examples, one or more blocks may be executed in various orders, so that the results are stored and utilized in subsequent or other blocks that are not immediately following. It should also be noted that individual blocks of the block diagrams and / or flowcharts, and combinations of blocks of the block diagrams and / or flowcharts, may be implemented by special-purpose hardware-based systems that perform specific functions or operations or perform combinations of special-purpose hardware and computer instructions.
[0337] All means or steps + corresponding structures, materials, operations, and equivalents of which may be found in the claims below are intended to include any structure, material, or operation for performing a function in combination with other claimed elements as specifically claimed.
[0338] The method may include the step of generating, using a processor, a logical architecture for an application and a first interface solution that specifies the mapping of logical resources to the hardware of an interface circuit block between a DPE array and programmable logic, in the case of an application that specifies a software part for implementation within a DPE array of a device and a hardware part for implementation within a PL of the device. The method includes the step of constructing a block diagram of the hardware part based on the logical architecture and the first interface solution and the step of performing an implementation flow on the block diagram using a processor. The method includes the step of compiling the software part of the application for implementation in one or more DPEs of the DPE array using a processor.
[0339] In another aspect, the step of constructing a block diagram includes the step of adding at least one IP core to the block diagram for implementation within programmable logic.
[0340] In another aspect, during the implementation flow, the hardware compiler constructs the block diagram and performs the implementation flow by exchanging design data with the DPE compiler configured to compile the software parts.
[0341] In another aspect, the hardware compiler exchanges additional design data with the NoC compiler. The hardware compiler receives a first NoC solution configured to implement roots through the device's NoC, which couples the DPE array to the device's PL.
[0342] In another aspect, the step of executing the implementation flow is performed based on the exchanged design data.
[0343] In another aspect, the step of compiling the software part is performed based on the implementation of the hardware part of the application for implementation in the PL, which is generated from the implementation flow.
[0344] In another aspect, the method includes the step of providing constraints on an interface circuit block to a DPE compiler configured to compile a software part in response to a hardware compiler configured to build a block diagram and perform an implementation flow determining that the implementation of the block diagram does not satisfy design metrics for the hardware part. The hardware compiler receives from the DPE compiler a second interface solution generated by the DPE compiler based on the constraints.
[0345] In another aspect, the step of performing the implementation flow is performed based on the second interface solution.
[0346] In another aspect, the hardware compiler provides constraints on the NoC to the NoC compiler in response to a determination that the implementation of the block diagram does not satisfy design metrics using a first NoC solution for the NoC. The hardware compiler receives from the NoC compiler a second NoC solution generated by the NoC compiler based on the constraints on the NoC.
[0347] The system includes a processor configured to initiate operations. The operations include generating a logical architecture for an application and a first interface solution that specifies the mapping of logical resources to the hardware of an interface circuit block between a DPE array and a PL, for an application that specifies a software portion for implementation within a DPE array of the device and a hardware portion for implementation within a PL of the device. The operations include constructing a block diagram of the hardware portion based on the logical architecture and the first interface solution, performing an implementation flow on the block diagram, and compiling the software portion of the application for implementation in one or more DPEs of the array of DPEs.
[0348] In another aspect, the operation of constructing a block diagram includes the operation of adding at least one IP core to the block diagram for implementation within the PL.
[0349] In another aspect, the operations include executing a hardware compiler that performs the implementation flow by exchanging design data with a DPE compiler configured to build block diagrams and compile software parts during the implementation flow.
[0350] In another aspect, operations include the hardware compiler exchanging additional design data with the NoC compiler and receiving a first NoC solution configured to enable the hardware compiler to implement roots through the device's NoC, which couples the DPE array to the device's PL.
[0351] In another aspect, the operation performing the implementation flow is carried out based on the exchanged design data.
[0352] In another aspect, the operation of compiling the software part is performed based on the hardware design for the hardware part of the application for implementation in the PL, which is generated from the implementation flow.
[0353] In another aspect, the operations include providing constraints on the interface circuit block to a DPE compiler configured to compile the software part in response to a hardware compiler configured to build the block diagram and perform the implementation flow determining that the implementation of the block diagram does not satisfy design constraints on the hardware part. The hardware compiler receives from the DPE compiler a second interface solution generated by the DPE compiler based on the constraints.
[0354] In another aspect, the operation performing the implementation flow is performed based on the second interface solution.
[0355] In another aspect, the hardware compiler provides constraints on the NoC to the NoC compiler in response to a determination that the implementation of the block diagram does not satisfy design metrics using a first NoC solution for the NoC. The hardware compiler receives from the NoC compiler a second NoC solution generated by the NoC compiler based on the constraints on the NoC.
[0356] The method comprises, for an application having a software part for implementation in a DPE array of a device and a hardware part for implementation in a PL of the device, using a processor running a hardware compiler, performing an implementation flow for the hardware part based on an interface block solution that maps logical resources used by the software part to the hardware of the interface block coupling the DPE array to the PL. The method comprises, in response to failure to satisfy design metrics during the implementation flow, providing interface block constraints to a DPE compiler using a processor running a hardware compiler. The method also comprises, in response to receiving interface block constraints, generating an updated interface block solution using a processor running a DPE compiler, and providing the updated interface block solution from the DPE compiler to the hardware compiler.
[0357] In another aspect, interface block constraints map the logical resources used by the software part to the physical resources of the interface block.
[0358] In another aspect, the hardware compiler continues the implementation flow using the updated interface block solution.
[0359] In another aspect, the hardware compiler repeatedly provides interface block constraints to the DPE compiler in response to the failure to satisfy design constraints for the hardware part.
[0360] In another aspect, interface block constraints include hard constraints and soft constraints. In this case, the method includes the step of a DPE compiler routing the software portion of an application using both hard constraints and soft constraints to generate an updated interface block solution.
[0361] In another aspect, the method includes the step of routing a software part of an application using only hard constraints to generate an updated interface block in response to a failure to generate an updated interface block solution using both hard constraints and soft constraints.
[0362] In another aspect, the method includes the step of generating an updated interface block solution by mapping the software part using both hard and soft constraints and routing the software part using only hard constraints in response to the failure to generate an updated mapping using only hard constraints.
[0363] In another aspect, each of the interface block solution and the updated interface block solution has a score, and the method includes the steps of comparing the scores and, in response to a determination that the score for the interface block solution exceeds the score for the updated interface block solution, relaxing the interface block constraints and submitting the relaxed interface block constraints to the DPE compiler to obtain an additional updated interface block solution.
[0364] In another aspect, the interface block solution and the updated interface block solution each have a score. The method includes the step of comparing the scores, and the step of performing an implementation flow using the updated interface block solution in response to a determination that the score for the updated interface block solution exceeds the score for the interface block solution.
[0365] The system includes a processor configured to initiate operations. The operations include, for an application having a software part for implementation in a device's DPE array and a hardware part for implementation in a device's PL, using a hardware compiler, performing an implementation flow for the hardware part based on an interface block solution that maps logical resources used by the software part to the hardware of the interface block coupling the DPE array to the PL. The operations include, using the hardware compiler, providing interface block constraints to the DPE compiler in response to failure to satisfy design metrics during the implementation flow. The operations further include, using the DPE compiler, generating an updated interface block solution in response to receiving interface block constraints, and providing the updated interface block solution from the DPE compiler to the hardware compiler.
[0366] In another aspect, interface block constraints map the logical resources used by the software part to the physical resources of the interface block.
[0367] In another aspect, the hardware compiler continues the implementation flow using the updated interface block solution.
[0368] In another aspect, the hardware compiler repeatedly provides interface block constraints to the DPE compiler in response to the failure to satisfy design constraints for the hardware part.
[0369] In another aspect, interface block constraints include hard constraints and soft constraints. In this case, the processor is configured to initiate actions including the DPE compiler routing the software portion of the application using both hard constraints and soft constraints to generate an updated interface block solution.
[0370] In another aspect, actions include routing the software portion of the application using only hard constraints to generate an updated interface block solution in response to a failure to generate an updated mapping using both hard and soft constraints.
[0371] In another aspect, the operations include the step of generating an updated interface block solution by mapping the software part using both hard and soft constraints and routing the software part using only hard constraints in response to the failure to generate an updated mapping using only hard constraints.
[0372] In another aspect, the interface block solution and the updated interface block solution each have a score. The processor is configured to initiate operations including an operation to compare the scores, and, in response to a determination that the score for the interface block solution exceeds the score for the updated interface block solution, an operation to relax the interface block constraint and submit the relaxed interface block constraint to the DPE compiler to obtain an additional updated interface block solution.
[0373] In another aspect, the interface block solution and the updated interface block solution each have a score. The processor is configured to initiate operations including an operation to compare the scores, and an operation to perform an implementation flow using the updated interface block solution in response to a determination that the score for the updated interface block solution exceeds the score for the interface block solution.
[0374] The method comprises, for an application specifying a software part for implementation within a DPE array of a device and a hardware part having HLS kernels for implementation within a PL of the device, using a processor to generate a first interface solution that maps logical resources used by the software part to hardware resources of an interface block coupling the DPE array and the PL. The method comprises, using a processor, generating a connectivity graph that specifies connectivity between nodes of the software part to be implemented in the DPE array and HLS kernels, and using a processor to generate a block diagram based on the connectivity graph and HLS kernels, wherein the block diagram is synthesizable. The method further comprises, using a processor, performing an implementation flow on the block diagram based on the first interface solution and using a processor to compile the software part of the application for implementation in one or more DPEs of the DPE array.
[0375] In another aspect, the step of generating a block diagram includes the step of performing HLS on HLS kernels to generate synthesizable versions of HLS kernels and the step of constructing a block diagram using the synthesizable versions of HLS kernels.
[0376] In another aspect, synthesizable versions of HLS kernels are designated as RTL blocks.
[0377] In another aspect, the step of generating a block diagram is performed based on a description of the architecture of the SoC on which the application will be implemented.
[0378] In another aspect, the step of generating a block diagram includes the step of connecting the block diagram to a base platform.
[0379] In another aspect, the step of performing the implementation flow includes the step of synthesizing a block diagram for implementation in the PL, and the step of placing and routing the synthesized block diagram based on the first interface solution.
[0380] In another aspect, the method includes the step of executing a hardware compiler that performs the implementation flow by exchanging design data with a DPE compiler configured to build block diagrams and compile software parts during the implementation flow.
[0381] In another aspect, the method includes the step of a hardware compiler exchanging additional design data with a NoC compiler and receiving a first NoC solution configured such that the hardware compiler implements roots through the NoC of the device, which couples a DPE array to the PL of the device.
[0382] In another aspect, the method includes the step of providing constraints on an interface circuit block to a DPE compiler configured to compile a software part in response to a hardware compiler configured to build a block diagram and perform an implementation flow determining that the implementation of the block diagram does not satisfy design metrics for the hardware part. The method also includes the step of the hardware compiler receiving, from the DPE compiler, a second interface solution generated by the DPE compiler based on the constraints.
[0383] In another aspect, the step of performing the implementation flow is performed based on the second interface solution.
[0384] The system includes a processor configured to initiate operations. The operations include generating a first interface solution that maps logical resources used by the software part to the hardware resources of an interface block coupling the DPE array and the PL, for an application specifying a software part for implementation within the device's DPE array and a hardware part having HLS kernels for implementation within the device's PL. The operations include generating a connection graph that specifies the connectivity between the nodes of the software part to be implemented in the DPE array and the HLS kernels, and generating a block diagram based on the connection graph and the HLS kernels, wherein the block diagram is synthesizable. The operations further include the step of performing an implementation flow on the block diagram based on the first interface solution and compiling the software part of the application for implementation in one or more DPEs of the DPE array.
[0385] In another aspect, the operation of generating a block diagram includes the operation of performing HLS on HLS kernels to generate synthesizable versions of HLS kernels and the operation of constructing a block diagram using the synthesizable versions of HLS kernels.
[0386] In another aspect, synthesizable versions of HLS kernels are designated as RTL blocks.
[0387] In another aspect, the operation of generating a block diagram is performed based on a description of the architecture of the SoC on which the application will be implemented.
[0388] In another aspect, the operation of generating a block diagram includes the operation of connecting the block diagram to a base platform.
[0389] In another aspect, the operation of performing the implementation flow includes the operation of synthesizing a block diagram for implementation in the PL, and the operation of placing and routing the synthesized block diagram based on the first interface solution.
[0390] In another aspect, the operations include executing a hardware compiler that performs the implementation flow by exchanging design data with a DPE compiler configured to build block diagrams and compile software parts during the implementation flow.
[0391] In another aspect, operations include receiving a first NoC solution configured such that the hardware compiler exchanges additional design data with the NoC compiler and the hardware compiler implements roots through the device's NoC, which couples the DPE array to the device's PL.
[0392] In another aspect, the operations include providing constraints on an interface circuit block to a DPE compiler configured to compile a software part in response to a hardware compiler configured to build a block diagram and perform an implementation flow determining that the implementation of the block diagram does not satisfy design metrics for the hardware part. The method also includes the step of the hardware compiler receiving, from the DPE compiler, a second interface solution generated by the DPE compiler based on the constraints.
[0393] In another aspect, the implementation flow is performed based on the second interface solution.
[0394] One or more computer program products comprising a computer-readable storage medium storing program code are disclosed herein. The program code is executable by computer hardware to initiate various operations described within this disclosure.
[0395] The description of the arrangements of the invention provided herein is for illustrative purposes only and is not intended to be comprehensive or limited to the disclosed forms and examples. The terms used herein are chosen to describe the principles of the arrangements of the invention, practical applications of technologies found in the market, or technical improvements, and / or to enable those skilled in the art to understand the arrangements of the invention disclosed herein. Variations and modifications may be obvious to those skilled in the art without departing from the scope and spirit of the described arrangements of the invention. Accordingly, reference should be made to the following claims defining the scope of such features and embodiments rather than to the foregoing disclosure.
[0396] Example 1 illustrates an exemplary schema for a logical architecture derived from an application. Example 1
[0397] Example 2 illustrates an exemplary schema for an SoC interface block solution for an application to be implemented in a DPE array (202). Example 2
[0398] Example 3 illustrates an exemplary schema for a NoC solution for an application to be implemented in NoC (208). Example 3
[0399] Example 4 illustrates an exemplary schema for specifying SoC interface block constraints and / or NoC constraints. Example 4
[0400] Example 5 illustrates an exemplary schema for specifying NoC traffic. Example 5
Claims
Claim 1 A method comprising: using a processor to generate a first interface solution for an application that specifies a software portion for implementation within a data processing engine (DPE) array of a device and a hardware portion for implementation within programmable logic of the device, the first interface solution specifying a logical architecture for the application and a mapping of logical resources for hardware implementations of a plurality of stream channels of an interface circuit block between the DPE array and the programmable logic; building a block diagram of the hardware portion based on the logical architecture and the first interface solution; using the processor to perform an implementation flow on the block diagram; and using the processor to compile the software portion of the application for implementation in one or more DPEs of the DPE array, wherein the interface circuit block includes a plurality of tiles including the plurality of stream channels and is capable of communicating with any DPE of the DPE array by communicating with one or more selected DPEs of the DPE array directly connected to the interface circuit block, and the mapping further specifies specific stream channels of specific tiles among the plurality of tiles of the interface circuit block. Claim 2 A method according to claim 1, further comprising the step of constructing the block diagram and performing the implementation flow by exchanging design data with a DPE compiler configured to compile the software part of the hardware compiler during the implementation flow. Claim 3 A method according to claim 2, further comprising: a step in which the hardware compiler exchanges additional design data with a Network-on-Chip (NoC) compiler; and a step in which the hardware compiler receives a first NoC solution configured to implement routes through the NoC of the device that couple the DPE array to the programmable logic of the device. Claim 4 In claim 2, the step of compiling the software part is performed based on the implementation of the hardware part of the application for implementation in the programmable logic generated from the implementation flow. Claim 5 A method according to claim 1, further comprising: providing constraints on the interface circuit block to a DPE compiler configured to compile the software part in response to a hardware compiler configured to construct the block diagram and perform the implementation flow determining that the implementation of the block diagram does not satisfy design metrics for the hardware part; and the hardware compiler receiving a second interface solution generated by the DPE compiler based on the constraints from the DPE compiler. Claim 6 In claim 5, the step of performing the implementation flow is performed based on the second interface solution. Claim 7 A method according to claim 5, wherein the hardware compiler provides a constraint on the NoC to the NoC compiler in response to determining that the implementation of the block diagram does not satisfy the design metric using a first NoC solution for the network-on-chip (NoC); and the hardware compiler receives from the NoC compiler a second NoC solution generated by the NoC compiler based on the constraint on the NoC. Claim 8 A system comprising a processor configured to initiate operations; wherein the operations include: generating a first interface solution that specifies a logical architecture for an application and a mapping of logical resources for hardware implementations of a plurality of stream channels of an interface circuit block between the DPE array and the programmable logic, for an application that specifies a software portion for implementation within a data processing engine (DPE) array of a device and a hardware portion for implementation within programmable logic of the device; constructing a block diagram of the hardware portion based on the logical architecture and the first interface solution; performing an implementation flow for the block diagram; and compiling the software portion of the application for implementation in one or more DPEs of the DPE array, wherein the interface circuit block comprises a plurality of tiles including the plurality of stream channels and is capable of communicating with any DPE of the DPE array by communicating with one or more selected DPEs of the DPE array directly connected to the interface circuit block, and the mapping further specifies specific stream channels of specific tiles among the plurality of tiles of the interface circuit block. Claim 9 A system according to claim 8, wherein the operation of constructing the block diagram includes the operation of adding at least one IP (Intellectual Property) core to the block diagram for implementation within the programmable logic. Claim 10 A system according to claim 8, wherein the processor is configured to initiate operations further comprising, during the implementation flow, an operation of executing a hardware compiler that constructs the block diagram and performs the implementation flow by exchanging design data with a DPE compiler configured to compile the software part. Claim 11 A system configured to disclose operations, wherein, in claim 10, the processor further comprises: an operation in which the hardware compiler exchanges additional design data with a network-on-chip (NoC) compiler; and an operation in which the hardware compiler receives a first NoC solution configured to implement routes through the NoC of the device that couple the DPE array to the programmable logic of the device. Claim 12 In claim 10, the operation of compiling the software part is performed based on a hardware design for the hardware part of the application for implementation in the programmable logic generated from the implementation flow, in a system. Claim 13 A system configured to disclose operations, wherein the processor, in response to a hardware compiler configured to construct the block diagram and perform the implementation flow determining that the implementation of the block diagram does not satisfy design constraints for the hardware part, provides constraints for the interface circuit block to a DPE compiler configured to compile the software part; and the hardware compiler further comprises the operation of receiving from the DPE compiler a second interface solution generated by the DPE compiler based on the constraints. Claim 14 In claim 13, the operation of performing the above implementation flow is performed based on the second interface solution, in a system. Claim 15 A system according to claim 13, wherein the hardware compiler provides a constraint on the NoC to the NoC compiler in response to determining that the implementation of the block diagram does not satisfy the design metric using a first NoC solution for the network-on-chip (NoC); and the hardware compiler receives from the NoC compiler a second NoC solution generated by the NoC compiler based on the constraint on the NoC.
Citation Information
Patent Citations
Heterogeneous multiprocessor program compilation targeting programmable integrated circuits
KR1020170084206A
Management of memory resources in a programmable integrated circuit
KR1020180008625A
IP cores in reconfigurable three dimensional integrated circuits
US20090070728A1
Computing system with hardware reconfiguration mechanism and method of operation thereof
US20120284501A1
Shared memory interface in a programmable logic device using partial reconfiguration
US7546572B1