Process Convergence in the Hardware-Software Design Process for Heterogeneous Programmable Devices
By generating a mapping solution for logical architecture and interface circuit blocks, the problem of separation of hardware and software design processes in heterogeneous programmable ICs is solved, and the coordinated design of hardware and software is realized, and the design efficiency and quality are improved.
Patent Information
- Application Number
- CN202080038360.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2019-05-23
- Filing Date
- 2020-05-12
- Publication Date
- 2025-09-02
- Estimated Expiration
- 2040-05-12
AI Technical Summary
In the prior art, there is a separation of the hardware and software design process of heterogeneous programmable integrated circuits, resulting in the inability to effectively converge. Hardware design relies on the platform created by hardware designers and cannot meet the software requirements and the interface layout instructions between subsystems.
By generating a mapping scheme of logic architecture and interface circuit blocks, using the processor to build a block diagram of the hardware part, and compiling the software part to implement in the DPE array and programmable logic, using the hardware compiler and the DPE compiler to interact, adjust interface block constraints in response to design indicators to achieve collaborative design of hardware and software.
It realizes efficient integration of hardware and software parts in heterogeneous programmable ICs, shortens design time, improves design quality and feasibility, and meets design indicators such as timing, area and power requirements.
Smart Images

Figure CN113874834B_ABST
Abstract
Description
[0001] Reservation of rights in copyright material
[0002] A portion of the disclosure of this patent document contains material which is subject to copyright protection. The copyright owner has no objection to the facsimile reproduction by anyone of the patent document or the patent disclosure, as it appears in the Patent and Trademark Office patent file or records, but otherwise reserves all copyright rights whatsoever. Technical Field
[0003] The present disclosure relates to integrated circuits (ICs), and more particularly, to implementing an application including hardware and software portions within a heterogeneous and programmable IC. Background Art
[0004] A programmable integrated circuit (IC) is an IC that includes programmable logic. An example of a programmable IC is a field-programmable gate array (FPGA). FPGAs are characterized by their inclusion of programmable circuit blocks. Examples of programmable circuit blocks include, but are not limited to, input / output blocks (IOBs), configurable logic blocks (CLBs), application-specific random access memory (BRAM), digital signal processing blocks (DSPs), processors, clock managers, and delay-locked loops (DLLs).
[0005] Modern programmable ICs have evolved to include programmable logic combined with one or more other subsystems. For example, some programmable ICs have evolved into systems on a chip, or "SoCs," which include programmable logic and a hardwired processor system. Other types of programmable ICs include other and / or different subsystems. The increasing diversity of subsystems included in programmable ICs presents challenges for implementing applications in these devices. The traditional design process for ICs with hardware-based and software-based subsystems (e.g., programmable logic circuits and processors) relies on hardware designers to first create a monolithic hardware design for the IC. This hardware design serves as a platform for subsequently creating, compiling, and executing the software design. This approach is often overly restrictive.
[0006] In other cases, the software and hardware design processes may be separate. However, the separation of the hardware and software design processes does not provide any indication of the software requirements or the layout of interfaces between the various subsystems in the IC. As a result, the hardware and software design processes may not converge on a feasible implementation for the application in the IC. Summary of the Invention
[0007] In one aspect, a method includes, for an application having a software portion designated for implementation in a data processing engine (DPE) array of a device and a hardware portion designated for implementation in programmable logic of the device, generating, using a processor, a logical architecture for the application and a first interface plan that specifies a mapping of logic resources to hardware interface circuit blocks between the DPE array and the programmable logic. The method may include constructing a block diagram for the hardware portion based on the logical architecture and the first interface plan, and executing an implementation flow on the block diagram using the processor. The method may also include compiling, using the processor, the software portion of the application for implementation in one or more DPEs of the DPE array.
[0008] In another aspect, a system includes a processor configured to initiate operations. The operations include generating, for an application specifying a software portion implemented in a DPE array of a device and a hardware portion implemented in a PL of the device, a logical architecture and a first interface plan for the application that specifies a mapping of logical resources to hardware interface circuit blocks between the DPE array and the PL. The operations may include constructing a block diagram of the hardware portion based on the logical architecture and the first interface plan, executing an implementation process on the block diagram, and compiling the software portion of the application for implementation in one or more DPEs of the DPE array.
[0009] In another aspect, a computer program product includes a computer-readable storage medium having program code stored thereon. The program code is executable by computer hardware to initiate operations. The operations may include, for an application that specifies a software portion for implementation within a DPE array of a device and a hardware portion for implementation within a PL of the device, generating a first interface plan for the application that specifies a logical architecture for the application and a hardware mapping of logic resources to interface circuit blocks between the DPE array and the PL. The operations may include constructing a block diagram of the hardware portion based on the logical architecture and the first interface plan, executing an implementation process on the block diagram, and compiling the software portion of the application for implementation in one or more DPEs of the DPE array.
[0010] In another aspect, a method may include, for an application having a software portion for implementation in a DPE array of a device and a hardware portion for implementation in a PL of the device, executing an implementation flow on the hardware portion using a processor executing a hardware compiler based on an interface block plan, the interface block plan mapping logic resources used by the software portion to hardware of an interface block coupling the DPE array to the PL. The method may include, in response to design specifications not being met during the implementation flow, providing interface block constraints to the DPE compiler using the processor executing the hardware compiler. The method may also include, in response to receiving the interface block constraints, generating an updated interface block plan using the processor executing the DPE compiler and providing the updated interface block plan from the DPE compiler to the hardware compiler.
[0011] In another aspect, a system includes a processor configured to initiate operations. The operations include, for an application having a software portion for implementation in a DPE array of a device and a hardware portion for implementation in a PL of the device, executing an implementation process on the hardware portion using a hardware compiler based on an interface block plan that maps logic resources used by the software portion to hardware of an interface block coupling the DPE array to the PL. The operations may include, in response to design specifications not being met during the implementation process, providing interface block constraints to a DPE compiler using the hardware compiler. The operations may also include, in response to receiving the interface block constraints, generating an updated interface block plan using the DPE compiler and providing the updated interface block plan from the DPE compiler to the hardware compiler.
[0012] In another aspect, a computer program product includes a computer-readable storage medium having program code stored thereon. The program code is executable by computer hardware to initiate operations. The operations may include, for an application having a software portion for implementation in a DPE array of a device and a hardware portion for implementation in a PL of the device, performing an implementation process on the hardware portion using a hardware compiler based on an interface block solution that maps logic resources used by the software portion to hardware interface blocks coupling the DPE array to the PL. The operations may include, in response to design specifications not being met during the implementation process, providing, using the hardware compiler, interface block constraints to a DPE compiler. The operations may also include, in response to receiving the interface block constraints, generating, using the DPE compiler, an updated interface block solution and providing the updated interface block solution from the DPE compiler to the hardware compiler.
[0013] In another aspect, a method may include, for an application that specifies a software portion to be implemented within a DPE array of a device and a hardware portion having an HLS kernel to be implemented within a PL of the device, generating, using a processor, a first interface scheme that maps logical resources used by the software portion to hardware resources of an interface block coupling the DPE array and the PL. The method may include generating, using the processor, a connectivity graph that specifies connectivity between nodes of the software portion to be implemented in the DPE array and the HLS kernel, and generating, using the processor, a block diagram based on the connectivity graph and the HLS kernel, wherein the block diagram is synthesizable. The method may also include executing, using the processor, an implementation flow based on the block diagram based on the first interface scheme, and compiling, using the processor, the software portion of the application for implementation in one or more DPEs of the DPE array.
[0014] In another aspect, a system includes a processor configured to initiate operations. The operations may include, for an application that specifies a software portion to be implemented within a DPE array of a device and a hardware portion having an HLS kernel to be implemented within a PL of the device, generating a first interface scheme that maps logical resources used by the software portion to hardware resources of an interface block coupling the DPE array and the PL. The operations may include generating a connectivity graph that specifies connectivity between nodes of the software portion to be implemented in the DPE array and the HLS kernel, and generating a block diagram based on the connectivity graph and the HLS kernel, wherein the block diagram is synthesizable. The operations may also include executing an implementation process on the block diagram based on the first interface scheme and compiling the software portion of the application for implementation in one or more DPEs of the DPE array.
[0015] In another aspect, a computer program product includes a computer-readable storage medium having program code stored thereon. The program code is executable by computer hardware to initiate operations. The operations may include, for an application specifying a software portion for implementation within a DPE array of a device and a hardware portion having an HLS kernel for implementation within a PL of the device, generating a first interface scheme that maps logical resources used by the software portion to hardware resources of an interface block coupling the DPE array and the PL. The operations may include generating a connection graph that specifies connectivity between nodes of the software portion to be implemented in the DPE array and the HLS kernel, and generating a block diagram based on the connection graph and the HLS kernel, wherein the block diagram is synthesizable. The operations may also include executing an implementation process on the block diagram based on the first interface scheme and compiling the software portion of the application for implementation in one or more DPEs of the DPE array.
[0016] This summary is provided only to introduce certain concepts and is not intended to identify any key or essential features of the claimed subject matter. Other features of the inventive arrangement will be apparent from the accompanying drawings and the following detailed description. BRIEF DESCRIPTION OF THE DRAWINGS
[0017] The arrangement of the present invention is shown by way of example in the accompanying drawings. However, the accompanying drawings should not be interpreted as limiting the arrangement of the present invention to only the specific embodiments shown. Various aspects and advantages will become apparent by reading the following detailed description and referring to the accompanying drawings.
[0018] Figure 1 An example of a computing node for use with one or more embodiments described herein is shown;
[0019] Figure 2 An example architecture of a system-on-chip (Soc) type of integrated circuit (IC) is shown;
[0020] Figure 3 Shown Figure 2 Example architecture of a data processing engine (DPE) of a DPE array;
[0021] Figure 4 Shown Figure 3 Other aspects of the example architecture;
[0022] Figure 5 Another example architecture of a DPE array is shown;
[0023] Figure 6 An example architecture of a unit block of a SoC interface block of a DPE array is shown;
[0024] Figure 7 Shown Figure 1 An example implementation of a Network-on-Chip (NoC)
[0025] Figure 8 Shown by NoC Figure 1 A block diagram of the connections between endpoint circuits in a system on a chip (Soc);
[0026] Figure 9 A block diagram of a NoC according to another example is shown;
[0027] Figure 10 An exemplary method of programming Noc is shown;
[0028] Figure 11 Another exemplary method of programming Noc is shown;
[0029] Figure 12 shows an exemplary data path through the Noc between endpoint circuits;
[0030] Figure 13 An exemplary method of processing read / write requests and responses associated with a NoC is shown;
[0031] Figure 14 An exemplary implementation of a Noc master unit is shown;
[0032] Figure 15 An exemplary implementation of a Noc slave unit is shown;
[0033] Figure 16 Shows the combination Figure 1 An example of the software architecture described as being executed by the system;
[0034] Figure 17A and 17B shows that by using the combination Figure 1 The described system maps applications to an example of a SoC;
[0035] Figure 18shows an example implementation of another application that has been mapped to the SoC;
[0036] Figure 19 Shows the combination Figure 1 Another example of a software architecture executed by the system described;
[0037] Figure 20 An example method of performing a design flow to implement an application in a SoC is shown;
[0038] Figure 21 Another example method of performing a design flow to implement an application in a SoC is shown;
[0039] Figure 22 An example method of communication between a hardware compiler and a DPE compiler is shown;
[0040] Figure 23 An exemplary method of processing a SoC interface block solution is shown;
[0041] Figure 24 Another example of an application for implementation in a SoC is shown;
[0042] Figure 25 shows an example of a SoC interface block solution generated by the DPE compiler;
[0043] Figure 26 shows an example of a routable SoC interface block constraint received by a DPE compiler;
[0044] Figure 27 shows an example of a non-routable SoC interface block constraint;
[0045] Figure 28 shows that the DPE compiler ignores Figure 27 Example of soft type SoC interface block constraints;
[0046] Figure 29 Another example of a non-routable SoC interface block constraint is shown;
[0047] Figure 30 Shown Figure 29 Example of mapping DPE nodes;
[0048] Figure 31 Another example of a non-routable SoC interface block constraint is shown;
[0049] Figure 32 Shown Figure 31 Example of mapping DPE nodes;
[0050] Figure 33 Shown by Figure 1Another example software architecture executable by the system;
[0051] Figure 34 Another example method of performing a design flow to implement an application in a SoC is shown;
[0052] Figure 35 Another example method of performing a design flow to implement an application in a SoC is shown. DETAILED DESCRIPTION
[0053] Although the present disclosure is summarized by claims defining novel features, it is believed that the various features described in this disclosure will be better understood by considering the embodiments in conjunction with the accompanying drawings. The processes, machines, manufactures, and any variations thereof described herein are provided for illustrative purposes. The specific structural and functional details described in this disclosure should not be interpreted as limiting, but merely as a basis for the claims and a representative basis for teaching one skilled in the art to use the described features in various ways in virtually any appropriately detailed structure. Furthermore, the terms and phrases used in this disclosure are not intended to be limiting, but rather to provide an understandable description of the described features.
[0054] The present disclosure relates to integrated circuits (ICs), and more particularly, to implementing applications including hardware and software portions in heterogeneous programmable ICs. An example of a heterogeneous programmable IC is a device (e.g., an integrated circuit) that includes programmable circuitry, referred to herein as "programmable logic" or "PL," and a plurality of hardwired and programmable data processing engines (DPEs). The plurality of DPEs may be arranged in an array that is communicatively linked to the PL of the IC via a system-on-chip (SoC) interface block. As defined in the present disclosure, a DPE is a hardwired and programmable circuit block that includes a core capable of executing program code and a memory module coupled to the core. As described in more detail in the present disclosure, the DPEs are capable of communicating with each other.
[0055] An application intended to be implemented in the described device includes a hardware portion implemented using the device's PL and a software portion implemented in and executed by the device's DPE array. The device may also include a hardwired processor system or "PS" capable of executing additional program code (e.g., another software portion of the application). For example, the PS includes a central processing unit or "CPU" or other hardwired processor capable of executing program code. Thus, the application may also include additional software portions intended to be executed by the CPU of the PS.
[0056] According to the inventive arrangements described in this disclosure, a design flow that can be executed by a data processing system is provided. The design flow can implement the hardware and software portions of an application within a heterogeneous programmable integrated circuit including a programmable logic controller (PL), a DPE array, and / or a power supply (PS). The IC can also include a programmable network-on-chip (NOC).
[0057] In some embodiments, an application is specified as a dataflow graph comprising a plurality of interconnected nodes. Nodes of the dataflow graph are specified for implementation within an array of DPEs or within the PL. For example, a node implemented in a DPE is ultimately mapped to a specific DPE in the array of DPEs. Object code is generated to be executed by each DPE in the array for the application to implement the node. For example, a node implemented in the PL can be synthesized and implemented in the PL, or implemented using a pre-built core (e.g., a register transfer level or "RTL" core).
[0058] This inventive arrangement provides an example design flow that can coordinate the construction and integration of different parts of an application for implementation in different heterogeneous subsystems of an IC. Different stages in the example design flow are targeted at specific subsystems. For example, one or more stages of the design flow are targeted at implementing the hardware portion of the application in the PL, while one or more other stages of the design flow are targeted at implementing the software portion of the application in the DPE array. However, one or more other stages of the design flow are targeted at implementing another software portion of the application in the PS. Still other stages of the design flow are targeted at implementing routing or data transfer between different subsystems and / or circuit blocks through the NoC.
[0059] Different stages of the example design flow corresponding to different subsystems can be performed by different compilers specific to the subsystem. For example, the software part can be implemented using a DPE compiler and / or a PS compiler. The hardware part to be implemented in the PL can be implemented by a hardware compiler. The routing of the NoC can be implemented by a NoC compiler. The various compilers are able to communicate and interact with each other while implementing the respective subsystems specified by the application in order to converge to a solution that can feasibly implement the application in the IC. For example, the compilers are able to exchange design data during operation to converge to a solution that meets the design indicators specified for the application. In addition, the implemented solution (e.g., the implementation of the application in the device) is a solution that maps the various parts of the application to the various subsystems in the device and the interfaces between the different subsystems are consistent and mutually agreed upon.
[0060] Using the example design flows described in this disclosure, a system is able to implement an application within a heterogeneous programmable IC in a shorter time (e.g., less execution time) than would otherwise be the case, for example, where all parts of the application are implemented together on the device. Furthermore, the example design flows described in this disclosure achieve feasibility and quality (e.g., convergence of design metrics such as timing, area, power, etc.) of the final implementation of the application within the heterogeneous programmable IC that is generally superior to results obtained using other conventional techniques in which each part of the application is mapped completely independently and then spliced or combined together. The example design flows achieve these results at least in part through the loosely-coupled co-convergence techniques described herein that rely on shared interface constraints between different subsystems.
[0061] Other aspects of the present invention are described in more detail below with reference to the accompanying drawings. For simplicity and clarity of illustration, the elements shown in the figures are not necessarily drawn to scale. For example, the dimensions of some elements may be exaggerated relative to other elements for clarity. In addition, where appropriate, reference numerals are repeated in multiple figures to indicate corresponding, similar, or analogous features.
[0062] Figure 1 An example of a computing node 100 is shown. Computing node 100 may include a host data processing system (host system) 102 and a hardware accelerator board 104. Computing node 100 is only one example implementation of a computing environment that may be used with a hardware accelerator board. In this regard, computing node 100 may be used independently, as a bare metal server, as part of a computing cluster, or as a cloud computing node within a cloud computing environment. Figure 1 It is not intended to imply any limitation on the scope of use or functionality of the examples described herein. Compute node 100 is an example of a system and / or computer hardware capable of performing the various operations described in this disclosure in connection with implementing applications within SoC 200. For example, compute node 100 may be used to implement an electronic design automation (EDA) system.
[0063] The host system 102 can operate with numerous other general-purpose or special-purpose computing system environments or configurations. Examples of computing systems, environments, and / or configurations that may be suitable for use with the host system 102 include, but are not limited to, personal computer systems, server computer systems, thin clients, thick clients, handheld or laptop devices, multiprocessor systems, microprocessor-based systems, set-top boxes, programmable consumer electronics, network PCs, minicomputer systems, mainframe computer systems, and distributed cloud computing environments that include any of the foregoing.
[0064] As shown, host system 102 is illustrated in the form of a computing device, such as a computer or server. Host system 102 can be implemented as a standalone device, implemented in a cluster, or implemented in a distributed cloud computing environment (wherein tasks are performed by a remote processing device linked by a communication network). In a distributed cloud computing environment, program modules can be located in a local and remote computer system storage medium comprising a memory storage device. The components of host system 102 may include, but are not limited to, one or more processors 106 (e.g., central processing unit), memory 108, and the bus 110 that couples the various system components comprising memory 108 to processor 106. Processor 106 may include any one of a variety of processors capable of executing program code. Example processor types include, but are not limited to, processors with x86 type architecture (IA-32, IA-64, etc.), Power architecture, ARM processors, etc.
[0065] Bus 110 represents one or more of a variety of communication bus structures, including a memory bus or memory controller, a peripheral bus, an accelerated graphics port, and a processor or local bus using any of a variety of available bus architectures. By way of example and not limitation, such architectures include an Industry Standard Architecture (ISA) bus, a Micro Channel Architecture (MCA) bus, an Enhanced ISA (EISA) bus, a Video Electronics Standards Association (VESA) local bus, a Peripheral Component Interconnect (PCI) bus, and a PCI Express (PCIe) bus.
[0066] The host system 102 typically includes a variety of computer-readable media. Such media can be any available media that can be accessed by the host system 102 and can include any combination of volatile media, non-volatile media, removable media, and / or non-removable media.
[0067] The memory 108 may include computer-readable media in the form of volatile memory, such as random access memory (RAM) 112 and / or cache memory 114. The host system 102 may also include other removable / non-removable, volatile / non-volatile computer system storage media. For example, a storage system 116 may be provided for reading from and writing to a non-removable non-volatile magnetic medium (not shown and typically referred to as a "hard drive"). Although not shown, a magnetic disk drive for reading from and writing to a removable non-volatile magnetic disk (e.g., a "floppy disk") and an optical disk drive for reading from or writing to a removable non-volatile optical disk (e.g., a CD-ROM, DVD-ROM, or other optical media) may be provided. In this case, each may be connected to the bus 110 via one or more data media interfaces. As will be further depicted and described below, the memory 108 may include at least one computer program product having a set (e.g., at least one) program module (e.g., program code) configured to perform the functions and / or operations described in the present disclosure.
[0068] A program / utility 118 having a set (at least one) of program modules 120 may be stored, by way of example and not limitation, in the memory 108, as well as an operating system, one or more application programs, other program modules, and program data. The program modules 120 generally implement the functions and / or methods of the embodiments of the present invention as described herein. For example, the program modules 120 may include one or more applications and a driver or protection program for communicating with the hardware accelerator board 104 and / or the SoC 200.
[0069] Program / utility 118 is executable by processor 106. Program / utility 118 and any data items used, generated, and / or manipulated by processor 106 are functional data structures that impart functionality when used by processor 106. As defined herein, a "data structure" is a physical implementation of the data organization of a data model within physical memory. Thus, a data structure is formed by specific electrical or magnetic structural elements in memory. When used by an application executed using the processor, the data structure imposes a physical organization on the data stored in memory.
[0070] The host system 102 may include one or more input / output (I / O) interfaces 128 communicatively linked to the bus 110. The I / O interfaces 128 allow the host system 102 to communicate with external devices, to couple to external devices that allow a user to interact with the host system 102, to couple to external devices that allow the host system 102 to communicate with other computing devices, and the like. For example, the host system 102 may be communicatively linked to a display 130 and a hardware accelerator board 104 via the I / O interfaces 128. The host system 102 may be coupled to other external devices, such as a keyboard (not shown), via the I / O interfaces 128. Examples of the I / O interfaces 128 may include, but are not limited to, a network card, a modem, a network adapter, a hardware controller, and the like.
[0071] In an exemplary implementation, the I / O interface 128 through which the host system 102 communicates with the hardware accelerator board 104 is a PCIe adapter. The hardware accelerator board 104 can be implemented as a circuit board, such as a card, coupled to the host system 102. The hardware accelerator board 104 can be inserted into a card slot, such as an available bus and / or PCIe slot of the host system 102.
[0072] The hardware acceleration board 104 includes a SoC 200. The SoC 200 is a heterogeneous programmable IC and therefore has multiple heterogeneous subsystems. Figure 2 An example architecture of the SoC 200 is described in more detail. The hardware accelerator board 104 also includes a volatile memory 134 coupled to the SoC 200 and a non-volatile memory 136 also coupled to the SoC 200. The volatile memory 134 can be implemented as RAM and is considered "local memory" of the SoC 200, while the memory 108 within the host system 102 is not considered local to the SoC 200, but rather local to the host system 102. In some embodiments, the volatile memory 134 can include multiple gigabytes of RAM, such as 64GB of RAM. Examples of non-volatile memory 136 include flash memory.
[0073] exist Figure 1 In the example of FIG, the computing node 100 is capable of operating an application of the SoC 200 and implementing the application within the SoC 200. The application may include hardware and software components corresponding to the different heterogeneous subsystems available in the SoC 200. Generally, the computing node 100 is capable of mapping the application to the SoC 200 for execution by the SoC 200.
[0074] Figure 2 An example architecture of SoC 200 is shown. SoC 200 is an example of a programmable IC and an integrated programmable device platform. Figure 2In the example shown, the various subsystems or areas of the SoC 200 can be implemented on a single die provided within a single integrated package. In other examples, the different subsystems can be implemented on multiple interconnected dies provided as a single integrated package.
[0075] In this example, SoC 200 includes multiple regions of circuits with different functions. In this example, SoC 200 optionally includes a data processing engine (DPE) array 202. SoC 200 includes a programmable logic (PL) region 214 (hereinafter referred to as the PL region or PL), a processing system (PS) 212, a network on chip (NoC) 208, and one or more hardwired circuit blocks 210. DPE array 202 is implemented as multiple interconnected, hardwired, and programmable processors that have interfaces to other regions of SoC 200.
[0076] PL 214 is a circuit that can be programmed to perform a specified function. As an example, PL 214 can be implemented as a field programmable gate array type circuit. PL 214 may include an array of programmable circuit blocks. Examples of programmable circuit blocks within PL 214 include, but are not limited to, configurable logic blocks (CLBs), dedicated random access memory blocks (BRAM and / or UltraRAM or URAM), digital signal processing blocks (DSPs), clock managers, and / or delay locked loops (DLLs).
[0077] Each programmable circuit block within the PL 214 typically includes programmable interconnect circuitry and programmable logic circuitry. The programmable interconnect circuitry typically includes a large number of interconnect lines of varying lengths interconnected by programmable interconnect points (PIPs). Typically, the interconnect lines are configured (e.g., on a per-line basis) to provide connectivity on a bit-by-bit basis (e.g., where each line carries one bit of information). The programmable logic circuitry implements user-designed logic using programmable elements, which may include, for example, lookup tables, registers, arithmetic logic, and the like. The programmable interconnect and programmable logic circuitry can be programmed by loading configuration data into internal configuration memory cells that define how the programmable elements are configured and operate.
[0078] PS 212 is implemented as a hard-wired circuit, which is manufactured as a part for SoC 200. PS 212 can be implemented as, or include, any of a variety of different processor types, each of which can execute program code. For example, PS 212 can be implemented as a separate processor, such as a single core that can execute program code. In another example, PS 212 can be implemented as a multi-core processor. In another example, PS 212 can include one or more cores, modules, coprocessors, interfaces and / or other resources. PS 212 can use any of a variety of different types of architectures to implement. The example architectures that can be used to implement PS 212 can include but are not limited to ARM processor architecture, x86 processor architecture, GPU architecture, mobile processor architecture, DSP architecture, other suitable architectures that can execute computer-readable instructions or program code, and / or the combination of different processors and / or processor architectures.
[0079] NoC 208 includes an interconnect network for sharing data between endpoint circuits in SoC 200. Endpoint circuits can be arranged in DPE array 202, PL area 214, PS 212, and / or hardwired circuit blocks 210. NoC 208 can include high-speed data paths with dedicated switches. In one example, NoC 208 includes horizontal paths, vertical paths, or both horizontal and vertical paths. Figure 1 The arrangement and number of regions shown are examples only. NoC 208 is an example of a common infrastructure within SoC 200 that can be used to connect selected components and / or subsystems.
[0080] NoC 208 provides connectivity to PL 214, PS 212, and selected circuit blocks in hardwired circuit block 210. NoC 208 is programmable. In cases where a programmable NoC is used with other programmable circuits, the network and / or data transmissions to be routed through NoC 208 are unknown before a user circuit design is created for implementation within SoC 200. NoC 208 can be programmed by loading configuration data into internal configuration registers that define how elements within NoC 208 (e.g., switches and interfaces) are configured and operate to pass data from switch to switch and between NoC interfaces.
[0081] The NoC 208 is manufactured as part of the SoC 200 and, while not physically modifiable, can be programmed to establish connections between different master circuits and different slave circuits in a user circuit design. For example, the NoC 208 can include a plurality of programmable switches capable of establishing a packet-switched network connecting user-specified master circuits and slave circuits. In this regard, the NoC 208 can accommodate different circuit designs, each of which has a different combination of master circuits and slave circuits implemented at different locations in the SoC 200 and that can be coupled by the NoC 208. The NoC 208 can be programmed to route data, such as application data and / or configuration data, between the master circuits and slave circuits in the user circuit design. For example, the NoC 208 can be programmed to couple different user-specified circuits implemented within the PL 214 to the PS 212 and / or the DPE array 202, to different hardwired circuit blocks, and / or to different circuits and / or systems external to the SoC 200.
[0082] Hardwired circuit blocks 210 may include input / output (I / O) blocks and / or transceivers for sending and receiving signals to circuits and / or systems external to SoC 200, memory controllers, and the like. Examples of different I / O blocks may include single-ended and pseudo-differential I / O, as well as high-speed differential clock transceivers. Furthermore, hardwired circuit blocks 210 may be implemented to perform specific functions. Other examples of hardwired circuit blocks 210 include, but are not limited to, cryptographic engines, digital-to-analog converters, analog-to-digital converters, and the like. Hardwired circuit blocks 210 within SoC 200 may be referred to herein from time to time as application-specific blocks.
[0083] exist Figure 2 In the example shown, PL 214 is shown in two separate regions. In another example, PL 214 can be implemented as a unified region of programmable circuitry. In yet another example, PL 214 can be implemented as more than two distinct regions of programmable circuitry. The specific organization of PL 214 is not intended to be limiting. In this regard, SoC 200 includes one or more PL regions 214, PS 212, and NoC 208.
[0084] In other example embodiments, the SoC 200 may include two or more DPE arrays 202 located in different areas of the IC. In still other examples, the SoC 200 may be implemented as a multi-die IC. In that case, each subsystem may be implemented on a different die. The different dies may be communicatively linked using any of a variety of available multi-die IC technologies, such as stacking the dies side-by-side on an interposer, using a stacked die architecture in which the IC is implemented as a multi-chip module (MCM), and the like. In the multi-die IC example, it should be understood that each die may include a single subsystem, two or more subsystems, a subsystem and another partial subsystem, or any combination thereof.
[0085] The DPE array 202 is implemented as a two-dimensional array of DPEs 204 including SoC interface blocks 206. The DPE array 202 may be implemented using any of a variety of different architectures described in more detail below. For purposes of illustration and not limitation, Figure 2 The DPEs 204 are shown arranged in aligned rows and aligned columns. However, in other embodiments, the DPEs 204 may be arranged so that the DPEs in a selected row and / or column are horizontally inverted or flipped relative to the DPEs in adjacent rows and / or columns. In one or more other embodiments, the rows and / or columns of DPEs may be offset relative to adjacent rows and / or columns. One or more or all of the DPEs 204 may be implemented to include one or more cores, each capable of executing program code. The number of DPEs 204, the specific arrangement of the DPEs 204, and / or the orientation of the DPEs 204 are not intended to be limiting.
[0086] The SoC interface block 206 can couple the DPE 204 to one or more other subsystems of the SoC 200. In one or more embodiments, the SoC interface block 206 is coupled to adjacent DPEs 204. For example, the SoC interface block 206 can be directly coupled to each DPE 204 in the bottom row of DPEs in the DPE array 202. In the illustration, the SoC interface block 206 can be directly connected to DPEs 204-1, 204-2, 204-3, 204-4, 204-5, 204-6, 204-7, 204-8, 204-9, and 204-10.
[0087] Figure 2The SoC interface block 206 is provided for illustrative purposes. In other embodiments, the SoC interface block 206 may be located at the top of the DPE array 202, to the left of the DPE array 202 (e.g., as a column), to the right of the DPE array 202 (e.g., as a column), or at multiple locations in and around the DPE array 202 (e.g., as one or more intervening rows and / or columns within the DPE array 202). Depending on the layout and location of the SoC interface block 206, the specific DPEs coupled to the SoC interface block 206 may vary.
[0088] For purposes of illustration, if the SoC interface block 206 is located on the left side of the DPE 204, the SoC interface block 206 can be directly coupled to the left column of DPEs, including DPE 204-1, DPE 204-11, DPE 204-21, and DPE 204-31. If the SoC interface block 206 is located on the right side of the DPE 204, the SoC interface block 206 can be directly coupled to the right column of DPEs, including DPE 204-10, DPE 204-20, DPE 204-30, and DPE 204-40. If SoC interface block 206 is located at the top of DPEs 204, SoC interface block 206 may be coupled to the top row of DPEs, including DPEs 204-31, DPE 204-32, DPE 204-33, DPE 204-34, DPE 204-35, DPE 204-36, DPE 204-37, DPE 204-38, DPE 204-39, and DPE 204-40. If SoC interface block 206 is located in multiple locations, the specific DPEs directly connected to SoC interface block 206 may vary. For example, if SoC interface blocks are implemented as rows and / or columns in DPE array 202, the DPEs directly coupled to SoC interface block 206 may be those DPEs adjacent to SoC interface block 206 on one, multiple, or each side of SoC interface block 206.
[0089] The DPEs 204 are interconnected via DPE interconnects (not shown), which, when considered collectively, form a DPE interconnect network. Thus, the SoC interface block 206 is able to communicate with any DPE 204 of the DPE array 202 by communicating with one or more selected DPEs 204 of the DPE array 202 directly connected to the SoC interface block 206 and utilizing the DPE interconnect network formed by the DPE interconnects implemented within each respective DPE 204.
[0090] The SoC interface block 206 can couple each DPE 204 within the DPE array 202 to one or more other subsystems of the SoC 200. For example, the SoC interface block 206 can couple the DPE array 202 to the NoC 208 and the PL 214. Thus, the DPE array 202 can communicate with circuit blocks implemented in the PL 214, the PS 212, and / or any hardwired circuit blocks 210. For example, the SoC interface block 206 can establish a connection between a selected DPE 204 and the PL 214. The SoC interface block 206 can also establish a connection between a selected DPE 204 and the NoC 208. Through the NoC 208, the selected DPE 204 can communicate with the PS 212 and / or the hardwired circuit blocks 210. The selected DPE 204 can communicate with the hardwired circuit blocks 210 via the SoC interface block 206 and the PL 214. In certain embodiments, SoC interface block 206 may be directly coupled to one or more subsystems of SoC 200. For example, SoC interface block 206 may be directly coupled to PS 212 and / or hardwired circuit block 210.
[0091] In one or more embodiments, the DPE array 202 includes a single clock domain. Other subsystems such as the NoC 208, PL 214, PS 212, and various hardwired circuit blocks 210 may be in one or more separate or distinct clock domains. Nevertheless, the DPE array 202 may include other clocks that can be used to interact with other subsystems within the DPE array. In certain embodiments, the SoC interface block 206 includes a clock signal generator capable of generating one or more clock signals that can be provided or distributed to the DPEs 204 of the DPE array 202.
[0092] The DPE array 202 can be programmed by loading configuration data into internal configuration memory cells (also referred to herein as "configuration registers") that define the connectivity between the DPEs 204 and the SoC interface block 206 and how the DPEs 204 and SoC interface block 206 operate. For example, for a particular DPE 204 or group of DPEs 204 to communicate with a subsystem, the DPE(s) 204 and SoC interface block 206 are programmed to do so. Similarly, for one or more specific DPEs 204 to communicate with one or more other DPEs 204, the DPEs are programmed to do so. The DPEs 204 and SoC interface block 206 can be programmed by loading configuration data into configuration registers within the DPEs 204 and SoC interface block 206, respectively. In another example, a clock signal generator that is part of the SoC interface block 206 can be programmable using configuration data to change the clock frequency provided to the DPE array 202.
[0093] Figure 3 Shown Figure 2 Example architecture of DPE array 202 and DPE 204. Figure 3 In the example of , DPE 204 includes core 302 , memory module 304 , and DPE interconnect 306 . Each DPE 204 is implemented as hardwired and programmable circuit blocks on SoC 200 .
[0094] Core 302 provides the data processing capabilities of DPE 204. Core 302 can be implemented as any of a variety of different processing circuits. Figure 3 In the example of FIG, core 302 includes an optional program memory 308. In an example implementation, core 302 is implemented as a processor capable of executing program code (e.g., computer readable instructions). In that case, program memory 308 is included and can store instructions executed by core 302. Core 302 can be implemented as a CPU, GPU, DSP, vector processor, or other type of processor capable of executing instructions. Core 302 can be implemented using any of the various CPU and / or processor architectures described herein. In another example, core 302 is implemented as a very long instruction word (VLIW) vector processor or a DSP.
[0095] In certain embodiments, program memory 308 is implemented as dedicated program memory dedicated to core 302 (e.g., exclusively accessed by core 302). Program memory 308 can only be used by cores of the same DPE 204. Thus, program memory 308 can only be accessed by core 302 and is not shared with any other DPE or components of another DPE. Program memory 308 may include a single port for read and write operations. Program memory 308 may support program compression and may be addressed using a memory-mapped network portion of DPE interconnect 306, described in more detail below. For example, program memory 308 may be loaded with program code that may be executed by core 302 via the memory-mapped network of DPE interconnect 306.
[0096] Core 302 may include configuration registers 324. Configuration registers 324 may be loaded with configuration data to control the operation of core 302. In one or more embodiments, core 302 may be activated and / or deactivated based on the configuration data loaded into configuration registers 324. Figure 3 In the example of , configuration registers 324 are addressable (eg, readable and / or writable) via a memory-mapped network of DPE interconnect 306, described in greater detail below.
[0097] In one or more embodiments, the memory module 304 can store data used by and / or generated by the core 302. For example, the memory module 304 can store application data. The memory module 304 can include read / write memory, such as random access memory (RAM). Thus, the memory module 304 can store data that can be read and consumed by the core 302. The memory module 304 can also store data written by the core 302 (e.g., results).
[0098] In one or more other embodiments, the memory module 304 can store data, such as application data, that can be used and / or generated by one or more other cores of other DPEs within the DPE array. One or more other cores of the DPE can also read from and / or write to the memory module 304. In certain embodiments, the other cores that can read from and / or write to the memory module 304 can be cores of one or more adjacent DPEs. Another DPE that shares a boundary or demarcation with the DPE 204 (e.g., adjacent) is referred to as a "neighboring" DPE relative to the DPE 204. By allowing the core 302 and one or more other cores from the adjacent DPE to read from and / or write to the memory module 304, the memory module 304 implements shared memory that supports communication between different DPEs and / or cores that can access the memory module 304.
[0099] refer to Figure 2 For example, DPEs 204-14, 204-16, 204-5, and 204-25 are considered to be neighboring DPEs of DPE 204-15. In one example, the core within each of DPEs 204-16, 204-5, and 204-25 can read and write to the memory modules within DPE 204-15. In certain embodiments, only those neighboring DPEs that are adjacent to the memory modules can access the memory modules of DPE 204-15. For example, DPE 204-14, while adjacent to DPE 204-15, may not be adjacent to the memory modules of DPE 204-15 because the core of DPE 204-15 may be located between the core of DPE 204-14 and the memory modules of DPE 204-15. Therefore, in certain embodiments, the core of DPE 204-14 may not be able to access the memory modules of DPE 204-15.
[0100] In certain embodiments, whether a core of a DPE can access a memory module of another DPE depends on the number of memory interfaces included in the memory module and whether the cores are connected to available memory interfaces among the memory interfaces of the memory module. In the above example, the memory module of DPE 204-15 includes four memory interfaces, wherein the cores of each of DPEs 204-16, 204-5, and 204-25 are connected to such memory interfaces. Core 302 within DPE 204-15 is itself connected to the fourth memory interface. Each memory interface may include one or more read and / or write channels. In certain embodiments, each memory interface includes multiple read channels and multiple write channels, enabling a particular core coupled thereto to simultaneously read and / or write to multiple banks within memory module 304.
[0101] In other examples, more than four memory interfaces may be available. These additional memory interfaces may be used to allow DPEs diagonally opposite DPE 204-15 to access the memory modules of DPE 204-15. For example, if cores in a DPE (e.g., DPEs 204-14, 204-24, 204-26, 204-4, and / or 204-6) are also coupled to available memory interfaces of the memory modules in DPE 204-15, such other DPEs will also be able to access the memory modules of DPE 204-15.
[0102] The memory module 304 may include configuration registers 336. The configuration registers 336 may be loaded with configuration data to control the operation of the memory module 304. Figure 3 In the example of , configuration registers 336 (and 324) are addressable (eg, readable and / or writable) via a memory-mapped network of DPE interconnect 306, described in greater detail below.
[0103] exist Figure 3 In the example shown, DPE interconnect 306 is dedicated to DPE 204. DPE interconnect 306 facilitates various operations, including communication between DPE 204 and one or more other DPEs of DPE array 202 and / or with other subsystems of SoC 200. DPE interconnect 306 further enables configuration, control, and debugging of DPE 204.
[0104] In certain embodiments, the DPE interconnect 306 is implemented as an on-chip interconnect. An example of an on-chip interconnect is an Advanced Microcontroller Bus Architecture (AMBA) eXtensible Interface (AXI) bus (e.g., or switch). The AMBA AXI bus is an embedded microcontroller bus interface for establishing on-chip connections between circuit blocks and / or systems. The AXI bus is provided herein as an example of an interconnect circuit that may be used with the inventive arrangements described in this disclosure and is therefore not intended to be limiting. Other examples of interconnect circuits may include other types of buses, crossbars, and / or other types of switches.
[0105] In one or more embodiments, DPE interconnect 306 includes two distinct networks. The first network can exchange data with other DPEs in DPE array 202 and / or other subsystems of SoC 200. For example, the first network can exchange application data. The second network can exchange data for the DPEs (e.g., configuration, control, and / or debug data).
[0106] exist Figure 3 In the example shown, the first network of DPE interconnect 306 is formed by a stream switch 326 and one or more stream interfaces (not shown). For example, stream switch 326 includes a stream interface for connecting to each of core 302, memory module 304, memory mapping switch 332, the upper DPE, the left DPE, the right DPE, and the lower DPE. Each stream interface may include one or more masters and one or more slaves.
[0107] Stream switch 326 can allow non-adjacent DPEs and / or DPEs not coupled to the memory interfaces of memory module 304 to communicate with core 302 and / or memory module 304 via a DPE interconnect network formed by the DPE interconnects of the individual DPEs 204 of DPE array 202 .
[0108] Reference again Figure 2, and using DPE 204-15 as a reference point, stream switch 326 is coupled to and can communicate with another stream switch located in the DPE interconnect of DPE 204-14. Stream switch 326 is coupled to and can communicate with another stream switch located in the DPE interconnect of DPE 204-25. Stream switch 326 is coupled to and can communicate with another stream switch located in the DPE interconnect of DPE 204-16. Stream switch 326 is coupled to and can communicate with another stream switch located in the DPE interconnect of DPE 204-5. In this way, core 302 and / or memory module 304 can also communicate with any DPE within DPE array 202 via the DPE interconnect in the DPE.
[0109] Flow switch 326 can also be used to interact with subsystems (e.g., PL 214 and / or NoC 208). Typically, flow switch 326 is programmed to operate as a circuit-switched flow interconnect or a packet-switched flow interconnect. Circuit-switched flow interconnects enable point-to-point dedicated flows suitable for high-bandwidth communication between DPEs. Packet-switched flow interconnects allow shared flows to time-division multiplex multiple logical flows onto a single physical flow for medium-bandwidth communication.
[0110] The stream switch 326 may include configuration registers (in Figure 3 Configuration registers 334 may be written to configuration registers 334 by way of a memory-mapped network of DPE interconnect 306. The configuration data loaded into configuration registers 334 indicates which other DPEs and / or subsystems (e.g., NoC 208, PL 214, and / or PS 212) DPE 204 will communicate with, and whether such communications are to be established as circuit-switched point-to-point connections or as packet-switched connections.
[0111] The second network of DPE interconnects 306 is formed by a memory-mapped switch 332. Memory-mapped switch 332 includes a plurality of memory-mapped interfaces (not shown). Each memory-mapped interface may include one or more masters and one or more slaves. For example, memory-mapped switch 332 includes a memory-mapped interface for connecting to each of core 302, memory module 304, a memory-mapped switch in a DPE above DPE 204, and a memory-mapped switch in a DPE below DPE 204.
[0112] The memory mapped switch 332 is used to transfer configuration, control and debug data of the DPE 204. Figure 3, the memory-mapped switch 332 can receive configuration data for configuring the DPE 204. The memory-mapped switch 332 can receive the configuration data from a DPE located below the DPE 204 and / or from the SoC interface block 206. The memory-mapped switch 332 can forward the received configuration data to one or more other DPEs above the DPE 204, to the core 302 (e.g., to program memory 308 and / or to configuration registers 324), to the memory module 304 (e.g., to memory within the memory module 304 and / or to configuration registers 336), and / or to the configuration registers 334 within the stream switch 326.
[0113] The DPE interconnect 306 is coupled to the DPE interconnect of each adjacent DPE and / or SoC interface block 206 based on the location of the DPE 204. In general, the DPE interconnects of the DPEs 204 form a DPE interconnect network (which may include a stream network and / or a memory-mapped network). The configuration registers of the stream switches of each DPE can be programmed by loading configuration data via the memory-mapped switches. Through configuration, the stream switches and / or stream interfaces are programmed to establish connections (whether packet-switched or circuit-switched) with other endpoints (whether in one or more other DPEs 204 and / or in the SoC interface block 206).
[0114] In one or more embodiments, DPE array 202 is mapped into the address space of a processor system, such as PS 212. Thus, any configuration registers and / or memory within DPE 204 can be accessed via a memory-mapped interface. For example, memory in memory module 304, program memory 308, configuration registers 324 in core 302, configuration registers 336 in memory module 304, and / or configuration registers 334 can be read and / or written via memory-mapped switch 332.
[0115] exist Figure 3 In the example of , memory mapped switch 332 can receive configuration data for DPE 204. The configuration data can include program code to be loaded into program memory 308 (if included), configuration data to be loaded into configuration registers 324, 334, and / or 336, and / or data to be loaded into memory (e.g., memory banks) of memory module 304. Figure 3 In the example of FIG. 3 , configuration registers 324 , 334 , and 336 are shown as being located within the specific circuit structures (eg, core 302 , stream switch 326 , and memory module 304 ) that the configuration registers are intended to control. Figure 3The examples are for illustration purposes only and illustrate that elements within the core 302, memory module 304, and / or stream switch 326 can be programmed by loading configuration data into corresponding configuration registers. In other embodiments, the configuration registers may be consolidated within specific areas of the DPE 204, while controlling the operation of components distributed throughout the DPE 204.
[0116] Thus, the stream switch 326 can be programmed by loading configuration data into the configuration registers 334. The configuration data programs the stream switch 326 to operate in a circuit-switched mode between two different DPEs and / or other subsystems, or in a packet-switched mode between selected DPEs and / or other subsystems. Thus, the connections established by the stream switch 326 to other stream interfaces and / or switches are programmed by loading the appropriate configuration data into the configuration registers 334 to establish the actual connections or application data paths within the DPE 204 with the other DPEs and / or with the subsystems of the other integrated circuit 300.
[0117] Figure 4 Shown Figure 3 Further aspects of the example architecture. Figure 4 In the example shown, details related to the DPE interconnect 306 are not shown. Figure 4 The connection of core 302 to other DPEs via shared memory is shown. Figure 4 Other aspects of the memory module 304 are also shown. For illustration purposes, Figure 4 Refer to DPE 204-15.
[0118] As shown, the memory module 304 includes a plurality of memory interfaces 402, 404, 406, and 408. Figure 4 , memory interfaces 402 and 408 are abbreviated as "MI". Memory module 304 also includes a plurality of memory banks 412-1 through 412-N. In a particular embodiment, memory module 304 includes eight memory banks. In other embodiments, memory module 304 may include fewer or more memory banks 412. In one or more embodiments, each memory bank 412 is single-ported, thereby allowing a maximum of one access to each memory bank per clock cycle. Where memory module 304 includes eight memory banks 412, this configuration supports eight parallel accesses per clock cycle. In other embodiments, each memory bank 412 is dual-ported or multi-ported, thereby allowing a greater number of parallel accesses per clock cycle.
[0119] exist Figure 4In the example of FIG, each of the memory banks 412-1 to 412-N has a respective arbiter 414-1 to 414-N. Each arbiter 414 can generate a stall signal in response to detecting a conflict. Each arbiter 414 can include arbitration logic. In addition, each arbiter 414 can include a crossbar. Therefore, any subject can write to any particular one or more memory banks 412. Figure 3 As noted, the memory module 304 is coupled to the memory-mapped switch 332, thereby facilitating the reading and writing of data to the memory bank 412. Thus, as part of a configuration, control, and / or debugging process, specific data stored in the memory module 304 may be controlled (e.g., written) via the memory-mapped switch 332.
[0120] The memory module 304 also includes a direct memory access (DMA) engine 416. In one or more embodiments, the DMA engine 416 includes at least two interfaces. For example, one or more interfaces can receive an input data stream from the DPE interconnect 306 and write the received data to the memory bank 412. One or more other interfaces can read data from the memory bank 412 and send the data out through a stream interface (e.g., a stream switch) of the DPE interconnect 306. For example, the DMA engine 416 can include a processor for accessing the memory bank 412. Figure 3 The flow interface of the flow switch 326.
[0121] The memory module 304 can operate as a shared memory accessible by multiple different DPEs. Figure 4 In the example of FIG, memory interface 402 is coupled to core 302 via core interface 428 included in core 302. Memory interface 402 provides core 302 with access to memory group 412 via arbiter 414. Memory interface 404 is coupled to the core of DPE 204-25. Memory interface 404 provides the core of DPE 204-25 with access to memory group 412. Memory interface 406 is coupled to the core of DPE 204-16. Memory interface 406 provides the core of DPE 204-16 with access to memory group 412. Memory interface 408 is coupled to the core of DPE 204-5. Memory interface 408 provides the core of DPE 204-5 with access to memory group 412. Thus, in FIG. Figure 4 In the example shown, each DPE having a shared boundary with the memory module 304 of the DPE 204-15 can read from and write to the memory bank 412. Figure 4 In the example of FIG. 3 , the core of DPE 204 - 14 cannot directly access the memory module 304 of DPE 204 - 15 .
[0122] The core 302 can access the memory modules of other adjacent DPEs via core interfaces 430, 432, and 434. Figure 4 In the example shown, core interface 434 is coupled to a memory interface of DPE 204-25. Thus, core 302 can access the memory modules of DPE 204-25 via core interface 434 and the memory interfaces contained within the memory modules of DPE 204-25. Core interface 432 is coupled to a memory interface of DPE 204-14. Thus, core 302 can access the memory modules of DPE 204-14 via core interface 432 and the memory interfaces contained within the memory modules of DPE 204-14. Core interface 430 is coupled to a memory interface within DPE 204-5. Thus, core 302 can access the memory modules of DPE 204-5 via core interface 430 and the memory interfaces contained within the memory modules of DPE 204-5. As discussed, core 302 can access the memory modules 304 within DPE 204-15 via core interface 428 and memory interface 402.
[0123] exist Figure 4 In the example shown in FIG. 2 , core 302 can read and write to any of the memory modules of DPEs (e.g., DPEs 204-25, 204-14, and 204-5) that share a boundary with core 302 in DPE 204-15. In one or more embodiments, core 302 can view the memory modules within DPEs 204-25, 204-15, 204-14, and 204-5 as a single, contiguous memory (e.g., as a single address space). Thus, the process by which core 302 reads and / or writes to the memory modules of such DPEs is the same as the process by which core 302 reads and / or writes to memory module 304. Assuming this contiguous memory model, core 302 can generate addresses for reading and writing. Core 302 can direct read and / or write requests to the appropriate core interface 428, 430, 432, and / or 434 based on the generated addresses.
[0124] As described above, core 302 can map read and / or write operations in the correct direction based on the addresses of such operations through core interfaces 428, 430, 432, and / or 434. When core 302 generates an address for a memory access, core 302 can decode the address to determine the direction (e.g., the specific DPE to be accessed) and forward the memory operation to the correct core interface in the determined direction.
[0125] Thus, core 302 can communicate with the core of DPE 204-25 via shared memory, which can be a memory module within DPE 204-25 and / or memory module 304 of DPE 204-15. Core 302 can communicate with the core of DPE 204-14 via shared memory, which can be a memory module within DPE 204-14. Core 302 can communicate with the core of DPE 204-5 via shared memory, which can be a memory module within DPE 204-5 and / or memory module 304 of DPE 204-15. Furthermore, core 302 can communicate with the core of DPE 204-16 via shared memory, which can be memory module 304 within DPE 204-15.
[0126] As discussed, the DMA engine 416 may include one or more stream-to-memory interfaces. Through the DMA engine 416, application data may be received from other sources within the SoC 200 and stored in the memory module 304. For example, data may be received from other DPEs that share a boundary with the DPE 204-15 and / or that do not share a boundary with the DPE 204-15 via the stream switch 326. Data may also be received from other subsystems of the SoC (e.g., the NoC 208, the hardwired circuit block 210, the PL 214, and / or the PS 212) via the DPE's stream switch via the SoC interface block 206. The DMA engine 416 is capable of receiving such data from the stream switch and writing the data to one or more appropriate memory banks 412 within the memory module 304.
[0127] The DMA engine 416 may include one or more memory-to-stream interfaces. Through the DMA engine 416, data can be read from one or more memory banks 412 of the memory module 304 and sent to other destinations via the stream interface. For example, the DMA engine 416 can read data from the memory module 304 and send such data to other DPEs that share and / or do not share a boundary with the DPE 204-15 through the stream switch. The DMA engine 416 can also send such data to other subsystems (e.g., the NoC 208, the hardwired circuit block 210, the PL 214, and / or the PS 212) through the stream switch and the SoC interface block 206.
[0128] In one or more embodiments, the DMA engine 416 is programmable by the memory-mapped switch 332 within the DPE 204-15. For example, the DMA engine 416 may be controlled by configuration registers 336. The configuration registers 336 may be written to using the memory-mapped switch 332 of the DPE interconnect 306. In certain embodiments, the DMA engine 416 may be controlled by the stream switch 326 within the DPE 204-15. For example, the DMA engine 416 may include control registers that are writable by the stream switch 326 connected thereto. Based on the configuration data loaded into the configuration registers 324, 334, and / or 336, streams received via the stream switch 326 within the DPE interconnect 306 may be connected to the DMA engine 416 in the memory module 304 and / or directly to the core 302. Based on the configuration data loaded into the configuration registers 324, 334, and / or 336, streams may be sent from the DMA engine 416 (e.g., the memory module 304) and / or the core 302.
[0129] The memory module 304 may also include a hardware synchronization circuit 420 (in Figure 4 In general, the hardware synchronization circuit 420 is capable of synchronizing different cores (e.g., cores of adjacent DPEs), Figure 4 304, and other external entities (e.g., PS 212) that may communicate via DPE interconnect 306. As illustrative and non-limiting examples, hardware synchronization circuitry 420 can synchronize the operations of two different cores, stream switches, memory mapping interfaces, and / or DPEs 204-15 and / or DMAs in different DPEs that access the same (e.g., shared) buffers in memory module 304.
[0130] In the case where the two DPEs are not adjacent, the two DPEs cannot access a common memory module. In that case, application data can be transferred via a data stream (the terms "data stream" and "stream" may be used interchangeably from time to time in this disclosure). Therefore, the local DMA engine can convert the transfer from a local memory-based transfer to a stream-based transfer. In that case, the core 302 and the DMA engine 416 can be synchronized using the hardware synchronization circuit 420.
[0131] The PS 212 can communicate with the core 302 via the memory-mapped switch 332. For example, the PS 212 can access the memory module 304 and the hardware synchronization circuit 420 by initiating memory reads and writes. In another embodiment, the hardware synchronization circuit 420 can also send an interrupt to the PS 212 when the state of the lock changes to avoid polling of the PS 212 by the hardware synchronization circuit 420. The PS 212 can also communicate with the DPE 204-15 through a stream interface.
[0132] In addition to communicating with adjacent DPEs through shared memory modules, and with adjacent and / or non-adjacent DPEs via DPE interconnect 306, core 302 may also include a cascade interface. Figure 4 In the example of FIG, core 302 includes cascade interfaces 422 and 424 (in Figure 4 Cascade interfaces 422 and 424 can provide direct communication with other cores. As shown, cascade interface 422 of core 302 receives input data streams directly from the core of DPE 204-14. The data streams received via cascade interface 422 can be provided to data processing circuitry within core 302. Cascade interface 424 of core 302 can send output data streams directly to the core of DPE 204-16.
[0133] exist Figure 4 In the example of , each of the cascade interface 422 and the cascade interface 424 may include a first-in-first-out (FIFO) interface for buffering. In certain embodiments, the cascade interfaces 422 and 424 are capable of transmitting data streams that may be hundreds of bits wide. The specific bit widths of the cascade interfaces 422 and 424 are not intended to be limiting. Figure 4 In the example of FIG, the cascade interface 424 is coupled to the accumulator register 436 within the core 302 ( Figure 4 Cascade interface 424 can output the contents of accumulator register 436 and can do so every clock cycle. Accumulator register 436 can store data generated and / or operated on by data processing circuitry within core 302.
[0134] exist Figure 4 In the example of FIG, cascade interfaces 422 and 424 can be programmed based on configuration data loaded into configuration register 324. For example, cascade interface 422 can be activated or deactivated based on configuration register 324. Similarly, cascade interface 424 can be activated or deactivated based on configuration register 324. Cascade interface 422 can be activated and / or deactivated independently of cascade interface 424.
[0135] In one or more other embodiments, cascade interfaces 422 and 424 are controlled by core 302. For example, core 302 may include instructions to read / write cascade interfaces 422 and / or 424. In another example, core 302 may include hardwired circuitry capable of reading and / or writing to cascade interfaces 422 and / or 424. In certain embodiments, cascade interfaces 422 and 424 may be controlled by an entity external to core 302.
[0136] In the embodiments described in this disclosure, DPE 204 does not include cache memory. By omitting cache memory, DPE array 202 is able to achieve predictable (e.g., deterministic) performance. Furthermore, since there is no need to maintain coherence between cache memories located in different DPEs, significant processing overhead is avoided.
[0137] According to one or more embodiments, the core 302 of the DPE 204 does not have input interrupts. Therefore, the core 302 of the DPE 204 can run uninterrupted. Omitting input interrupts to the core 302 of the DPE 204 also allows the DPE array a02 to achieve predictable (eg, deterministic) performance.
[0138] Figure 5 Another example architecture of a DPE array is shown. Figure 5 In the example of FIG, SoC interface block 206 provides an interface between DPE 204 and other subsystems of SoC 200. SoC interface block 206 integrates the DPE into the device. SoC interface block 206 can communicate configuration data to DPE 204, communicate events from DPE 204 to other subsystems, communicate events from other subsystems to DPE 204, generate and communicate interrupts to entities external to DPE array 202, communicate application data between other subsystems and DPE 204, and / or communicate trace and / or debug data between other subsystems and DPE 204.
[0139] exist Figure 5 In the example of , SoC interface block 206 includes a plurality of interconnected unit blocks (tiles). For example, SoC interface block 206 includes unit blocks 502, 504, 506, 508, 510, 512, 514, 516, 518 and 520. Figure 5 In the example of FIG, the cell blocks 502-520 are organized in a row. In other embodiments, the cell blocks may be arranged in columns, in a grid, or in another layout. For example, the SoC interface block 206 may be implemented as a column of cell blocks to the left of the DPE 204, to the right of the DPE 204, between the columns of DPE 204, etc. In another embodiment, the SoC interface block 206 may be located above the DPE array 202. The SoC interface block 206 may be implemented such that the cell blocks are located below the DPE array 202, to the left of the DPE array 202, to the right of the DPE array 202, and / or above the DPE array 202 in any combination. In this regard, Figure 5 It is provided for purposes of illustration and not limitation.
[0140] In one or more embodiments, the cell blocks 502-520 have the same architecture. In one or more other embodiments, the cell blocks 502-520 may be implemented using two or more different architectures. In certain embodiments, the cell blocks within the SoC interface block 206 may be implemented using different architectures, where each different cell block architecture supports communication with a different type of subsystem or combination of subsystems of the SoC 200.
[0141] exist Figure 5 In the example shown in FIG. 5 , cell blocks 502-520 are coupled so that data can propagate from one cell block to another. For example, data can propagate from cell block 502 through cell blocks 504, 506, and down the line of cell blocks to cell block 520. Similarly, data can propagate in the opposite direction from cell block 520 to cell block 502. In one or more embodiments, each of cell blocks 502-520 can operate as an interface to multiple DPEs. For example, each of cell blocks 502-520 can operate as an interface to a subset of DPEs 204 of DPE array 202. The subset of DPEs to which each cell block provides an interface can be mutually exclusive, such that no single DPE is interfaced by more than one cell block of SoC interface block 206.
[0142] In one example, each of the cell blocks 502-520 provides an interface to a column of DPEs 204. For purposes of illustration, cell block 502 provides an interface to the DPEs in column A. Cell block 504 provides an interface to the DPEs in column B, and so on. In each case, the cell block includes a direct connection to the adjacent DPE in the column of DPEs, which in this example is the bottom DPE. For example, referring to column A, cell block 502 is directly connected to DPE 204-1. Other DPEs in column A can communicate with cell block 502, but do so through the DPE interconnects of intervening DPEs in the same column.
[0143] For example, cell block 502 can receive data from another source such as PS 212, PL 214, and / or another hardwired circuit block 210 (e.g., a dedicated circuit block). Cell block 502 can provide those portions of the data addressed to the DPEs in rank A to those DPEs while sending data addressed to the DPEs in other ranks (e.g., DPEs to which cell block 502 is not interfaced) to cell block 504. Cell block 504 can perform the same or similar processing, wherein data received from cell block 502 addressed to the DPEs in rank B is provided to those DPEs, while data addressed to the DPEs in other ranks is sent to cell block 506, and so on.
[0144] In this manner, data can propagate from cell block to cell block of the SoC interface block 206 until it reaches a cell block that serves as an interface to the DPE to which the data is addressed (e.g., a "target DPE"). The cell block operating as an interface to the target DPE can direct the data to the target DPE using the DPE's memory-mapped switch and / or the DPE's stream switch.
[0145] As described above, the use of columns is an example implementation. In other embodiments, each cell block of SoC interface block 206 can provide an interface to a row of DPEs in DPE array 202. This configuration can be used when SoC interface block 206 is implemented as a column of cell blocks, whether to the left, right, or between columns of DPEs 204. In other embodiments, the subset of DPEs to which each cell block provides an interface can be any combination of fewer than all DPEs in DPE array 202. For example, DPEs 204 can be assigned to cell blocks of SoC interface block 206. The specific physical layout of such DPEs can vary depending on the connectivity of the DPEs established by the DPE interconnect. For example, cell block 502 can provide interfaces to DPEs 204-1, 204-2, 204-11, and 204-12. Another cell block of SoC interface block 206 can provide interfaces to four other DPEs, and so on.
[0146] Figure 6 FIG. 2 shows an example architecture of a unit block of the SoC interface block 206. Figure 6 In the example of FIG, two different types of cell blocks are shown for the SoC interface block 206. Cell block 602 is configured to serve as an interface between the DPE and only the PL 214. Cell block 610 is configured to serve as an interface between the DPE and the NoC 208 and between the DPE and the PL 214. The SoC interface block 206 may include a combination of cell blocks using both architectures as shown for cell blocks 602 and cell block 610, or, in another example, only a cell block having the architecture shown for cell block 610.
[0147] exist Figure 6 In the example shown, cell block 602 includes a flow switch 604 connected to a PL interface 606 and a DPE (e.g., DPE 204-1 located immediately above). PL interface 606 is connected to a boundary logic interface (BLI) circuit 620 and a BLI circuit 622 located in PL 214, respectively. Cell block 610 includes a flow switch 612 connected to a NoC and PL interface 614 and to a DPE (e.g., DPE 204-5 located immediately above). NoC and PL interface 614 is connected to BLI circuits 624 and 626 in PL 214, and also to a NoC master unit (NMU) 630 and a NoC slave unit (NSU) 632 of NoC 208.
[0148] exist Figure 6 In the example shown, each stream interface 604 can output six different 32-bit data streams to the DPE to which it is coupled, and receive four different 32-bit data streams from it. PL interface 606 and each of NoC and PL interface 614 can provide six different 64-bit data streams to PL 214 via BLI 620 and BLI 624, respectively. Generally, each of BLIs 620, 622, 624, and 626 provides an interface or connection point within PL 214 to which PL interface 606 and / or NoC and PL interface 614 connect. PL interface 606 and each of NoC and PL interface 614 can receive eight different 64-bit data streams from PL 214 via BLI 622 and BLI 624, respectively.
[0149] NoC and PL interface 614 is also connected to NoC 208. Figure 6 In the example of FIG208 , the NoC and PL interface 614 is connected to one or more NMUs 630 and one or more NSUs 632. In one example, the NoC and PL interface 614 can provide two different 128-bit data streams to the NoC 208, where each data stream is provided to a different NMU 630. The NoC and PL interface 614 can receive two different 128-bit data streams from the NoC 208, where each data stream is received from a different NSU 632.
[0150] The stream switches 604 in adjacent cell blocks are connected. In one example, the stream switches 604 in adjacent cell blocks can communicate via four different 32-bit data streams in each of the left and right directions (e.g., as long as the cell block is on the right or left, as appropriate).
[0151] Each of the cell blocks 602 and 610 can include one or more memory-mapped switches to communicate configuration data. For illustrative purposes, the memory-mapped switches are not shown. For example, the memory-mapped switches can be vertically connected to the memory-mapped switches of the DPE immediately above, connected to memory-mapped switches in other adjacent cell blocks in the SoC interface block 206 in the same or similar manner as the stream switch 604, connected to configuration registers (not shown) in the cell blocks 602 and 610, and / or connected to the PL interface 608 or the NoC and PL interface 614, as appropriate.
[0152] The various bit widths and numbers of data flows described in conjunction with the various switches included in the unit blocks 602 and / or 610 of the DPE 204 and / or SoC interface block 206 are provided for illustrative purposes and are not intended to limit the inventive arrangements described in this disclosure.
[0153] Figure 7 An example implementation of a NoC 208 is shown. The NoC 208 includes an NMU 702, an NSU 704, a network 714, a NoC peripheral interconnect (NPI) 710, and registers 712. Each NMU 702 is an ingress circuit that connects endpoint circuits to the NoC 208. Each NSU 704 is an egress circuit that connects the NoC 208 to endpoint circuits. The NMU 702 is connected to the NSU 704 via a network 714. In one example, the network 714 includes a NoC packet switch 706 (NPS) and routing 708 between the NPSs 706. Each NPS 706 performs NoC packet switching. The NPSs 706 are connected to each other and to the NMUs 702 and NSUs 704 via routing 708 to implement multiple physical channels. The NPSs 706 also support multiple virtual channels for each physical channel.
[0154] The NPI 710 includes circuitry for programming the NMU 702, NSU 704, and NPS 706. For example, the NMU 702, NSU 704, and NPS 706 may include registers 712 that determine their functionality. The NPI 710 includes a peripheral interconnect coupled to the registers 712 for programming them to set functionality. The registers 712 in the NoC 208 support interrupts, quality of service (QoS), error handling and reporting, transaction control, power management, and address mapping control. The registers 712 can be initialized to a usable state before being reprogrammed, for example, by writing to the registers 712 using a write request. Configuration data for the NoC 208 can be stored in non-volatile memory (NVM), for example, as part of a programming device image (PDI), and provided to the NPI 710 for programming the NoC 208 and / or other endpoint circuitry.
[0155] NMU 702 is the traffic entry point. NSU 704 is the traffic exit point. The endpoint circuits coupled to NMU 702 and NSU 704 can be hardened circuits (e.g., hardwired circuit block 210) or circuits implemented in PL 214. A given endpoint circuit can be coupled to more than one NMU 702 or more than one NSU 704.
[0156] Figure 82 is a block diagram illustrating connections between endpoint circuits in SoC 200 through NoC 208, according to an example. In this example, endpoint circuit 802 is connected to endpoint circuit 804 through NoC 208. Endpoint circuit 802 is a master circuit that is coupled to NMU 702 of NoC 208. Endpoint circuit 804 is a slave circuit that is coupled to NSU 704 of NoC 208. Each of endpoint circuits 802 and 804 can be a circuit in PS 212, a circuit in PL region 214, or a circuit in another subsystem (e.g., hardwired circuit block 210).
[0157] The network 714 includes multiple physical channels 806. The physical channels 806 are implemented by programming the NoC 208. Each physical channel 806 includes one or more NPSs 706 and associated routes 708. The NMU 702 is connected to the NSU 704 via at least one physical channel 806. A physical channel 806 may also have one or more virtual channels 808.
[0158] The connections over network 714 use a master-slave arrangement. In an example, the most basic connection on network 714 includes a single master connected to a single slave. However, in other examples, more complex structures can be implemented.
[0159] Figure 9 is a block diagram illustrating a NoC 208 according to another example. In this example, the NoC 208 includes a vertical portion 902 (VNoC) and a horizontal portion 904 (HNoC). Each VNoC 902 is disposed between PL regions 214. The HNoC 904 is disposed between the PL regions 214 and an I / O group 910 (e.g., an I / O block and / or transceiver corresponding to the hardwired circuit block 210). The NoC 208 is connected to a memory interface 908 (e.g., the hardwired circuit block 210). The PS 212 is coupled to the HNoC 904.
[0160] In this example, PS 212 includes multiple NMUs 702 coupled to an HNoC 904. VNoC 902 includes NMUs 702 and NSUs 704 arranged in PL area 214. Memory interface 908 includes NSU 704 coupled to HNoC 904. Both HNoC 904 and VNoC 902 include NPSs 706 connected by routes 708. In VNoC 902, routes 708 extend vertically. In HNoC 904, routes extend horizontally. In each VNoC 902, each NMU 702 is coupled to an NPS 706. Similarly, each NSU 704 is coupled to an NPS 706. NPSs 706 are coupled to each other to form a switch matrix. Some NPSs 706 in each VNoC 902 are coupled to other NPSs 706 in the HNoC 904.
[0161] Although only a single HNoC 904 is shown, in other examples, the NoC 208 may include more than one HNoC 904. Furthermore, although two VNoCs 902 are shown, the NoC 208 may include more than two VNoCs 902. Although a memory interface 908 is shown by way of example, it should be understood that hardwired circuit blocks 210, other hardwired circuit blocks 210 may be used in place of, or in addition to, the memory interface 908.
[0162] Figure 10 An example method 1000 is shown for programming the NoC 208. Although described independently of other subsystems of the SoC 200, the method 1000 may be included and / or used as part of a larger boot-up or programming process for the SoC 200.
[0163] At block 1002, a platform management controller (PMC) implemented in SoC 200 receives NoC programming data at boot time. The NoC programming data may be part of the PDI. The PMC is responsible for managing SoC 200. The PMC is capable of maintaining a secure and reliable environment, booting SoC 200, and managing SoC 200 during normal operation.
[0164] At block 1004, the PMC loads the NoC programming data into registers 712 via the NPI 710 to create the physical channel 806. In an example, the programming data may also include information for configuring the routing table in the NPS 706. At block 1006, the PMC boots the SoC 200. In this manner, the NoC 208 includes at least configuration information for the physical channel 806 between the NMU 702 and the NSU 704. As further described below, the remaining configuration information for the NoC 208 may be received during runtime. In another example, all or part of the configuration information described below that is received during runtime may be received at boot time.
[0165] Figure 11 An example method 1100 for programming the NoC 208 is shown. At block 1102, the PMC receives NoC programming data during execution time. At block 1104, the PMC loads the programming data into the NoC registers 712 via the NPI 710. In the example, at block 1106, the PMC configures the routing table in the NPS 706. At block 1108, the PMC configures the QoS path 806 on the physical channel. At block 1110, the PMC configures the address space mapping. At block 1112, the PMC configures the ingress / egress interface protocol, width, and frequency. The QoS path, address space mapping, routing table, and ingress / egress configuration are discussed further below.
[0166] Figure 12 An example data path 1200 through the NoC 208 between endpoint circuits is shown. The data path 1200 includes an endpoint circuit 1202, an AXI master circuit 1204, an NMU 1206, an NPS 1208, an NSU 1210, an AXI slave circuit 1212, and an endpoint circuit 1214. The endpoint circuit 1202 is coupled to the AXI master circuit 1204. The AXI master circuit 1204 is coupled to the NMU 1206. In another example, the AXI master circuit 1204 is part of the NMU 1206.
[0167] The NMU 1206 is coupled to an NPS 1208. The NPSs 1208 are coupled to each other to form a chain of NPSs 1208 (e.g., a chain of five NPSs 1208 in this example). Generally, there is at least one NPS 1208 between the NMU 1206 and the NSU 1210. The NSU 1210 is coupled to one of the NPSs 1208. An AXI slave circuit 1212 is coupled to the NSU 1210. In another example, the AXI slave circuit 1212 is part of the NSU 1210. An endpoint circuit 1214 is coupled to the AXI slave circuit 1212.
[0168] Endpoint circuits 1202 and 1214 can each be a hardened circuit (e.g., a PS circuit, hardwired circuit 210, or one or more DPEs 204) or a circuit configured in PL 214. Endpoint circuit 1202 acts as a master and sends read / write requests to NMU 1206. In this example, endpoint circuits 1202 and 1214 communicate with NoC 208 using the AXI protocol. Although AXI is described in this example, it should be understood that NoC 208 can be configured to receive communications from endpoint circuits using other types of protocols known in the art. For clarity of example, NoC 208 is described herein as supporting the AXI protocol. NMU 1206 forwards the request through a set of NPSs 1208 to the destination NSU 1210. NSU 1210 passes the request to an attached AXI slave circuit 1212 for processing and distribution of the data to endpoint circuit 1214. AXI slave circuit 1212 can send read / write responses back to NSU 1210. The NSU 1210 may forward the response to the NMU 1206 through a set of NPSs 1208. The NMU 1206 transmits the response to the AXI master 1204, which distributes the data to the endpoint circuits 1202.
[0169] Figure 13 An example method 1300 for processing read / write requests and responses is shown. Method 1300 begins at block 1302, where endpoint circuit 1202 sends a request (e.g., a read request or a write request) to NMU 1206 via AXI master circuit 1204. At block 1304, NMU 1206 processes the response. In the example, NMU 1206 performs asynchronous crossing and rate matching between the clock domain of endpoint circuit 1202 and NoC 208. NMU 1206 determines the destination address of NSU 1210 based on the request. If virtualization is employed, NMU 1206 may perform address remapping. NMU 1206 also performs AXI translation of the request. NMU 1206 further packages the request into a data packet stream.
[0170] At block 1306, the NMU 1206 sends the requested packet to the NPS 1208. Each NPS 1208 performs a table lookup for the target output port based on the destination address and routing information. At block 1308, the NSU 1210 processes the requested packet. In one example, the NSU 1210 unpacks the request, performs AXI translation, and performs asynchronous crossing and rate matching from the NoC clock domain to the clock domain of the endpoint circuit 1214. At block 1310, the NSU 1210 sends the request to the endpoint circuit 1214 via the AXI slave circuit 1212. The NSU 1210 may also receive a response from the endpoint circuit 1214 via the AXI slave circuit 1212.
[0171] At block 1312, the NSU 1210 processes the response. In one example, the NSU 1210 performs asynchronous crossing and rate matching from the clock domain of the endpoint circuit 1214 and the clock domain of the NoC 208. The NSU 1210 also packages the response into a packet stream. At block 1314, the NSU 1210 sends the packet through the NPS 1208. Each NPS 1208 performs a table lookup for the target output port based on the destination address and routing information. At block 1316, the NMU 1206 processes the packet. In one example, the NMU 1206 unpacks the response, performs AXI translation, and performs asynchronous crossing and rate matching from the NoC clock domain to the clock domain of the endpoint circuit 1202. At block 1318, the NMU 1206 sends the response to the endpoint circuit 1202 via the AXI master circuit 1204.
[0172] Figure 14An example implementation of an NMU 702 is shown. The NMU 702 includes an AXI master interface 1402, a packetization circuit 1404, an address mapping table 1406, an unpacking circuit 1408, a QoS circuit 1410, a VC mapping circuit 1412, and a clock management circuit 1414. The AXI master interface 1402 provides an AXI interface to the NMU 702 for endpoint circuits. In other examples, different protocols may be used, and thus the NMU 702 may have different master interfaces that conform to the selected protocol. The NMU 702 routes inbound traffic to the packetization circuit 1404, which generates packets from the inbound data. The packetization circuit 1404 determines the destination ID from the address mapping table 1406, which is used to route the packet. The QoS circuit 1410 may provide ingress rate control to control the rate at which packets are injected into the NoC 208. The VC mapping circuit 1412 manages QoS virtual channels on each physical channel. The NMU 702 may be configured to select the virtual channel to which a packet is mapped. Clock management circuitry 1414 performs rate matching and asynchronous data crossing to provide an interface between the AXI clock domain and the NoC clock domain. Unpacking circuitry 1408 receives return data packets from NoC 208 and is configured to unpack the data packets for output by AXI master interface 1402 .
[0173] Figure 15 An example implementation of an NSU 704 is shown. NSU 704 includes an AXI slave interface 1502, clock management circuitry 1504, packetization circuitry 1508, depacketization circuitry 1506, and QoS circuitry 1510. AXI slave interface 1502 provides an AXI interface to NSU 704 for endpoint circuits. In other examples, different protocols may be used, and thus NSU 704 may have different slave interfaces conforming to the selected protocol. NSU 704 routes inbound traffic from NoC 208 to depacketization circuitry 1506, which generates depacketized data. Clock management circuitry 1504 performs rate matching and asynchronous data crossing to provide an interface between the AXI clock domain and the NoC clock domain. Packetization circuitry 1508 receives return data from slave interface 1502 and is configured to packetize the return data for transmission through NoC 208. QoS circuitry 1510 may provide ingress rate control to control the rate at which data packets are injected into NoC 208.
[0174] Figure 16 Shows that the combination Figure 1 The described system implements an example software architecture. For example, Figure 16 The architecture can be implemented as Figure 1 One or more program modules 120. Figure 16 The software architecture includes a DPE compiler 1602 , a NoC compiler 1604 , and a hardware compiler 1606 . Figure 16Examples of various types of design data that may be exchanged between compilers during operation (eg, executing a design flow for implementation in SoC 200 ) are shown.
[0175] The DPE compiler 1602 can generate one or more binary files from an application that can be loaded into one or more DPEs of the DPE array 202 and / or a subset of the DPEs 204. Each binary file can include object code that can be executed by the DPE's core, optional application data, and configuration data for the DPE. The NoC compiler 1604 can generate a binary file that includes configuration data that is loaded into the NoC 208 to create a data path for the application. The hardware compiler 1606 can compile the hardware portion of the application to generate a configuration bitstream for implementation in the PL 214.
[0176] Figure 16 An example of how DPE compiler 1602, NoC compiler 1604, and hardware compiler 1606 communicate with each other during operation is shown. The compilers communicate in a coordinated manner by exchanging design data to converge on a solution. This solution is an implementation applied within SoC 200 that meets the design specifications and constraints and includes a common interface through which the various heterogeneous subsystems of SoC 200 communicate.
[0177] As defined in this disclosure, the term "design metric" defines a goal or requirement for an application to be implemented in the SoC 200. Examples of design metrics include, but are not limited to, power consumption requirements, data throughput requirements, timing requirements, etc. Design metrics may be provided by user input, a file, or otherwise to define higher or system-level requirements for an application. As defined in this disclosure, a "design constraint" is a requirement that an EDA tool may or may not follow to achieve a design metric or requirement. Design constraints may be specified as compiler directives and typically specify lower-level requirements or recommendations for an EDA tool (e.g., a compiler) to follow. Design constraints may be specified by user input, a file containing one or more design constraints, command line input, or the like.
[0178] In one aspect, the DPE compiler 1602 can generate a logical architecture and SoC interface block solution for an application. For example, the DPE compiler 1602 can generate a logical architecture based on high-level, user-defined metrics for the software portion of the application to be implemented in the DPE array 202. Examples of metrics can include, but are not limited to, data throughput, latency, resource utilization, and power consumption. Based on the metrics and the application (e.g., a specific node to be implemented in the DPE array 202), the DPE compiler 1602 can generate a logical architecture.
[0179] A logical architecture is a file or data structure that can specify the hardware resource block information required for various parts of an application. For example, the logical architecture can specify the number of DPEs 204 required to implement the software portion of the application, any intellectual property (IP) cores required to communicate with the DPE array 202 in the PL 214, the connections that need to be routed through the NoC 208, and the port information for the IP cores in the DPE array 202, NoC 208, and PL 214. An IP core is a reusable block or portion of logic, cell, or IC layout design that can be used in a circuit design as a reusable circuit block capable of performing a specific function or operation. An IP core can be specified in a format that can be incorporated into a circuit design for implementation within the PL 214. Although this disclosure refers to various types of cores, the term "core" without any other modifiers is generally intended to refer to such different types of cores.
[0180] Example 1 of this disclosure at the end of the detailed description shows an example schema that can be used to specify the logical architecture of an application. Example 1 illustrates the various types of information included in the logical architecture of an application. In one aspect, hardware compiler 1606 can implement the hardware portion of an application based on or using the logical architecture and SoC interface block scheme, as opposed to using the application itself.
[0181] The port information for DPE array 202 and the IP cores in NoC 208 and PL 214 may include the logical configuration of the ports, such as whether each port is a streaming data port, a memory-mapped port, or a parameter port, and whether the port is a master or a slave. Other examples of IP core port information include the data width and operating frequency of the port. Connectivity between DPE array 202, NoC 208, and IP cores in PL 214 may be specified as logical connections between ports of various hardware resource blocks specified in the logical architecture.
[0182] The SoC interface block schema is a data structure or file that specifies the mapping of physical data paths (e.g., physical resources) that connect to the SoC interface block 206 and enter and exit the DPE array 202. For example, the SoC interface block schema maps specific logical connections used to transfer data into and out of the DPE array 202 to specific flow channels of the SoC interface block 206, e.g., to specific cell blocks, flow switches, and / or flow switch interfaces (e.g., ports) of the SoC interface block 206. Example 2, located after Example 1 and near the end of the detailed description, illustrates an example mode of applying the SoC interface block schema.
[0183] In one aspect, the DPE compiler 1602 can analyze or simulate data traffic on the NoC 208 based on the application and logical architecture. The DPE compiler 1602 can provide the data transmission requirements of the software portion of the application, such as "NoC traffic," to the NoC compiler 1604. The NoC compiler 1604 can generate routes for data paths through the NoC 208 based on the NoC traffic received from the DPE compiler 1602. The results from the NoC compiler 1604 (shown as "NoC plan") can be provided to the DPE compiler 1602.
[0184] In one aspect, the NoC scheme may be an initial NoC scheme that specifies only the entry and / or exit points of the NoC 208 to which nodes of the application connected to the NoC 208 will be connected. For example, for compiler convergence purposes, more detailed routing and / or configuration data for data paths within the NoC 208 (e.g., between entry and exit points) may be excluded from the NoC scheme. Example 3, following Example 2 at the end of the detailed description, illustrates an example schema for the NoC scheme for this application.
[0185] Hardware compiler 1606 can operate on the logic architecture to implement the hardware portion of the application in PL 214. In the event that hardware compiler 1606 is unable to generate an implementation of the hardware portion of the application that meets established design constraints (e.g., timing, power, data throughput, etc.) (e.g., using the logic architecture), hardware compiler 1606 can generate one or more SoC interface block constraints and / or receive one or more user-specified SoC interface block constraints. Hardware compiler 1606 can provide the SoC interface block constraints as a request to DPE compiler 1602. The SoC interface block constraints effectively remap one or more portions of the logic architecture to different flow paths of the SoC interface block 206. The SoC interface block constraints provided by hardware compiler 1606 further facilitate hardware compiler 1606 in generating the hardware portion of the application in PL 214 that meets the design specifications. Example 4, located after Example 3 toward the end of the detailed description, illustrates example constraints for the SoC interface block and / or NoC of an application.
[0186] In another aspect, the hardware compiler 1606 can also generate NoC traffic based on the application and the logical architecture and provide the NoC traffic to the NoC compiler 1604. For example, the hardware compiler 1606 can analyze or simulate the hardware portion of the application to determine the data traffic generated by the hardware portion of the design, which will be transmitted through the NoC 208 to the PS 212, the DPE array 202, and / or other portions of the SoC 200. The NoC compiler 1604 can generate and / or update the NoC scheme based on information received from the hardware compiler 1606. The NoC compiler 1604 can provide the NoC scheme or an updated version thereof to the hardware compiler 1606 and the DPE compiler 1602. In this regard, the DPE compiler 1602 can update the SoC interface block scheme and provide the updated scheme to the hardware compiler 1606 in response to the NoC scheme or the updated NoC scheme received from the NoC compiler 1604, and / or in response to one or more SoC interface block constraints received from the compiler 1606. The DPE compiler 1602 generates an updated SoC interface block plan based on the SoC interface block constraints received from the hardware compiler 1606 and / or the updated NoC plan received from the NoC compiler 1604 .
[0187] It should be understood that Figure 16 The flow of data between compilers in the examples in the present disclosure is for illustrative purposes only. In this regard, information exchange between compilers can be performed at various stages of the example design flow described in this disclosure. In other aspects, the exchange of design data between compilers can be performed in an iterative manner, such that each compiler can continuously improve the implementation of the application portion processed by the compiler based on information received from other compilers to converge on a solution.
[0188] In a specific example, after receiving the logic architecture and SoC interface block plan from DPE compiler 1602 and the NoC plan from NoC compiler 1604, hardware compiler 1606 may determine that it is not possible to generate an implementation of the hardware portion of the application that meets the established design specifications. The initial SoC interface block plan generated by DPE compiler 1602 is generated based on DPE compiler 1602's understanding of a portion of the application to be implemented in DPE array 202. Similarly, the initial NoC plan generated by NoC compiler 1604 is generated based on initial NoC traffic provided by DPE compiler 1602 to NoC compiler 1604. Example 5, located after Example 4 at the end of the detailed description, illustrates an example pattern of NoC traffic for an application. It should be understood that while patterns are used in Examples 1-5, other formats and / or data structures may be used to specify the illustrated information.
[0189] The hardware compiler 1606 attempts to execute the implementation process on the hardware portion of the application, including synthesis (if necessary), placement, and routing of the hardware portion. Consequently, the initial SoC interface block solution and the initial NoC solution may result in placement and / or routing that does not meet established timing constraints within the PL 214. In other cases, the SoC interface block solution and the NoC solution may not have a sufficient number of physical resources (e.g., wires) to accommodate the data that must be transmitted, resulting in congestion within the PL 214. In such cases, the hardware compiler 1606 can generate one or more different SoC interface block constraints and / or receive one or more user-specified SoC interface block constraints and provide the SoC interface block constraints to the DPE compiler 1602 as a request to regenerate the SoC interface block solution. Similarly, the hardware compiler 1606 can generate one or more different NoC constraints and / or receive one or more user-specified NoC constraints and provide the NoC constraints to the NoC compiler 1604 as a request to regenerate the NoC solution. In this manner, the hardware compiler 1606 invokes the DPE compiler 1602 and / or the NoC compiler 1604.
[0190] The DPE compiler 1602 can take the SoC interface block constraints received from the hardware compiler 1606 and update the SoC interface block solution using the received SoC interface block constraints, if possible, and provide the updated SoC interface block solution back to the hardware compiler 1606. Similarly, the NoC compiler 1604 can take the NoC constraints received from the hardware compiler 1606 and update the NoC solution using the received NoC constraints, if possible, and provide the updated NoC solution back to the hardware compiler 1606. The hardware compiler 1606 can then continue the implementation flow to generate the hardware portion of the application for implementation within the PL 214 using the updated SoC interface block solution received from the DPE compiler 1602 and the updated NoC solution received from the NoC compiler 1604.
[0191] In one aspect, the hardware compiler 1606 invoking the DPE compiler 1602 and / or the NoC compiler 1604 by providing one or more SoC interface block constraints and one or more NoC constraints, respectively, is part of a verification process. For example, the hardware compiler 1606 is seeking verification from the DPE compiler 1602 and / or the NoC compiler 1604 that the SoC interface block constraints and NoC constraints provided by the hardware compiler 1606 can be used or integrated into a routable SoC interface block solution and / or NoC solution.
[0192] Figure 17A Shows the use of Figure 1The described system maps to an example of an application 1700 on SoC 200. For illustrative purposes, only a subset of the different subsystems of SoC 200 are shown. Application 1700 includes nodes A, B, C, D, E, and F with the connectivity shown. Example 6 below illustrates example source code that may be used to specify application 1700.
[0193] Example 6
[0194]
[0195]
[0196] In one aspect, application 1700 is specified as a data flow graph comprising a plurality of nodes. Each node represents a computation, corresponding to a function rather than a single instruction. Nodes are interconnected by edges representing data flows. The hardware implementation of a node may execute only in response to receiving data from each input to that node. Nodes are typically executed in a non-blocking manner. The data flow graph specified by application 1700 represents a parallel specification to be implemented in SoC 200, rather than a sequential program. The system is capable of operating on application 1700 (e.g., in the form of a graph as shown in Example 1) to map various nodes to appropriate subsystems of SoC 200 for implementation therein.
[0197] In one example, application 1700 is specified in a high-level programming language (HLL), such as C and / or C++. As noted, while specified in an HLL typically used to create sequential programs, application 1700, as a dataflow graph, is a parallel specification. The system can provide a class library for building dataflow graphs and applications 1700 like these. Dataflow graphs are user-defined and compiled into the architecture of SoC 200. The class library can be implemented as a helper library with predefined classes and constructors that can be used to build graphs, nodes, and edges for application 1700. Application 1700 effectively executes on SoC 200 and includes delegate objects that execute in PS 212 of SoC 200. Objects of application 1700 executed in PS 212 can be used to direct and monitor actual computations running on SoC 200, for example, in PL 214, DPE array 202, and / or hardwired circuit blocks 210.
[0198] According to the inventive arrangements described in this disclosure, accelerators (e.g., PL nodes) can be represented as objects in a dataflow graph (e.g., an application). The system can automatically synthesize the PL nodes and connect the synthesized PL nodes for implementation in PL 214. In contrast, in traditional EDA systems, users specify hardware-accelerated applications that utilize sequential semantics. Hardware-accelerated functions are specified through function calls. The interface of the hardware-accelerated function (e.g., the PL node in this example) is defined by the function call and the various parameters provided in the function call, rather than by connections on the dataflow graph.
[0199] As shown in the source code of Example 6, nodes A and F are specified for implementation in PL 214, while nodes B, C, D, and E are specified for implementation in DPE array 202. The connectivity of the nodes is specified by data transfer edges in the source code. The source code of Example 6 also specifies a top-level testbench and control program that executes in PS 212.
[0200] refer to Figure 17A , application 1700 is mapped to SoC 200. As shown, nodes A and F are mapped to PL 214. Shaded DPEs 204-13 and 204-14 represent the DPEs 204 to which nodes B, C, D, and E are mapped. For example, nodes B and C are mapped to DPE 204-13, while nodes D and E are mapped to DPE 204-4. Nodes A and F are implemented in PL 214 and connected to DPEs 204-13 and 204-44 via PL 214, specific unit blocks and switches in SoC interface block 206, switches in the DPE interconnect of intervening DPEs 204, and routing using specific memory of selected adjacent DPEs 204.
[0201] The binary file generated for DPE 204-13 includes the object code required for DPE 204-13 to implement the computations corresponding to nodes B and C, and configuration data for establishing data paths between DPE 204-13 and DPE 204-14, and between DPE 204-13 and DPE 204-3. The binary file generated for DPE 204-4 includes the object code required for DPE 204-4 to implement the computations corresponding to nodes D and E, and configuration data for establishing data paths with DPE 204-14 and DPE 204-5.
[0202] Other binary files are generated for other DPEs 204 (e.g., DPEs 204-3, 204-5, 204-6, 204-7, 204-8, and 204-9) to connect DPE 204-13 and DPE 204-4 to SoC interface block 206. Obviously, if such other DPEs 204 implement other computations (nodes with applications assigned thereto), such binary files will include any object code.
[0203] In this example, due to the long route connecting DPE 204-14 and node F, the hardware compiler 1606 is unable to generate an implementation of the hardware portion that meets the timing constraints. In the present disclosure, a specific state of the implementation of the hardware portion of an application can be referred to as a state of the hardware design, wherein the hardware design is generated and / or updated throughout the implementation process. For example, the SoC interface block solution can assign the signal crossover of node F to the cell block of the SoC interface block below DPE 204-9. In that case, the hardware compiler 1606 can provide the requested SoC interface block constraints to the DPE compiler 1602 to request that the crossover of node F through the SoC interface block 206 be moved closer to DPE 204-4. For example, the requested SoC interface block constraints from the hardware compiler 1606 can request that the logical connection of DPE 204-4 be mapped to the cell block immediately below DPE 204-4 within the SoC interface block 206. This remapping will allow the hardware compiler to place node F closer to DPE 204-4 to improve timing.
[0204] Figure 17B Another example mapping of application 1700 to SoC 200 is shown. Figure 17B Shows an alternative and Figure 17A A more detailed example is shown in . For example, Figure 17B 17. The diagram illustrates the mapping of nodes of application 1700 to specific DPEs 204 of DPE array 202, the connections established between the DPEs 204 to which the nodes of application 1700 are mapped, the allocation of memory in the memory modules of the DPEs 204 to the nodes of application 1700, the transfer of data to the memory and core interfaces (e.g., 428, 430, 432, 434, 402, 404, 406, and 408) of the DPEs 204 (represented by double-headed arrows), and / or the mapping to the stream switches of the DPE interconnect 306, as performed by DPE compiler 1602.
[0205] exist Figure 17B, memory modules 1702, 1706, 1710, 1714, and 1718 are shown along with cores 1704, 1708, 1712, 1716, and 1720. Cores 1704, 1708, 1712, 1716, and 1720 include program memories 1722, 1724, 1726, 1728, and 1730, respectively. In the top row, core 1704 and memory module 1706 form DPE 204, while core 1708 and memory module 1710 form another DPE 204. In the bottom row, memory module 1714 and core 1716 form DPE 204, while memory 1718 and core 1720 are used for another DPE 204.
[0206] As shown, nodes A and F are mapped to PL 214. Node A is connected to a memory bank (e.g., the shaded portion of the memory bank) in memory module 1702 via a stream switch and an arbiter in memory module 1702. Nodes B and C are mapped to core 1704. Instructions for implementing nodes B and C are stored in program memory 1722. Nodes D and E are mapped to core 1716, with instructions for implementing nodes D and E stored in program memory 1728. Node B is assigned and has access to the shaded portion of the memory bank in memory module 1702 via a core-memory interface, while node C is assigned and has access to the shaded portion of the memory bank in memory module 1706 via a core-memory interface. Nodes B, C, and E are assigned and have access to the shaded portion of the memory bank in memory module 1714 via a core-memory interface. Node D has access to the shaded portion of the memory bank in memory module 1718 via a core-memory interface. Node F is connected to memory module 1718 via an arbiter and a stream switch.
[0207] Figure 17B Connectivity between nodes for applications that may use memory and / or core interfaces that share memory between cores and that are implemented using the DPE interconnect 306 is shown.
[0208] Figure 18 2 shows an example implementation of another application that has been mapped onto SoC 200. For illustrative purposes, only a subset of the different subsystems of SoC 200 are shown. In this example, connections to nodes A and F (each of which is implemented in PL 214) are routed through NoC 208. NoC 208 includes ingress / egress points 1802, 1804, 1806, 1808, 1810, 1812, 1814, and 1816 (e.g., NMUs / NSUs). Figure 18The example of FIG20 shows a situation where node A is placed relatively close to entry / exit point 1802, while node F, which accesses volatile memory 134, has a long path through PL 214 to entry / exit point 1816. If hardware compiler 1606 is unable to place node F closer to entry / exit point 1816, hardware compiler 1606 can request an updated NoC plan from NoC compiler 1604. In that case, hardware compiler 1606 can invoke NoC compiler 1604 with NoC constraints to generate an updated NoC plan that specifies different entry / exit points for node F, such as entry / exit point 1812. The different entry / exit points for node F will allow hardware compiler 1606 to place node F closer to the newly specified entry / exit points specified in the updated NoC plan and take advantage of the faster data paths available in NoC 208.
[0209] Figure 19 Shown by combining Figure 1 Another example software architecture 1900 that can be executed by the described system. For example, the architecture 1900 can be implemented as Figure 1 One or more program modules 120. Figure 19 In the example of , application 1902 is intended for implementation within SoC 200 .
[0210] exist Figure 19 In the example of , the user can interact with a user interface 1906 provided by the system. When interacting with the user interface 1906, the user can specify or provide an application 1902, performance and partition constraints 1904 of the application 1902, and a base platform 1908.
[0211] Application 1902 may include multiple different parts, each corresponding to a different subsystem available in SoC 200. For example, application 1902 may be specified as described in conjunction with Example 6. Application 1902 includes a software portion to be implemented in DPE array 202 and a hardware portion to be implemented in PL 214. Application 1902 may optionally include other software portions to be implemented in PS 212 and portions to be implemented in NoC 208.
[0212] Partitioning constraints (of performance and partitioning constraints 1904) optionally specify the location or subsystem in which various nodes of application 1902 are to be implemented. For example, a partitioning constraint may indicate, on a per-node basis for application 1902, whether the node is to be implemented in DPE array 202 or in PL 214. In other examples, location constraints can provide more specific or detailed information to DPE compiler 1602 to perform mapping of cores to DPEs, networks or data streams to stream switches, and buffers to memory modules and / or groups of memory modules of a DPE.
[0213] As an illustrative example, the implementation of an application may require a specific mapping. For example, in an application where multiple copies of a kernel are to be implemented in a DPE array, and each copy of the kernel operates on a different data set simultaneously, it may be desirable to locate the data set at the same relative address (location in memory) for each copy of the kernel executed in a different DPE in the DPE array. This can be achieved by using location constraints. If the DPE compiler 1602 does not support this condition, each copy of the kernel must be programmed separately or independently, rather than replicating the same programming across multiple different DPEs in the DPE array.
[0214] Another illustrative example is placing location constraints on applications that utilize cascade interfaces between DPEs. Because cascade interfaces flow in one direction within each row, it is preferable that a chain of DPEs coupled using cascade interfaces not begin at a DPE with a missing cascade interface (e.g., a corner DPE) or at a location that cannot be easily replicated elsewhere in the DPE array (e.g., the last DPE in a row). Location constraints can force the start of an application's DPE chain to begin at a specific DPE.
[0215] The performance constraints (of performance and partitioning constraints 1904 ) may specify various metrics (eg, power requirements, latency requirements, timing, and / or data throughput) to be achieved by the implementation of a node, whether in the DPE array 202 or the PL 214 .
[0216] Base platform 1908 is a description of infrastructure circuitry to be implemented in SoC 200 that interacts and / or interfaces with circuitry on a circuit board to which SoC 200 is coupled. Base platform 1908 may be synthesizable. For example, base platform 1908 may specify circuitry to be implemented within SoC 200 that receives signals from outside SoC 200 (e.g., external to SoC 200) and provides signals to systems and / or circuitry external to the SoC. For example, base platform 1908 may specify circuit resources, such as those used to interface with the SoC. Figure 1 100, one or more memory controllers for accessing volatile memory 134 and / or non-volatile memory 136, and / or other resources such as internal interfaces coupling DPE array 202 and / or PL 214 to the PCIe nodes. The circuitry specified by base platform 1908 can be used for any application that can be implemented in SoC 200 given a particular type of circuit board. In this regard, base platform 1908 is specific to the particular circuit board to which SoC 200 is coupled.
[0217] In one example, a partitioner 1910 can separate different portions of an application 1902 based on the subsystems of the SoC 200 in which each portion of the application 1902 is to be implemented. In an example implementation, the partitioner 1910 is implemented as a user-directed tool, where a user provides input indicating which of the different portions (e.g., nodes) of the application 1902 corresponds to each of the different subsystems of the SoC 200. For example, the input provided may be performance and partitioning constraints 1904. For illustrative purposes, the partitioner 1910 divides the application 1902 into a PS portion 1912 to be executed on the PS 212, a DPE array portion 1914 to be executed on the DPE array 202, a PL portion 1916 to be implemented in the PL 214, and a NoC portion 1936 to be implemented in the NoC 208. In one aspect, partitioner 1910 can generate each of PS portion 1912, DPE array portion 1914, PL portion 1916, and NoC portion 1936 as a separate file or separate data structure.
[0218] As shown, each different portion corresponding to a different subsystem is processed by a different compiler specific to that subsystem. For example, PS compiler 1918 can compile PS portion 1912 to generate one or more binary files comprising object code executable by PS 212. DPE compiler 1602 can compile DPE array portion 1914 to generate one or more binary files comprising object code executable by different DPEs, application data, and / or configuration data. Hardware compiler 1606 can perform an implementation flow on PL portion 1916 to generate a configuration bitstream that can be loaded into SoC 200 to implement PL portion 1916 in PL 214. As defined herein, the term "implementation flow" refers to the process in which placement and routing, and optionally synthesis, are performed. NoC compiler 1604 can generate a binary file specifying configuration data for NoC 208, which, when loaded into NoC 208, creates data paths connecting the various masters and slaves of application 1902. These various outputs generated by compilers 1918 , 1602 , 1604 , and / or 1606 are shown as binary files and a configuration bitstream 1924 .
[0219] In certain embodiments, some of compilers 1918, 1602, 1604, and / or 1606 are capable of communicating with each other during operation. By communicating at various stages during the design flow operating on application 1902, compilers 1918, 1602, 1604, and / or 1606 are able to converge on a solution. Figure 19In the example shown in FIG. 1 , the DPE compiler 1602 and the hardware compiler 1606 can communicate during operation while respectively compiling portions 1914 and 1916 of the application 1902. The hardware compiler 1606 and the NoC compiler 1604 can communicate during operation while respectively compiling portions 1916 and 1936 of the application 1902. The DPE compiler 1602 can also call the NoC compiler 1604 to obtain a NoC routing plan and / or an updated NoC routing plan.
[0220] The generated binary files and configuration bitstream 1924 can be provided to any of a variety of different targets. For example, the generated binary files and configuration bitstream 1924 can be provided to a simulation platform 1926, a hardware emulation platform 1928, an RTL simulation platform 1930, and / or a target IC 1932. In the case of the RTL simulation platform 1930, the hardware compiler 1922 can be configured to output RTL for the PL portion 1916 that can be simulated in the RTL simulation platform 1930.
[0221] Results obtained from simulation platform 1926, emulation platform 1928, RTL simulation platform 1930, and / or from implementation of application 1902 in target IC 1932 may be provided to performance analyzer and debugger 1934. Results from performance analyzer and debugger 1934 may be provided to user interface 1906, where a user may view the results of executing and / or simulating application 1902.
[0222] Figure 20 An example method 2000 of performing a design flow to implement an application in SoC 200 is shown. The method 2000 may be combined with Figure 1 The system described can be executed in conjunction with Figure 16 and Figure 19 The software architecture described.
[0223] In block 2002 , the system receives an application. The application may specify a software portion for implementation within the DPE array 202 of the SoC 200 and a hardware portion for implementation within the PL 214 of the SoC 200 .
[0224] In block 2004, the system can generate a logical architecture for the application. For example, the DPE compiler 1602 executed by the system can generate the logical architecture based on the software portion of the application to be implemented in the DPE array 202 and any high-level, user-specified metrics. The DPE compiler 1602 can also generate a SoC interface block schema that specifies the mapping of physical data paths in and out of the DPE array 202 connected to the SoC interface blocks 206.
[0225] In another aspect, when generating the logic architecture and SoC interface block solution, the DPE compiler 1602 can generate an initial mapping of nodes (referred to as "DPE nodes") of the application to be implemented in the DPE array 202 to specific DPEs 204. The DPE compiler 1602 can optionally generate an initial mapping and routing of the application's global memory data structures to global memory (e.g., volatile memory 134) by providing NoC traffic for the global memory to the NoC compiler 1604. As discussed, the NoC compiler 1604 can generate a NoC solution from the received NoC traffic. Using the initial mapping and routing, the DPE compiler 1602 can simulate the DPE portion to verify the initial implementation of the DPE portion. The DPE compiler 1602 can output the data generated by the simulation to the hardware compiler 1606 corresponding to each flow channel used in the SoC interface block solution.
[0226] In one aspect, the generated logic architecture as performed by the DPE compiler 1602 implements the previously combined Figure 19 The various example modes illustrate the partitioning of Figure 19 16. The example modes further illustrate how decisions and / or constraints can be logically shared across different subsystems of SoC 200 by different compilers in the SoC (DPE compiler 1602, hardware compiler 1606, and NoC compiler 1604) when compiling the portion of the application assigned to each compiler.
[0227] At block 2006, the system can construct a block diagram for the hardware portion. For example, a hardware compiler 1606 executed by the system can generate a block diagram. The block diagram combines the hardware portion of the application, as specified by the logical architecture, with the base platform of the SoC 200. For example, the hardware compiler 1606 can connect the hardware portion and the base platform when generating the block diagram. In addition, the hardware compiler 1606 can generate a block diagram based on the SoC interface block solution to connect the IP core corresponding to the hardware portion of the application to the SoC interface block.
[0228] For example, each node in the hardware portion of the application, as specified by the logic architecture, can be mapped to a specific RTL core (e.g., a specified portion of user-provided or customized RTL) or an available IP core. By mapping user-specified nodes to cores, hardware compiler 1606 can construct a block diagram to specify various circuit blocks of the base platform, any IP cores of PL 214 that are required to interact with DPE array 202 according to the logic architecture, and / or any other user-specified IP cores and / or RTL cores to be implemented in PL 214. Examples of other IP cores and / or RTL cores that can be manually inserted by the user include, but are not limited to, data width conversion blocks, hardware buffers, and / or clock domain logic. In one aspect, each block of the block diagram can correspond to a specific core (e.g., a circuit block) to be implemented in PL 214. The block diagram specifies the connectivity of the core to be implemented in the PL and the connectivity of the core with the physical resources of NoC 208 and / or SoC interface block 206 as determined by the SoC interface block scheme and the logic architecture.
[0229] In one aspect, the hardware compiler 1606 can also create logical connections between the cores of the PL 214 and global memory (e.g., volatile memory 134) by creating NoC traffic according to the logic architecture and executing the NoC compiler 1604 to obtain the NoC solution. In one example, the hardware compiler 1606 can route the logical connections to verify the ability of the PL 214 to implement the block diagram and the logical connections. In another aspect, the hardware compiler 1606 can use SoC interface block traces (e.g., described in more detail below) in conjunction with one or more data traffic generators as part of the simulation to verify the functionality of the block diagram using actual data traffic.
[0230] The system performs an implementation flow on the block diagram in block 2008. For example, a hardware compiler can perform an implementation flow on the block diagram involving synthesis (if necessary), placement, and routing to generate a configuration bitstream that can be loaded into the SoC 200 to implement the hardware portion 214 of the application in the PL.
[0231] The hardware compiler 1606 can execute an implementation flow on the block diagram using the SoC interface block plan and the NoC plan. For example, because the SoC interface block plan specifies specific flow channels of the SoC interface block 206 through which a particular DPE 204 communicates with the PL 214, the placer can place blocks of the block diagram that are connected to the DPE 204 via the SoC interface block 206 so that they are close (e.g., within a specific distance) to the specific flow channels of the SoC interface block 206 to which these blocks are connected. For example, ports of a block can be associated with flow channels specified by the SoC interface block plan. The hardware compiler 1606 can also route connections between ports of blocks of the block diagram that are connected to the SoC interface block 206 by routing signals input to and / or output from the ports to the BLI of the PL 214, where the BLI of the PL 214 is connected to the specific flow channels coupled to these ports, as determined from the SoC interface block plan.
[0232] Similarly, because the NoC scheme specifies specific entry / exit points to which circuit blocks in the PL 214 are to be connected, the placer can place blocks of the block diagram that are connected to the NoC 208 so that they are close to (e.g., within a specific distance) the specific entry / exit points to which these blocks are connected. For example, ports of a block can be associated with entry / exit points of the NoC scheme. Hardware compiler 1606 can also route connections between ports of blocks of the block diagram that are connected to entry / exit points of the NoC 208 by routing signals input to and / or output from the ports to the entry / exit points of the NoC 208, where the entry / exit points of the NoC 208 are logically coupled to these ports, as determined from the NoC scheme. Hardware compiler 1606 can also route any signals that connect ports of blocks in the PL 214 to each other. However, in some applications, the NoC 208 may not be used to transfer data between the DPE array 202 and the PL 214.
[0233] In block 2010, during the implementation flow, the hardware compiler optionally exchanges design data with the DPE compiler 1602 and / or the NoC compiler 1604. For example, the hardware compiler 1606, the DPE compiler 1602, and the NoC compiler 1604 can be combined as shown in FIG. Figure 16 The design data may be exchanged one-time, as needed, or on an iterative or repetitive basis as described. Block 2010 may be performed selectively. For example, hardware compiler 1606 may exchange design data with DPE compiler 1602 and / or NoC compiler 1604 before or during construction of the block diagram, before and / or during placement, and / or before and / or during routing.
[0234] In block 2012, the system exports the final hardware design generated by the hardware compiler 1606 as a hardware package. The hardware package contains the configuration bitstream used to program the PL 214. The hardware package is generated based on the hardware portion of the application.
[0235] In block 2014, the user configures a new platform using the hardware package. The user initiates the generation of a new platform based on the configuration provided by the user. The platform generated by the system using the hardware package is used to compile the software portion of the application.
[0236] In block 2016, the system compiles the software portion of the application for implementation in the DPE array 202. For example, the system executes the DPE compiler 1602 to generate one or more binary files that can be loaded into each of the DPEs 204 in the DPE array 202. The binary files for the DPEs 204 can include object code, application data, and configuration data for the DPEs 204. Once the configuration bitstream and binary files are generated, the system can load the configuration bitstream and binary files into the SoC 200 to implement the application therein.
[0237] On the other hand, the hardware compiler 1606 can provide the hardware implementation to the DPE compiler 1602. The DPE compiler 1602 can extract the final SoC interface block solution that the hardware compiler 1606 relies on when executing the implementation process. The DPE compiler 1602 uses the same SoC interface block solution used by the hardware compiler 1606 to perform compilation.
[0238] exist Figure 20 In the example, each part of the application is solved by a subsystem-specific compiler. The compiler is able to communicate design data, such as constraints and / or suggested solutions, to ensure that the interfaces between the various subsystems implemented for the application (e.g., SoC interface blocks) are compliant and consistent. Although not in Figure 20 As specifically shown in FIG, if used in an application, a NoC compiler 1604 may also be called to generate a binary file for programming the NoC 208.
[0239] Figure 21 Another example method 2100 of performing a design flow to implement an application in SoC 200 is shown. The method 2100 may be performed by combining Figure 1 The system described herein performs the following operations: Figure 16 or Figure 19The method 2100 may begin at block 2102, where the system receives an application. The application may be specified as a data flow graph to be implemented in the SoC 200. The application may include a software portion for implementation in the DPE array 202, a hardware portion for implementation in the PL 214, and a hardware portion for implementing data transfers in the NoC 208 of the SoC 200. The application may also include a further software portion for implementation in the PS 212.
[0240] In block 2104, the DPE compiler 1602 can generate a logical architecture, a SoC interface block schema, and a SoC interface block trace from the application. The logical architecture can be based on the DPE 204 required to implement the software portion of the application designated for implementation within the DPE array 202, as well as any IP cores required to interact with the DPE 204 to be implemented in the PL 214. As described above, the DPE compiler 1602 can generate an initial DPE schema, where the DPE compiler 1602 performs an initial mapping of nodes (of the software portion of the application) to the DPE array 202. The DPE compiler 1602 can also generate a schema for the initial SoC interface block, which maps logical resources to physical resources (e.g., flow channels) of the SoC interface block 206. In one aspect, the SoC interface block schema can be generated using the initial NoC schema generated by the NoC compiler 1604 from the data transfer. The DPE compiler 1602 can also simulate the initial DPE schema with the SoC interface block schema to simulate the data flow through the SoC interface block 206. The DPE Compiler 1602 is capable of capturing data transfers through SoC interface blocks during simulation as "SoC interface block traces" for use in Figure 21 Subsequent use during the design flow is shown.
[0241] In block 2104, hardware compiler 1606 generates a block diagram for the hardware portion of the application to be implemented in PL 214. Hardware compiler 1606 generates the block diagram based on the logic architecture and SoC interface block scheme, as well as optional user-specified additional IP cores to be included in the block diagram with the circuit blocks specified in the logic architecture. In one aspect, the user manually inserts such additional IP cores and connects the IP cores to the other circuit blocks of the hardware description specified in the logic architecture.
[0242] In block 2106 , the hardware compiler 1606 optionally receives one or more user-specified SoC interface block constraints and provides the SoC interface block constraints to the DPE compiler 1602 .
[0243] In one aspect, before implementing the hardware portion of an application, hardware compiler 1606 can evaluate the physical connections defined between NoC 208, DPE array 202, and PL 214 based on the block diagram and logical architecture. Hardware compiler 1606 can perform an architectural simulation of the block diagram to evaluate the connections between the block diagram (e.g., the PL portion of the design) and the DPE array 202 and / or NoC 208. For example, hardware compiler 1606 can perform the simulation using the SoC interface block traces generated by DPE compiler 1602. As an illustrative and non-limiting example, hardware compiler 1606 can perform a SystemC simulation of the block diagram. In the simulation, using the SoC interface block traces, data traffic is generated for the block diagram and for the flow channels (e.g., physical connections) between PL 214 and DPE array 202 (via SoC interface block 206) and / or NoC 208, using the SoC interface block routing. The simulation generates system performance and / or debug information that is provided to hardware compiler 1606.
[0244] Hardware compiler 1606 can evaluate system performance data. For example, if hardware compiler 1606 determines, based on the system performance data, that one or more design specifications of the hardware portion of the application are not met, hardware compiler 1606 can generate one or more SoC interface block constraints under user guidance. Hardware compiler 1606 provides the SoC interface block constraints as a request to DPE compiler 1602.
[0245] The DPE compiler 1602 can perform an updated mapping of the DPE portion of the application to the DPEs 204 of the DPE array 202, using the SoC interface block constraints provided by the hardware compiler 1606. For example, if the hardware portion of the application in the PL 214 is implemented with the hardware portion connected directly to the DPE array 202 via the SoC interface block 206 (e.g., not through the NoC 208), the DPE compiler 1602 can generate an updated SoC interface block solution for the hardware compiler 1606 without involving the NoC compiler 1604.
[0246] In block 2108, hardware compiler 1606 optionally receives one or more user-specified NoC constraints and provides the NoC constraints to the NoC compiler for verification. Hardware compiler 1606 may also provide NoC traffic to NoC compiler 1604. NoC compiler 1604 can use the received NoC constraints and / or NoC traffic to generate an updated NoC solution. For example, if an application is implemented where the hardware portion of PL 214 is connected to DPE array 202, PS 212, hardwired circuit block 210, or volatile memory 134 via NoC 208, hardware compiler 1606 can invoke NoC compiler 1604 by providing the NoC constraints and / or NoC traffic to NoC compiler 1604. NoC compiler 1604 can update routing information for data paths through NoC 208 as an updated NoC solution. The updated routing information may specify updated routes and specific ingress / egress points for the routes. The hardware compiler 1606 may obtain the updated NoC solution and, in response, generate updated SoC interface block constraints that are provided to the DPE compiler 1602. This process may be iterative in nature. The DPE compiler 1602 and the NoC compiler 1604 may operate concurrently, as shown in blocks 2106 and 2108.
[0247] In block 2110, the hardware compiler 1606 can perform synthesis on the block diagram. In block 2112, the hardware compiler 1606 performs placement and routing on the block diagram. In block 2114, while performing placement and / or routing, the hardware compiler can determine whether the implementation of the block diagram (e.g., the current state of the implementation of the hardware portion (e.g., the hardware design) at any of these different stages of the implementation process) meets the design specifications for the hardware portion of the application. For example, the hardware compiler 1606 can determine whether the current implementation meets the design specifications before placement, during placement, before routing, or during routing. In response to determining that the current implementation of the hardware portion of the application does not meet the design specifications, the method 2100 continues to block 2116. Otherwise, the method 2100 continues to block 2120.
[0248] In block 2116, the hardware compiler can provide one or more user-specified SoC interface block constraints to the DPE compiler 1602. The hardware compiler 1606 can optionally provide one or more NoC constraints to the NoC compiler 1604. As discussed, the DPE compiler 1602 generates an updated SoC interface block solution using the SoC interface block constraints received from the hardware compiler 1606. The NoC compiler 1604 optionally generates an updated NoC solution. For example, if one or more data paths between the DPE array 202 and the PL 214 flow through the NoC 208, the NoC compiler 1604 can be invoked. In block 2118, the hardware compiler 1606 receives the updated SoC interface block solution and, optionally, the updated NoC solution. After block 2118, the method 2100 continues to block 2112, where the hardware compiler 1606 continues to perform placement and / or routing using the updated SoC interface block solution and, optionally, the updated NoC solution.
[0249] Figure 21 It is shown that the exchange of design data between compilers can be performed in an iterative manner. For example, at any of a number of different points during the placement and / or routing phases, the hardware compiler 1606 can determine whether the current state of the implementation of the hardware portion of the application meets the established design criteria. If not, the hardware compiler 1606 can initiate an exchange of design data as described to obtain an updated SoC interface block solution and an updated NoC solution, which the hardware compiler 1606 uses for placement and routing purposes. It should be understood that the hardware compiler 1606 only needs to call the NoC compiler 1604 in situations where the configuration of the NoC 208 is to be updated (e.g., data from the PL 214 is provided to and / or received from other circuit blocks via the NoC 208).
[0250] In block 2120, if the hardware portion of the application meets the design specifications, the hardware compiler 1606 generates a configuration bitstream that specifies the implementation of the hardware portion within the PL 214. The hardware compiler 1606 can also provide the final SoC interface block solution (e.g., the SoC interface block solution for placement and routing) to the DPE compiler 1602 and provide the final NoC solution, which may have been used for placement and routing, to the NoC compiler 1604.
[0251] In block 2122, the DPE compiler 1602 generates a binary file for programming the DPEs 202 of the DPE array 204. The NoC compiler 1604 generates a binary file for programming the NoC 208. For example, throughout blocks 2106, 2108, and 2116, the DPE compiler 1602 and the NoC compiler 1604 may perform an incremental verification function, where the SoC interface block solution and the NoC solution used are generated based on a verification program. This verification process may be performed in less execution time than determining the complete solution for the SoC interface block and the NoC. In block 2122, the DPE compiler 1602 and the NoC compiler 1604 may generate final binary files for programming the DPE array 202 and the NoC 208, respectively.
[0252] In block 2124, the PS compiler 1918 generates a PS binary file. The PS binary file includes object code to be executed by the PS 212. For example, the PS binary file implements a control program executed by the PS 212 to monitor the operation of the SoC 200 and the applications implemented therein. The DPE compiler 1602 may also generate a DPE array driver, which may be compiled by the PS compiler 1918 and executed by the PS 212 to read and / or write to the DPE 204 of the DPE array 202.
[0253] In block 2126, the system can deploy configuration bitstreams and binaries in the SoC 200. For example, the system can combine the various binaries and configuration bitstream groups into a PDI that can be provided to and loaded into the SoC 200 to implement applications therein.
[0254] Figure 22 An example method 2200 is shown for communication between the hardware compiler 1606 and the DPE compiler 1602. The method 2200 presents an example method for communicating between the hardware compiler 1606 and the DPE compiler 1602. Figure 16 、 19 , 20, and 21. Method 2200 illustrates an example implementation of verification calls (e.g., verification procedures) made between the hardware compiler 1606 and the DPE compiler 1602. The example of method 2200 provides an alternative to performing full placement and routing on the DPE array 202 and / or NoC 208 to generate updated SoC interface block plans in response to SoC interface block constraints provided from the hardware compiler 1606. Method 2200 illustrates an incremental approach in which rerouting is attempted before initiating mapping and routing of the software portion of the application.
[0255] The method 2200 may begin at block 2202, where the hardware compiler 1606 provides one or more SoC interface block constraints to the DPE compiler 1602. For example, the hardware compiler 1606 may receive one or more user-specified SoC interface block constraints and / or generate one or more SoC interface block constraints during an implementation process and in response to determining that a design specification of a hardware portion of an application is not or will not be met. The SoC interface block constraints may specify a preferred mapping of logical resources to physical flow channels of the SoC interface block 206, which is expected to result in an improved quality of results (QoS) for the hardware portion of the application.
[0256] Hardware compiler 1606 provides SoC interface block constraints to DPE compiler 1602. The SoC interface block constraints provided by hardware compiler 1606 can fall into two different categories. The first category of SoC interface block constraints are hard constraints. The second category of SoC interface block constraints are soft constraints. Hard constraints are design constraints that must be met to implement an application within SoC 200. Soft constraints are design constraints that may be violated during application implementation within SoC 200.
[0257] In one example, hard constraints are user-specified constraints on the hardware portion of an application to be implemented in the PL 214. Hard constraints can include any available constraint type, such as location, power, timing, etc., which are user-specified constraints. Soft constraints can include any available constraints generated by the hardware compiler 1606 and / or the DPE compiler 1602 throughout the implementation flow, such as constraints that specify a specific mapping of logic resources to flow channels of the described SoC interface block 206.
[0258] In response to receiving the SoC interface block constraints, the DPE compiler 1602 initiates a verification process to merge the received SoC interface block constraints into generating an updated SoC interface block solution in block 2204. In block 2206, the DPE compiler 1602 can distinguish between hard constraints and soft constraints related to the hardware portion of the application received from the hardware compiler 1606.
[0259] At block 2208, the DPE compiler 1602 routes the software portion of the application while complying with both the hard and soft constraints provided by the hardware compiler. For example, the DPE compiler 1602 can route connections between the DPEs 204 of the DPE array 202 and data paths between the DPEs 204 and the SoC interface block 206 to determine which flow channels (e.g., cell blocks, flow switches, and ports) of the SoC interface block 206 are used for data paths crossing between the DPE array 202 and the PL 214 and / or the NoC 208. If the DPE compiler 1602 successfully routes the software portion of the application for implementation in the DPE array 202 while complying with both the hard and soft constraints, the method 2200 continues to block 2218. If the DPE compiler 1602 cannot generate a route for the software portion of the application in the DPE array while complying with both the hard and soft constraints, e.g., the constraint is not routable, the method 2200 continues to block 2210.
[0260] In block 2210, the DPE compiler 1602 routes the software portion of the application while respecting only the hard constraints. In block 2210, the DPE compiler 1602 ignores the soft constraints for the purpose of routing operations. If the DPE compiler 1602 successfully routes the software portion of the application for implementation in the DPE array 202 while respecting only the hard constraints, the method 2200 continues to block 2218. If the DPE compiler 1602 cannot generate a route for the software portion of the application in the DPE array 202 while respecting only the hard constraints, the method 2200 continues to block 2212.
[0261] Blocks 2208 and 2210 illustrate a method for verification operations that seeks to create an updated SoC interface block solution in a shorter time than would be required to perform a complete mapping (e.g., location) and routing of DPE nodes using the SoC interface block constraints provided from the hardware compiler 1606. Thus, blocks 2208 and 2210 only involve routing and do not attempt to map (e.g., remap) or "place" DPE nodes to the DPEs 204 of the DPE array 202.
[0262] In the event that routing alone using the SoC interface block constraints from the hardware compiler cannot reach the updated SoC interface block solution, the method 2200 continues to block 2212. In block 2212, the DPE compiler 1602 can map the software portion of the application to the DPEs in the DPE array 202 using both hard and soft constraints. The DPE compiler 1602 is also programmed with the architecture (e.g., connectivity) of the SoC 200. The DPE compiler 1602 performs the actual allocation of logical resources to the physical channels (e.g., to flow channels) of the SoC interface block 206 and can also model the architectural connections of the SoC 200.
[0263] As an example, consider DPE node A communicating with PL node B. Each block of the block diagram may correspond to a specific core (e.g., a circuit block) to be implemented in PL 214. PL node B communicates with DPE node A via physical channel X in SoC interface block 206. Physical channel X carries the data flow between DPE node A and PL node B. DPE compiler 1602 can map DPE node A to a specific DPE Y such that the distance between DPE Y and physical channel X is minimized.
[0264] In some implementations of the SoC interface block 206, one or more cell blocks included therein are not connected to the PL 214. The unconnected cell blocks may be a result of the layout of specific hard-wired circuit blocks 210 in and / or around the PL 214. This architecture (e.g., having unconnected cell blocks in the SoC interface block 206) complicates routing between the SoC interface block 206 and the PL 214. Connection information about the unconnected blocks is modeled in the DPE compiler 1602. As part of the execution mapping, the DPE compiler 1602 can select DPE nodes that have connections to the PL 214. As part of the execution mapping, the DPE compiler 1602 can minimize the number of selected DPE nodes that are mapped to the DPE 204 in the column of the DPE array 202 that is immediately above the unconnected cell blocks of the SoC interface block 206. The DPE compiler 1602 maps DPE nodes that are not connected (eg, directly connected) to the PL 214 (eg, nodes connected to other DPEs 204 ) to columns of the DPE array 202 located above the unconnected cell blocks of the SoC interface block 206 .
[0265] In block 2214, the DPE compiler 1602 routes the remapped software portion of the application while respecting only the hard constraints. If the DPE compiler 1602 successfully routes the remapped software portion of the application for implementation in the DPE array 202 while respecting only the hard constraints, the method 2200 proceeds to block 2218. If the DPE compiler 1602 cannot generate a route for the software portion of the application in the DPE array 202 while respecting only the hard constraints, the method 2200 proceeds to block 2216. In block 2216, the DPE compiler 1602 indicates that the verification operation failed. The DPE compiler 1602 may output a notification and may provide the notification to the hardware compiler 1606.
[0266] In block 2218, the DPE compiler 1602 generates an updated SoC interface block solution and a score for the updated SoC interface block solution. The DPE compiler 1602 generates the updated SoC interface block solution based on the updated routing or updated mapping and routing determined in block 2208, block 2210, or blocks 2212 and 2214.
[0267] The score generated by the DPE compiler 1602 indicates the quality of the SoC interface block solution based on the mapping and / or routing operations performed. In one example implementation, the DPE compiler 1602 determines a score based on how many soft constraints were not satisfied and a distance between the flow channels requested in the soft constraints and the actual channels allocated in the updated SoC interface block solution. For example, both the number of unsatisfied soft constraints and the distance may be inversely proportional to the score.
[0268] In another example implementation, the DPE compiler 1602 determines a score based on the quality of the updated SoC interface block solution using one or more design cost metrics. These design cost metrics may include the number of data moves supported by the SoC interface block solution, memory conflict costs, and routing delays. In one aspect, the number of data moves in the DPE array 202, in addition to those required to transfer data across the SoC interface block 206, can be quantified by the number of DMA transfers used in the DPE array 202. The memory conflict costs can be determined based on the number of concurrent access circuits (e.g., DPEs or DMAs) per memory bank. The routing delay can be quantified by the minimum number of cycles required to transfer data between the SoC interface block 206 port and a separate source or destination DPE 204. When the design cost metrics are lower (e.g., the sum of the design cost metrics is lower), the DPE compiler 1602 determines a higher score.
[0269] In another example implementation, the total score for the updated SoC interface block solution is calculated as a score (e.g., 80 / 100) where the numerator is subtracted from 100 the sum of the number of other DMA transfers, the number of concurrent access circuits per memory bank when there are more than two, and the number of hops required for routing between the SoC interface block 206 ports and the DPE 204 core.
[0270] In block 2220, the DPE compiler 1602 provides the updated SoC interface block solution and the score to the hardware compiler 1606. The hardware compiler 1606 can evaluate the various SoC interface block solutions received from the DPE compiler 1602 based on the scores of each corresponding SoC interface block solution. In one aspect, for example, the hardware compiler 1606 can retain the previous SoC interface block solution. The hardware compiler 1606 can compare the score of the updated SoC interface block solution with the score of the previous (e.g., the immediately preceding SoC interface block solution) and use the updated SoC interface block solution if the score of the updated SoC interface block solution exceeds the score of the previous SoC interface block solution.
[0271] In another example, hardware compiler 1606 receives a SoC interface block solution with a score of 80 / 100 from DPE compiler 1602. Hardware compiler 1606 cannot implement the hardware portion of the application within PL 214 and provides one or more SoC interface block constraints to DPE compiler 1602. The updated SoC interface block solution received by hardware compiler 1606 from DPE compiler 1602 has a score of 20 / 100. In that case, in response to determining that the score of the newly received SoC interface block solution does not exceed (e.g., is lower than) the score of the previous SoC interface block solution, hardware compiler 1606 relaxes one or more interface block constraints (e.g., soft constraints) in the SoC and provides the SoC interface block constraints (including the relaxed constraints) to DPE compiler 1602. DPE compiler 1602 attempts to generate another SoC interface block solution that, given the relaxed design constraints, has a score higher than 20 / 100 and / or 80 / 100.
[0272] In another example, the hardware compiler 1606 may select to use a previous SoC interface block solution with a higher or highest score. The hardware compiler 1606 may revert to an earlier SoC interface block solution at any time, for example, in response to receiving a SoC interface block solution with a lower score than an immediately previous SoC interface block solution or in response to receiving an interface block solution with a lower score than a previous SoC interface block solution after one or more SoC interface block constraints have been relaxed.
[0273] Figure 23 An example method 2300 for processing SoC interface block solutions is shown. The method 2300 can be executed by the hardware compiler 1606 to evaluate received SoC interface block solutions and select a SoC interface block solution, referred to as a current best SoC interface block solution, for executing an implementation flow on the hardware portion of the application.
[0274] In block 2302, the hardware compiler 1606 receives a SoC interface block solution from the DPE compiler 1602. The SoC interface block solution received in block 2302 may be an initial or first SoC interface block solution provided by the DPE compiler 1602. When providing the SoC interface block solution to the hardware compiler 1606, the DPE compiler 1602 further provides scores for the SoC interface block solutions. At least initially, the hardware compiler 1606 selects the first SoC interface block solution as the current best SoC interface block solution.
[0275] In block 2304, the hardware compiler 1606 optionally receives one or more hard SoC interface block constraints from the user. In block 2306, the hardware compiler can generate one or more soft SoC interface block constraints to implement the hardware portion of the application. The hardware compiler generates the soft SoC interface block constraints in an effort to meet the hardware design specifications.
[0276] In block 2308, the hardware compiler 1606 sends the SoC interface block constraints (e.g., both hard and soft constraints) to the DPE compiler 1602 for verification. In response to receiving the SoC interface block constraints, the DPE compiler can generate an updated SoC interface block solution based on the SoC interface block constraints received from the hardware compiler 1606. The DPE compiler 1602 provides the updated SoC interface block solution to the hardware compiler 1606. Thus, in block 2310, the hardware compiler receives the updated SoC interface block solution.
[0277] In block 2312, the hardware compiler 1606 compares the score of the updated SoC interface block solution (eg, the most recently received SoC interface block solution) with the score of the first (eg, previously received) SoC interface block solution.
[0278] In block 2314, hardware compiler 1606 determines whether the score of the updated (e.g., most recently received) SoC interface block solution exceeds the score of the previously received (e.g., first) SoC interface block solution. In block 2316, hardware compiler 1606 selects the most recently received (e.g., updated) SoC interface block solution as the current best SoC interface block solution.
[0279] In block 2318, the hardware compiler 1606 determines whether the improvement goal has been achieved or whether the time budget has been exceeded. For example, the hardware compiler 1606 can determine whether the current implementation state of the hardware portion of the application meets a larger number of design specifications and / or is close to meeting one or more design specifications. The hardware compiler 1606 can also determine whether the time budget has been exceeded based on the amount of processing time spent on placement and / or routing and whether that time exceeds the maximum placement time, the maximum routing time, or the maximum time for placement and routing. In response to determining that the improvement goal has been achieved or the time budget has been exceeded, the method 2300 continues to block 2324. If not, the method 2300 continues to block 2320.
[0280] In block 2324, the hardware compiler 1606 implements the hardware portion of the application using the current best SoC interface block solution.
[0281] Continuing with block 2320, the hardware compiler 1606 relaxes one or more of the SoC interface block constraints. The hardware compiler 1606 may relax, for example, or change one or more of the soft constraints. An example of relaxing or changing a soft SoC interface block constraint includes removing (e.g., deleting) the soft SoC interface block constraint. Another example of relaxing or changing a soft SoC interface block constraint includes replacing the soft SoC interface block constraint with a different SoC interface block constraint. The replaced soft SoC interface block constraint may not be as strict as the original constraint being replaced.
[0282] In block 2322, the hardware compiler 1606 can send the SoC interface block constraints (including the relaxed SoC interface block constraints) to the DPE compiler 1602. After block 2322, the method 2300 loops back to block 2310 to continue processing as described. For example, the DPE compiler generates a further updated SoC interface block solution based on the SoC interface block constraints received from the hardware compiler in block 2322. In block 2310, the hardware compiler receives the further updated SoC interface block solution.
[0283] Method 2300 illustrates an example process for selecting a SoC interface block solution from the DPE compiler 1602 for use in executing an implementation flow, and the circumstances under which SoC interface block constraints may be relaxed. It should be understood that the hardware compiler 1606 may provide the SoC interface block constraints to the DPE compiler 1602 at any of a variety of different points during the implementation flow to obtain an updated SoC interface block solution as part of a coordination and / or verification process. For example, at any point where the hardware compiler 1606 determines (e.g., based on timing, power, or other checks or analyses) that the implementation of the hardware portion of the application in its current state does not meet or will not meet the design specifications of the application, the hardware compiler 1606 may request an updated SoC interface block solution by providing the updated SoC interface block constraints to the DPE compiler 1602.
[0284] Figure 24 Another example of an application 2400 for implementation in SoC 200 is shown. Application 2400 is specified as a directed flow graph. Nodes are shaded and shaped differently to distinguish between PL nodes, DPE nodes, and I / O nodes. In the example shown, I / O nodes can be mapped to SoC interface block 206. PL nodes are implemented in the PL. DPE nodes are mapped to specific DPEs. Although not shown in full, application 2400 includes 36 cores (e.g., nodes) to be mapped to DPE 204, 72 PLs with data flows to the DPE arrays, and 36 DPE arrays with data flows to the PLs.
[0285] Figure 25 is an example diagram of a SoC interface block solution generated by the DPE compiler 1602. Figure 25 The SoC interface block solution can be generated by the DPE compiler 1602 and provided to the hardware compiler 1606. Figure 25 The example of FIG. 1 shows a scenario in which the DPE compiler 1602 generates an initial mapping of DPE nodes to the DPEs 204 of the DPE array 202. Furthermore, the DPE compiler 1602 successfully routes the initial mapping of the DPE nodes. Figure 25 In the example shown, only columns 6-17 of the DPE array 202 are shown.
[0286] Figure 25 The mapping of DPE nodes to DPE 204 of DPE array 202 and the routing of data flows to SoC interface block 206 hardware are shown. The mapping of DPE nodes 0-35 of application 2400 to DPE 204 as determined by DPE compiler 1602 is shown with reference to DPE array 202. The routing of data flows between the DPE and specific unit blocks of SoC interface block 206 is shown as a set of arrows. Figure 25-30 For the purpose, Figure 25The buttons displayed in are used to distinguish between data flows controlled by soft constraints, hard constraints, and data flows with no constraints applied.
[0287] refer to Figure 25-30 , soft constraints correspond to routes determined by the DPE compiler 1602 and / or the hardware compiler 1606, while hard constraints may include user-specified SoC interface block constraints. Figure 25 All constraints shown in are soft constraints. Figure 25 , where the DPE compiler 1602 successfully determines an initial SoC interface block solution. In one aspect, the DPE compiler 1602 can be configured to at least initially attempt to use vertical routing for the SoC interface block solution before attempting to use other routing from one column to another along (e.g., from left to right) a row of DPEs 204.
[0288] Figure 26 16 shows an example of a routable SoC interface block constraint received by the DPE compiler 1602. The DPE compiler 1602 is capable of generating an updated SoC interface block solution that specifies updated routing in the form of updated SoC interface block constraints. Figure 26 In the example shown, a large number of SoC interface block constraints are hard constraints. In this example, the DPE compiler 1602 successfully routes the data flow of the DPE array 202 while observing each type of constraint shown.
[0289] Figure 27 An example of a non-routable SoC interface block constraint that will be observed by the DPE compiler 1602 is shown. The DPE compiler 1602 cannot generate an observation Figure 27 The constraints shown in the SoC interface block scheme.
[0290] Figure 28 An example is shown where the DPE compiler 1602 ignores the Figure 27 Soft type SoC interface block constraints. Figure 28 In the example shown, DPE compiler 1602 uses only hard constraints to successfully route the software portion of the application for implementation in DPE array 202. Those data flows that are not controlled by constraints can be routed in any way that DPE compiler 1602 deems appropriate or is able to do so.
[0291] Figure 29 Another example of a non-routable SoC interface block constraint is shown. Figure 29 In the example of , there are only hard constraints. Therefore, the DPE compiler 1602 cannot start the mapping (or remapping) operation without ignoring the hard constraints.
[0292] Figure 30 Shown Figure 29In this example, after remapping, the DPE compiler 1602 is able to successfully route the DPE nodes to generate an updated SoC interface block solution.
[0293] Figure 31 Another example of a non-routable SoC interface block constraint is shown. Figure 31 In the example of , there are only hard constraints. Therefore, the DPE compiler 1602 that cannot ignore the hard constraints initiates the mapping operation. For illustrative purposes, the DPE array 202 includes only three rows of DPEs (eg, 3 DPEs per column).
[0294] Figure 32 Shown Figure 31 Example of DPE node mapping. Figure 32 Shown from the combination of Figure 31 The result obtained by the initiated remapping operation. In this example, after remapping, the DPE compiler 1602 is able to successfully route the applied software solution to generate an updated SoC interface block solution.
[0295] In one aspect, the system is able to perform the Figure 25-32 The mapping shown. The ILP formulation can include multiple different variables and constraints that define the mapping problem. The system is capable of solving the ILP formulation while also minimizing the cost. The cost can be determined at least in part based on the number of DMA engines used. In this way, the system is capable of mapping the DFG onto the DPE array.
[0296] In another aspect, the system can sort the nodes of the DFG in descending order of priority. The system can determine the priority based on one or more factors. Examples of factors may include, but are not limited to, the height of the node in the DFG graph, the total degree of the node (e.g., the sum of all edges entering and leaving the node), and / or the type of edges connected to the node, such as memory, flow, and cascade. The system can place the node on the best available DPE based on affinity and availability. The system can determine availability based on whether all resource requirements of the node (e.g., compute resources, memory buffers, flow resources) can be met on a given DPE. The system can determine affinity based on one or more other factors. Examples of affinity factors may include placing the node on the same DPE or on a neighboring DPE (where the node's neighbors are already placed) to minimize DMA communication, architectural constraints (e.g., whether the node is part of a cascade chain), and / or finding the DPE with the most available resources. If the node is placed while meeting all constraints, the system can increase the priority of the neighboring nodes of the placed node so that they can be processed next. If there are no valid placements available for the current node, the system may attempt to de-place some other nodes from their best candidate DPEs to make room for the node. The system may place the de-placed nodes back into the priority queue for further placement. The system can limit the total effort spent on finding a good solution by tracking the total number of placements and de-placements performed. However, it should be understood that other mapping techniques can be used and the examples provided here are not intended to be limiting.
[0297] Figure 33 Shows that the combination Figure 1 Another example software architecture 3300 for system implementation is described. For example, Figure 33 The architecture 3300 can be composed of Figure 1 The system is implemented by one or more program modules 120. Figure 33 The example software architecture 3300 may be used in a situation where an application (e.g., a data flow graph) specifies one or more high-level synthesis (HLS) kernels for implementation in the PL 214. For example, a PL node of the application references an HLS kernel that requires HLS processing. In one aspect, the HLS kernel is specified in a high-level language (HLL) such as C and / or C++.
[0298] exist Figure 33 In the example of FIG, the software architecture 3300 includes a DPE compiler 1602, a hardware compiler 1606, an HLS compiler 3302, and a system linker 3304. The NoC compiler 1604 can be included and used in conjunction with the DPE compiler 1602 to perform verification checks 3306 as previously described.
[0299] As shown, DPE compiler 1602 receives application 3312, SoC architecture description 3310, and optional test bench 3314. As discussed, application 3312 can be specified as a dataflow graph including parallel execution semantics. Application 3312 can include interconnected PL nodes and DPE nodes and specify execution time parameters. In this example, the PL node references an HLS kernel. SoC architecture description 3310 can be a data structure or file that specifies information such as the size and dimensions of DPE array 202, the size of PL 214 and the various programmable circuit blocks available therein, the type of PS 212 (e.g., the type of processors and other devices included in PS 212), and other physical characteristics of the circuitry in SoC 200 in which application 3312 will be implemented. SoC architecture description 3310 can also specify connectivity (e.g., interfaces) between the subsystems included therein.
[0300] The DPE compiler 1602 can output the HLS kernel to the HLS compiler 3302. The HLS compiler 3302 converts the HLS kernel specified in the HLL into HLS IP that can be synthesized by the hardware compiler. For example, the HLS IP can be specified as a register transfer level (RTL) block. For example, the HLS compiler 3302 generates an RTL block for each HLS kernel. As shown in the figure, the HLS compiler 3302 outputs the HLS IP to the system linker 3304.
[0301] The DPE compiler 1602 generates other outputs, such as an initial SoC interface block plan and a connectivity graph. The DPE compiler 1602 outputs the connectivity graph to the system linker 3304 and outputs the SoC interface block plan to the hardware compiler 1606. The connectivity graph specifies the connectivity between the nodes corresponding to the HLS kernel to be implemented in the PL 214 (now converted to HLS IP) and the nodes to be implemented in the DPE array 202.
[0302] As shown, the system linker 3304 receives the SoC architecture description 3310. The system linker 3304 can also receive one or more HLS and / or RTL blocks directly from the application 3312, which are not processed by the DPE compiler 1602. The system linker 3304 can automatically generate a block diagram corresponding to the hardware portion of the application using the received HLS and / or RTL blocks, HLS IP, and a connection diagram of the connectivity between the specified IP cores and the connectivity between the IP cores and the DPE nodes. In one aspect, the system linker 3304 can integrate the block diagram with a base platform (not shown) of the SoC 200. For example, the system linker 3304 can connect the block diagram to the base platform to produce an integrated block diagram. The block diagram and the connected base platform can be referred to as a synthesizable block diagram.
[0303] On the other hand, the HLS IP and RTL IP referred to as kernels within an SDF graph (e.g., application 3312) may be compiled into IP outside of the DPE compiler 1602. The compiled IP may be provided directly to the system linker 3304. The system linker 3304 can automatically generate a block diagram corresponding to the hardware portion of the application using the provided IP.
[0304] In one aspect, the system linker 3304 can include other hardware-specific details derived from the original SDF (e.g., application 3312) and the generated connection diagram within the block diagram. For example, because the application 3312 includes software models that are actual HLS models, which can be converted to IP or associated (e.g., matched) with IP in a database of such IP using some mechanism (e.g., by name or other matching / correlation techniques), the system linker 3304 can automatically generate the block diagram (e.g., without user intervention). In this example, custom IP may not be used. In the automatically generated block diagram, the system linker 3304 can automatically insert one or more other circuit blocks, such as data width conversion blocks, hardware buffers, and / or clock domain crossing logic, which in other cases described herein are manually inserted and connected by the user. For example, the system linker 3304 can analyze the data type and software model to determine that one or more other circuit blocks are required, as described, to create the connections specified by the connection diagram.
[0305] The system linker 3304 outputs the block diagram to the hardware compiler 1606. The hardware compiler 1606 receives the block diagram and the initial SoC interface block plan generated by the DPE compiler 1602. The hardware compiler 1606 can initiate verification checks 3306 with the DPE compiler 1602 and the optional NoC compiler 1604, as previously described in conjunction with Figure 20 Block 2010, Figure 21 Blocks 2106, 2108, 2112, 2114, 2116, and 2118, Figure 22 and Figure 23 Verification may be an iterative process in which the hardware compiler provides design data, such as various types of constraints (which may include relaxed / modified constraints in an iterative approach), to the DPE compiler 1602 and, optionally, to the NoC compiler 1604, and, in return, receives an updated SoC interface block solution from the DPE compiler 1602 and, optionally, an updated NoC solution from the NoC compiler 1604.
[0306] The hardware compiler 1606 can generate a hardware package that includes a configuration bitstream for the hardware portion of the application 3312 implemented in the PL 214. The hardware compiler 1606 can output the hardware package to the DPE compiler 1602. The DPE compiler 1602 can generate DPE array configuration data (e.g., one or more binary files) that programs the software portion of the application 3312 intended to be implemented in the DPE array 202 therein.
[0307] Figure 34 Another example method 3400 for performing a design flow to implement an application in SoC 200 is shown. The method 3400 may be combined with Figure 1 The system described can be executed in conjunction with Figure 33 The software architecture described. Figure 34 As shown in the example of , the application being processed includes a node that specifies an HLS kernel to be implemented in PL 214.
[0308] In block 3402, the DPE compiler 1602 receives the application, the SoC architecture description of the SoC 200, and an optional testbench. In block 3404, the DPE compiler 1602 can generate a connectivity diagram and provide the connectivity diagram to the system linker. In block 3406, the DPE compiler 1602 generates an initial SoC interface block plan and provides the initial SoC interface block plan to the hardware compiler 1606. The initial SoC interface block plan can specify the initial mapping of the application's DPE nodes to the DPEs 204 of the DPE array 202 and the mapping of the physical data paths of the DPE array 202 to and from the SoC interface block 206.
[0309] In block 3408, the HLS compiler 3302 can perform HLS on the HLS kernel to generate a synthesizable IP core. For example, the DPE compiler 1602 provides the HLS kernel specified by the application node to the HLS compiler 3302. The HLS compiler 3302 generates an HLS IP for each HLS kernel received. The HLS compiler 3302 outputs the HLS IP to the system linker.
[0310] In block 3410, the system linker can automatically generate a block diagram corresponding to the hardware portion of the application by using the connection diagram, the SoC architecture description, and the HLS IP. In block 3412, the system linker can integrate the block diagram with the base platform for the SoC 200. For example, the hardware compiler 1606 can connect the block diagram to the base platform to produce an integrated block diagram. In one aspect, the block diagram and the connected base platform are referred to as a synthesizable block diagram.
[0311] In block 3414, the hardware compiler 1606 can execute an implementation flow on the integrated block diagram. During the implementation flow, the hardware compiler 1606 can cooperate with the DPE compiler 1602 and the optional NoC compiler 1604 to perform verification as described herein to converge on an implementation of the hardware portion of the application for implementation in the PL. For example, as discussed, the hardware compiler 1606 can invoke the DPE compiler 1602 and, optionally, the NoC compiler 1604 in response to determining that the current implementation state of the hardware portion of the application does not meet one or more design criteria. The hardware compiler 1606 can invoke the DPE compiler 1602 and, optionally, the NoC compiler 1604 before placement, during placement, before routing, and / or during routing.
[0312] In block 3416, the hardware compiler 1606 outputs the hardware implementation to the DPE compiler 1602. In one aspect, the hardware implementation can be output as a device support archive (DSA) file. The DSA file can include platform metadata, simulation data, one or more configuration bitstreams generated by the hardware compiler 1606 from the implementation flow, and the like. The hardware implementation can also include a final SoC interface block solution and an optional final NoC solution, which the hardware compiler 1606 uses to create an implementation of the hardware portion of the application.
[0313] In block 3418, the DPE compiler 1602 completes software generation of the DPE array. For example, the DPE compiler 1602 generates binary files for programming the DPEs used in the application. When generating the binary files, the DPE compiler 1602 can perform an implementation flow using the final SoC interface block solution and, optionally, the final NoC solution used by the hardware compiler 1606. In one aspect, the DPE compiler can determine the SoC interface block solution used by the hardware compiler by examining the configuration bitstream and / or metadata included in the DSA.
[0314] In block 3420, the NoC compiler 1604 generates one or more binary files for programming the NoC 208. In block 3422, the PS compiler 1918 generates the PS binary files. In block 3424, the system can deploy the configuration bitstream and binary files in the SoC 200.
[0315] Figure 35 Another example method 3500 for executing a design flow to implement an application in SoC 200 is shown. The method 3500 may be performed by combining Figure 1 The described system performs an application. An application may be specified as a dataflow graph as described herein and includes a software portion for implementation within DPE array 202 and a hardware portion for implementation within PL 214.
[0316] In block 3502, the system can generate a first interface plan that maps logic resources used by the software portion to hardware resources of an interface block coupling the DPE array 202 and the PL 214. For example, the DPE compiler 1602 can generate an initial or first SoC interface block plan.
[0317] In block 3504, the system can generate a connectivity graph that specifies connectivity between the HLS kernel and the nodes of the software portion to be implemented in the DPE array. In one aspect, the DPE compiler 1602 can generate the connectivity graph.
[0318] In block 3506, the system can generate a block diagram based on the connectivity graph and the HLS kernel. The block diagram is synthesizable. For example, the system linker can generate a synthesizable block diagram.
[0319] At block 3508, the system can execute an implementation flow on the block diagram using the first interface solution. As discussed, the hardware compiler 1606 can exchange design data with the DPE compiler 1602 and, optionally, the NoC compiler 1604 during the implementation flow. The hardware compiler 1606 and the DPE compiler 1602 can iteratively exchange data, with the DPE compiler 1602 providing updated SoC interface block solutions to the hardware compiler 1606 in response to being invoked by the hardware compiler 1606. The hardware compiler 1606 can invoke the DPE compiler by providing one or more constraints for its SoC interface blocks. The hardware compiler 1606 and the NoC compiler 1604 can iteratively exchange data, with the NoC compiler 1604 providing updated NoC solutions to the hardware compiler 1606 in response to being invoked by the hardware compiler 1606. The hardware compiler 1606 can invoke the NoC compiler 1604 by providing one or more constraints to the NoC 208.
[0320] In block 3510, the system can compile the software portion of the application using the DPE compiler 1602 for implementation in one or more DPEs 204 of the DPE array 202. The DPE compiler 1602 can receive the results of the implementation flow so that a consistent interface is used between the DPE array 202 and the PL 214 (e.g., the same SoC interface block scheme used by the hardware compiler 1606 during the implementation flow).
[0321] For purposes of explanation, specific nomenclature is set forth to provide a thorough understanding of the various inventive concepts of the present disclosure. However, the terminology used herein is for the purpose of describing particular aspects of the present arrangements only and is not intended to be limiting.
[0322] As defined herein, the singular forms "a," "an" and "the" are intended to include the plural forms as well, unless the context clearly indicates otherwise.
[0323] As defined herein, unless expressly stated otherwise, the terms "at least one," "one or more," and "and / or" are open-ended expressions that are both conjunctive and disjunctive in operation. For example, the expressions "at least one of A, B, and C," "at least one of A, B, or C," "one or more of A, B, and C," "one or more of A, B, or C," and "A, B, and / or C" refer to A alone, B alone, C alone, A and B together, A and C together, B and C together, or A, B, and C together.
[0324] As defined herein, the term "automatically" means without user intervention. As defined herein, the term "user" refers to a human being.
[0325] As defined herein, the term "computer-readable storage medium" refers to a storage medium that contains or stores program code for use by or in connection with an instruction execution system, apparatus, or device. As defined herein, a "computer-readable storage medium" is not itself a transient, propagating signal. A computer-readable storage medium may be, but is not limited to, an electronic storage device, a magnetic storage device, an optical storage device, an electromagnetic storage device, a semiconductor storage device, or any suitable combination of the foregoing. As described herein, various forms of memory are examples of computer-readable storage media. A non-exhaustive list of more specific examples of computer-readable storage media may include: a portable computer disk, a hard disk, RAM, read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), electronically erasable programmable read-only memory (EEPROM), static random access memory (SRAM), a portable compact disk read-only memory (CD-ROM), a digital versatile disk (DVD), a floppy disk, and the like.
[0326] As defined herein, the term "if" means "when" or "on" or "in response to" or "in response to," depending on the context. Thus, the phrase "if it is determined" or "if [stated condition or event] is detected" may be interpreted to mean "when determining" or "in response to determining" or "when [stated condition or event] is detected," or "in response to detecting [stated condition or event]," or "in response to detecting [stated condition or event]," depending on the context.
[0327] As defined herein, the term "high-level language" or "HLL" refers to a programming language or instruction set used to program a data processing system, wherein the instructions have a strong abstraction from the details of the data processing system, such as machine language. For example, HLLs can automate or hide operational aspects of the data processing system, such as memory management. Although referred to as HLLs, these languages are generally classified as "efficiency-level languages." HLLs directly expose programming models supported by the hardware. Examples of HLLs include, but are not limited to, C, C++, and other suitable languages.
[0328] HLL can be contrasted with hardware description languages (HDLs) such as Verilog, SystemVerilog, and VHDL, which are used to describe digital circuits. HDLs allow designers to create a definition of a digital circuit design that can be compiled into a register transfer level (RTL) netlist that is generally technology-independent.
[0329] As defined herein, the term "in response to" and similar language as described above, such as "if," "when," or "upon," refers to the tendency to respond or react to an action or event. The response or reaction is automatically performed. Thus, if a second action is performed "in response to" a first action, a causal relationship exists between the occurrence of the first action and the occurrence of the second action. The term "in response to" indicates a causal relationship.
[0330] As defined herein, the terms "one embodiment," "an embodiment," "one or more embodiments," "a specific embodiment," or similar language mean that a particular feature, structure, or characteristic described in connection with the embodiment is included in at least one embodiment described within the present disclosure. Thus, throughout this disclosure, appearances of the phrases "in one embodiment," "in an embodiment," "in one or more embodiments," "in a specific embodiment," and similar language may, but do not necessarily, all refer to the same embodiment. Throughout this disclosure, the terms "embodiment" and "arrangement" are used interchangeably.
[0331] As defined herein, the term "output" means stored in a physical memory element (eg, a device), written to a display or other peripheral output device, sent or transmitted to another system, exported, and the like.
[0332] As defined herein, the term "substantially" means that the stated feature, parameter or value need not be achieved precisely, but rather deviations or variations, including, for example, tolerances, measurement errors, measurement accuracy limitations and other factors known to those skilled in the art, may occur in amounts that do not negate the effect that the characteristic is intended to provide.
[0333] The terms "first," "second," etc. may be used herein to describe various elements. These elements should not be limited by these terms because, unless otherwise specified or the context clearly indicates otherwise, these terms are only used to distinguish one element from another.
[0334] The computer program product may include a computer-readable storage medium having computer-readable program instructions thereon, and the computer-readable program instructions are used to cause a processor to perform various aspects of the invention arrangement described herein. In this disclosure, the term "program code" and the term "computer-readable program instructions" are used interchangeably. The computer-readable program instructions described herein can be downloaded from the computer-readable storage medium to the corresponding computing / processing device, or downloaded to an external computer or external storage device via a network (e.g., the Internet, LAN, WAN and / or wireless network). The network may include copper transmission cables, optical transmission fibers, wireless transmissions, routers, firewalls, switches, gateway computers and / or edge devices including edge servers. The network adapter card or network interface in each computing / processing device receives the computer-readable program instructions from the network, and forwards the computer-readable program instructions to be stored in the computer-readable storage medium in the corresponding computing / processing device.
[0335] The computer-readable program instructions for performing the operation of the invention arrangement as described herein can be assembler instructions, instruction set architecture (ISA) instructions, machine instructions, machine-related instructions, microcode, firmware instructions or source code or the object code written in any combination of one or more programming languages, including object-oriented programming languages and / or procedural programming languages.The computer-readable program instructions can include state setting data.The computer-readable program instructions can be fully on the user's computer, partly on the user's computer, as an independent software package, partly on the user's computer and partly on a remote computer or fully on a remote computer or server.In the latter case, the remote computer can be connected to the user's computer through any type of network (including LAN or WAN), or can be connected (for example, using an internet service provider to be connected through the internet) with an external computer.In some cases, the electronic circuit comprising, for example, a programmable logic circuit, FPGA or PLA can be personalized electronic circuit by utilizing the state information of the computer-readable program instructions, execute the computer-readable program instructions, thereby perform the work of the invention arrangement described below.
[0336] The flow charts and / or block diagrams of reference methods, devices (systems) and computer program products are herein described to describe certain aspects of the present invention. It should be understood that each block of the flow charts and / or block diagrams, and the combination of the blocks in the flow charts and / or block diagrams, can be implemented by computer-readable program instructions, such as program codes.
[0337] These computer-readable program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, or other programmable data processing apparatus to produce a machine, such that the instructions are executed in the apparatus via the processor of the computer or other programmable data processing apparatus, creating an apparatus for implementing the functions / actions specified in the flowchart and / or block diagram blocks. These computer-readable program instructions can also be stored in a computer-readable storage medium that can direct the computer, programmable data processing apparatus, and / or other device to function in a particular manner, such that the computer-readable storage medium having the instructions stored therein comprises an article of manufacture that includes instructions for implementing various aspects of the operations specified in the flowchart and / or block diagram or blocks.
[0338] The computer-readable program instructions may also be loaded onto a computer, other programmable data processing apparatus, or other device to cause a series of operations to be performed on the computer, other programmable apparatus, or other device to produce a computer-implemented process, such that the instructions executed on the computer, other programmable apparatus, or other device implement the functions / actions specified in the flowchart and / or block diagram blocks.
[0339] The flowcharts and block diagrams in the figures illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various aspects of the present invention. In this regard, each block in the flowchart or block diagram may represent a module, segment, or portion of instructions, which includes one or more executable instructions for implementing the specified operation.
[0340] In some alternative embodiments, the operations indicated in the blocks may not occur in the order indicated in the figures. For example, two blocks shown in succession may be performed substantially simultaneously, or may sometimes be performed in reverse order, depending on the functions involved. In other examples, blocks may generally be performed in increasing numerical order, while in other examples, one or more blocks may be performed in a varying order, with the results stored and utilized in subsequent or non-immediately following other blocks. It should also be noted that each block of the block diagrams and / or flow chart illustrations, and the combination of blocks in the block diagrams and / or flow chart illustrations, may be implemented by a dedicated hardware-based system that performs a specified function or action, or implements a combination of hardware and computer instructions for a special purpose.
[0341] The corresponding structures, materials, acts, and equivalents of all means or step plus function elements found in the claims below are intended to include any structure, material, or act for performing the function in combination with other claimed elements as specifically claimed.
[0342] A method includes, for an application having a software portion for implementation within a DPE array of a device and a hardware portion for implementation within a programmable logic controller (PL) of the device, generating, using a processor, a logical architecture and a first interface scheme for the application, the first interface scheme specifying a mapping of logic resources to hardware interface circuit blocks between the DPE array and programmable logic. The method includes constructing a block diagram of the hardware portion based on the logical architecture and the first interface scheme, and executing an implementation process on the block diagram using the processor. The method includes compiling, using the processor, the software portion of the application for implementation in one or more DPEs of the DPE array.
[0343] In another aspect, constructing the block diagram includes adding at least one IP core to the block diagram for implementation within the programmable logic.
[0344] On the other hand, during the implementation flow, the hardware compiler builds the block diagram and executes the implementation flow by exchanging design data with the DPE compiler configured to compile the software portion.
[0345] In another aspect, the hardware compiler exchanges further design data with the NoC compiler.The hardware compiler receives a first NoC solution configured to implement routing through a NoC of the device, the NoC of the device coupling the DPE array to the PL of the device.
[0346] On the other hand, executing the implementation process is executing the implementation process based on the exchanged design data.
[0347] In another aspect, compiling the software portion is performed based on the implementation of the hardware portion of the application for implementation in the PL generated by the implementation flow.
[0348] In another aspect, in response to a hardware compiler configured to construct the block diagram and execute the implementation process determining that the implementation of the block diagram does not meet the design specifications of the hardware portion, constraints for the interface circuit block are provided to a DPE compiler configured to compile the software portion. The hardware compiler receives from the DPE compiler a second interface solution generated by the DPE compiler based on the constraints.
[0349] On the other hand, the implementation process is performed based on the second interface solution.
[0350] In another aspect, the hardware compiler provides constraints on the NoC to the NoC compiler in response to determining that the implementation of the block diagram does not meet the design criteria using the first NoC solution for the NoC. The hardware compiler receives from the NoC compiler a second NoC solution generated by the NoC compiler based on the constraints on the NoC.
[0351] A system includes a processor configured to initiate operations. The operations include, for an application specifying a software portion to be implemented within a DPE array of a device and a hardware portion to be implemented within a PL of the device, generating a logical architecture and a first interface scheme for the application, the first interface scheme specifying a mapping of logical resources to hardware interface circuit blocks between the DPE array and the PL. The operations include constructing a block diagram of the hardware portion based on the logical architecture and the first interface scheme, executing an implementation process on the block diagram, and compiling the software portion of the application for implementation in one or more DPEs of the DPE array.
[0352] In another aspect, constructing the block diagram includes adding at least one IP core to the block diagram for implementation within the PL.
[0353] In another aspect, the operations include executing, during an implementation flow, a hardware compiler that constructs a block diagram and performs the implementation flow by exchanging design data with a DPE compiler configured to compile a software portion.
[0354] In another aspect, the operations include the hardware compiler exchanging further design data with the NoC compiler, and the hardware compiler receiving a first NoC scheme configured to implement routing through a NoC of a device coupling the DPE array to a PL of the device.
[0355] On the other hand, the implementation process is performed based on the exchanged design data.
[0356] On the other hand, based on the hardware design of the hardware portion of the application, a compilation software portion is performed for implementation in the PL generated by the implementation flow.
[0357] In another aspect, the operation includes, in response to a hardware compiler configured to construct a block diagram and execute an implementation flow determining that an implementation of the block diagram does not satisfy design constraints of the hardware portion, providing constraints of the interface circuit block to a DPE compiler configured to compile the software portion. The hardware compiler receives from the DPE compiler a second interface solution generated by the DPE compiler based on the constraints.
[0358] On the other hand, the implementation process is performed based on the second interface solution.
[0359] In another aspect, the hardware compiler provides constraints for the NoC to the NoC compiler in response to determining that the implementation of the block diagram does not meet the design metrics using the first NoC solution for the NoC. The hardware compiler receives from the NoC compiler a second NoC solution generated by the NoC compiler based on the constraints of the NoC.
[0360] A method includes, for an application having a software portion for implementation in a device's DPE array and a hardware portion for implementation in a program logic program (PL) of the device, executing an implementation flow on the hardware portion using a processor executing a hardware compiler based on an interface block plan, wherein the interface block plan maps logic resources used by the software portion to hardware of an interface block coupling the DPE array to the program logic program (PL). The method includes, in response to design specifications not being met during the implementation flow, providing interface block constraints to the DPE compiler using the processor executing the hardware compiler. The method also includes, in response to receiving the interface block constraints, generating an updated interface block plan using the processor executing the DPE compiler and providing the updated interface block plan from the DPE compiler to the hardware compiler.
[0361] On the other hand, the interface block constraints map the logical resources used by the software portion to the physical resources of the interface block.
[0362] On the other hand, the hardware compiler continues the implementation flow using the updated interface block scheme.
[0363] In another aspect, in response to design constraints of the hardware portion not being satisfied, the hardware compiler iteratively provides interface block constraints to the DPE compiler.
[0364] In another aspect, the interface block constraints include hard constraints and soft constraints. In that case, the method includes the DPE compiler routing the software portion of the application using both the hard constraints and the soft constraints to generate an updated interface block solution.
[0365] In another aspect, the method includes, in response to failing to generate an updated interface block plan using both hard and soft constraints, routing the software portion of the application using only the hard constraints to generate an updated interface block plan.
[0366] In another aspect, the method includes, in response to failing to generate an updated mapping using only hard constraints, mapping the software portion using both hard and soft constraints and routing the software portion using only hard constraints to generate an updated interface block solution.
[0367] In another aspect, wherein the interface block solution and the updated interface block solution each have a score, the method includes comparing the scores, and in response to determining that the score of the interface block solution exceeds the score of the updated interface block solution, relaxing the interface block constraints, and submitting the relaxed interface block constraints to the DPE compiler to obtain a further updated interface block solution.
[0368] In another aspect, the interface block solution and the updated interface block solution each have a score. The method includes comparing the scores and, in response to determining that the score of the updated interface block solution exceeds the score of the interface block solution, performing the implementation process using the updated interface block solution.
[0369] A system includes a processor configured to initiate operations. The operations include, for an application having a software portion for implementation in a device's DPE array and a hardware portion for implementation in a program program (PL) of the device, executing an implementation flow on the hardware portion using a hardware compiler based on an interface block solution that maps logic resources used by the software portion to hardware interface blocks coupling the DPE array to the program program (PL). The operations include, in response to design specifications not being met during the implementation flow, providing, using the hardware compiler, interface block constraints to a DPE compiler. The operations also include, in response to receiving the interface block constraints, generating, using the DPE compiler, an updated interface block solution and providing the updated interface block solution from the DPE compiler to the hardware compiler.
[0370] On the other hand, the interface block constraints map the logical resources used by the software portion to the physical resources of the interface block.
[0371] On the other hand, the hardware compiler continues the implementation flow using the updated interface block scheme.
[0372] In another aspect, in response to design constraints of the hardware portion not being satisfied, the hardware compiler iteratively provides interface block constraints to the DPE compiler.
[0373] In another aspect, the interface block constraints include hard constraints and soft constraints. In that case, the processor is configured to initiate operations including the DPE compiler routing the software portion of the application using both the hard constraints and the soft constraints to generate an updated interface block solution.
[0374] In another aspect, the operations include, in response to failing to generate an updated mapping using both the hard and soft constraints, routing the software portion of the application using only the hard constraints to generate an updated interface block solution.
[0375] In another aspect, the operations include, in response to failing to generate an updated mapping using only hard constraints, mapping the software portion using both hard and soft constraints and routing the software portion using only hard constraints to generate an updated interface block solution.
[0376] In another aspect, the interface block plan and the updated interface block plan each have a score. The processor is configured to initiate operations including comparing the scores and, in response to determining that the score of the interface block plan exceeds the score of the updated interface block plan, relaxing the interface block constraints and submitting the relaxed interface block constraints to the DPE compiler to obtain a further updated interface block plan.
[0377] In another aspect, the interface block solution and the updated interface block solution each have a score. The processor is configured to initiate operations including comparing the scores and, in response to determining that the score of the updated interface block solution exceeds the score of the interface block solution, performing the implementation process using the updated interface block solution.
[0378] A method includes, for an application that specifies a software portion to be implemented within a DPE array of a device and a hardware portion having an HLS kernel to be implemented within a PL of the device, generating, using a processor, a first interface scheme that maps logical resources used by the software portion to hardware resources of an interface block coupling the DPE array and the PL. The method includes generating, using the processor, a connectivity graph that specifies connectivity between nodes of the software portion to be implemented in the DPE array and the HLS kernel, and generating, using the processor, a block diagram based on the connectivity graph and the HLS kernel, wherein the block diagram is synthesizable. The method also includes executing, using the processor, an implementation flow based on the block diagram based on the first interface scheme, and compiling, using the processor, the software portion of the application for implementation in one or more DPEs of the DPE array.
[0379] In another aspect, generating the block graph includes executing HLS on the HLS kernel to generate a synthesizable version of the HLS kernel and constructing the block graph using the synthesizable version of the HLS kernel.
[0380] On the other hand, the synthesizable version of the HLS kernel is specified as RTL blocks.
[0381] In another aspect, generating the block diagram is performed based on a description of the architecture of a SoC in which the application is to be implemented.
[0382] In another aspect, generating the block diagram includes connecting the block diagram with a base platform.
[0383] In another aspect, performing the implementation flow includes synthesizing a block diagram for implementation in the PL, and placing and routing the synthesized block diagram based on the first interface scheme.
[0384] In another aspect, the method includes executing a hardware compiler during an implementation flow that constructs a block diagram and performs the implementation flow by exchanging design data with a DPE compiler configured to compile a software portion.
[0385] In another aspect, the method includes the hardware compiler exchanging further design data with the NoC compiler, and the hardware compiler receiving a first NoC scheme configured to implement routing through a NoC of a device coupling the DPE array to a PL of the device.
[0386] In another aspect, the method includes providing constraints of the interface circuit block to a DPE compiler configured to compile the software portion in response to a hardware compiler configured to construct the block diagram and execute an implementation flow determining that the implementation of the block diagram does not meet design specifications for the hardware portion. The method also includes the hardware compiler receiving, from the DPE compiler, a second interface solution generated by the DPE compiler based on the constraints.
[0387] On the other hand, the implementation process is performed based on the second interface solution.
[0388] A system includes a processor configured to initiate operations. The operations include, for an application having a software portion designated for implementation within a DPE array of a device and a hardware portion for an HLS kernel implemented within a PL of the device, generating a first interface scheme that maps logical resources used by the software portion to hardware resources of an interface block coupling the DPE array and the PL. The operations include generating a connectivity graph that specifies connectivity between nodes of the software portion to be implemented in the DPE array and the HLS kernel, and generating a block diagram based on the connectivity graph and the HLS kernel, wherein the block diagram is synthesizable. The operations also include executing an implementation process on the block diagram based on the first interface scheme and compiling the software portion of the application for implementation in one or more DPEs of the DPE array.
[0389] In another aspect, generating the block graph includes executing HLS on the HLS kernel to generate a synthesizable version of the HLS kernel and constructing the block graph using the synthesizable version of the HLS kernel.
[0390] On the other hand, the synthesizable version of the HLS kernel is specified as RTL blocks.
[0391] In another aspect, generating the block diagram is performed based on a description of the architecture of a SoC in which the application is to be implemented.
[0392] In another aspect, generating the block diagram includes connecting the block diagram with a base platform.
[0393] In another aspect, performing the implementation flow includes synthesizing a block diagram for implementation in the PL, and placing and routing the synthesized block diagram based on the first interface scheme.
[0394] In another aspect, the operations include executing a hardware compiler during an implementation flow that constructs a block diagram and performs the implementation flow by exchanging design data with a DPE compiler configured to compile a software portion.
[0395] In another aspect, the operations include the hardware compiler exchanging further design data with the NoC compiler, and the hardware compiler receiving a first NoC scheme configured to implement routing through a NoC of a device coupling the DPE array to a PL of the device.
[0396] In another aspect, the operations include, in response to a hardware compiler configured to construct a block diagram and execute an implementation flow determining that an implementation of the block diagram does not meet design specifications for the hardware portion, providing constraints of the interface circuit block to a DPE compiler configured to compile the software portion. The method also includes the hardware compiler receiving, from the DPE compiler, a second interface solution generated by the DPE compiler based on the constraints.
[0397] On the other hand, the execution process is performed based on the second interface solution.
[0398] One or more computer program products are disclosed herein, which include a computer-readable storage medium having program code stored thereon. The program code can be executed by computer hardware to initiate various operations described in this disclosure.
[0399] The description of the inventive arrangements provided herein is for illustrative purposes and is not intended to be exhaustive or limited to the disclosed forms and examples. The terminology used herein is selected to explain the principles of the inventive arrangements, practical applications, or technical improvements to marketed technologies, and / or to enable one of ordinary skill in the art to understand the inventive arrangements disclosed herein. Modifications and variations will be apparent to one of ordinary skill in the art without departing from the scope and spirit of the described inventive arrangements. Therefore, reference should be made to the following claims, rather than to the foregoing disclosure, to indicate the scope of such features and embodiments.
[0400] Example 1 shows an example pattern derived from the logical architecture of an application.
[0401] Example 1
[0402]
[0403]
[0404]
[0405]
[0406]
[0407] Example 2 shows an example schema of a SoC interface block solution for an application to be implemented in the DPE array 202 .
[0408] Example 2
[0409]
[0410]
[0411] Example 3 shows an example pattern of a NoC solution for an application to be implemented in the NoC 208 .
[0412] Example 3
[0413]
[0414]
[0415]
[0416] Example 4 shows an example schema for specifying SoC interface block constraints and / or NoC constraints.
[0417] Example 4
[0418]
[0419]
[0420] Example 5 shows an example pattern for specifying NoC traffic.
[0421] Example 5
[0422]
[0423]
[0424]
Claims
1. A method for process convergence in a hardware-software design process for heterogeneous programmable devices, the method comprising: For an application having a software portion for implementation in a data processing engine (DPE) array of the heterogeneous programmable device and a hardware portion for implementation in programmable logic of the heterogeneous programmable device, executing an implementation flow on the hardware portion by using a processor that executes a hardware compiler based on an interface block scheme, wherein the interface block scheme maps logic resources used by the software portion to hardware of an interface block that couples the DPE array to the programmable logic; In response to not meeting design specifications in the implementation flow executed based on the interface block solution, providing interface block constraints to a DPE compiler by using a processor executing the hardware compiler; in response to receiving the interface block constraints, generating an updated interface block solution using the interface block constraints by using a processor executing the DPE compiler; and The updated interface block solution is provided from the DPE compiler to the hardware compiler.
2. The method according to claim 1, characterized in that In response to design constraints for the hardware portion not being satisfied, the hardware compiler iteratively provides interface block constraints to the DPE compiler.
3. The method according to claim 1, characterized in that The interface block constraints include hard constraints and soft constraints, and the method further includes: The DPE compiler routes the software portion of the application by using both the hard constraints and the soft constraints to generate the updated interface block solution.
4. The method according to claim 3, characterized in that The method further comprises: In response to a failure to generate the updated interface block solution by using both the hard constraints and the soft constraints, routing the software portion of the application by using only the hard constraints to generate the updated interface block solution.
5. The method according to claim 4, characterized in that The method further comprises: In response to failing to generate the updated mapping by using only the hard constraints, mapping the software portion by using both the hard constraints and the soft constraints, and routing the software portion by using only the hard constraints to generate the updated interface block solution.
6. The method according to claim 1, characterized in that The interface block solution and the updated interface block solution each have a score, and the method further comprises: comparing the scores; and In response to determining that the score of the interface block solution exceeds the score of the updated interface block solution, the interface block constraints are relaxed and the relaxed interface block constraints are submitted to the DPE compiler to obtain a further updated interface block solution.
7. The method according to claim 1, characterized in that The interface block solution and the updated interface block solution each have a score, and the method further comprises: comparing the scores; and In response to determining that the score of the updated interface block solution exceeds the score of the interface block solution, the implementation process is performed using the updated interface block solution.
8. A system for process convergence in a hardware-software design process for heterogeneous programmable devices, the system comprising: A processor configured to initiate operations comprising: For an application having a software portion for implementation in a data processing engine (DPE) array of the heterogeneous programmable device and a hardware portion for implementation in programmable logic of the heterogeneous programmable device, executing an implementation flow on the hardware portion by using a hardware compiler based on an interface block solution, wherein the interface block solution maps logic resources used by the software portion to hardware of an interface block coupling the DPE array to the programmable logic; In response to not meeting design specifications in the implementation flow executed based on the interface block solution, providing interface block constraints to a DPE compiler by using the hardware compiler; In response to receiving the interface block constraints, generating an updated interface block solution using the interface block constraints by using the DPE compiler; and The updated interface block solution is provided from the DPE compiler to the hardware compiler.
9. The system according to claim 8, characterized in that The hardware compiler continues the implementation flow by using the updated interface block solution.
10. The system according to claim 8, wherein: In response to design constraints of the hardware portion not being satisfied, the hardware compiler iteratively provides interface block constraints to the DPE compiler.
11. The system according to claim 8, wherein: The interface block constraints include hard constraints and soft constraints, wherein the processor is configured to initiate operations including: The DPE compiler routes the software portion of the application by using both the hard constraints and the soft constraints to generate the updated interface block solution.
12. The system according to claim 11, wherein: The processor is configured to initiate operations including: In response to failing to generate the updated mapping by using both the hard constraints and the soft constraints, routing the software portion of the application by using only the hard constraints to generate the updated interface block solution.
13. The system according to claim 12, wherein: The processor is configured to initiate operations including: In response to failing to generate the updated mapping by using only the hard constraints, mapping the software portion by using both the hard constraints and the soft constraints, and routing the software portion by using only the hard constraints to generate an updated interface block solution.
14. The system according to claim 8, wherein: The interface block scheme and the updated interface block scheme each have a score, wherein the processor is configured to initiate operations comprising: comparing the scores; and In response to determining that the score of the interface block solution exceeds the score of the updated interface block solution, the interface block constraints are relaxed and the relaxed interface block constraints are submitted to the DPE compiler to obtain a further updated interface block solution.
15. The system according to claim 8, wherein: The interface block scheme and the updated interface block scheme each have a score, wherein the processor is configured to initiate operations comprising: comparing said scores; and; In response to determining that the score of the updated interface block solution exceeds the score of the interface block solution, the implementation process is performed using the updated interface block solution.
Citation Information
Patent Citations
Software and hardware cooperation based cross-linking simulation test method for programmable logic device
CN105302950A