Host Endpoint Adaptive Computing Configurability

The computing system dynamically reallocates core complex chiplets between CPU and accelerator functions to optimize resource utilization and computational efficiency by addressing the inefficiencies in existing DPU power and space usage.

JP2025527690APending Publication Date: 2025-08-22XILINX INC
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
JP2025511590
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2022-08-22
Filing Date
2023-04-20
Publication Date
2025-08-22

AI Technical Summary

Technical Problem

Current acceleration devices, such as data processing units (DPUs), often have either excessive or insufficient computational power, leading to inefficiencies in power and space usage due to independent sizing of host and embedded processing resources.

Method used

A computing system that dynamically allocates core complex chiplets between CPU and accelerator functions using a composable agent, allowing flexible assignment of chiplets to form integrated IO devices based on workload demands.

Benefits of technology

This approach optimizes resource utilization by adapting hardware resources to specific application needs, reducing power waste and improving computational efficiency by dynamically reallocating chiplets between CPU and accelerator roles.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2025527690000001_ABST
    Figure 2025527690000001_ABST
Patent Text Reader

Abstract

Embodiments herein describe a processor system including an integrated adaptive accelerator. In one embodiment, the processor system includes multiple core complex chiplets, each including one or more processing cores for a host CPU. In addition, the processor system includes accelerator chiplets. The processor system can assign one or more of the core complex chiplets to the accelerator chiplet to form an IO device, with the remaining core complex chiplets forming a CPU for the host. In this way, rather than the accelerator and CPU having independent computer resources, the accelerator can be integrated into the host's processor system, allowing hardware resources to be divided between the CPU and the accelerator according to the needs of the particular application being executed by the host.
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] Examples of the present disclosure generally relate to processor systems that include processing chiplets that can be allocated to a host central processing unit (CPU) or accelerator chiplets to form integrated IO devices. [Background technology]

[0002] Current acceleration devices (e.g., input / output (IO) devices) such as data processing units (DPUs) include different components, such as an I / O gateway, a processor subsystem, a network-on-chip (NoC), a storage and data accelerator, a data processing engine, and programmable logic (PL). Currently, DPUs connect to a host's processor complex using a PCIe connection. The host processing power and the embedded processing power of the DPU are sized independently. For some workloads, the DPU may have much more computational power than is required to perform the task, thereby wasting power and space in the computing system. For other workloads, the DPU may not have enough computational power required to perform the task and becomes a bottleneck. Summary of the Invention

[0003] One embodiment describes a computing system including a substrate, a processor system within a host including a plurality of core complex chiplets each including at least one processor core, an accelerator chiplet, and a composable agent configured to assign at least one of the plurality of core complex chiplets to the accelerator chiplet to form an IO device, and the remaining plurality of core complex chiplets to form a central processing unit (CPU) for the host.

[0004] Another embodiment described herein is a processor system including a plurality of core complex chiplets, each including at least one processor core, an accelerator chiplet, and an interconnect connecting the plurality of core complex chiplets to each other and to the accelerator chiplet, The interconnect includes a composable agent configured to assign at least one of the plurality of core complex chiplets to the accelerator chiplet to form an IO device, and the remaining plurality of core complex chiplets to form a central processing unit (CPU) for a host.

[0005] Another embodiment described herein is a method that includes selecting and assigning at least one of a plurality of core complex chiplets to an IO device while assigning remaining core complex chiplets of the plurality of core complex chiplets to a CPU of a host; removing the selected core complex chiplet as a peer of the remaining core complex chiplets of the plurality of core complex chiplets; and adding the selected core complex chiplet to the IO device, wherein the plurality of core complex chiplets are located on the same substrate as accelerator chiplets also assigned to the IO device.

[0006] In a manner in which the above-recited features may be understood in detail, a more particular description briefly summarized above may be made by reference to exemplary implementations, some of which are illustrated in the accompanying drawings. It should be noted, however, that the accompanying drawings illustrate only typical example implementations and therefore should not be considered limiting of the scope thereof. [Brief explanation of the drawings]

[0007] [Figure 1] 1 illustrates a processor system including an integrated adaptive accelerator, according to one embodiment. [Figure 2] 1 is a flowchart for adding a processing core to an integrated accelerator, according to one embodiment. [Figure 3]FIG. 1 is a block diagram of a processor system, according to one embodiment. [Figure 4] 1 illustrates a processor system including an integrated adaptive accelerator, according to one embodiment. [Figure 5] 1 illustrates a processor system including an integrated adaptive accelerator, according to one embodiment. [Figure 6] 1 illustrates a processor system including multiple integrated adaptive accelerators, according to one embodiment. DETAILED DESCRIPTION OF THE INVENTION

[0008] For ease of understanding, wherever possible, identical reference numbers have been used to indicate identical elements common to the figures. It is contemplated that elements of one embodiment may be beneficially incorporated in other embodiments.

[0009] Various features are described below with reference to the drawings. It should be noted that the drawings may or may not be drawn to scale, and that elements of similar structure or function are represented by similar reference numerals throughout the drawings. It should be noted that the drawings are intended only to facilitate the description of features. They are not intended as an exhaustive description of the specification or as limitations on the scope of the claims. In addition, the illustrated example need not have all the aspects or advantages shown. An aspect or advantage described in connection with a particular embodiment is not necessarily limited to that embodiment and may be implemented in any other embodiment, even if not so illustrated or explicitly described.

[0010] Embodiments herein describe a processor system including an integrated adaptive accelerator. In one embodiment, the processor system includes multiple core complex chiplets, each including one or more processing cores for a host CPU. In addition, the processor system includes accelerator chiplets (e.g., a SmartNIC, a database accelerator, an artificial intelligence (AI) accelerator, a graphics processor unit (GPU), etc.). The processor system can assign one or more of the core complex chiplets to the accelerator chiplets to form an IO device, and the remaining core complex chiplets form a CPU for the host. For example, the number of core complex chiplets assigned to the accelerator can be determined depending on whether an application executed by the host requires more computer resources from the CPU or more computer resources from the accelerator. In accelerator-intensive applications, more core complex chiplets may be assigned to the accelerator and the IO device, while in CPU-intensive applications, fewer core complex chiplets (or none at all) are assigned to the accelerator.

[0011] In this way, rather than the accelerator and CPU having independent computational resources, the accelerator can be integrated into the host's processor system, so that hardware resources can be divided between the CPU and the accelerator according to the needs of the particular application being executed by the host.

[0012] 1 illustrates a processor system 100 including an integrated adaptive accelerator, according to one embodiment. In this example, the processor system 100 includes a substrate 105 on which multiple core complex chiplets 110, interconnects 120, and accelerator chiplets 115 are disposed. In one embodiment, the core complex chiplets 110, interconnects 120, and accelerator chiplets 115 are individual integrated circuits (ICs). Each of the core complex chiplets 110 and accelerator chiplets 115 may be connected to the interconnects 120 using traces within the substrate 105. That is, the interconnects 120 enable the core complex chiplets 110 to communicate with each other and with the accelerator chiplets 115. However, in other embodiments, the core complex chiplets 110 may have direct connections to each other.

[0013] The interconnect 120 allows core complex chiplets 110 to be assigned to other chiplets 110 to form CPUs 140 for hosts or to accelerator chiplets 115 to form IO devices 130. That is, each core complex chiplet 110 can be assigned as a peer with other chiplets 110 to form CPUs 140, or disconnected from other chiplets 110 and assigned to IO devices 130. When a core complex chiplet 110 is assigned to an IO device 130, the coherent connection between it and the other chiplets 110 is severed and it is no longer treated as a peer. Instead, the core complex chiplet 110 may no longer share the same memory space as the other chiplets 110 in CPU 140, or run the same operating system, or rely on different interrupt semantics. This is described in more detail in FIG. 3.

[0014] In one embodiment, core complex chiplet 110 is a replicated IC, e.g., having the same number of processing cores and other hardware circuits. However, accelerator complex chiplet 115 may be different from core complex chiplet 110. For example, accelerator complex chiplet 115 may include dedicated circuitry for performing accelerator tasks, such as a network interface (if accelerator complex chiplet 115 is a DPU or database accelerator), a dedicated data processing engine (e.g., for performing network or AI tasks), programmable logic for additional user configurability, a host interface for managing memory requests and interrupts using interconnect 120, etc. In one embodiment, accelerator complex chiplet 115 does not include an embedded processor (e.g., a processor core). That is, in systems where accelerator complex chiplet 115 and IO device 130 are not integrated into host processor system 100, these accelerators may include their own processor cores (e.g., general-purpose processing units) as well as dedicated accelerator engines. However, because processing system 100 can allocate one or more of core complex chiplets 110 to accelerator chiplet 115, accelerator chiplet 115 no longer requires its own processor system, thereby saving space and power within processor system 100. However, in other embodiments, it may be advantageous to still include processor cores within accelerator chiplet 115.

[0015] Interconnect 120 includes composable agent 125, which assigns core complex chiplets 110 to accelerator chiplets 115. In this embodiment, agent 125 assigns core complex chiplets 110A and 110C to accelerator chiplet 115 to form IO device 130. In contrast, composable agent 125 assigns core complex chiplets 110B and 110D-G to form CPU 140. In one embodiment, composable agent 125 makes this assignment when the computing system is booting.

[0016] In one embodiment, this allocation can be changed. For example, if the workload changes and IO device 130 no longer needs more computational resources to perform its acceleration tasks, composable agent 125 can remove core compute chiplet 110C from IO device 130 and reassign it to CPU 140. In this manner, each of core complex chiplets 110 can be assigned to either integrated IO device 130 or CPU 140.

[0017] Furthermore, it is not necessary to have a separate interconnect 120 IC as shown. In other embodiments, the functions performed by interconnect 120 may be performed using hardware on core complex chiplets 110 and accelerator chiplets 115. Substrate 105 may include a network of interconnects for selectively connecting core complex chiplets 110 and accelerator chiplets 115. For example, portions of interconnect 120 may be distributed across each core complex chiplet 110, with substrate connections 105 achieving multi-core complex connectivity to create CPU 140.

[0018] In another embodiment, accelerator chiplet 115 or the circuitry within a chiplet is part of interconnect 120. That is, accelerator chiplet 115 may be integrated into the same IC as interconnect 120, rather than having a separate IC as shown in Figure 1. In other words, interconnect 120 (which may include composable agent 125 and other circuitry shown in Figure 1) may be part of the same IC as accelerator chiplet 115.

[0019] FIG. 2 is a flowchart of a method 200 for adding processing cores to an integrated accelerator, according to one embodiment. In block 205, CPU and accelerator workloads are determined. This determination may be performed by a user application running on the host, a system administrator, or the processor system. For example, a composable agent may rely on historical data to determine CPU and accelerator workloads. Alternatively, the composable agent may determine what types of user applications are running or will be running on the host and use this to estimate the CPU and accelerator workloads when running these applications. Alternatively, a silicon manufacturer may create different sets of products, each with a different mix of processor cores as part of I / O device 130, and static configurations of composable agent 125, each targeting a different workload with different computational resources for CPU 140.

[0020] For example, a control plane heavy workload, such as an unconsolidated storage workload, may use significantly more processing than a data plane heavy workload, where the accelerator performs most of the calculations and the processor cores in the CPU are lightly loaded. By determining (or estimating) the workload on the CPU and accelerator, the composable agent can determine how much computing resource each requires.

[0021] In block 210, the composable agent selects at least one of the core complex chiplets to assign to the IO device. That is, the composable agent selects at least one core complex chiplet to add to the accelerator chiplet to form the IO device. The number of core complex chiplets to assign to the IO device may depend on the workload of the IO device. For heavier workloads, additional core complex chiplets may be assigned to the IO device. For lighter workloads, fewer core complex chiplets (or none) are assigned to the IO device.

[0022] In block 215, the composable agent removes the selected core complex chiplet as a peer of the remaining core complex chiplets. For example, the agent may deactivate connections between the selected core complex chiplet and the chiplets used to form the CPU. These connections may be cache coherent connections and inter-processor interrupt connections between the selected core complex chiplets. Furthermore, hardware within the selected core complex chiplet may be reconfigured to operate as a processor core in an IO device rather than a peer processor in a CPU. For example, the selected core complex chiplet may issue read and write memory requests to an IO memory management unit (MMU) in the interconnect. Furthermore, the selected core complex chiplet may use message signaling interrupts (MSIs) to signal interrupts to the host processing system (e.g., the interconnect) even if the selected core complex chiplet is physically part of the host processing system. The MSIs may be generated by the core complex, or the interconnect may translate the core complex interrupts to MSIs to the host processing system. However, MSIs are just one example of a suitable IO interrupt protocol that may be used.

[0023] In block 220, the composable agent adds the selected core complex chiplet to the IO device. For example, the interconnect may establish a communication path between the selected core complex chiplet and the accelerator chiplet such that the core complex chiplet functions like a processor subsystem integrated into the accelerator. Using the established communication path, the core complex chiplet can run an operating system that is independent from the operating system running on the CPU. Here, the operating system resources may include the accelerator chiplet that uses the established communication path between the selected core complex chiplet and the accelerator chiplet.

[0024] In block 225, method 400 determines whether the workload has changed. This may be the workload of a CPU or the workload on an IO device. For example, the workload on a CPU may have increased such that the composable agent determines to move a core complex chiplet previously assigned to an IO device to the CPU. Alternatively, the workload on an IO device may have increased such that the composable agent moves a core complex chiplet previously assigned to a CPU to the IO device.

[0025] In block 230, the composable agent adjusts the number of core complex chiplets allocated to the IO device. This may include adding more core complex chiplets to the IO device or removing core complex chiplets from the IO device. This adjustment of core complex chiplets may occur when the computing system is rebooted, or it may be possible to adjust the number of core complex chiplets allocated to the IO device without having to reboot the computing system.

[0026] 3 is a block diagram of a processor system 300, according to one embodiment. Similar to processor system 100, processor system 300 includes core complex chiplet 110, interconnect 120, and accelerator chiplet 115. In this example, core complex chiplet 110 has both a coherent interconnect and an IO interconnect to interconnect 120. In one embodiment, the coherent interconnect is used when core complex chiplet 110 is part of a CPU, while the IO interconnect is used when core complex chiplet 110 is part of an IO device.

[0027] Core complex chiplet 110 includes processing cores 305, coherent hardware 310, and IO hardware 315. Coherent hardware 310 includes circuitry or firmware used when core complex chiplet 110 is a peer within a CPU, while IO hardware 315 includes circuitry or firmware used when core complex chiplet 110 is part of an IO device. For example, when part of a CPU, coherent hardware 310 may perform memory read and write requests as a cache coherent peer to other core complex chiplets 110 that form the CPU. However, when part of the IO domain, IO hardware 315 submits requests to read and write data to IO MMU 320 in interconnect 120, similar to accelerator chiplets 115 and other IO devices. Furthermore, rather than maintaining page tables with other core complex chiplets 110 that form the CPU, core complex chiplet 110 may receive page table translations from IO MMU 320 to access CPU memory when part of an IO device. Compositable agent 125 can allocate at least one of memory ports 330 as IO domain memory. When part of the IO domain, core composite chiplet 110 can run an operating system independent of the operating system running on the CPU and maintain its own page tables mapped to IO domain memory. Additionally, IO hardware 315 can support IO device interrupt semantics (e.g., MSI) that are subordinate to the interrupts used by coherent hardware 310. The interrupt semantics can send interrupts to an interrupt manager 325 in interconnect 120, which handles interrupts received from chiplets in the IO device.

[0028] Interconnect 120 includes composable agent 125, IO MMU 320, interrupt manager 325, and memory port 330. As described above, composable agent 125 selects and assigns core complex chiplets 110 to CPUs or IO devices. Additionally, composable agent 125 can inform IO MMU 320, interrupt manager 325, and memory port 330 which one of core complex chiplets 110 is part of a CPU and which is part of an IO device. These circuits can then use appropriate techniques to transmit data and handle interrupts.

[0029] In one embodiment, interconnect 120 may include programmable logic for managing transitions of core complex chiplet 110 between a coherent interconnect (or path) and an IO interconnect (or path), transitions between homogeneous core interrupts and IO device semantics interrupts (e.g., MSI), and transitions between homogeneous core MMUs and IO device semantics IO translation lookaside buffering (IOTLB) performed by IO MMU 320. For example, the programmable logic may implement a bridge between core interrupt semantics used when core complex chiplet 110 is part of a CPU and IO message interrupt semantics used with core complex chiplet 110 when it is part of an IO device. The programmable logic may also include a bridge for translating between a core page table walk used when core complex chiplet 110 is part of a CPU and, for example, an IO PCIe address translation service (ATS) used when core complex chiplet 110 is part of an IO device. This bridge can translate core page table accesses into PCIe ATS messages, and the data structures received in the ATS responses are then modified to appear as native to-core Instruction Set Architecture (ISA) page tables.

[0030] Accelerator chiplet 115 includes acceleration circuitry 335, host interface 340, and network interface 345. Acceleration circuitry 335 can include an acceleration engine (e.g., a data processing engine, a cryptographic engine, a compression engine, an AI engine, etc.), programmable logic, or a combination thereof. Acceleration circuitry 335 can be any circuitry that performs acceleration tasks assigned by a host (e.g., a CPU). For example, core complex chiplet 110 forming a CPU (or software running on chiplet 110) can offload acceleration tasks to accelerator chiplet 115 using interconnect 120.

[0031] Host interface 340 allows accelerator chiplets 115 to communicate with interconnect 120 via the IO interconnect. Host interface 340 can use a PCIe or Universal Interconnect Express (UCIe) connection to communicate with interconnect 120. Host interface 340 may use a cache coherent protocol, such as Compute Express Link (CXL®) or Cache Coherent Interconnect for Accelerators (CCIX®), to communicate with interconnect 120 and core complex chiplets 110 assigned to CPUs. These protocols allow accelerators 120 and any core complex chiplets 110 assigned to IO devices to be cache coherent with core complex chiplets 110 that form CPUs, but are subordinate to those core complex chiplets 110.

[0032] Network interface 345 enables accelerator chiplet 115 to communicate with a network (e.g., a local area network (LAN) or a wide area network (WAN)). In this example, accelerator chiplet 115 may be a DPU or a database accelerator. However, in embodiments in which the accelerator is an AI or cryptographic accelerator that does not communicate with a network, network interface 345 may be omitted.

[0033] Figure 4 illustrates a processor system including an integrated adaptive accelerator, according to one embodiment. For simplicity, the processor system of Figure 4 illustrates only core complex chiplets 110 and accelerator chiplets 115, but does not illustrate interconnections that may also be present.

[0034] In this example, the processor system includes two accelerator chiplets 115A and 115B. In one embodiment, the accelerator chiplets 115 are the same type of accelerator (e.g., two DPU chiplets or two database accelerators). In one embodiment, the accelerator chiplets 115 may be the same integrated circuit. For example, the processor system may be designed to be used to execute applications that require specific accelerator tasks to be performed by the accelerator chiplets 115. Thus, the processor system may include two accelerator chiplets 115 instead of just one as shown in FIG. 1 .

[0035] As described above, one or more of core complex chiplets 110 can be assigned to accelerator chiplets 115 to form IO device 405. In this example, core complex chiplet 110A is assigned to accelerator chiplets 115A and 115B, and core complex chiplets 110B-D form CPU 410 for the host. However, in other embodiments, two or more of core complex chiplets 110 can be assigned to IO device 405, or none of core complex chiplets 110 are assigned to IO device 405, in which case accelerator tasks are performed only by accelerator chiplets 115A and 115B.

[0036] Figure 5 illustrates a processor system including an integrated adaptive accelerator, according to one embodiment. For simplicity, the processor system of Figure 5 illustrates only core complex chiplets 110 and accelerator chiplets 115, but does not illustrate interconnections that may also be present.

[0037] Similar to FIG. 4 , here the processor system includes two accelerator chiplets 115A and 115B. In one embodiment, the accelerator chiplets 115 are the same type of accelerator (e.g., two DPU chiplets or two database accelerators). In one embodiment, the accelerator chiplets 115 may be the same integrated circuit. However, unlike FIG. 4 , FIG. 5 illustrates that one of the accelerator chiplets may be assigned to be part of the CPU 510. Thus, FIG. 5 illustrates that in addition to assigning the core complex chiplet 110 to the IO device 505, the accelerator chiplet 115 may be assigned to the CPU 510. For example, the current workload of the IO device 505 may be much less than the workload on the CPU 510. Thus, the composable agent may reassign the accelerator chiplet 115B to the CPU 510 while leaving the accelerator chiplet 115A performing the functions of the IO device 505.

[0038] Accelerator chiplet 115B can function as a peer processor to other core complex chiplets 110 in the processor system. To do so, accelerator chiplet 115 may also include a coherent interconnect to core complex chiplet 110, as well as hardware to support the interrupt and memory management protocols used by core complex chiplet 110 when part of CPU 510.

[0039] 5, if the workload of IO device 505 and CPU 510 changes, the composable agent can reassign accelerator chiplet 115B to IO device 505. Thus, FIG. 5 shows that accelerator chiplet 115, like core complex chiplet 110, can be designed to support both coherent and IO modes of operation.

[0040] Figure 6 illustrates a processor system including multiple integrated adaptive accelerators, according to one embodiment. For simplicity, the processor system of Figure 6 illustrates only core complex chiplets 110 and accelerator chiplets 115, and does not illustrate interconnections that may also be present.

[0041] The processor system of Figure 6 includes accelerator chiplet 605 and accelerator chiplet 610, which are different accelerators (i.e., chiplets 605 and 610 are different). For example, accelerator chiplet 605 may be a different type of accelerator than accelerator chiplet 610. For example, accelerator chiplet 605 may be a DPU, while accelerator chiplet 610 is an AI accelerator. Thus, Figure 6 illustrates that a processor system can include multiple different types of accelerator chiplets.

[0042] Additionally, accelerator chiplets can be used to form different IO devices. In this example, accelerator chiplet 605 and core complex chiplet 110A form IO device 615, while accelerator chiplet 610 and core complex chiplet 110B form IO device 620. The remaining core complex chiplets 110 form CPU 625 for the host. The composable agent can assign core complex chiplets 110 according to the individual workloads of the IO devices. For example, if IO device 615 has (or is expected to have) a larger workload than IO device 620, the composable agent may assign two of the core complex chiplets 110 to IO device 620. Thus, FIG. 6 illustrates that a processor system can include any number of different types of accelerators to support any number of different types of IO devices. Furthermore, these IO devices can also include any number of core complex chiplets 110 according to their workloads.

[0043] In the foregoing, reference is made to the embodiments presented in this disclosure. However, the scope of the disclosure is not limited to the specific described embodiments. Instead, any combination of the described features and elements, whether associated with different embodiments or not, is contemplated for implementing and practicing the contemplated embodiments. Moreover, while the embodiments disclosed herein may achieve advantages over other possible solutions or prior art, whether or not a particular advantage is achieved by a given embodiment does not limit the scope of the disclosure. Accordingly, the foregoing aspects, features, embodiments, and advantages are merely illustrative and are not considered elements or limitations of the appended claims unless expressly recited in the claims.

[0044] As will be appreciated by one skilled in the art, embodiments disclosed herein may be embodied as a system, method, or computer program product. Accordingly, aspects may take the form of an entirely hardware embodiment, an entirely software embodiment (including firmware, resident software, microcode, etc.), or an embodiment combining software and hardware aspects, all of which may be generally referred to herein as a "circuit," "module," or "system." Furthermore, aspects may take the form of a computer program product embodied in one or more computer-readable medium(s) having computer-readable program code embodied therein.

[0045] Any combination of one or more computer-readable media may be utilized. The computer-readable medium may be a computer-readable signal medium or a computer-readable storage medium. The computer-readable storage medium may be, for example, but not limited to, an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any suitable combination of the foregoing. More specific examples (non-exhaustive list) of computer-readable storage media include an electrical connection having one or more wires, a portable computer floppy disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing. In the context of this specification, a computer-readable storage medium is any tangible medium that can contain or store a program for use by or in connection with an instruction execution system, apparatus, or device.

[0046] A computer-readable signal medium may include a propagated data signal in which computer-readable program code is embodied, for example, in baseband or as part of a carrier wave. Such a propagated signal may take any of a variety of forms, including, but not limited to, electromagnetic, optical, or any suitable combination thereof. A computer-readable signal medium is not a computer-readable storage medium but may be any computer-readable medium that can communicate, propagate, or transport a program for use by or in connection with an instruction execution system, apparatus, or device.

[0047] The program code embodied on the computer readable medium may be transmitted using any suitable medium, including but not limited to wireless, wireline, fiber optic cable, RF, etc., or any suitable combination of the foregoing.

[0048] Computer program code for carrying out operations of aspects of the present disclosure may be written in any combination of one or more programming languages, including, for example, object-oriented programming languages ​​such as Java, Smalltalk, C++, and conventional procedural programming languages ​​such as the "C" programming language or similar programming languages. The program code may execute entirely on the user's computer, partially on the user's computer as a standalone software package, partially on the user's computer, partially on a remote computer, or entirely on a remote computer or server. In the latter scenario, the remote computer may be connected to the user's computer via any type of network, including a local area network (LAN) or a wide area network (WAN), or may be connected to an external computer (e.g., via the Internet using an Internet Service Provider).

[0049] Aspects of the present disclosure are described below with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments presented in the present disclosure. It will be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, or other programmable data processing apparatus, such that the instructions, executed by the processor of the computer or other programmable data processing apparatus, create means for implementing the functions / acts specified in the flowchart and / or block diagram blocks.

[0050] These computer program instructions may also be stored on a computer-readable storage medium, and the instructions may direct a computer, programmable data processing apparatus, and / or other device to function in a particular manner to produce an article of manufacture including instructions that implement the functions / acts specified in the flowchart and / or block diagram blocks.

[0051] Computer program instructions may also be loaded into a computer, other programmable data processing apparatus, or other device to cause a series of operational steps to be performed on the computer, other programmable apparatus, or other device to create a computer-implemented process, such that the instructions executing on the computer or other programmable apparatus provide a process for implementing the functions / acts specified in the flowchart and / or block diagram blocks.

[0052] The flowcharts and block diagrams in the figures illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of the present invention. In this regard, each block in the flowcharts or block diagrams may represent a module, segment, or portion of instructions, including one or more executable instructions for implementing a specified logical function. In some alternative implementations, the functions noted in the blocks may occur out of the order noted in the figures. For example, two blocks shown in succession may in fact be executed substantially concurrently, or the blocks may be executed in the reverse order, depending on the functionality involved. It should also be noted that each block of the block diagrams and / or flowchart illustrations, and combinations of blocks in the block diagrams and / or flowchart illustrations, may be implemented by a dedicated hardware-based system that performs the specified functions or acts or a combination of dedicated hardware and computer instructions.

[0053] The technology disclosed herein can be illustrated by the following non-limiting examples.

[0054] Example 1. A processor system within a host, the processor system comprising: a substrate; a plurality of core complex chiplets each including at least one processor core; an accelerator chiplet; and a composable agent configured to assign at least one of the plurality of core complex chiplets to the accelerator chiplet to form an IO device, and the remaining plurality of core complex chiplets to form a central processing unit (CPU) for the host.

[0055] Example 2. The processor system of example 1, wherein the multiple core composite chiplets are replicated integrated circuits and the accelerator chiplets are implemented using integrated circuits that are different from the replicated integrated circuits.

[0056] Example 3. The processor system of example 2, wherein the accelerator chiplet does not include a processor core.

[0057] Example 4. The processor system of example 3, wherein the accelerator chiplet is combined with at least one of the plurality of core complex chiplets to form a data processing unit (DPU).

[0058] Example 5. The processor system of example 1, further comprising: an interconnect disposed on the substrate and implemented using an integrated circuit separate from the plurality of core complex chiplets and the accelerator chiplet, the interconnect configured to enable each of the plurality of core complex chiplets to communicate with each other and to enable the plurality of core complex chiplets to communicate with the accelerator chiplet.

[0059] Example 6. The processor system of example 5, wherein the interconnect includes: a first circuit that supports interrupt semantics used by the plurality of core complex chiplets forming the CPU and interrupt semantics used by at least one of the plurality of core complex chiplets assigned to the IO device; and a second circuit that supports memory accesses used by the plurality of core complex chiplets forming the CPU and memory accesses used by at least one of the plurality of core complex chiplets assigned to the IO device.

[0060] Example 7. The processor system of example 1, further comprising an interconnect configured to enable each of the plurality of core complex chiplets to communicate with each other and to enable the plurality of core complex chiplets to communicate with an accelerator chiplet, wherein the interconnect and the accelerator chiplet are part of the same integrated circuit.

[0061] Example 8. The processor system of example 1, wherein after assigning at least one of the plurality of core complex chiplets to an accelerator chiplet, the composable agent is configured to reallocate at least one of the plurality of core complex chiplets to a CPU such that at least one of the plurality of core complex chiplets is no longer part of an IO device.

[0062] Example 9. The processor system of example 1, wherein the composable agent is configured to assign additional core complex chiplets of the plurality of core complex chiplets to the IO device such that the IO device includes more than one core complex chiplet of the plurality of core complex chiplets.

[0063] Example 10. A processor system comprising: a plurality of core complex chiplets, each including at least one processor core; an accelerator chiplet; and an interconnect connecting the plurality of core complex chiplets to each other and to the accelerator chiplet, the interconnect including a composable agent configured to assign at least one of the plurality of core complex chiplets to the accelerator chiplet to form an IO device, while the remaining plurality of core complex chiplets form a central processing unit (CPU) for a host.

[0064] Example 11. The processor system of example 10, wherein the multiple core complex chiplet is a replicated integrated circuit, the accelerator chiplet is implemented using an integrated circuit that is different from the replicated integrated circuit, and the interconnect is implemented using an integrated circuit that is separate from the multiple core complex chiplet.

[0065] Example 12. The processor system of example 10, wherein the accelerator chiplet does not include a processor core.

[0066] Example 13. The processor system of example 10, wherein the interconnect includes: a first circuit that supports interrupt semantics used by the multiple core complex chiplets forming the CPU and interrupt semantics used by at least one of the multiple core complex chiplets assigned to the IO device; and a second circuit that supports memory accesses used by the multiple core complex chiplets forming the CPU and memory accesses used by the multiple core complex chiplets assigned to the IO device.

[0067] Example 14. A method comprising: selecting and assigning at least one of a plurality of core complex chiplets to an IO device while assigning remaining core complex chiplets of the plurality of core complex chiplets to a CPU of a host; removing the selected core complex chiplet as a peer of the remaining core complex chiplets of the plurality of core complex chiplets; and adding the selected core complex chiplet to the IO device, wherein the plurality of core complex chiplets are located on the same substrate as accelerator chiplets that are also assigned to the IO device.

[0068] Example 15. The method of example 14, wherein the multiple core complex chiplet is a replicated integrated circuit and the accelerator chiplet is implemented using an integrated circuit that is different from the replicated integrated circuit.

[0069] Example 16. The method of example 14, wherein each of the plurality of core complex chiplets includes at least one processor core, and the accelerator chiplet does not include a processor core.

[0070] Example 17. The method of example 14, wherein the accelerator chip is combined with at least one of the multiple core complex chiplets to form a DPU.

[0071] Example 18. The method of example 14, wherein the interconnect is disposed on the substrate and is implemented using an integrated circuit separate from the plurality of core composite chiplets and the accelerator chiplet.

[0072] Example 19. The method of Example 14, further comprising determining a workload of the CPU and the IO device, wherein the workload determines the number of multiple core complex chiplets to allocate to the IO device and the CPU.

[0073] Example 20. The method of example 19, further comprising: determining that the workload of at least one of the CPU and the IO device has changed; and adjusting the number of multiple core complex chiplets assigned to the IO device.

[0074] While the above is directed to particular examples, other and further examples may be devised without departing from the basic scope thereof, which scope is determined by the following claims.

Claims

1. A processor system in a host, comprising: A substrate; a plurality of multi-core chiplets each including at least one processor core; an accelerator chiplet; and a composable agent configured to assign at least one of the plurality of core complex chiplets to the accelerator chiplet to form an IO device, while the remaining plurality of core complex chiplets form a central processing unit (CPU) for the host.

2. The processor system of claim 1 , wherein the plurality of core complex chiplets are replicated integrated circuits, and the accelerator chiplet is implemented using an integrated circuit different from the replicated integrated circuit.

3. The processor system of claim 2 , wherein the accelerator chiplet does not include a processor core.

4. 10. The processor system of claim 1, further comprising: an interconnect disposed on the substrate and implemented using an integrated circuit separate from the plurality of core complex chiplets and the accelerator chiplet, the interconnect configured to enable each of the plurality of core complex chiplets to communicate with each other and the plurality of core complex chiplets to communicate with the accelerator chiplet.

5. 10. The processor system of claim 1, further comprising an interconnect configured to enable each of the plurality of core complex chiplets to communicate with each other and to enable the plurality of core complex chiplets to communicate with the accelerator chiplet, wherein the interconnect and the accelerator chiplet are part of the same integrated circuit.

6. 2. The processor system of claim 1, wherein after assigning the at least one of the plurality of core complex chiplets to the accelerator chiplet, the composable agent is configured to reassign the at least one of the plurality of core complex chiplets to the CPU such that the at least one of the plurality of core complex chiplets is no longer part of the IO device.

7. 2. The processor system of claim 1, wherein the composable agent is configured to assign an additional one of the plurality of core complex chiplets to the IO device such that the IO device includes multiple core complex chiplets of the plurality of core complex chiplets.

8. 1. A processor system, comprising: a plurality of multi-core chiplets each including at least one processor core; an accelerator chiplet; and an interconnect connecting the plurality of core complex chiplets to each other and to the accelerator chiplet, the interconnect comprising: an interconnect including a composable agent configured to assign at least one of the plurality of core complex chiplets to the accelerator chiplet to form an IO device, while the remaining plurality of core complex chiplets form a central processing unit (CPU) for a host.

9. 9. The processor system of claim 8, wherein the multiple core complex chiplets are replicated integrated circuits, the accelerator chiplet is implemented using an integrated circuit different from the replicated integrated circuit, and the interconnect is implemented using an integrated circuit separate from the multiple core complex chiplets.

10. The processor system of claim 8 , wherein the accelerator chiplet does not include a processor core.

11. 9. The processor system of claim 8, wherein the interconnect includes: a first circuit that supports interrupt semantics used by the plurality of core complex chiplets forming the CPU and interrupt semantics used by the at least one of the plurality of core complex chiplets assigned to the IO device; and a second circuit that supports memory accesses used by the plurality of core complex chiplets forming the CPU and memory accesses used by the plurality of core complex chiplets assigned to the IO device.

12. 1. A method comprising: Selecting and allocating at least one of the plurality of core complex chiplets to an IO device, while allocating the remaining core complex chiplets to a CPU of a host; removing the selected core complex chiplet as a peer of the remaining core complex chiplets of the plurality of core complex chiplets; adding the selected core complex chiplet to the IO device, wherein the plurality of core complex chiplets are located on the same substrate as an accelerator chiplet also assigned to the IO device.

13. the plurality of core complex chiplets are replicated integrated circuits, and the accelerator chiplet is implemented using an integrated circuit different from the replicated integrated circuit; each of the plurality of core complex chiplets includes at least one processor core, and the accelerator chiplet does not include a processor core; the accelerator chip is combined with the at least one of the plurality of core complex chiplets to form a DPU; The method of claim 12 , wherein the interconnect is implemented using an integrated circuit disposed on the substrate and separate from the plurality of core complex chiplets and the accelerator chiplet.

14. 13. The method of claim 12, further comprising determining a workload of the CPU and the IO device, the workload determining a number of the multiple core complex chiplets to allocate to the IO device and the CPU.

15. determining that the workload of at least one of the CPU and the IO device has changed; 15. The method of claim 14, further comprising: adjusting a number of the multiple core complex chiplets assigned to the IO device.