Job Target Aliasing in Unintegrated Computer Systems

By dynamically composing computing units from a pool of physical components based on job requirements, the method addresses inefficiencies in resource allocation in large-scale computing clusters, enhancing utilization and reducing space and cost requirements.

JP7706096B2Active Publication Date: 2025-07-11リキッド インコーポレイテッド
View PDF 8 Cites 0 Cited by

Patent Information

Application Number
JP2023553070
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Priority Date
2021-03-03
Filing Date
2022-02-18
Publication Date
2025-07-11
Estimated Expiration
2042-02-18

AI Technical Summary

Technical Problem

Existing workload managers in large-scale computing clusters face challenges in managing resource demand due to fixed server configurations, leading to inefficiencies in space and cost requirements as jobs are deployed across multiple servers regardless of actual resource needs.

Method used

A method and apparatus that dynamically compose computing units from a pool of physical components via a communication fabric, determining resource requirements for each job and forming composite machines to efficiently allocate resources, allowing for flexible and efficient use of computing resources.

Benefits of technology

This approach enhances resource utilization by dynamically allocating computing resources based on job needs, improving efficiency and reducing the need for excess servers, thereby optimizing space and cost in data centers.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007706096000001
    Figure 0007706096000001
  • Figure 0007706096000002
    Figure 0007706096000002
  • Figure 0007706096000003
    Figure 0007706096000003
Patent Text Reader

Abstract

Presented herein is a deployment of an arrangement of physical computing components coupled via a communications fabric. In one example, the method includes presenting a target machine to a workload manager that can receive an execution job from the workload manager. The target machine has a network state and includes a selection of computing components. The method also includes receiving a job issued by the workload manager that is directed to the target machine. Based on characteristics of the job, the method includes determining resource requirements for processing the job, forming a composite machine with physical computing components that support the resource requirements of the job, transferring the network state of the target machine to the composite machine to indicate the network state of the composite machine to the workload manager, and commencing execution of the job on the composite machine.
Need to check novelty before this filing date? Find Prior Art

Description

Background Art

[0001] Clustered computing systems are becoming increasingly popular because of the growing demand for data storage, data processing, and communication processing. Data centers generally include large rack-mounted and network-connected data storage and data processing systems. These data centers can receive data for storage from external users via network links and can also receive data generated from applications executed by processing elements within the data center. Often, data centers and associated computing equipment are used to execute jobs for multiple simultaneous users or applications. Jobs can use a central processing unit (CPU) or a graphics processing unit (GPU) to process data and to route data related to these resources between a temporary storage device and a long-term storage device or between various network locations using the resources of the data center. GPU-based processing is becoming increasingly popular for use in artificial intelligence (AI) and machine learning regimes. In these regimes, a computing system such as a blade server can include one or more GPUs along with an associated CPU for processing large data sets.

[0002] A workload manager has been developed that can receive and deploy computing jobs for execution by a server, such as in a large-scale cloud system and computing cluster. Exemplary workload managers include the Slurm workload manager, OpenStack, Kubernetes, and other common workload and cloud orchestration / deployment services. These workload managers generally have a list of servers that can be selected for job processing. When a server is selected for a job, the job can be deployed for execution or other types of processing by the selected server. However, it can be difficult to manage the demand for these workload managers across a large-scale computing cluster with servers whose configurations can change over time.

[0003] In a large-scale computing cluster, density limitations can occur. Specifically, each server generally includes a fixed configuration among a CPU, a GPU, and storage elements housed in a common enclosure or chassis. When an incoming job is deployed within the data center, the granularity of computing resources is limited to individual servers. Thus, a deployed job generally occupies one or more servers along with all of the corresponding CPU, GPU, and storage elements of each server, regardless of whether the entire resources of the server are actually needed to execute the job. To compensate, data center operators generally deploy an increasingly large amount of servers to handle the increasing traffic from jobs. This strategy can face barriers to the physical space required for rack-mounted servers, as well as large space and cost requirements. SUMMARY OF THE INVENTION

[0004] This specification presents the deployment of an arrangement of physical computing components coupled via a communication fabric. In one example, the method includes presenting to a workload manager a target machine that can receive execution jobs from the workload manager. The target machine has a network state and includes a selection of computing components. The method also includes receiving a job issued by the workload manager directed to the target machine. Based on the characteristics of the job, the method includes determining resource requirements for processing the job, forming a composite machine comprising physical computing components that support the resource requirements of the job, transferring the network state of the target machine to the composite machine to indicate the network state of the composite machine to the workload manager, and initiating execution of the job on the composite machine.

[0005] In other examples, the apparatus includes program instructions stored on one or more computer-readable storage media that, based on being executed by a processing system, direct the processing system to present to the workload manager a target machine that can receive execution jobs, where the target machine has a network state and includes a selection of computing components. The program instructions direct the processing system to receive a job issued by the workload manager and directed to the target machine. Based on the characteristics of the job, the program instructions direct the processing system to determine resource requirements for processing the job and to form a composite machine comprising physical computing components that support the resource requirements of the job. The program instructions direct the processing system to transfer the network state of the target machine to the composite machine, indicate the network state of the composite machine to the workload manager, and initiate execution of the job on the composite machine.

[0006] Other examples include a system comprising a job interface configured to present computing targets as virtual targets for job execution, where each computing target has an associated set of advertised computing components and a corresponding network addressing. The job interface receives a job for execution directed to a selected computing target having a corresponding network address. The system also includes a controller configured to select a set of physical computing components necessary to support execution of the job from a pool of physical computing components, synthesize a physical computing node comprising the set of physical computing components, configure the physical computing node to communicate via the corresponding network address in place of the selected computing target, and deploy the job to the physical computing node for processing.

[0007] This summary is provided to introduce a selection of concepts in a simplified form that are further described below in the disclosure. This summary is not intended to identify key features or essential features of the claimed subject matter, nor is it intended to be used to limit the scope of the claimed subject matter.

Brief Description of the Drawings

[0008] Many aspects of the present disclosure can be better understood in relation to the following drawings. The components of the drawings are not necessarily to scale, and instead emphasis is placed on clearly illustrating the principles of the present disclosure. Further, in the drawings, like reference numerals designate corresponding parts throughout several views. Although several embodiments are described in relation to these drawings, the present disclosure is not limited to the embodiments disclosed herein. Rather, the intention is to cover all alternative forms, modifications, and equivalents.

[0009]

Figure 1

[0010]

Figure 2

[0011]

Figure 3

[0012]

Figure 4

[0013]

Figure 5

[0014]

Figure 6

Modes for Carrying Out the Invention

[0015] Using a data center with associated computing devices, it is possible to process the execution jobs of multiple simultaneous users or simultaneous data applications. Jobs can utilize the resources of the data center to process data and transfer data related to these resources between temporary storage and long-term storage, or between various network destinations. Data center processing resources can include a central processing unit (CPU) along with various types of coprocessing units (CoPUs) such as graphics processing units (GPUs), tensor processing units (TPUs), field programmable gate arrays (FPGAs), or application-specific integrated circuits (ASICs). Processing of the coprocessing unit type is becoming increasingly popular for use in artificial intelligence (AI) and machine learning systems. In the examples herein, the limitations of blade server-based data systems can be overcome using a distributed computing system that can dynamically synthesize groups of computing on the fly according to the requirements of each incoming execution job. Groupings herein referred to as computing units, computing nodes, or bare metal machines can include resources that meet the needs of various execution jobs and are adjusted to such jobs. Instead of having a fixed arrangement among the CPUs, CoPUs, and memory elements housed in a common enclosure or chassis, the examples herein can flexibly include any number of CPUs, CoPUs, and memory elements formed logically on a communication fabric across any number of enclosures / chassis. Computing units can be further grouped into sets or clusters of many computing units / machines to achieve greater parallelism and throughput. Thus, the data system can better utilize resources by not having idle or wasted portions of blade servers that are not required for a particular job or a particular part of a job. The operator of the data center can achieve a very high utilization level of the data center, higher than can be achieved using fixed-configuration servers.

[0016] This specification presents the deployment of an arrangement of physical computing components coupled via a communication fabric. An execution job directed to a computing cluster is received. The cluster includes at least one “machine” or computing unit, and the computing unit includes at least one processor element (e.g., a CPU). The computing unit may also include a CoPU (such as a GPU), a network interface element (e.g., a NIC), or a data storage element (e.g., an SSD), although these elements are not required for the computing unit. The computing unit or cluster is formed from a pool of computing components coupled via one or more communication fabrics. Based on the characteristics of the execution job, the control system can determine the resources required for the job and the resource scheduling for processing the execution job. When a job is scheduled to be executed, the control system facilitates the composition of a computing unit for processing the execution job. The computing unit is composed from among the computing components that form the pool of computing components. Logical partitions are established within the communication fabric to form the computing units and separate each computing unit from one another. In response to completion of the execution job, the computing unit is broken down and returned to the pool of computing components.

[0017] This specification describes various individual physical computing components coupled via one or more shared communication fabrics. Various communication fabric types can be used in this specification. For example, a Peripheral Component Interconnect Express (PCIe) fabric can be used, including various versions such as, among others, 3.0, 4.0, or 5.0. Instead of a PCIe fabric, other point-to-point communication fabrics or communication buses with associated physical layers, electrical signaling, protocols, and layered communication stacks can be used. These can include, among others, Gen-Z, Ethernet, InfiniBand, NVMe, Internet Protocol (IP), Serial Attached SCSI (SAS), Fibre Channel, Thunderbolt, Serial ATA Express (SATAExpress), NVLink, Cache Coherent Interconnect for Accelerators (CCIX), Compute Express Link (CXL), Open Coherent Accelerator Processor Interface (OpenCAPI), Wireless Ethernet or Wi-Fi (802.11x), or cellular wireless technology. Ethernet can refer to any of the various network communication protocol standards and bandwidths available, such as 10BASE-T, 100BASE-TX, 1000BASE-T, 10GBASE-T (10 Gigabit Ethernet), 40GBASE-T (40 Gigabit Ethernet), Gigabit (GbE), Terabit (TbE), 200GbE, 400GbE, 800GbE, or other various wired and wireless Ethernet formats and speeds. Cellular wireless technology can include various wireless protocols and networks built around the 3rd Generation Partnership Project (3GPP (registered trademark)) standards, including, among others, 4G Long Term Evolution (LTE), 5GNR (New Radio), and related 5G standards.

[0018] Some of the foregoing signaling or protocol types are built on top of PCIe and thus add additional functionality to the PCIe interface. Parallel, serial, or combined parallel / serial type interfaces can also be applied to the examples herein. In the following examples, PCIe is used as an exemplary fabric type, but it should be understood that other alternatives can be used instead. PCIe is a high-speed serial computer expansion bus standard that generally has point-to-point connections between a host and component devices, or between peer devices. PCIe generally has individual serial links that connect all devices to a root complex, also called a host. The PCIe communication fabric can be established using the various switching circuits and control architectures described herein.

[0019] The components of the various computing systems herein can be included in one or more physical enclosures, such as rack-mountable modules that can further be included in a shelf or rack unit. A modular framework where modules can be inserted and removed according to the needs of a particular end user allows a number of components to be inserted or installed in a physical enclosure. A sealed modular system can include a physical support structure and enclosure that includes circuitry, printed circuit boards, semiconductor systems, and structural elements. Modules that include components such as computing system 100 can be insertable and removable in a rack-mount type or rack unit (U) type enclosure. It should be understood that the components of FIG. 1 can be included in any physical mounting environment and do not necessarily include associated enclosures or rack-mount elements.

[0020] As a first exemplary system, FIG. 1 is presented. FIG. 1 is a system diagram showing a computing system 100 that uses a workload-based hardware synthesis technique. The computing system 100 includes a computing cluster 101 having physical computing components coupled via a communication fabric (not shown). The computing system 100 includes a management system 110 having a job interface 111 and a job queue 112. The computing system 100 includes a workload manager 120 having a user interface 121. The management system 110 and the workload manager 120 can communicate via one or more network links such as the link 150 in FIG. 1.

[0021] During operation, several computing unit aliases 130, namely computing unit aliases 131 - 133, are provided by the management system 110, although a different number can also be provided. Each computing unit alias includes an indication of a set of virtual computing components that includes the corresponding computing unit alias. These virtual computing components describe what types and amounts of computing components are available for processing jobs, such as for job execution. The computing unit aliases 130 are presented to the workload manager 120 via one or more network links 150 via the job interface 111, along with an indication of the set of virtual computing components that includes each computing unit alias. This presentation and indication can be referred to as the advertisement of the computing unit aliases by the management system via the job interface 111.

[0022] The workload manager 120 can start a job for execution via the user interface 121, or receive an instruction for a job, and transfer a request via the job interface 111 for processing the job by any of the compute unit aliases 130. The job interface 111 can include, among other interfaces, a network interface, a user interface, a terminal interface, an application programming interface (API), a representational state transfer (REST) interface, or a RestAPI. In some examples, the workload manager 120 establishes a front end (user interface 121) for a user or operator that can create, schedule, and transfer jobs for execution or processing by the system computing cluster 101. These execution jobs have characteristics that describe the nature of the execution, operation, and processing processes of each job. For example, a job can have an associated set of metadata indicating the resources required for job execution, or a minimum set of system / computing requirements is necessary to support job execution. Job requirements can be indicated as component type, processing capacity, storage usage, maximum job completion time frame, or other instruction specifications.

[0023] When a job is received by the management system 110, the job is not expanded to calculate the unit alias 130. Instead, the management system 110 processes the characteristics of the incoming job to determine which physical computing components are actually required to support the job. Next, the management system 110 dynamically composes physical computing units to execute the job. These physical computing units, such as the physical computing unit 135, include a set of physical computing components logically coupled by partitions configured within a communication fabric that combines the physical computing components. Each physical computing unit, such as the physical computing unit 135 shown in FIG. 1, can be composed of any number of job-defined quantities of CPUs, CoPUs, NICs, GPUs, or storage units selected from one or more pools 160 of physical computing components, including 0 of several types of components. Network states, such as network addressing, ports, sockets, or other network state information, are assigned from the compute unit alias selected for the job by the workload manager 120 to the physical computing unit 135. Thereby, the workload manager 120 can continue to communicate with the elements of the computing cluster 101 for job execution, status, and processing without interruption.

[0024] Initially, the physical computing unit is not formed or established to support the execution or processing of various jobs. Instead, a pool 160 of physical components is established, and the computing unit can be formed on-the-fly from the components within these pools to match the specific requirements of the execution job. To determine the components that need to be included within the computing unit for a particular execution job, the management system 110 processes the aforementioned characteristics of the execution job to determine which resources are required to support the execution or processing of the job and establishes a computing unit for processing the job. Thus, the total resources of the computing cluster 101 can be dynamically subdivided as needed to support the execution of various execution jobs received via the job interface 111. The computing unit is formed at a specific time, called synthesized or composed, and the software of the job is deployed to the elements of the computing unit to execute / process according to the nature of the job. When a particular job is completed on a particular computing unit, that computing unit can be disassembled, and it is added to the pool 160 of physical components containing individual physical components for use in creating additional computing units for additional jobs. As described herein, various techniques are used to synthesize and disassemble these computing units.

[0025] In addition to the hardware or physical components that are combined into the physical computing unit, when the computing unit is combined, the software components of the job are deployed. The job can include software components to be deployed for execution, such as user applications, user datasets, models, scripts, or other job-providing software. Other software such as an operating system, virtualization system, hypervisor, device driver, bootstrap software, BIOS elements and configurations, status information, or other software components may be provided by the management system 110. For example, the management system 110 can determine that a specific operating system, such as a version of Linux (registered trademark), should be deployed to the combined computing unit to support the execution of a specific job. The indication of the type or version of the operating system may be included in the characteristics associated with the incoming job, or in other metadata of the job. The operating system in the form of an operating system image can be deployed to the data storage element included in the combined computing unit, together with the device drivers necessary to support the other physical computing components of the computing unit. The job can include one or more datasets to be processed by the computing unit, together with one or more applications that execute data processing. Various monitoring or telemetry components, such as utilization levels, job execution status indicating completeness levels, watchdog monitors, or other elements, can be deployed to monitor the activities of the computing unit. In other examples, a catalog of available applications and operating systems can be provided by the computing cluster 101, and the computing cluster can be selected by the job for inclusion in the associated computing unit. Finally, when the hardware and software components are combined / deployed to form a computing unit, the job can be executed on the computing unit.

[0026] To compose computing units, the management system 110 issues commands or control instructions for controlling elements of a communication fabric that couple physical computing components. These physical computing components can be logically separated into any number of distinct, arbitrarily defined arrangements (computing units). The communication fabric can be configured by the management system 110 to selectively route traffic among components of a particular computing unit while maintaining the logical separation between different computing units. In this way, a flexible “bare metal” configuration can be established among the physical components of the computing cluster 101. Individual computing units can be associated with external users or client machines that can utilize the computing, storage, network, or graphics processing resources of the computing unit. Further, any number of computing units can be grouped into “clusters” of computing units for greater parallelism and capacity. Although not shown in FIG. 1 for clarity, various power modules as well as associated power and control distribution links may also be included in each of the components.

[0027] In one example of a communication fabric, a PCIe fabric is employed. The PCIe fabric is formed from a plurality of PCIe switch circuits that can be referred to as PCIe cross-point switches. The PCIe switch circuits can be configured to logically interconnect various PCIe links, at least based on the traffic conveyed by each PCIe link. In these examples, a domain-based PCIe signaling distribution can be included that enables separation of PCIe ports of the PCIe switch according to operator-defined groups. The operator-defined groups can be managed by a management system 110, which logically assembles components into associated computing units and logically separates components of different computing units. The management system 110 can control the PCIe switch circuits via a fabric interface coupled to the PCIe fabric, change the logical partitioning or separation between PCIe ports, and thus change the composition of the grouping of physical components. In addition to or as an alternative to domain-based separation, each PCIe switch port can be a non-transparent (NT) port or a transparent port. An NT port can enable some logical separation between endpoints, such as a bridge, while a transparent port does not enable logical separation and has the effect of connecting endpoints in a purely switched configuration. Access via one or more NT ports can include additional handshakes between the PCIe switch and the initiating endpoint to select a particular NT port or to enable visibility via the NT port. Preferably, this domain-based separation (NT port-based separation) enables physical components (i.e., CPUs, CoPUs, storage units, NICs) to be coupled to a shared or common fabric, but can only have visibility to the components included via separation / partitioning into computing units. Thus, logical partitioning between PCIe fabrics can be used to achieve grouping between multiple physical components. This partitioning is inherently scalable and can be dynamically changed as needed by the management system 110 or other control elements.

[0028] Returning to the description of the elements of FIG. 1, the management system 110 can include one or more microprocessors and other processing circuitry that retrieves and executes software such as a job interface 111 and fabric management software from an associated storage system (not shown). The management system 110 can be implemented within a single processing device, but can also be distributed across multiple processing devices or subsystems that cooperate when executing program instructions. Examples of the management system 110 include general-purpose central processing units, application-specific processors, and logic devices, as well as any other type of processing device, combinations thereof, or variations. In some examples, the management system 110 includes an Intel® microprocessor, an Apple® microprocessor, an AMD® microprocessor, an ARM® microprocessor, a field-programmable gate array (FPGA), an application-specific integrated circuit (ASIC), an application-specific processor, or other microprocessor or processing element. The management system 110 includes or provides a job interface 111. These elements can include various software components executed by the processor elements of the management system 110, or alternatively can include circuitry.

[0029] In FIG. 1, the management system 110 includes a fabric interface. The fabric interface includes a communication link between the management system 110 and any component coupled to a related communication fabric that may include one or more PCIe links. In some examples, the fabric interface can use Ethernet traffic transferred via a PCIe link or other link. Further, each CPU included in the computing unit of FIG. 1 can be configured with a driver or emulation software that can provide Ethernet communication transmitted via a PCIe link. Thus, any of the CPUs in pool 160 (when deployed in the computing unit) and the management system 110 can communicate over Ethernet transferred via the PCIe fabric. However, the implementation is not limited to Ethernet over PCIe, and other communication interfaces including PCIe traffic on the PCIe interface may be used.

[0030] FIG. 2 is included to illustrate an exemplary operation of the elements of FIG. 1. In operation 201, the job interface 111 of the management system 110 advertises a computing unit alias via link 150 monitored by the workload manager 120. These computing unit aliases, shown as computing unit aliases 131-133 in FIG. 1, include a predetermined set of computing components. Computing unit alias 131 represents a set of examples of computing components that can include any number of CPUs, NICs, GPUs, CoPUs, storage units, or other components. Some exemplary computing unit aliases can have a set of components corresponding to physical components, but the examples herein also have the option of presenting computing unit aliases with a set of fictional components. That is, a set of computing components that do not correspond to the physical placement, physical quantity, or current availability of a job. Thus, a computing unit alias can include placeholder entities that are advertised by the management system 110 and are available for receiving jobs for execution. FIG. 1 shows one such job 140, which is initiated by the workload manager 120 in response to a job request that takes over the user interface 121. When a job is dispatched by the management system 110, the computing unit alias is not used to execute the job; instead, physical computing units are synthesized on-the-fly to accommodate the job. In this way, computing unit aliases enable more flexibility with respect to the start and processing of jobs in a distributed computing cluster such as shown herein.

[0031] In some examples, computing unit aliases 131-133 are not only virtual but also over-provisioned. Over-provisioning refers to having a large amount or high availability of computing components that may be available for job execution in computing cluster 101. In some examples, all computing unit aliases represent the entire amount of physical computing components of computing cluster 101 as being available for processing jobs, regardless of current availability or physical presence. Thus, computing unit aliases represent a virtual number of computing components. Due in part to the disaggregated fabric coupling characteristics of computing cluster 101, rapid and dynamic reconfiguration of physically present computing components can occur, allowing for a higher capacity or amount of components as perceived from the perspective of workload manager 120. A user or operator of computing cluster 101 can receive an over-provisioning instruction (211) and enable or disable over-provisioning of computing cluster 101 and the computing unit aliases. This instruction can also indicate the rate or amount of over-provisioning of the computing unit aliases, such as a fixed or dynamic margin that exceeds the physically present amount of computing components or data processing capabilities at any given time. In response, management system 110 can form computing unit aliases based on the indicated amount of over-provisioning (212), and these over-provisioned computing unit aliases can be advertised to workload manager 120 via job interface 111.

[0032] After the workload manager 120 starts operating to process jobs, various jobs can be transferred to the job interface 111 via the link 150. Specifically, the computing unit aliases 131-133 can receive jobs and are advertised or presented via the job interface 111 as including corresponding sets of computing components. The workload manager 120 can select one or more of the computing unit aliases 131-133 to process a particular job and transfer the instructions for those jobs together with an instruction as to which of the computing unit aliases 131-133 should process the job. Further, the job instructions can include the amount or set of resources required to process the job, such as the amount of CPU, GPU, memory unit, or NIC, among other quantities, which can be indicated in granularity units corresponding to the amount of components or alternatively as minimum performance requirements of the job. These minimum performance requirements can include, among other specifications, the required CPU processing capabilities (various units of processing capabilities such as the number of cores or operations per second), memory space (in bytes), NIC capacity in communication bandwidth, and GPU capacity (in operations per second). Each of the computing unit aliases 131-133 also has a related network state indicated by a network address specification, network port, network socket, or other network identifier indicated by the workload manager 120 within the job request, used by the workload manager 120 to start the job, obtain the status regarding the execution or processing of the job, and receive result data or completion status related to the job.

[0033] In operation 202, the job interface 111 of the management system 110 receives an execution job targeting the computing unit aliases 131 to 133 of the computing cluster 101. As described above, the management system 110 analyzes the job characteristics included in the job request to determine the computing resources required to execute the job. In operation 203, the management system 110 can then add the job to the job queue 112 together with the computing unit composition scheduling information. Based on the characteristics of the job, the management system 110 determines a resource scheduling for processing the job, and the resource scheduling indicates the temporal allocation of the resources of the computing cluster 101. The resource scheduling can include one or more data structures regarding the identifier of the job, the representation of the set of computing components required to execute each job, the time frame to start the composition and decomposition of the computing units, and the representation of the software to be deployed on the computing units to execute the job on the computing units that execute the job.

[0034] In FIG. 1, an example execution job 140 is highlighted. When job 140 is received, the characteristics of job 140 are analyzed by management system 110 to determine which physical computing components are required to execute job 140. The physical computing components can include CPUs, CoPUs, storage units, and NICs selected from pool 160 to support job 140. The selected computing components are combined into physical computing unit 135. This combination process can correspond to the scheduling or time allocation of the resources of computing cluster 101 of job 140. Similarly, other jobs received by job interface 111 can have a different set of physical computing components allocated from pool 160 based on the characteristics of the jobs. The physical computing units can use the same computing components, but the scheduled times are different. This reuse of the same physical computing components across various jobs is made possible in part by the dynamic combination, decomposition, and recombining of physical computing devices according to incoming jobs, job completion states, and job performance requirements. Preferably, when jobs are scheduled and executed on different physical computing units, computing unit aliases 131-133 remain constant, thus presenting a consistent set of computing units over time to workload manager 120. As will be described below, various handshakes and transitions are facilitated, and jobs dispatched by workload manager 120 to any of virtual computing unit aliases 131-133 can be actually deployed to the physically computed units that are synthesized and decomposed on the fly.

[0035] In operation 205, the management system 110 synthesizes computing units to support jobs according to the schedules and characteristics shown in queue 112. The management system 110 starts an execution job on the computing cluster 101 according to resource scheduling by at least instructing a communication fabric associated with pool 160 to synthesize a computing unit that includes a set of physical computing components selected from pool 160 to process the execution job. Instructing the communication fabric to synthesize a computing unit includes instructing the communication fabric to form a logical separation within the communication fabric that communicatively couples the set of physical computing components. Each logical separation enables the physical computing components within each set to communicate only via the communication fabric within the corresponding logical separation. When each physical computing unit is formed, the management system 110 controls the communication fabric to deploy software components to the computing unit to execute the job. Next, in operation 207, the synthesized computing unit executes the corresponding job on the synthesized computing unit.

[0036] However, as described above, the workload manager 120 deploys jobs to the compute unit aliases 131-133, rather than directly to the physical compute units. Before a job is deployed and executed by a physical compute unit such as the physical compute unit 135, the network state is transferred from the specific compute unit alias indicated by the workload manager 120 for the job to the physical compute unit (operation 206). The network state can include various network addressing and related information. In one example, a media access control address (MAC address) corresponds to the compute unit alias 131, and a different MAC address corresponds to the network interface of the physical compute unit 135. The MAC address of the compute unit alias 131 is associated with other network addressing specifications such as the IP address presented to the workload manager 120. Since the physical compute unit 135 is synthesized on the fly and can include one or more NICs, the MAC address of the NIC of the physical compute unit 135 is determined in operation 213. Next, a reconfiguration process is used to change which MAC address corresponds to the network addressing specification associated with the compute unit alias 131. Specifically, in operation 214, an address resolution protocol (ARP) reconfiguration or rebroadcast process is executed to associate the IP address of the compute unit alias 131 with the MAC address (e.g., Ethernet address) of the NIC of the physical compute unit 135. This has the effect of assigning the network state of the compute unit alias 131 to the physical compute unit 135. Here, the workload manager 120 can continue to communicate using the same IP address, socket, or port as before, but instead, the communication can be made to correspond to the physical compute unit 135 instead of the compute unit alias 131. The interruption of communication between the workload manager 120 and the compute unit alias 131 during the migration to the physical compute unit 135 is handled by the TCP / IP flow control and error correction functions.

[0037] Once the network state is transferred after the composition of the physical computing unit 135, the corresponding job can be executed by the physical computing unit 135. Regarding the deployment, execution, or completion of the job, various states can be provided to the workload manager 120. This state can be provided by the physical computing unit 135 using the relevant network state (i.e., network addressing) for delivery to the workload manager 120, and the workload manager 120 can communicate with the physical computing unit 135 using such a network state. This network state can be shown to the workload manager 120 as element 142 in FIG. 1.

[0038] Finally, when the execution job is completed, in operation 208, the management system 110 decomposes and returns the resources of the physical computing unit 135 to the pool 160. Before decomposition, various results from the job can be transferred to the workload manager 120. These results can include the processed dataset, the state at completion, or other data executed, processed, operated on, or deployed to the physical computing unit 135. To perform the decomposition, the management system 110 instructs the communication fabric to remove the corresponding logical separation of the physical computing unit so that the computing components of the computing unit become available for composition into additional physical computing units. As part of the decomposition, the network state transferred to the physical computing unit 135 is returned to the computing unit alias 131. This return of the network state can include changing the MAC address of the computing unit alias 131 associated with the IP address that was used for the job and originally transferred to the physical computing unit 135. An ARP reconfiguration or re - broadcast process (operation 215) can perform this change of the MAC address or Ethernet address to accommodate the change in the MAC address of the computing unit alias 131.

[0039] The operations of FIG. 2 and other examples in this specification illustrate job-based initiation of the composition and decomposition of computing units, along with the aliasing of computing units and the transfer of network states to physical computing units. Thus, the initiation of an execution job and related job characteristics can trigger the creation of the physical computing unit that executes the job and the transfer of network states from the aliased computing unit to the physical computing unit.

[0040] To modify or change the physical computing unit, various triggers can be used, either separately or in combination with the composition based on the aforementioned jobs. In the first trigger, event-based triggers are employed. These event-based triggers can change or modify the computing unit or add additional computing units to support a job or a work unit containing the job. Based on the observation by the management system 110 of the dynamic events or patterns indicated by the job, the management system 110 can initiate changes to the configuration of the computing unit and resources assigned to it. Examples of such events or patterns include observed resource shortages in a process, specific character strings identified by a function, specific signals identified by intelligent infrastructure algorithms, or other factors that can be monitored by the management system 110. The analysis of the telemetry of a running job or the characteristics of a job before or during execution can notify the management system 110 to initiate dynamic changes to the computing unit. Thus, the management system 110 can change the composition of the computing unit to add or remove resources (e.g., physical computing components) of the computing unit according to an event or pattern. Preferably, the computing unit can be better optimized to support the current resource needs of each job, and when it becomes unnecessary for the current job, the resources can be intelligently returned to the pool for use by other future jobs.

[0041] Another alternative trigger includes a time trigger based on a machine learning type algorithm or a user-defined time frame. In this example, the pattern or behavior of the synthetic computing unit can be determined or learned over time such that a particular type of job exhibits a particular type of behavior. Based on these behaviors, changes to the computing unit can be made dynamically to support the workload pattern. For example, the management system 110 can determine whether more / less storage resources or more / less coprocessing resources are required at a particular stage of the execution of a particular type of job. The management system 110 can predictively or proactively change the composition of the computing unit, which can include adding or removing or including resources, to better optimize the current resources allocated to the computing unit in the state in which the work unit is being executed by the job. The time characteristics can be determined by the management system 110 based on explicit user input or based on a machine learning process to determine the time frame for adding or removing resources from the computing unit. The management system 110 can include a resource scheduler element that can determine which resource changes are needed and when these changes are needed to support current and future job needs. The changes to the computing unit described herein may, in some examples, require recomposition and restart of the computing unit and the associated operating system, such as when adding or removing a particular physical component or resource. However, other changes, such as adding / removing storage or network interface resources, can be achieved on-the-fly without restarting or recomposing a particular computing unit.

[0042] Figure 3 shows further techniques and structures for deploying a computing unit to process incoming jobs. Figure 3 includes a system 300 that includes a computing cluster 301, a management controller 310, and a workload manager 320. The management controller 310 controls and manages the operation and configuration of the computing cluster 301 and presents a plurality of virtual targets or target aliases (330) to the workload manager 320 via one or more API-style interfaces presented over a network link. During operation, the management controller 310 receives jobs for processing or execution by elements of the computing cluster 301, interprets the requirements of the jobs, and dynamically composes a computing unit for processing the jobs from among various pools of physical computing components. As described below, the jobs target virtual targets or target aliases shown as targets 331-333 in Figure 3, which include a set of over-provisioned components. In contrast, the physical pools of components include actual hardware that can be reconfigured into various groups or sets called computing units or physical computing units since the set includes actual hardware. Figure 3 shows several pools within the computing cluster 301, namely a CPU pool 341, a CoPU pool 342, a storage pool 343, and a NIC pool 344. All components within each pool are communicatively coupled via a common communication fabric such as fabric 340. Fabric 340 includes any of the communication fabric types described herein such as PCIe. The management controller 310 can interface with the switching elements of the fabric 340 to form a computing unit by reconfiguring the logical separation and partitioning within the fabric 340.

[0043] The management controller 310 adopts application programming interfaces (APIs) compliant with various interface specifications such as the Representational State Transfer (REST) interface specification. An API that conforms to the REST architectural constraints is called a RESTful API. Thus, the management controller 310 can present one or more interfaces, including a RESTful API also called a RestAPI that standardizes the definitions and protocols for communication between the workload manager 320 (or any other workload management or orchestration software entity) and the elements of the computing cluster 301 managed by the management controller 310. As part of this API, the management controller 310 provides one or more configuration files to the workload manager 320 that identifies various targets that can receive jobs for execution or other data processing. At startup, the workload manager 320 can read these configuration files and determine, along with the network addressing associated with each target, which targets are available and which resources each target has available.

[0044] In FIG. 3, three exemplary targets, namely targets 331-333, are presented to the workload manager 320. Each of these targets represents a set of computing components that can process jobs, and the workload manager 320 can select individual targets to process individual jobs as needed. However, as described herein, these targets 331-333 are fictitious or aliased and do not correspond to the actual hardware of the computing cluster 301. Thus, a more flexible configuration and quantity of computing components can be included in each target. One such arrangement is to over-provision these targets to have more components physically available within the computing cluster 301. This physical availability can be related to the presence of computing components within a physical modular chassis and rack-mount system, or it can be related to the current availability to accept workloads (i.e., idle). Thus, targets 331-333 present more computing components to the workload manager 320 and abstract the physical components of the computing cluster 301 from the workload manager 320. This over-provisioning is made possible in part due to the speed at which physical components can be synthesized and disassembled on-the-fly to appear to have a larger capacity than expected. Since the management controller 310 can arbitrarily assemble and disassemble the logical relationships between physical computing components, the computing unit can process a large number of simultaneous jobs issued to targets 331-333. The amount of over-provisioning can be specified by the user or operator, such as a ratio, percentage, or absolute value of over-provisioning with respect to bandwidth, processing capacity, or the granularity instance of the physical components. For example, a 20% over-provisioning could be specified that presents each of the targets 331-333 as having 20% more computing components than the computing components physically available within the computing cluster 301.Overprovisioning can be applied to all types of components, or alternatively may relate only to specific components such as a CPU or GPU.

[0045] One advantage of the offload targets 331 - 333 is that when the characteristics of the target change, a particular workload manager cannot reconfigure on the fly (i.e., without a restart / restart). Thus, the management controller 310 can present the maximum number or overprovisioned number of computing components available in the computing cluster 301 for all targets so that the workload manager 320 does not need to be restarted. The workload manager 320 can use any or all components of the computing cluster 301 for any and all targets 331 - 333 at any given time, and even when physical computing units are combined and decomposed as needed to support jobs, the workload manager 320 does not need to be restarted. The overprovisioning aspect allows for more concurrency of jobs deployed to the target and shifts the burden of orchestration of physical components to the management controller 310. The workload manager 310 can blindly dispatch jobs to the targets 331 - 333 as needed, and the management controller 310 can handle the details of execution on the physical hardware. Further, any job initiated by the workload manager 320 can access the full (or overprovisioned) resources of the computing cluster 301 having each target 331 - 333, even if the hardware components are currently being used by other jobs. Status inquiries to the targets 331 - 333 regarding busy / idle status can be answered by the management controller 310 to always present an idle target and maintain the illusion of available physical targets to the workload manager 320. The management controller 310 does this using network state spoofing or network addressing for each of the targets 331 - 333.

[0046] Each of the targets 331 to 333 has a corresponding network state associated therewith, such as a network socket defined by an IP address and a network port, among other network characteristics. When the workload manager 320 desires to communicate with a target, the workload manager 320 dispatches traffic to the network socket associated with each target. As shown in FIG. 3, target 331 has network socket "A", target 332 has network socket "B", and target 333 has network socket "C", each of which includes a unique IP address. Initially, the IP addresses of targets 331 to 333 are also associated with a MAC address assignment or Ethernet address assignment unique to the target. However, these IP addresses route traffic to the management controller 310 that presents the target objects 331 to 333 as virtual entities.

[0047] In response to a job directed at an individual target received by the management controller 310, the management controller 310 can queue the job and dispatch it to the actual physical hardware for completion instead of the virtual targets 331-333. In this example, the management controller 310 receives a job request via the RestAPI, interprets the job request to determine the hardware necessary to support the job, and composes a machine or computing unit to process the job. For the first job, the management controller 310 composes a physical computing device 350 that includes a set of physical computing components such as the CPU 351, NIC 352, GPU 353, and storage device 354. For the second job, the management controller 310 composes a physical computing unit 360 that includes a set of physical computing components such as the CPU 361, NIC 362, GPU 363, and storage unit 364. Further, the network state is transferred from each of the virtual targets to the physical computing unit so that the workload manager 320 can communicate with the physical computing unit during the execution and completion of the job. In FIG. 3, this is shown as a change in the relationship established between the IP address specification of the target and the MAC address specification of the target, and instead corresponds to the MAC address specification associated with the physical computing unit (at the time of composition). This can be achieved in various ways such as an ARP reconfiguration or re-broadcast process. Thus, the workload manager 320 still uses the same network socket as that used for the job request when the physical computing unit is dispatched to process the job. From here, the job is executed or processed by the hardware included in the physical computing unit. Job execution can include data processing operations, machine learning processes, data storage operations, data transfer operations, data conversion operations, graphics rendering operations, or any other data or processing operations that can be processed by the included hardware. Data, state, and other information can be transferred between the workload manager 320 and the physical computing unit during the execution of the job.After the job is completed, the physical computing units are broken down and returned to the various pools of the computing cluster 301. This breakdown is performed by the management controller 310 removing the various partitions or logical associations within the fabric 340 between the hardware components and logging the state of each hardware component as being available for composition into further computing units.

[0048] As described above, the components of the computing cluster 301 include the communication fabric 340, CPUs, CoPUs, and storage devices. Various other devices can be included, such as NICs, FPGAs, RAM, or programmable read-only memory (PROM) devices. Each CPU in the CPU pool 341 includes other processing circuitry for retrieving and executing software, such as user applications, from a microprocessor, system-on-chip device, or associated storage system. Each CPU can be implemented within a single processing device, but can also be distributed across multiple processing devices or subsystems that cooperate when executing program instructions. Examples of each CPU include general-purpose central processing units, application-specific processors, and logic devices, as well as any other type of processing device, combinations thereof, or variations. In some examples, each CPU includes an Intel®, AMD®, Apple®, or ARM® microprocessor, graphics core, compute core, ASIC, FPGA portion, or other microprocessor or processing element. Each CPU includes one or more fabric communication interfaces, such as PCIe, that couple the CPU to the switch elements of the communication fabric 340. The CPU may comprise a PCIe endpoint device or a PCIe host device that may or may not have a root complex.

[0049] Each CoPU in CoPU pool 342 includes a coprocessing element for special processing of a dataset. For example, CoPU pool 342 can include graphics processing resources that can be assigned to one or more computing units. A GPU can include a graphics processor, shader, pixel rendering element, frame buffer, texture mapper, graphics core, graphics pipeline, graphics memory, or other graphics processing and processing elements. In some examples, each GPU includes a graphics "card" that includes circuitry to support the GPU chip. Exemplary GPU cards include NVIDIA® or AMD® graphics cards that include graphics processing elements along with various support circuits, connectors, and other elements. In further examples, other styles of coprocessing units or coprocessing assemblies can be used, such as machine learning processing units, tensor processing units (TPUs), FPGAs, ASICs, or other dedicated processors.

[0050] Each storage unit in storage pool 343 includes one or more data storage drives, such as solid state storage drives (SSDs) or magnetic hard disk drives (HDDs), along with associated enclosures and circuitry. Each storage unit also includes a fabric interface (such as a PCIe interface), a control processor, and power system elements. In still other examples, each storage device includes an array of one or more separate data storage devices along with associated enclosures and circuitry. In some examples, a fabric interface circuit is added to the storage drive to form the storage device. Specifically, the storage drive can include a storage interface, such as SAS, SATAExpress, NVMe, or other storage interface, and the storage interface is coupled to communication fabric 340 using communication conversion circuitry included in the storage unit to convert the communication to PCIe communication or other fabric interface.

[0051] Each NIC in the NIC pool 344 includes circuitry for communicating via a packet network such as an Ethernet and TCP / IP (Transmission Control Protocol / Internet Protocol) network. Some examples transmit other traffic over Ethernet or TCP / IP such as iSCSI (Internet Small Computer System Interface). Each NIC includes an Ethernet interface device and can communicate via a wired, optical, or wireless link. External access to the components of the computing cluster 301 can be provided via the packet network links provided by the NICs, which can include presenting iSCSI, Network File System (NFS), Server Message Block (SMB), or Common Internet File System (CIFS) shares over the network links. In some examples, a fabric interface circuit is added to a storage drive to form a storage device. Specifically, the NIC can include communication conversion circuitry included in the NIC to couple the NIC to the communication fabric 340 using PCIe communication or other fabric interfaces.

[0052] The communication fabric 340 includes a plurality of fabric links coupled by communication switch circuits. In an example where PCIe is used, the communication fabric 340 comprises a plurality of PCIe switches that communicate with members of the compute cluster 301 via associated PCIe links. Each PCIe switch comprises a PCIe cross-connect switch for establishing a switching connection between any PCIe interfaces processed by each PCIe switch. The communication fabric 340 can enable a plurality of PCIe hosts to exist on the same fabric while being communicatively coupled only to associated PCIe endpoints. Thus, many hosts (e.g., CPUs) can communicate independently with many endpoints using the same fabric. The PCIe switches can be used to transfer data between CPUs, CoPUs, and storage units within the compute units, and between compute units when host-to-host communication is used. The PCIe switches described herein can be configured to logically interconnect various ones of the associated PCIe links, based at least on the traffic carried by each PCIe link. In these examples, domain-based PCIe signaling distribution can be included that enables separation of the PCIe ports of the PCIe switches according to user-defined groups. The user-defined groups can be managed by a management controller 310 that logically integrates components into associated compute units and logically separates components from different compute units. In addition to or instead of domain-based separation, each PCIe switch port can be a non-transparent (NT) or a transparent port. An NT port can enable some logical separation between endpoints, like a bridge, while a transparent port does not enable logical separation and has the effect of connecting endpoints in a pure circuit-switching configuration. Access via one or more NT ports can include an additional handshake between the PCIe switch and the initiating endpoint to select a particular NT port or to enable visibility via the NT port.In some examples, each PCIe switch comprises a PLX / Broadcom / Avago PEX series chip such as a PEX8796 24-port, 96-lane PCIe switch chip, a PEX8725 10-port, 24-lane PCIe switch chip, a PEX97xx chip, a PEX9797 chip, or other PEX87xx / PEX97xx chips.

[0053] FIG. 4 is a system diagram showing a computing platform 400. The computing platform 400 can include elements of the computing cluster 101 or 301 of FIGS. 1 and 3, but variations are possible. The computing platform 400 comprises a rack-mounted apparatus of a plurality of modular chassis. One or more physical enclosures such as modular chassis can be further included in a shelf or a rack unit. Chassis 410, 420, 430, 440, and 450 are included in the computing platform 400 and may be attached to a common rack-mounted configuration within one or more data centers or may span multiple rack-mounted configurations. Within each chassis, the modules are attached to a shared PCIe switch along with various power systems, structural supports, and connector elements. A predetermined number of components of the computing platform 400 can be inserted or installed in a physical enclosure such as a modular framework in which modules can be inserted and removed according to the needs of a particular end user. A sealed modular system can include a physical support structure and enclosure including circuits, printed circuit boards, semiconductor systems, and structural elements. The modules comprising the components of the computing platform 400 are insertable and removable in a rack-mounted enclosure. In some examples, the elements of FIG. 4 are included in a "U"-style chassis for attachment within a larger rack-mounted environment. It should be understood that the components of FIG. 4 can be included in any physical attachment environment and need not include associated enclosures or rack-mounted elements.

[0054] The chassis 410 includes a management module or a top-of-rack (ToR) switch chassis, and includes a management processor 411 and a PCIe switch 460. The management processor 411 includes a management operating system (OS) 412, a user interface 413, and a job interface 414. The management processor 411 is connected to the PCIe switch 460 via one or more PCIe links including one or more PCIe lanes.

[0055] The PCIe switch 460 is coupled to PCIe switches 461 - 464 in other chassis within the computing platform 400 via one or more PCIe links. These one or more PCIe links are represented by the PCIe module interconnect 465. The PCIe switches 460 - 464 and the PCIe module interconnect 465 form a PCIe fabric that communicatively couples all of the various physical computing elements of FIG. 4. In some examples, the management processor 411 can communicate with elements of the PCIe fabric via a special management PCIe link or sideband signaling (not shown), such as an inter-integrated circuit (I2C) interface, to control the operation and partitioning of the PCIe fabric. These control operations can include the composition and decomposition of compute units, changes to logical partitioning within the PCIe fabric, monitoring of PCIe fabric telemetry, control of power-up / down operations of modules on the PCIe fabric, updating of firmware for various circuits comprising the PCIe fabric, and other operations.

[0056] Chassis 420 includes a plurality of CPUs 421-425, each coupled to a PCIe fabric via a PCIe switch 461 and an associated PCIe link (not shown). Chassis 430 includes a plurality of GPUs 431-435, each coupled to a PCIe fabric via a PCIe switch 462 and an associated PCIe link (not shown). Chassis 440 includes a plurality of SSDs 441-445, each coupled to a PCIe fabric via a PCIe switch 463 and an associated PCIe link (not shown). Chassis 450 includes a plurality of NICs 451-455, each coupled to a PCIe fabric via a PCIe switch 464 and an associated PCIe link (not shown). Each chassis 420, 430, 440, and 450 can include various modular bays for attaching modules that include corresponding elements of each CPU, GPU, SSD, or NIC. A power system, monitoring elements, internal / external ports, attachment / detachment hardware, and other related features can be included in each chassis. Further description of the individual elements of chassis 420, 430, 440, and 450 is included below.

[0057] When the various CPU, GPU, SSD, or NIC components of computing platform 400 are installed in a related chassis or enclosure, the components are coupled via a PCIe fabric and can be logically separated into any number of separately and arbitrarily defined configurations called "machines" or compute units. Each compute unit can be composed of a selected number of CPUs, GPUs, SSDs, and NICs, including zero of any type of module, but generally, at least one CPU is included in each compute unit. An example of a physical compute unit 401 is shown in FIG. 4, which includes CPU 421, GPUs 431-432, SSD 441, and NIC 451. Compute unit 401 is synthesized using a logical partition within the PCIe fabric indicated by logical domain 470. The PCIe fabric can be configured by management processor 411 to selectively route traffic between the components of a particular compute unit while maintaining logical separation between components not included in that particular compute unit. In this way, a flexible "bare metal" configuration distributed among the components of platform 100 can be established. The individual compute units can be associated with external users, incoming jobs, or client machines that can utilize the computing, storage, network, or graphics processing resources of the compute units. Further, for greater parallelism and capacity, any number of compute units can be grouped into a "cluster" of compute units.

[0058] In some examples, the management processor 411 can result in the creation of a computing unit via one or more user interfaces or job interfaces. For example, the management processor 411 can provide a user interface 413 that can present a machine template for a computing unit that can specify the hardware components to be allocated and the software and configuration information for a computing unit created using the template. In some examples, the computing unit creation user interface can provide a machine template for the computing unit based on the use case or usage category of the computing unit. For example, the user interface can provide proposed machine templates or computing unit configurations for game server units, artificial intelligence learning computing units, data analysis units, and storage server units. For example, a game server unit template can specify additional processing resources compared to a storage server unit template. Further, the user interface can provide templates or computing unit configurations and options for customizing the creation of a computing unit template from component types arbitrarily selected by the user from a list or category of components.

[0059] In an additional example, the management processor 411 can provide policy-based dynamic adjustment to the computing unit during operation. In some examples, the user interface 413 can enable a user to define policies for adjusting the hardware and software assigned to the computing unit and its configuration information during operation. In one example, during operation, the management processor 411 can analyze the telemetry data of the computing unit to determine the current resource utilization rate. Based on the current utilization rate, the dynamic adjustment policy can specify whether processing resources, storage resources, networking resources, etc. are to be allocated to or removed from the computing unit. For example, the telemetry data indicates that the current usage level of the allocated storage resources of the storage computing unit is approaching 100%, and an additional storage device can be allocated to the computing unit.

[0060] In yet another example, the management processor 411 can bring about dynamic adjustment based on the execution job for the computing unit during operation. In some examples, the job interface 414 can receive instructions of execution jobs to be processed by the imaginary targets A - B presented for the computing platform 400. The management processor 411 can analyze these incoming jobs to determine the system requirements for executing / processing the jobs including the resources selected from among the CPU, GPU, SSD, NIC, and other resources. In FIG. 4, Table 490 shows some jobs received via the job interface 414 and queued in the job queue. Table 490 shows the unique job identifier (ID) and an indication of which target is associated with the job, followed by the system components of various granularities to be included within the computing unit formed to support the job. For example, job 491 has a job ID of 00001234, targets target A, and indicates that for the computing unit formed to execute job 491, there is 1 CPU, 2 GPUs, 1 SSD, and 1 NIC. Thus, when it is time to execute job 491, the management processor 411 constructs a computing unit 401 composed of CPU 421, GPUs 431 - 432, SSD 441, and NIC 451. The computing unit 401 is synthesized using logical partitioning within the PCIe fabric indicated by the logical domain 470. The logical domain 470 enables the CPU 421, GPUs 431 - 432, SSD 441, and NIC 451 to communicate via PCIe signaling, while at the same time separating the other components of other logical domains and other computing units from the computing unit 401 via PCIe communication, all sharing the same PCIe fabric. The computing unit 401 also has an IP address associated with the re - assigned target A such that the MAC address or Ethernet address of the NIC 451 is associated with the IP address initially associated with target A. Job 491 can be executed on the computing unit 401 when various software components are deployed to the computing unit 401.The computing unit 401 can be disassembled upon completion of a job and can return various network states to target A. Other targets A to C (or more) can be processed similarly.

[0061] Although the PCIe fabric is described in relation to FIG. 4, the management processor 411 can control and manage a plurality of protocol communication fabrics and communication fabrics different from PCIe. For example, the PCIe switch device of the management processor 411 and the PCIe fabric can bring about communication coupling of physical components using a plurality of different implementations or versions of PCIe and similar protocols. For example, different PCIe versions (e.g., 3.0, 4.0, 5.0, and later) may be adopted for different physical components within the same PCIe fabric. Further, next-generation interfaces such as Gen-Z, CCIX, CXL, OpenCAPI, or a wireless interface including a Wi-Fi interface or a cellular wireless interface can be used. Also, although PCIe is used in FIG. 4, PCIe may not be present, and it should be understood that different communication links or buses such as NVMe, Ethernet, SAS, FibreChannel, Thunderbolt, SATA Express, etc. can be used instead among other interconnections, networks, and link interfaces.

[0062] Referring to the components of the computing platform 400 here, the management processor 411 can include one or more microprocessors and other processing circuits that search for and execute software such as the management operating system 412, the user interface 413, and the job interface 414 from the associated storage system. The management processor 411 can be implemented within a single processing device, but can also be distributed across multiple processing devices or subsystems that cooperate when executing program instructions. Examples of the management processor 411 include general-purpose central processing units, application-specific processors, and logic devices, as well as any other type of processing device, combinations thereof, or variations. In some examples, the management processor 411 includes an Intel® or AMD® microprocessor, an Apple® microprocessor, an ARM® microprocessor, a field-programmable gate array (FPGA), an application-specific integrated circuit (ASIC), an application-specific processor, or other microprocessor or processing element.

[0063] The management operating system (OS) 412 is executed by the management processor 411 and manages the resources of the computing platform 400. This management includes, among other functions, computing unit aliasing, computing unit alias overprovisioning, computing unit composition, computing unit modification, computing unit decomposition, computing unit network state transfer, and monitoring of computing units. The management OS 412 provides the functions and operations described herein for the management processor 411. The user interface 413 can present a graphical user interface (GUI), an application programming interface (API), or a command line interface (CLI), a WebSocket interface to one or more users. The user interface 413 can be used by an end user or an administrator to establish computing units, allocate resources to the computing units, create clusters of computing units, and perform other operations. In some examples, the user interface 413 provides an interface that allows a user to determine one or more computing unit templates and dynamic adjustment policy sets for use or customization in creating computing units. The machine template can be managed, selected, and changed using the user interface 413. The policy of the computing unit can be managed, selected, and changed using the user interface 413. The user interface 413 can also provide the user with telemetry information for the operation of the computing platform 400 in one or more status interfaces or status views, etc. The status of various components or elements of the computing platform 400 can be monitored via the user interface 413, especially the CPU status, GPU status, NIC status, SSD status, PCIe switch / fabric status, etc. Various performance metrics, error situations can be monitored using the user interface 413.

[0064] A plurality of instances of elements 411 to 414 can be included in the computing platform 400. Each management instance can manage the resources of a predetermined number of clusters or computing units. User commands, such as those received via the GUI, can be received by any of the management instances and transferred to a handling management instance by the receiving management instance. Each management instance can have a unique or pre-assigned identifier that can help in the distribution of user commands to the appropriate management instance. Further, the management processors of each management instance can communicate with each other, such as by using a mailbox process or other data exchange techniques. This communication can be performed via a dedicated sideband interface such as an I2C interface, or via a PCIe or Ethernet interface that couples the management processors.

[0065] A plurality of CPUs 421 to 425 are included in the chassis 420. Each CPU can include a CPU module including one or more CPUs or microprocessors and other processing circuits that obtain and execute software such as an operating system, device drivers, and applications from an associated storage system. Each CPU can be implemented within a single processing device, but can also be distributed across multiple processing devices or subsystems that cooperate when executing program instructions. Examples of each CPU include general-purpose central processing units, application-specific processors, and logic devices, as well as any other type of processing device, combinations thereof, or variations. In some examples, each CPU includes an Intel® microprocessor, an Apple® microprocessor, an AMD® microprocessor, an ARM® microprocessor, a graphics processor, a computing core, a graphics core, an ASIC, an FPGA, or other microprocessor or processing element. Each CPU can also communicate with other computing units, such as those within the same storage assembly / enclosure or another storage assembly / enclosure, via one or more PCIe interfaces and the PCIe fabric.

[0066] A plurality of GPUs 431-435 are included in the chassis 430, which can represent any type of CoCPU. Each GPU can comprise a GPU module that includes one or more GPUs. Each GPU includes graphics processing resources that can be allocated to one or more computing units. A GPU can include a graphics processor, shader, pixel rendering elements, frame buffer, texture mapper, graphics core, graphics pipeline, graphics memory, or other graphics processing and processing elements. In some examples, each GPU comprises a graphics "card" that includes circuitry to support a GPU chip. Exemplary GPU cards include NVIDIA® or AMD® graphics cards that include graphics processing elements along with various support circuits, connectors, and other elements. In further examples, other styles of graphics processing units, graphics processing assemblies, or coprocessing elements can be used, such as other special processors including machine learning processing units, tensor processing units (TPUs), FPGAs, ASICs, or other special processing elements for concentrating processing and memory resources for the processing of special data sets.

[0067] A plurality of SSDs 441 - 445 are included in the chassis 440. Each SSD may comprise an SSD module that includes one or more SSDs. Each SSD includes one or more storage drives, such as solid state storage drives having a PCIe interface. Each SSD also includes a PCIe interface, a control processor, and power system elements. Each SSD can include a processor or control system for traffic statistics and status monitoring, among other operations. In yet other examples, each SSD may instead comprise a different data storage medium, such as a magnetic hard disk drive (HDD), a cross-point memory (e.g., an Optane™ device), a static random access memory (SRAM) device, a programmable read only memory (PROM) device, or other magnetic, optical, or semiconductor-based storage medium, along with an associated enclosure, control system, power system, and interface circuitry.

[0068] A plurality of NICs 451 - 455 are included within the chassis 450, each having an associated MAC address or Ethernet address. Each NIC can comprise a NIC module that includes one or more NICs. Each NIC can include a network interface controller card for communicating via a TCP / IP (Transmission Control Protocol (TCP) / Internet Protocol) network or for carrying user traffic such as iSCSI (Internet Small Computer System Interface) or NVMe (NVM Express) traffic. The NIC can comprise an Ethernet interface device and can communicate via a wired, optical, or wireless link. External access to the components of the computing platform 400 can be provided via the packet network link provided by the NIC. The NIC can communicate with other components of the associated computing unit via the associated PCIe link of the PCIe fabric. In some examples, the NIC is provided to communicate with the management processor 411 via an Ethernet link. In additional examples, the NIC is provided to communicate with one or more other chassis, rack - mount systems, data centers, computing platforms, communication fabrics, or other elements via an Ethernet link.

[0069] In addition to the CPU, GPU, SSD, and NIC, other dedicated devices may be employed in the computing platform. These other dedicated devices can include, among other circuits, dedicated coprocessing circuitry, fabric-coupled RAM devices, ASIC circuits, or FPGA circuits, as well as coprocessing modules comprising various memory components, storage components, and interface components. Each of the other dedicated devices can include a PCIe interface having one or more PCIe lanes. These PCIe interfaces communicate via the PCIe fabric and can be used to include other dedicated devices in one or more computing units. These other dedicated devices may comprise PCIe endpoint devices or PCIe host devices, which may or may not have a root complex.

[0070] An FPGA device can be employed as an example of other dedicated devices. The FPGA device can receive processing tasks from another PCIe device such as a CPU or GPU and offload those processing tasks to the FPGA programmable logic circuit. The FPGA is generally initialized to a program state using configuration data, which includes various logic configurations, memory circuits, registers, processing cores, special circuits, and other functions that provide special circuits or application-specific circuits. The FPGA device can reprogram to change the circuits implemented therein and execute different sets of processing tasks at different times. Using the FPGA device, machine learning tasks can be executed, artificial neural network circuits can be implemented, custom interfaces or glue logic can be implemented, encryption / decryption tasks can be executed, blockchain calculations and processing tasks can be executed, or other tasks can be executed. In some examples, the CPU provides data to the FPGA to be processed by the FPGA via a PCIe interface. The FPGA can process this data to generate a result and provide this result to the CPU via the PCIe interface. Two or more CPUs and / or FPGAs can be involved in parallelizing tasks or serially processing data via two or more devices. In some examples, the FPGA device can include locally stored configuration data that can be supplemented, replaced, or overwritten using the configuration data stored in the configuration data storage device. This configuration data can include firmware, programmable logic programs, bitstreams, or objects, PCIe device initial configuration data, among other configuration data described herein. The FPGA configuration can also include an SRAM device or a PROM device used to perform other functions for establishing boot programming, power-on configuration, or initial configuration of the FPGA device. In some examples, the SRAM or PROM device can be incorporated into the FPGA circuit or package.

[0071] The PCIe switches 460 - 464 communicate via the associated PCIe links. In the example of FIG. 4, the PCIe switches 460 - 464 can be used to carry user data between PCIe devices within each chassis and between each chassis. Each of the PCIe switches 460 - 464 includes a PCIe cross - connect switch to establish switching connections between any PCIe interfaces processed by each PCIe switch. The PCIe switches described herein can logically interconnect various ones of the associated PCIe links, at least based on the traffic carried by each PCIe link. In these examples, it can include domain - based PCIe signaling distribution that enables separating the PCIe ports of the PCIe switches according to user - defined groups. The user - defined groups can be managed by a management processor 411 that logically integrates components into associated computing units and logically separates the components and computing units from each other. In addition to, or instead of, domain - based separation, each PCIe switch port can be a non - transparent (NT) or transparent port. An NT port can enable some logical separation between endpoints, like a bridge, while a transparent port does not enable logical separation and has the effect of connecting endpoints in a purely switched configuration. Access via one or more NT ports can include an additional handshake between the PCIe switch and the initiating endpoint to select a particular NT port or to enable visibility via the NT port.

[0072] Preferably, this NT port-based separation or domain-based separation can enable visibility only for components in which physical components (i.e., CPU, GPU, SSD, NIC) are included via separation / division. Thus, using logical partitioning between PCIe fabrics, grouping among multiple physical components can be realized. This partitioning is inherently scalable and can be dynamically changed as needed by the management processor 411 or other control elements. The management processor 411 can control a PCIe switch circuit with a PCIe fabric to change the logical partitioning or separation between PCIe ports and thus change the composition of the grouping of physical components. In this specification, these groupings, called computing units, can form "machines" individually and can be further grouped into clusters of many computing units / machines. Among other considerations described in this specification, in accordance with user instructions received via the user interface, dynamically in response to the load / idle state, dynamically in response to incoming or queued execution jobs, or preemptively due to expected needs, physical components can be added to or removed from a computing unit.

[0073] In a further example, a memory-mapped direct memory access (DMA) conduit can be formed between individual CPU / PCIe device pairs. This memory mapping can be performed, among other configurations, on the PCIe fabric address space. To provide these DMA conduits on a shared PCIe fabric with multiple CPUs and GPUs, the logical partitioning described herein can be employed. Specifically, NT ports or domain-based partitioning on the PCIe switch can separate individual DMA conduits between the associated CPUs / GPUs. The PCIe fabric can have a 64-bit address space, which enables a 2^64-byte addressable space, resulting in at least 16 exabytes of byte-addressable memory. The 64-bit PCIe address space can be shared by all computing units or separated among various computing units that form a configuration for appropriate memory mapping to resources.

[0074] The PCIe interface can support multiple bus widths such as x1, x2, x4, x8, x16, and x32, and each multiple of the bus width includes additional "lanes" for data transfer. PCIe also supports the transfer of sideband signaling such as the System Management Bus (SMBus) interface and the Joint Test Action Group (JTAG) interface, as well as related clocks, power, and bootstrap, among other signaling. PCIe can also have different implementations or versions as used herein. For example, PCIe version 3.0 and later (e.g., 4.0, 5.0, or later) may be adopted. Further, next-generation interfaces such as Gen-Z, Cache Coherent CCIX, CXL, or OpenCAPI can be used. Also, although PCIe is used in FIG. 4, it should be understood that different communication links or buses such as NVMe, Ethernet, SAS, FibreChannel, Thunderbolt, SATA Express, etc. can be used instead among other interconnects, networks, and link interfaces. NVMe is an interface standard for mass storage devices such as hard disk drives and solid-state memory devices. NVMe can replace the SATA interface for interfacing with mass storage devices in personal computer and server environments. However, these NVMe interfaces are limited to a one-to-one host-drive relationship similar to SATA devices. In the examples described herein, the PCIe interface can be used to transfer NVMe traffic and present a multi-drive system with multiple storage drives as one or more NVMe Virtual Logical Unit Numbers (VLUNs) on the PCIe interface.

[0075] Any of the links in FIG. 4 can each use various communication media such as air, space, metal, optical fiber, or any other signal propagation path including combinations thereof. Any of the links in FIG. 4 can include any number of PCIe link or lane configurations. Any of the links in FIG. 4 can each be a direct link or can include various devices, intermediate components, systems, and networks. Any of the links in FIG. 4 can each be a common link, shared link, aggregated link, or can be composed of individual separate links.

[0076] Next, a detailed example of the formation and processing of a computing unit will be described. In FIG. 4, any CPU 421 - 425 has configurable logical visibility to any / all GPUs 431 - 435, SSDs 441 - 445, and NICs 451 - 455, or other physical components coupled to the PCIe fabric of computing platform 400, as if logically separated by the PCIe fabric. For example, any CPU 421 - 425 can transfer and retrieve storage data with any SSD 441 - 445 included in the same computing unit. Similarly, any CPU 421 - 425 can exchange data for processing by any GPU 431 - 435 included in the same computing unit. Thus, combining "m" SSDs or GPUs with "n" CPUs can enable a large - scale and scalable architecture with high levels of performance, redundancy, and density. In an example of graphics processing, NT partitioning or domain - based partitioning in the PCIe fabric can be brought about by one or more of the PCIe switches. This partitioning allows the GPU to interact with the desired one or more CPUs and associate two or more GPUs, such as eight GPUs, with a particular computing unit. Further, the relationship of the dynamic GPU computing units can be adjusted on - the - fly using partitioning across the entire PCIe fabric. Shared NIC resources can also be applied across the entire computing unit.

[0077] FIG. 5 is a system diagram including further details regarding the elements of FIG. 4, such as the formation of the computing unit and the deployment of software components thereto. System 500 includes a management processor 411 that communicates with the composite computing unit 401 via link 510. The composite computing unit 401 includes a CPU 421, GPUs 431-432, an SSD 441, and a NIC 451. The CPU 421 has deployed thereon software including an operating system 522, an application 524, a computing unit interface 525, and an execution job 491. Thus, the CPU 421 is shown as having several operating layers. The first layer 501 is the hardware layer or "metal" machine infrastructure of the computing unit 401 formed on the PCIe fabric using the logical domain 470. The second layer 502 provides the OS as well as the computing unit interface 525. Finally, the third layer 503 provides user-level applications and execution jobs.

[0078] The management OS 111 also includes a management interface 515 that communicates via link 510 with the computing unit interface 525 deployed on the computing unit 401. The management interface 515 enables communication with the computing unit to transfer software components to the computing unit and receive status, telemetry, and other data from the computing unit. The management interface 515 and the computing unit interface 525 provide a standardized interface for management traffic such as control commands, control responses, telemetry data, status information, or other data. The standardized interface can include one or more APIs.

[0079] In some examples, the compute unit interface includes an emulated network interface. This emulated network interface comprises a transport mechanism for transferring packet network traffic over one or more PCIe interfaces. The emulated network interface can emulate a network device, such as an Ethernet device, to the management processor 411, such that the management processor 411 can interact / interface with the CPU 421 of the compute unit 401 via the PCIe interface as if the management processor 411 and the CPU 421 were communicating via an Ethernet network interface. The emulated network interface enables the OS to interface using Ethernet-style commands and drivers, and can comprise kernel-level elements or modules that enable application or OS-level processes to communicate with the emulated network device without having the associated latency and processing overhead associated with a full network stack. The emulated network interface includes software components such as drivers, modules, kernel-level modules, or other software components that appear as network devices to application-level and system-level software executed by the CPU of the compute unit. Preferably, the emulated network interface does not require network stack processing to transfer communications. In the case of a compute unit such as compute unit 401, the emulated network interface does not use network stack processing and still appears as a network device to the operating system 522, such that the user software or operating system elements of the associated CPU can interact with the network interface and communicate via the PCIe fabric using communication methods that face existing networks such as Ethernet communications.The emulated network interface of the management processor 411 transfers communications as related traffic via a PCIe interface or PCIe fabric to another emulated network device located on the computing unit 401. The emulated network interface converts PCIe traffic to network device traffic and vice versa. Processing of communications transferred to the emulated network device via the network stack is omitted, and the network stack is generally used for the type of network device / interface presented. For example, the emulated network device may be presented as an Ethernet device to one or more operating systems or applications. Communications received from one or more operating systems are transferred by the emulated network device to one or more destinations. However, the emulated network interface does not include a network stack for processing communications from the application layer to the link layer. Instead, the emulated network interface converts payload data and a destination to PCIe traffic, such as by extracting the payload data and destination from communications received from one or more operating systems and encapsulating the payload data in a PCIe frame using the addressing associated with the destination.

[0080] The compute unit interface 525 can include an emulated network interface as described for the emulated network interfaces. Further, the compute unit interface 525 monitors the operation of the CPU 421 and the software executed by the CPU 421 and provides telemetry for this operation to the management processor 411. Thus, any user-provided software such as a user-provided operating system (Windows, Linux®, MacOS, Android, iOS, etc.), execution job 491, user application 524, or other software and drivers can be executed by the CPU 421. The compute unit interface 525 provides functions that enable the CPU 421 to participate in associated compute units and / or clusters and to provide telemetry data to the management processor 411 via the link 510. In an example where the compute unit includes physical components that utilize multiple or different communication protocols, the compute unit interface 525 can provide functions that enable protocol-to-protocol communication within the compute unit. Each CPU of the compute unit can also communicate with each other via an emulated network device that transmits network traffic via the PCIe fabric. The compute unit interface 525 can also provide APIs for user software and operating systems to interact with the compute unit interface 525, as well as APIs for exchanging control / telemetry signaling with the management processor 411.

[0081] Furthermore, the computing unit interface 525 can operate as an interface to the device drivers of the PCIe devices of the computing unit to facilitate protocol - to - protocol communication or peer - to - peer communication between the device drivers of the PCIe devices of the computing unit, for example, when the PCIe devices utilize different communication protocols. Further, the computing unit interface 525 can operate to facilitate continuous operation during dynamic adjustment to the computing unit based on a dynamics adjustment policy. Further, the computing unit interface 525 can operate to facilitate migration to alternative hardware in the computing platform based on a policy (e.g., migration from PCIe version 3.0 hardware to Gen - Z hardware based on a utilization or responsiveness policy). Control elements within the corresponding PCIe switch circuit may be configured to monitor PCIe communication between computing units that utilize different versions or communication protocols. As described above, within the computing platform and in some implementations within the computing unit, different versions or communication protocols can be utilized. In some examples, one or more PCIe switches or other devices within the PCIe fabric can operate to function as an interface between PCIe devices that utilize different versions or communication protocols. Detected data transfers can be "trapped" and converted or translated to the version or communication protocol utilized by the destination PCIe device by the PCIe switch circuit and then routed to the destination PCIe device.

[0082] FIG. 6 is a block diagram showing an implementation form of the management processor 600. The management processor 600 shows an example of any of the management processors described in this specification, such as the management system 110 of FIG. 1, the management controller 310 of FIG. 3, or the management processor 411 of FIGS. 4 and 5. The management processor 600 includes a communication interface 601, a job interface 602, a user interface 603, and a processing system 610. The processing system 610 includes a processing circuit 611 and a data storage system 612 that can include a random access memory (RAM) 613, but can include additional or different configured elements.

[0083] The processing circuit 611 can be implemented within a single processing device, but can also be distributed across a plurality of processing devices or subsystems that cooperate when executing program instructions. Examples of the processing circuit 611 include general-purpose central processing units, microprocessors, application-specific processors, and logic devices, as well as any other type of processing device. In some examples, the processing circuit 611 includes physically distributed processing devices such as a cloud computing system.

[0084] The communication interface 601 includes one or more communication and network interfaces for communicating via a communication link, a network such as a packet network, and the Internet. The communication interface can include a PCIe interface, an Ethernet interface, a serial interface, a serial peripheral interface (SPI) link, an inter-integrated circuit (I2C) interface, a universal serial bus (USB) interface, a UART interface, a wireless interface, or one or more local or wide area network communication interfaces capable of communicating via an Ethernet or Internet protocol (IP) link. The communication interface 601 can include a network interface configured to communicate using one or more network addresses that can be associated with different network links. Examples of the communication interface 601 include network interface card devices, transceivers, modems, and other communication circuits. The communication interface 601 can communicate with elements of a PCIe fabric or other communication fabric to establish logical partitions within the fabric via a management interface or a control interface of one or more communication switches of the communication fabric.

[0085] The job interface 602 comprises a network-based interface or other remote interface that receives execution jobs from one or more external systems and provides execution job results and status to such external systems. Jobs are received via the job interface 602 and placed in a job schedule 631 for execution or other types of processing by elements of the corresponding computing platform. The job interface 602 can include, among other interfaces, a network interface, a user interface, a terminal interface, an application programming interface (API), a representational state transfer (REST) interface, a RESTful interface, a RestAPI. In some examples, a workload manager software platform (not shown) establishes a front end for a user or operator that can create, schedule, and transfer jobs for execution or processing. The job interface 602 can receive instructions for these jobs from the workload manager software platform.

[0086] The user interface 603 can include a touch screen, keyboard, mouse, voice input device, voice input device, or other touch input device for receiving input from the user. Output devices such as a display, speaker, web interface, terminal interface, and other types of output devices may also be included in the user interface 603. The user interface 603 can provide output and receive input via a network interface such as the communication interface 601. In an example of a network, the user interface 603 can packetize display or graphics data for a remote display by a display system or computing system coupled via one or more network interfaces. Physical or logical elements of the user interface 603 can provide warnings or visual output to the user or other operator. The user interface 603 can also include associated user interface software executable by the processing system 610 that supports the various user input and output devices described above. Separately, or together with each other and other hardware and software elements, the user interface software and user interface devices can support a graphical user interface, natural user interface, or any other type of user interface.

[0087] The user interface 603 can present a graphical user interface (GUI) to one or more users. The GUI can be used by an end user or an administrator to establish clusters and assign assets (computing units / machines) to each cluster. In some examples, the GUI or other parts of the user interface 603 provide an interface that enables an end user to determine one or more computing unit templates and dynamic adjustment policy sets for use or customization in creating computing units. The user interface 603 can be used to manage, select, and change machine templates and to change the policies of computing units. The user interface 603 can also provide telemetry information, such as in one or more status interfaces or status views. The states of various components or elements can be monitored via the user interface 603, especially the processor / CPU state, network state, storage device state, PCIe element state, etc. Various performance metrics and error situations can be monitored using the user interface 603. The user interface 603 can provide other user interfaces other than the GUI, such as a command-line interface (CLI), an application programming interface (API), or other interfaces. A part of the user interface 603 can be provided via a WebSocket-based interface.

[0088] Both the storage system 612 and the RAM 613 can include a non - volatile data storage system, but modifications are also possible. The storage system 612 and the RAM 613 can each include any storage medium that is readable by the processing circuit 611 and can store software and OS images. The RAM 613 can include volatile and non - volatile, removable and fixed media implemented in any method or technology for storing information such as computer - readable instructions, data structures, program modules, or other data. The storage system 612 can include non - volatile storage media such as solid - state storage media, flash memory, phase - change memory, or magnetic memory, including combinations thereof. The storage system 612 and the RAM 613 can each be implemented as a single storage device, but can also be implemented across multiple storage devices or subsystems. The storage system 612 and the RAM 613 can each include additional elements such as a controller that can communicate with the processing circuit 611.

[0089] Software or data stored on storage system 612 or RAM 613 or within the storage system can include any other form of machine-readable processing instructions having a process that instructs processor 600 to operate as described herein when computer program instructions, firmware, or the processing system is executed. For example, software 620 can drive processor 600 to receive user commands for establishing computing units among a plurality of distributed physical computing components including, among other components, a CPU, GPU, SSD, and NIC. Software 620 can receive and monitor telemetry data, statistical information, operational data, and other data to provide telemetry to the user and drive processor 600 to change the operation of the computing units according to the telemetry data, policies, or other data and criteria. Software 620 can, among other things, manage cluster resources and computing unit resources, establish domain partitioning or NT partitioning among communication fabric elements, interface with individual communication switches, and drive processor 600 to control the operation of such communication switches. The software can also include user software applications, application programming interfaces (APIs), or user interfaces. The software can be implemented as a single application or multiple applications. Generally, when the software is loaded and executed on a processing system, it can convert the processing system from a general-purpose device to a dedicated device customized as described herein.

[0090] System software 620 shows a detailed diagram of an exemplary configuration of the RAM 613. It should be understood that different configurations are possible. System software 620 includes an application 621 and an operating system (OS) 622. Software applications 623 - 629 each include executable instructions that can be executed by the processor 600 to operate a computing system or a cluster controller according to the operations described herein, or to operate other circuits.

[0091] Specifically, the cluster management application 623, as shown in FIG. 1, establishes and maintains clusters and computing units among various hardware elements of the computing platform. The user interface application 624 provides one or more graphical or other user interfaces for an end user to manage associated clusters and computing units and monitor the operation of the clusters and computing units. The job processing application 625 receives execution jobs via the job interface 602 and analyzes the execution jobs for scheduling / queuing, along with a display of the computing components required for processing / executing the jobs within the composite computing unit. The job processing application 625 also indicates the job software or data that needs to be deployed to the composite computing unit for job execution, as well as the data, status, or results that need to be transferred to the originating system via the job interface 602 for the job. The module communication application 626 performs communication among other processor 600 elements such as I2C, Ethernet, emulated network devices, or PCIe interfaces. The module communication application 626 enables communication between the processor 600 and the composite computing unit, as well as with other elements.

[0092] The target aliasing handler 627 presents and manages job targets or placeholder entities to which one or more workload managers or other external entities can dispatch jobs. The target aliasing handler 627 provides a target machine that includes placeholder entities that need not correspond to physical computing components. The target machine can have associated network addressing or other network characteristics. The target aliasing handler 627 responds to status inquiries regarding the target machine by transferring a status response indicating that the corresponding selection or set of computing components is available for job execution, regardless of the availability state of the selection of computing components. When a job is dispatched to a synthetic computing unit, the target aliasing handler 627 can execute an ARP reconfiguration process to handle the transfer of network addressing to the network addressing of the NIC of the synthetic computing unit, such as associating an IP address from the initial MAC address of the target machine to a different MAC address of the synthetic machine or physical computing unit. The target aliasing handler 627 also handles the return of the network that addresses from the synthetic machine to the target machine after completion of the job and decomposition of the synthetic machine.

[0093] The user CPU interface 628 provides communication, APIs, and emulated network devices for communicating with the processor of the computing unit and its dedicated driver elements. The fabric interface 629 establishes various logical partitions or domains among communication fabric circuit elements such as PCIe switch elements of the PCIe fabric. The fabric interface 629 also controls the operation of the fabric switch element and receives telemetry from the fabric switch element. The fabric interface 629 also establishes an address trap or address redirect function within the communication fabric. The fabric interface 629 can interface with one or more fabric switch circuit elements to establish the address range to be monitored and redirected, thus forming an address trap within the communication fabric.

[0094] In addition to software 620, other data 630 can be stored by storage system 612 and RAM 613. The data 630 can include a job schedule 631 (or job queue), a template 632, a machine policy 633, a telemetry agent 634, telemetry data 635, fabric data 636, and a target aliasing configuration 637. The job schedule 631 can include a job identifier, job resources required for job execution, and a display of various other job information. This other job information can include timestamps for reception, start / end of execution, and other information. The job schedule 631 can include one or more data structures that hold a temporal representation of the execution jobs and related computing components necessary to include in a computing unit synthesized for the execution / processing of the execution jobs. The template 632 can include specifications or descriptions of various predefined hardware templates or machine templates. The template 632 can also include a list or data structure of components and component characteristics that can be used for template creation or template adjustment. The machine policy 633 can include specifications or descriptions of various previously defined machine policies. These machine policy specifications can include lists of criteria, triggers, thresholds, limits, or other information, as well as instructions for components or fabrics affected by the policy. The machine policy 633 can also include a list or data structure of policy factors, criteria, triggers, thresholds, limits, or other information that can be used for policy creation or policy adjustment. The telemetry agent 634 can include software elements that can be deployed to components within a computing unit to monitor the operation of the computing unit. The telemetry agent 634 can include hardware / software parameters, telemetry device addressing, or other information used for interfacing with monitoring elements such as IPMI-compliant hardware / software of the computing unit and communication fabric.Telemetry data 635 includes a data store of received data from telemetry elements of various computing units, and this received data can include telemetry data or monitoring data. The telemetry data 635 can be organized into a computing unit layout, a communication fabric layout, or other structures. The telemetry data 635 is cached as data 630 and then can be transferred to other elements of the computing system or used for presentation via a user interface. Fabric data 636 includes information and characteristics of various communication fabrics, including a pool of resources or a pool of components such as fabric type, protocol version, technology descriptors, header requirements, addressing information, and other data. The fabric data 636 can include the relationship between components and the specific fabric to which the components are connected.

[0095] The target aliasing configuration 637 receives and stores a selection or indication of an overprovisioning level indicating an overprovisioned number of computing components having a greater number of computing components than physically available computing components. The target aliasing configuration 637 can store an indication of the number of targets presented to an external entity and various configurations of such targets. For example, the target aliasing configuration 637 can store overprovisioning characteristics, component types, number of components, network addressing characteristics, or other characteristics.

[0096] Software 620 can exist in RAM 613 during the execution and operation of processor 600 and can also exist in the non-volatile portion of storage system 612, among other locations and states, such as during a power-off state. Software 620 can be loaded into RAM 613 during startup or boot procedures, as described for computer operating systems and applications. Software 620 can receive user input via user interface 603. This user input can include user commands and other inputs including combinations thereof.

[0097] Storage system 612 can include flash memory such as NAND flash or NOR flash memory, phase change memory, magnetic memory, among other solid state storage technologies. As shown in FIG. 6, storage system 612 includes software 620. As previously described, software 620 can be within non-volatile memory space for applications and the OS during periods when the power to processor 600 is off, among other operating software.

[0098] Processor 600 is generally intended to represent a computing system in which at least software 620 is deployed and executed to render or otherwise perform the operations described herein. However, processor 600 can also correspond to any computing system that can stage at least software 620 from which software 620 can be deployed and executed, or further distributed, transported, downloaded, or provided to yet another computing system for further distribution.

[0099] The systems and operations described herein perform dynamic allocation of computing resources (CPUs), graphics processing resources (GPUs), network resources (NICs), or storage resources (SSDs) to a computing cluster with computing units. The computing units exist within a pool of unused, unallocated, or free components until they are decomposed and allocated (assembled) to the computing units. The management processor can control the assembly and disassembly of the computing units and provide an interface to external users, job management software, or orchestration software. Processing resources and other elements (graphics processing, network, storage, FPGAs, or others) can be exchanged on-the-fly inside and outside the computing units and associated clusters, and these resources can be allocated to other computing units or clusters. In one example, the graphics processing resources can be dispatched / adjusted by a first computing resource / CPU and subsequently provide the graphics processing status / results to another computing unit / CPU. In another example, when a failure, hang, or overload condition occurs in a resource, additional resources can be introduced to the computing unit and cluster to supplement the resource.

[0100] Processing resources (e.g., CPUs) can be assigned unique identifiers for use by a management processor for identification and on the PCIe fabric. User-supplied software such as operating systems and applications can be deployed to the processing resources as needed when initialized after the CPUs are added to the compute units, and the user-supplied software can be removed from the CPUs when those CPUs are removed from the compute units. The user software can be deployed from a storage system that the management processor can access for deployment. Storage resources such as storage drives, storage devices, and other storage resources can be allocated and subdivided among compute units / clusters. These storage resources can span different or similar storage drives or devices and can have any number of logical units (LUNs), logical targets, partitions, or other logical arrangements. These logical arrangements can include one or more LUNs, iSCSI LUNs, NVMe targets, or other logical partitions. Arrays of storage resources such as mirroring, striping, redundant arrays of independent disks (RAID) arrays can be used, or other array configurations can be used across the storage resources. Network resources such as network interface cards can be shared among the compute units of a cluster using bridging or spanning techniques. Graphics resources (e.g., GPUs) or FPGA resources can be shared among multiple compute units of a cluster using NT partitioning or domain-based partitioning on the PCIe fabric and PCIe switches.

[0101] The functional block diagrams, operational scenarios and sequences, and flow diagrams provided in the figures represent exemplary systems, environments, and methodologies for implementing the novel aspects of the present disclosure. For the sake of simplicity in explanation, the methods included in this specification may be in the form of functional diagrams, operational scenarios or sequences, or flow diagrams, and may be described as a series of operations. However, it should be understood and recognized that since some operations may be performed in a different order than that shown and described herein, and / or concurrently with other operations, the methods are not limited by the order of operations. For example, those skilled in the art can understand and recognize that the method may alternatively be represented as a series of interrelated states or events such as a state diagram. Further, not all operations illustrated in the methodology are required for the novel implementation.

[0102] The descriptions and figures included herein illustrate specific embodiments for teaching those skilled in the art how to make and use the best options. For the purpose of teaching the principles of the present invention, some conventional aspects are simplified or omitted. Those skilled in the art can understand the variations from these implementations that fall within the scope of the present disclosure. Also, as will be appreciated by those skilled in the art, the foregoing features can be combined in various ways to form multiple implementations. As a result, the present invention is not limited to the specific embodiments described above, but is limited only by the claims and their equivalents.

Claims

A method implemented by a processing system, the method comprising, as steps executed by the processing system: Presenting to a workload manager a target machine capable of receiving an execution job, the target machine having a network state and including a selection of computing components; Receiving, by the workload manager, a job issued and directed to the target machine; Determining, based on characteristics of the job, resource requirements for processing the job and forming a composite machine comprising physical computing components that support the resource requirements of the job; Transferring the network state of the target machine to the composite machine and indicating the network state of the composite machine to the workload manager; Initiating execution of the job on the composite machine; A method comprising the above steps. Claim 2 The method according to claim 1, wherein the target machine comprises a placeholder entity that does not correspond to a physical computing component. Claim 3 The method according to claim 2, further comprising transferring a status response indicating that a selected computing component is available for execution of the job regardless of the availability status of the selection of computing components, in response to a status query regarding the placeholder entity. Claim 4 The method according to claim 1, wherein the network state comprises a network socket corresponding to the target machine. Claim 5 A step executed by the processing system: Transferring the network state by changing at least a media access control (MAC) address relationship such that an IP address initially corresponding to the target machine instead corresponds to a MAC address of the composite machine; The method according to claim 4, further comprising the above step. Claim 6 The method according to claim 1, wherein the selection of computing components comprises an over-provisioned number of computing components having a number of computing components greater than the physically available number. Claim 7 A step executed by the processing system: selecting the physical computing component from a pool of physical computing components; instructing at least one communication fabric that couples the pool of physical computing components to form a logical partition in the communication fabric to establish the composite machine, the logical partition separating the physical computing components of the composite machine from other physical computing components of the pool of physical computing components; The method according to claim 1, further comprising.

8. The pool of physical computing components comprises at least one of a central processing unit (CPU), a coprocessing unit, a graphics processing unit (GPU), a tensor processing unit (TPU), a field programmable gate array (FPGA), an application specific integrated circuit (ASIC), a storage drive, and a network interface controller (NIC) coupled to at least the communication fabric. The method according to claim 7.

9. As steps executed by the processing system, responding to completion of the job, decomposing the composite machine and transferring the network state of the composite machine to the target machine; The method according to claim 1, further comprising.

10. one or more computer-readable storage media; a processing system operably coupled to the one or more computer-readable storage media; program instructions stored on the one or more computer-readable storage media, which, when executed by the processing system, are based on at least presenting a target machine capable of receiving execution jobs to a workload manager, the target machine having a network state and including a selection of computing components; receiving a job issued by the workload manager and directed to the target machine; determining resource requirements for processing the job based on characteristics of the job and forming a composite machine comprising physical computing components that support the resource requirements of the job; Transfer the network state of the target machine to the composite machine, and indicate the network state of the composite machine to the workload manager, Start executing the job on the composite machine, Program instructions for instructing the processing system as follows, An apparatus comprising.

11. The apparatus according to claim 10, wherein the target machine comprises a placeholder entity that does not correspond to a physical computing component.

12. Based on being executed by the processing system, at least, In response to a status inquiry regarding the placeholder entity, transfer a status response indicating that the selected computing component is available for executing the job regardless of the availability status of the selection of the computing components. The apparatus according to claim 11, comprising program instructions for instructing the processing system as follows.

13. The apparatus according to claim 10, wherein the network state comprises a network socket corresponding to the target machine.

14. Based on being executed by the processing system, at least, Transfer the network state by changing at least the media access control (MAC) address relationship so that the IP address initially corresponding to the target machine instead corresponds to the MAC address of the composite machine. The apparatus according to claim 13, comprising program instructions for instructing the processing system as follows.

15. The apparatus according to claim 10, wherein the selection of the computing components comprises an over-provisioned number of computing components having a number of computing components greater than the physically available number.

16. Based on being executed by the processing system, at least, Select the physical computing component from a pool of physical computing components, Instruct to form a logical partition in the communication fabric to establish the composite machine in at least one communication fabric that couples the pool of physical computing components, the logical partition separating the physical computing components of the composite machine from other physical computing components of the pool of physical computing components. The apparatus according to claim 10, comprising program instructions for instructing the processing system as described above.

17. The pool of physical computing components of the apparatus according to claim 16 comprises at least one or more of a central processing unit (CPU), a coprocessing unit, a graphics processing unit (GPU), a tensor processing unit (TPU), a field programmable gate array (FPGA), an application specific integrated circuit (ASIC), a storage drive, and a network interface controller (NIC) coupled to at least the communication fabric.

18. A job interface, which presents computing targets as virtual targets for job execution, each of the computing targets having an associated set of advertised computing components and corresponding network addressing, receives a job for execution directed to a selected computing target having a corresponding network address, and is configured as such; A controller, selects a set of physical computing components necessary to support the execution of the job from a pool of physical computing components, synthesizes a physical computing node comprising the set of physical computing components, configures the physical computing node to communicate via a corresponding network address instead of the selected computing target, and deploys the job to the physical computing node for processing, and is configured as such; A computing system comprising the above.

19. Each of the associated sets of the advertised computing components of the system according to claim 18 comprises a greater number of computing components than the number available from the pool of physical computing components.

20. The controller, configured to instruct at least one of the communication fabrics to form a logical partition in the communication fabric and couple the pool of the physical computing components to establish the physical computing node, the logical partition separating the physical computing components of the physical computing node from other physical computing components of the pool of the physical computing components, The pool of the physical computing components comprises at least one or more of a central processing unit (CPU) coupled to at least the communication fabric, a coprocessing unit, a graphics processing unit (GPU), a tensor processing unit (TPU), a field programmable gate array (FPGA), an application specific integrated circuit (ASIC), a storage drive, and a network interface controller (NIC). The system according to claim 18.

Citation Information

Patent Citations

  • Information processing method, device, and program

    JP2015090675A

  • Method, apparatus, computer program product, and data center facility for implementing a disaggregated computing system

    JP2019511051A

  • Execution job computation unit synthesis in a computing cluster

    JP2023553213A

  • Resource manager for managing the sharing of resources among multiple workloads in a distributed computing environment

    US20080155100A1

  • Management of addresses in virtual machines

    US20150128245A1