Dynamic Server Rebalancing
The dynamic server rebalancing system addresses inefficiencies in data center resource utilization by virtually connecting peripheral devices across systems, optimizing space and cost through flexible resource allocation.
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- リキッド インコーポレイテッド
- Filing Date
- 2023-03-02
- Publication Date
- 2026-05-27
AI Technical Summary
Existing data center systems face limitations in resource utilization due to fixed server arrangements, leading to inefficiencies in space and cost requirements, as jobs often occupy entire servers regardless of actual resource needs, and existing technologies fail to dynamically rebalance peripheral devices across multiple systems.
Implementing a dynamic server rebalancing system that allows peripheral devices such as GPUs, FPGAs, and ASICs to be virtually connected and shared across multiple computer systems via a network, enabling flexible allocation and utilization of resources without physical constraints.
Enhances data center utilization by efficiently sharing and reallocating peripheral resources, optimizing space and cost by dynamically rebalancing computing resources across systems, supporting applications like AI, machine learning, and other data processing tasks.
Smart Images

Figure 0007866281000001 
Figure 0007866281000002 
Figure 0007866281000003
Abstract
Description
Background Art
[0001] [Cross - reference to Related Applications] This application claims the benefit and priority of U.S. Patent Application No. 17 / 831,518, titled "DYNAMIC SERVER REBALANCING", filed on June 3, 2022, and U.S. Provisional Patent Application No. 63 / 321,274, titled "DYNAMIC SERVER REBALANCING", filed on March 18, 2022.
[0002] Clustered computer systems are in high demand and are widespread for data storage, data processing, and communication processing. Data centers typically include large - scale rack - mounted and network - connected data storage and a data processing system. These data centers can receive data for storage from external users via network links, and in addition, can receive data generated from applications executed by processing elements within the data center. In many cases, data centers and related computer equipment can be utilized to execute jobs for multiple simultaneous users or applications. In addition to using a central processing unit (CPU) or a graphics processing unit (GPU) to process data, execution jobs can route data related to these resources between temporary storage and long - term storage or between various network locations. For example, GPU - based processing is becoming increasingly popular for use in artificial intelligence (AI) and machine learning regimes. In these regimes, a computer system such as a blade server can include one or more GPUs along with a related CPU for processing large - scale data sets.
[0003] Specifically, a server typically includes a fixed arrangement of a CPU, a GPU, and storage elements housed in a common enclosure or chassis. Once an incoming job is deployed within the data center, the granularity of computing resources is limited to individual servers. Therefore, regardless of whether the entire server resource is actually needed to run the job, a deployed job typically occupies one or more servers, along with all of the corresponding CPUs, GPUs, and storage elements of each server. To compensate, data center operators typically deploy a continuously increasing number of servers to handle the increasing traffic from jobs. This strategy can face significant space and cost requirements, in addition to the physical space required for rack-mount servers. [Overview of the project] [Means for solving the problem]
[0004] This paper presents improved devices, systems, and technologies for virtually connecting remotely located peripheral devices, physically coupled to a host computer system, to a client computer system. Even though the peripheral devices are located remotely from the host device and connected to the client device via a communication network link, they may be coupled to the client computer system as local devices. These improvements may provide peripheral devices such as GPUs, FPGAs, or ASICs (also known as co-processing units (CoPUs) or data processing accelerators) that are available on demand by the client computer system via a network link. These peripheral devices can be arbitrarily associated with and unassociated with various client devices (such as servers or other computer systems) as if they were local peripheral devices plugged into the client device. Therefore, a client device can add a greater number of peripheral devices for processing workloads or user data than would normally be possible if the peripheral devices were physically plugged into the client device's motherboard. The host device can share access to physically connected peripheral devices with remote client devices when the peripheral devices are not being used by the host, thereby efficiently utilizing all the resources of the computer cluster without the costs associated with adding additional server and space requirements.
[0005] In one exemplary implementation, the method may include receiving instructions for a peripheral device that is available for data processing and located in a first computer device; receiving a request from a second computer device to access the peripheral device; and, based on the request, instructing the second computer device to emulate the peripheral device as a local device installed in the second computer device. The method may further include routing data traffic from the second computer device to be processed by the peripheral device in the first computer device.
[0006] In another exemplary implementation, the system may include a first computer device including a network interface, which is configured to receive instructions from a server rebalancing system for peripheral devices available via the network interface for processing, and the peripheral devices are located in a second computer device. The first computer device may issue requests to access the peripheral devices to the server rebalancing system via the network interface and, based on the response from the server rebalancing system, emulate a local installation of the peripheral devices in the first computer system. The first computer device may also issue data traffic to the server rebalancing system via the network interface for processing by the peripheral devices.
[0007] In yet another exemplary implementation, the method may include issuing an instruction from the first computer device to the server rebalancing system via a network interface that a peripheral device located at the first computer device is available for processing, and receiving a second instruction from the server rebalancing system via a network interface at the first computer device that the peripheral device is assigned to the second computer device. The method may further include receiving data traffic from the second computer device for processing by the peripheral device at the first computer device via a network interface from the server rebalancing system, and providing the results of processing the data traffic by the peripheral device from the first computer device via a network interface.
[0008] This summary is provided to provide a simplified introduction to the selection of concepts further described below in this technical disclosure. It should be understood that this summary is not intended to identify any significant or essential features of the claimed subject matter, nor should it be used to limit the scope of the claimed subject matter. [Brief explanation of the drawing]
[0009] Many aspects of this disclosure can be preferably understood by reference to the following drawings. The elements of the drawings are not necessarily to scale, but rather the focus is on clearly illustrating the principles of this disclosure. Furthermore, in the drawings, similar reference numerals indicate corresponding parts through several figures. Although several embodiments are described in relation to these drawings, this disclosure is not limited to the embodiments disclosed herein. Rather, it is intended to cover all alternative forms, modifications, and equivalents.
[0010] [Figure 1] This is a diagram showing a computer system in one implementation configuration.
[0011] [Figure 2] This flowchart illustrates an example of how a computer system operates in one implementation configuration.
[0012] [Figure 3] This figure shows a host device in one implementation configuration.
[0013] [Figure 4] This figure shows a fabric control system in one implementation configuration.
[0014] [Figure 5] This is a diagram showing a computer system in one implementation configuration.
[0015] [Figure 6] This figure shows a control system in one implementation configuration. [Modes for carrying out the invention]
[0016] A data center containing related computer equipment is used to process jobs and data, and can also transfer data associated with jobs between temporary and long-term storage, or between various network destinations. A data center typically includes numerous rack-mount computer systems or servers, each containing independently packaged components, such as a corresponding set of processors, system memory, data storage devices, network interfaces, and other computer equipment coupled via an internal data bus. Once installed, these servers are usually not modified or altered, except for minor upgrades or repairs of individual components to replace existing ones. This relatively fixed arrangement of servers can be called a consolidated arrangement. Thus, each server represents a granularity of computer equipment where individual internal components, once housed by the server manufacturer and inserted into a rack by the system installer, are rarely modified.
[0017] The limitations of individual server-based computer systems can be overcome by using deaggregated physical components and peripheral devices that can be dynamically attached to client computer systems while not locally coupled to the local data bus of such client systems. Instead of having a fixed arrangement between computer devices and peripheral devices, with the entire computer system housed in a common enclosure or chassis, the examples herein can flexibly include any number of peripheral devices spanning any number of enclosures / chassis, dynamically formed into a logical arrangement via a communication fabric or network. Furthermore, in addition to deaggregated components that do not have conventional server motherboard relationships, the various exemplary integrated computer systems described herein can make unused locally connected data processing resources and peripheral devices available to other integrated computer devices. For example, even when a client is remotely accessing a device via a network, peripheral devices of the host computer system can be emulated as being locally mounted or locally installed on the client computer system for use. Thus, by not having idle or wasted portions of an integrated server that are not required for a particular task or a particular part of a task, and instead making those idle components available for use by other computer devices, the computer system can make better use of its resources. Data center operators can achieve extremely high levels of data center utilization that could not be achieved using fixed-location servers, and can augment existing servers with additional functionality via existing network connectivity. These operations and techniques can be called dynamic server rebalancing.
[0018] The systems and operations described herein provide dynamic rebalancing and allocation of peripheral resources of individual computer devices, such as computing resources (CPU), graphics processing resources (GPU), network interface resources (NIC), communication fabric interface resources, data storage resources (SSD), field-programmable gate arrays (FPGA), and system memory resources (RAM), across multiple computer devices, even when peripheral resources are not locally coupled to client devices utilizing those resources. Peripheral resources may also include coprocessing units (CoPUs) or data processing accelerators, such as GPUs, tensor processing units (TPUs), FPGAs, or application-specific integrated circuits (ASICs). Data processing of host device CPUs enhanced by CoPUs is gaining popularity for use in artificial intelligence (AI), machine learning systems, cryptocurrency mining and processing, advanced graphical visualization, biological system modeling, autonomous vehicle systems, and a variety of other tasks.
[0019] In one example, peripheral resources may be deaggregated and established as a pool of unused, unassigned, or free peripheral resources until they are allocated (configured) to a requesting client device using a communication fabric such as PCIe Express (Peripheral Component Interconnect Express) or Compute Express Link (CXL). A management processor, or dynamic server rebalancing system, may control the merging and dismerging of connected servers and computer systems and provide interfaces to external users, job management software, or orchestration software. Peripheral resources and other elements (graphics processing, networking, storage, FPGA, RAM, or others) are made available by the host device and can be attached / detached on the fly to various client devices. In another example, peripheral resources may reside within the enclosures of individual servers. These peripheral resources can be established in a pool as described above for deaggregated components, but instead are physically associated with individual servers. By using the improved techniques described herein, components located within a first server can be used for the activities of a second server as if those components were local devices to that second server. For example, graphics processing resources physically attached to a host device may be allocated to a first client device via a dynamic server rebalancing system, then removed from the first client device and allocated to a second client device. In another example, if a resource fails, malfunctions, or becomes overloaded, additional peripheral resources from other host devices may be introduced to the client device to supplement the existing resources.
[0020] As a first exemplary system, Figure 1 is presented. Figure 1 is a system diagram showing a computer system 100 that utilizes dynamic server rebalancing technology, which can encompass both deaggregated computer components and integrated servers. System 100 includes computer devices 110 and 140, a dynamic server rebalancing system 130, a network switch 131, and deaggregated peripheral devices 153. Computer devices 110 and 140 can communicate with the network switch 131 via network links 150-151. Peripheral devices 153 communicate with the network switch 131 via a communication fabric link 152. Although only two computer devices are shown in Figure 1, it should be understood that any number of computer devices can be included and coupled to the network switch 131 via the relevant network links.
[0021] Computer devices 110 and 140 include network interfaces 111 and 141, local peripheral device interconnection interfaces 112 and 142, and peripheral over fabric (PoF) systems 115 and 145. Network interfaces 111 and 141 may be coupled to network switch 131 via network links 150 to 151. Local interfaces 112 and 142 may be coupled to local peripheral devices 113 and 143 via local links. Computer device 110 and its associated components and connections are described below, but for the sake of clarity, the same functions may apply to computer device 140 unless otherwise specified. The PoF system 115 may be coupled to both network interface 111 and local interface 112 via software and hardware connections, such as via software interfaces to the associated protocol stacks or programming interfaces of network interface 111 and local interface 112.
[0022] During operation, the computer device 110 may use its onboard central processing unit (CPU), along with peripheral devices including a graphics processing device (e.g., GPU), a data storage device (e.g., SSD), a memory device (e.g., DRAM), a network interface (e.g., NIC), and a user interface device, to run system software and user applications for various tasks. The operator of the computer device 110, or the OS or other components of the computer device 110, may wish to add additional peripheral devices for use by the computer device 110. Conversely, the computer device 110 may indicate that it has local resources or peripheral devices that are idle or otherwise available for use by other computer devices in the system 100. To facilitate making local peripheral devices remotely available to other computer devices, and to facilitate adding remote peripheral devices to a computer device without physically connecting such peripheral devices to the computer device, various improved techniques and systems are presented herein. A peripheral device, including a data processing element (such as a CoPU) or other peripheral devices (such as data storage or memory devices), can be configured to be associated with the computer device 110 even if such a device or element is not physically local to the computer device 110. The computer device 110 and the elements included in the dynamic server rebalancing system 130 can enable remote sharing or addition of peripheral devices for use by the computer device 110, as if the remote peripheral device were a local device coupled via a local interface such as a PCIe interface. While PCI / PCIe connectivity is referred to herein as an example of a common peripheral device communication protocol, it should be understood that other peripheral device protocols, such as CXL, can be used without departing from the scope of this disclosure.By remotely sharing peripheral devices and other resources between computer devices, any association between any of peripheral devices 113, 143, and 140 and any of host devices 110, 140 can be created and changed on the fly. These associations are made via the network interface of computer device 110 and an optional communication interface coupled to peripheral device 140, as will be described in detail below.
[0023] Referring here to the explanation of the elements in Figure 1, the host device 110 comprises a computer system having processing elements, data storage and memory elements, and user interface elements. Examples of the host device 110 include servers, blade servers, desktop computers, laptop computers, tablet computers, smartphones, game systems, elements of distributed computer systems, customer equipment, access terminals, internet appliances, media players, or other computer systems. Typically, the host device 110 will include a motherboard or other system circuit board having a central processing unit (CPU) coupled with a memory device such as random access memory (RAM) or dynamic RAM (DRAM). The CPU may be a component in a processing system formed from one or more microprocessor elements, including Intel® microprocessors, Apple® microprocessors, AMD® microprocessors, ARM® microprocessors, field-programmable gate arrays (FPGAs), application-specific integrated circuits (ASICs), tensor processors, application-specific processors, or other microprocessors or processing elements. Various peripheral devices 113 may be locally and physically coupled to the peripheral device interconnect interface 112 of the host device 110 within the enclosure or chassis of the host device 110 via corresponding slots, connectors, or cabling. These peripheral devices 113 may include graphics cards housing graphics processing units (GPUs), data storage drives using various computer-readable media, network interface controllers (NICs) having physical layer elements for coupling to network links (e.g., Ethernet), or other devices including user interface devices. For example, a PCIe device may be included in the host device 110, which is coupled to a PCIe slot on the motherboard of the host device 110. Such devices 113 are referred to as local peripheral devices.Furthermore, the host device 110 also includes various software that runs on the processing system of the host device 110. This software typically includes an operating system, user applications, device drivers, user data, hypervisor software, telemetry software, or various other software elements.
[0024] The dynamic server rebalancing system 130, sometimes called a control system, control module, management system, management processor, or fabric control system, provides control and interaction elements for the network switch 131. The dynamic server rebalancing system 130 may be included in a top-of-rack (ToR) switch chassis or other network switching or routing elements. The dynamic server rebalancing system 130 comprises a management operating system (OS), an operator control interface, and various other elements as shown in Figures 4-6. The dynamic server rebalancing system 130 may include one or more microprocessors and other processing circuits that retrieve and execute software from an associated storage system (not shown). The dynamic server rebalancing system 130 may be implemented within a single processing device, but may be distributed across multiple processing devices or subsystems that cooperate in executing program instructions. Examples of the dynamic server rebalancing system 130 include general-purpose central processing units, application-specific processors, and logic devices, as well as any other types of processing devices, combinations thereof, or variations thereof. In some examples, the dynamic server rebalancing system 130 comprises an Intel® microprocessor, an Apple® microprocessor, an AMD® microprocessor, an ARM® microprocessor, an FPGA, an ASIC, an application-specific processor, or other microprocessor or processing element. The dynamic server rebalancing system 130 includes at least one network interface, and in some examples, at least one fabric interface. The network interface comprises a network stack and an associated physical layer interface used to communicate with the network switch 131, control elements of the network switch 131, and communicate with devices coupled to port 133 of the network switch 131. The fabric interface may include a communication link subsystem used to communicate with the network switch 131 and to communicate between peripheral devices coupled to port 134.The fabric interface may include one or more PCIe interfaces or other suitable fabric interfaces.
[0025] Network switch 131 includes network ports 133 that provide switched network connections to computer devices, as shown for network links 150-151. Network switch 131 includes various network switching circuits for communicatively linking individual ports to other ports based on traffic patterns, addressing, or other traffic properties. In one example, network switch 131 comprises an Ethernet or Wi-Fi (802.11xx) switch that supports wired or wireless connections, which may refer to any of the various network communication protocol standards and bandwidths available, e.g., 10BASE-T, 100BASE-TX, 1000BASE-T, 10GBASE-T (10 Gigabit Ethernet), 40GBASE-T (40 Gigabit Ethernet), Gigabit (GbE), Terabit (TbE), 200GbE, 400GbE, 800GbE, or other various wired and wireless formats and speeds.
[0026] The network switch 131 optionally includes a fabric port 134, which may be part of a separate fabric switch element that communicates with the dynamic server rebalancing system 130 via the fabric port. The fabric port 134 may be coupled to peripheral devices 153 via an associated fabric link 152, which typically comprises a point-to-point multi-lane serial link. Types of fabric ports and links include, among others, PCIe, Gen-Z, InfiniBand, NVMe, FibreChannel, NVLink, Cache Coherent Interconnect for Accelerators (CCIX), Compute Express Link (CXL), and Open Coherent Accelerator Processor Interface (OpenCAPI). In Figure 1, the fabric port 134 is shown coupled to a set or group of peripheral devices 153, the group including individual peripheral devices coupled to a backplane having a shared fabric connection to the fabric port 134 via link 152. While this shows a single group of uniform peripheral devices 153 (e.g., just a GPU box, a JBOG), it should be understood that any number of groups of peripheral devices in some arbitrary arrangement could be used instead, and individual peripheral devices could be coupled to the fabric port 134 without using a shared backplane.
[0027] Peripheral devices 113, 143, and 153 may comprise various coprocessing units (CoPUs) or data processing accelerators, such as graphics processing units (GPUs), tensor processing units (TPUs), field-programmable gate arrays (FPGAs), or application-specific integrated circuits (ASICs), as described above. Peripheral devices may also include data storage devices, such as solid-state storage devices (SSDs) containing flash memory or other media types, hard drives (HDDs) containing rotating magnetic media, magnetoresistive random access memory (MRAM), or other data storage devices having various media and interface types. Peripheral devices may also include fabric-coupled memory devices, such as dynamic random access memory (DRAM), static random access memory (SRAM), 3D XPoint memory, solid-state memory devices, magnetic random access memory (MRAM) devices, or various other memory devices. Peripheral devices may include a network interface controller (NIC) that contains various network interface elements such as physical layer elements (PHY), transport layer elements, TCP / IP traffic processing elements, routers, switches, or bridges, along with associated cable connectors.
[0028] Here, with reference to an exemplary set of the operation of the elements in Figure 1, operation 200 is shown in Figure 2. Operation 200 can be subdivided into operations performed by the dynamic server rebalancing system 130 (labeled as switch (SW) operations) and operations performed by the computer device 110 (labeled as server operations), although various operations may be performed by devices other than those listed in Figure 2.
[0029] In operation 201, the dynamic server rebalancing system 130 discovers a network-connected server device, such as a computer device 110. The dynamic server rebalancing system may receive initialization messaging provided by the computer device 110, including a device identifier (ID), vendor ID, addressing, and other information. This information may be automatically provided by the computer device 110 upon power-up or reset (for example, initiated as an operation of the server 110, not as an operation of the dynamic server rebalancing system 130), or it may be queried by the dynamic server rebalancing system 130 using various protocol-specific commands or control traffic. A communication protocol or handshake operation may be used between the server device 110 and the dynamic server rebalancing system 130, which provides the host IP address, port number, and login / password of the server device 110 so that the dynamic server rebalancing system 130 can issue attach / detach operations to the server device 110.
[0030] Regardless of the technique used to discover individual host devices via link 150, the dynamic server rebalancing system 130 may establish and maintain one or more data structures listing these properties and identifiers of host devices, along with their corresponding IP addresses, port numbers, and login / password parameters. One exemplary data structure is shown in Figure 1 as "Server Device 135". The dynamic server rebalancing system 130 may update this data structure over time to add or remove server devices in response to changes in the operational status, presence, or availability of the dynamic server rebalancing system. A similar discovery process may be performed to discover peripheral devices 153 via link 152.
[0031] In some examples, the dynamic server rebalancing system 130 discovers many network devices, but may only discover catalog computer devices 110 that support network-coupled peripheral device mode through the use of dedicated software, firmware, or hardware deployed on such computer devices 110, such as a PoF module 115. This special module, such as a driver module or kernel-level module, may report a status to the dynamic server rebalancing system 130 that includes a mode support instruction along with server-related network addresses, network ports, or network sockets.
[0032] In response to a discovery process between the dynamic server rebalancing system 130 and the server device 110, the server device 110 (or one or more modules thereof) may generate a list of available peripheral components physically located on the server device 110 and present it to the dynamic server rebalancing system (202). The list may include all peripheral components on the server device 110, or it may be limited to peripheral components that are idle or available for use by other computer systems. In some examples, the server device 110 may report resources, including storage capacity or processing resources, in addition to peripheral devices, without making components or entire peripheral devices available. For example, the server device 110 may make available a portion of its CPU processing power while retaining a portion of its own processing power, such as by using virtualization. The list of peripheral devices may include a device identifier (ID), a vendor ID, an addressing specification, a performance specification or requirement, and other information that the peripheral device can be identified for access. For example, the server device 110 may include multiple network ports, and different ports may be dedicated to one or more peripheral devices. In embodiments where multiple peripheral devices may be accessed through a single port, an access request may notify the server device 110 which peripheral device is being accessed, using a device identifier or other information.
[0033] In step 203, based on receiving a list of available peripheral devices from the server device 110 and optionally from discovery operations for additional peripheral devices 153, the dynamic server rebalancing system 130 may compile a data structure for the available peripheral devices. An exemplary data structure is shown in Figure 1 as “Peripheral Device 136”. The peripheral device data structure 136 may identify the type of resource or peripheral device, its specifications, the associated server device 110, addressing information and identifiers, port 133 or network connection 150, status (e.g., whether the resource is currently in use or assigned to the client-server device 110 or one of them), or other details. The peripheral device data structure 136 may be incorporated into the server device data structure 135 or maintained separately. When a resource comes online or becomes unavailable, or when a peripheral device is attached to / detached from the server system 110, the peripheral data structure may be updated by the dynamic server rebalancing system 130. The dynamic server rebalancing system may provide a list of available peripheral devices to the server device 110 connected to the network switch 131 via the network link 150. The list may be provided to the server device 110 automatically, for example, at regular intervals or whenever the list is updated. In some examples, the list may be provided to the server device 110 in response to a request from the server device to access resources. A subset of the total number of peripheral devices in “Peripheral Devices 136” may be provided based on the server device 110 requesting only specific types of devices. The PoF system 115 of the computer device 110 may then make a selection or request for peripheral devices shown to the dynamic server rebalancing system 130 via the network interface 111 (204). This selection may indicate an identification or identifier of a peripheral device that the server device 110 wishes to connect to via the network link 150, such as a device ID or vendor ID used to identify peripheral devices on the local peripheral device interconnection interface 112.By requesting the use of peripheral devices, server device 110 may be considered a “client” server or computer device for those peripheral devices, while the computer device on which the requested peripheral devices are physically located may be referred to as a “host” server or computer device. In some examples, a single server device 110 may simultaneously be a “host” for several peripheral devices and a “client” for other peripheral devices. For the purposes of the illustrative explanation in Figure 2, computer device 110 may be a “client” server, while computer device 140 may be a “host” server.
[0034] In 205, in response to the client server 110's selection of a peripheral device, the dynamic server rebalancing system 130 may provide one or more commands to instruct the PoF system 115 to remove the peripheral device from the available device pool (e.g., from “Peripheral Devices 136”) and to attach the peripheral device to the client system 110. The dynamic server rebalancing system 130 may also send an instruction to the host system 140 indicating that the peripheral device is occupied or unavailable while it is attached to the client system 110.
[0035] Considering peripheral device 147A of host server 140 as an exemplary selected peripheral device, in response to an instruction issued by the dynamic server rebalancing system 130, the PoF system 115 performs a hot-plug or install process to instantiate a virtual or emulated version of peripheral device 147A, shown as 147B in Figure 1, into the local peripheral device interconnect interface 112 of client device 110 via a virtual connection 116 (206). When the hardware device is not physically present in the slot or connector of client device 110, this hot-plug or install process may include emulating the physical presence of the new hardware device in the slot or connector of client device 110 in the local peripheral device interconnect interface 112 and triggering an initiation process. The PoF system 115 may invoke the initiation process in the local peripheral device interconnect interface 112 without using the corresponding hardware interrupt or hardware indicator that would normally occur from the physical plugging in or powering on of the peripheral device to client device 110. This may also involve modifications to the local peripheral device interconnect interface 112 to accept interrupts or attach commands from the PoF system 115, not just hardware changes. The PoF system 115 also emulates the hardware plugging process by providing an address space description to the local peripheral device interconnect interface 112 for accessing the peripheral device 147B as a local peripheral device coupled to the client system 110. In an example where the local peripheral device interconnect interface 112 has a PCIe subsystem for the client device 110, the PoF system 115 may issue an "attach" command that points the vendor ID and device ID to the PCIe subsystem. The PCIe subsystem may also request or provide a memory address location for memory-mapped access to the peripheral device 147B.The PoF system 115 can further emulate the behavior of the physical peripheral device 147A as a virtual peripheral device 147B by emulating these addressable locations and the structural properties of such locations (e.g., buffer sizing and formatting).
[0036] When instantiated on the local peripheral device interconnect interface 112 of the client device 110, the PoF system 115 can emulate the peripheral device 147B on the client server 110 (207). The device driver of the client device 110 can interface with the peripheral device 147B via the local peripheral device interconnect interface 112 to provide access to operating system processes, user applications, kernel resources, and various other interfaces. Any tools and libraries associated with the peripheral device driver function similarly for local peripheral devices physically coupled to the client device 110, or for peripheral devices mapped via a network link through the PoF system 115. Thus, the device driver for the peripheral device 147B is not normally aware that the actual peripheral device 147A is not locally connected to the client device 110. User applications, operating systems, kernel resources, hypervisors, etc., all interface with the peripheral device 147B as they would normally do when locally connected, and the behavior of the local hardware interface of the peripheral device 147B can be emulated on the local peripheral device interconnect interface 112 via the PoF system 115. This emulation may include the above behavior for instantiation and for handling subsequent communication between the local peripheral device interconnect interface 112 and the network interface 111. This communication may include configuration traffic, command and control handshakes, input / output traffic, and read / write traffic or data transfer between the client device 110 and the peripheral device 147A located on the host device 140. Thus, the PoF system 115 acts as an interaction unit for the traffic of peripheral devices 147A / B between the local peripheral device interconnect interface 112 and the network interface 111.The PoF system 115 interfaces with the network stack of the network interface 111 to send and receive this traffic to and from the actual peripheral device 147A via the network link 150. This may include intercepting client-derived traffic for the peripheral device 147B from the local peripheral device interconnect interface 112 and interpreting the client-derived traffic to convert or rebundle it from its native format (e.g., PCIe frames from the local peripheral device interconnect interface 112 or similar) into an IP packet format or Ethernet frame format suitable for forwarding via the network interface 111. The PoF system 115 then routes the host-derived traffic in packet format via the network interface 111 for distribution to the host server 140 and the physical peripheral device 147A via link 150.
[0037] The dynamic server rebalancing system 130 can receive traffic originating from client devices via link 150. Since network link 150 is coupled to client device 110 and network link 151 is coupled to host device 140, interaction operation between the two port connections 133 is established. If the peripheral device being accessed is between peripheral devices 153 rather than host 140, the dynamic server rebalancing system 130 provides interaction at least between network link 150 and network link 151, or in some examples, between network link 150 and fabric link 152 (208). The connection between network link 150 and fabric link 152 may include protocol conversion, which removes various network protocol-specific headers from network frames or IP packets and then repackages or encapsulates the payload data into frames or packets (or other fabric-native data crams) suitable for transmission over fabric link 152. Interaction traffic between network links of the same type, such as links 150-151, may not require format conversion. Various read, write, input / output, control, command, or other traffic may be processed in this manner to transfer transactions originating on the client device 110 to the peripheral device 147A on the host device 140. Similarly, the reverse operation may occur for transactions and responses originating from the host device 140 and peripheral device 147A for transfer to the client device 110.
[0038] Commands and other traffic sent from client 110 to host 140 may be received and processed at network interface 141. The PoF system 145 communicates with the network stack of network interface 141 and interprets these received network frames or packets to convert them into a native format suitable for the local peripheral device interconnect interface 142. The local peripheral device interconnect interface 142 provides the peripheral device 147A with communication in its native format for execution or processing. The processing results may be intercepted or passed from the peripheral device interconnect interface 142 to the PoF system 145, and then passed back to network interface 141 for transmission to client device 110 via network link 151 and dynamic server rebalancing system 130, similar to the transmission process described above for host device 110.
[0039] When received by the client system 110 via link 150 and network interface 111, these network frames or packets are processed by the network stack of network interface 111. The PoF system 115 communicates with the network stack of network interface 111 and interprets these received network frames or packets to convert them into a native format suitable for the local peripheral device interconnect interface 112, as if they were sent from an emulated peripheral device 147B. The local peripheral device interconnect interface 112 provides communication in native format to various software elements of the client device 110, such as device drivers that interface with user-level applications. Thus, the client device 110 can use the peripheral device 147A via the emulated peripheral device 148B provided by PoF 115 as if it were locally coupled to the local bus or connector of the client device 110.
[0040] In 209, when peripheral device 147A is no longer needed by client device 110, PoF system 115 may operate to remove peripheral device 147B from client device 110 and send a disconnection instruction to dynamic server rebalancing system 130 via link 150. In some examples, host device 140 may send an instruction to dynamic server rebalancing system 130 indicating that peripheral device 147A is no longer available, for example, based on the fact that host device 140 needs the peripheral device or that it is unavailable due to shutdown, forced termination, or other operation. In such a case, dynamic server rebalancing system 130 may send an instruction to client device 110 to disconnect peripheral device 147B. PoF system 115 may bring about the removal of the instance of peripheral device 147B from local peripheral device interconnect interface 112 by at least emulating a hardware device removal process or a "hot unplug" operation to local peripheral device interconnect interface 112. This hot unplug or detach process involves emulating in the local peripheral device interconnect interface 112 that the hardware device is no longer physically present in the slot or connector of the client device 110, and triggering a termination process. The PoF system 115 invokes the termination process in the local peripheral device interconnect interface 112 without using the corresponding hardware interrupt or hardware indicator that would normally result from the physical unplugging or power-off of the peripheral device to the client device 110. This may also involve modifications to the local peripheral device interconnect interface 112 to accept interrupts or detach commands from the PoF system 115, not just from hardware changes. For peripheral device 147B in the host device 110, the established address space description may be decomposed or removed in the local peripheral device interconnect interface 112.In another example, the "detach" command may be issued by the PoF system 115, which indicates the PCIe vendor ID and PCIe device ID of peripheral device 147B to the local peripheral device interconnect interface 112.
[0041] In step 210, the dynamic server rebalancing system 130 may update the data structure 136 to indicate that peripheral device 147A is no longer assigned to or attached to client system 110, and return peripheral device 147A to the available peripheral device pool (depending on the situation, move it to an unassigned, unavailable pool, or reassign it to host system 140). Peripheral device 147A may be returned to the available pool based on an instruction from client system 110 that peripheral device 147B has been disconnected, or the dynamic server rebalancing system 130 may disconnect peripheral device 147A from client 110 without requiring confirmation from client 110.
[0042] As described above, when peripheral device 147A is removed, it may be returned to the pool of peripheral devices, remaining in an inactive, removed state until required by the host device or another device. The installation process can then proceed as described above. A further description of the components or pool of peripheral devices is shown in Figure 5. Furthermore, other fabric-coupled devices, such as FPGAs, SSDs, NICs, memory devices, user interface devices, or other peripheral devices, may be installed / removed instead of the GPU or a device similar to peripheral device 147A, as described in the operation of Figure 2.
[0043] As described above, with respect to peripheral device 147A, computer device 140 may be the host and computer device 110 may be the client, but for other peripheral devices, the relationship may be reversed. For example, while computer device 110 is remotely accessing peripheral device 147A, computer device 140 may simultaneously remotely access peripheral device 148A from computer device 110 via the emulated peripheral device 148B shown in Figure 1.
[0044] In some exemplary embodiments, computer device 110 may have local access to its own peripheral devices 113 via peripheral device interconnection interface 112, and access only remote peripheral devices 143,153 via link 150 and dynamic server rebalancing system 130. In these cases, only available or idle peripheral devices or resources may be reported to the dynamic server rebalancing system 130, while peripheral devices 113 used by the computer system itself may not be reported as available. However, in some examples, a server 110 capable of dynamic server rebalancing (e.g., via PoF system 115) may report all local peripheral devices 113 to the dynamic server rebalancing system 130. In such embodiments, the computer system may not have direct access to its local peripheral devices, but instead may request the use of peripheral devices from an available pool of peripheral devices via the dynamic server rebalancing system 130 for all peripheral device needs. In this way, computer system 110 may function as both a host and a client to itself and access its own local peripheral devices "remotely". Such an implementation may simplify resource rebalancing between computer devices 110 without local resource contention between the host device and the remote computer system. For example, if a local peripheral device 148A of host 110 is idle and attached to a client system 140, and host 110 needs the peripheral device 148A, this could cause a contention or interruption of operation in client system 140. Such contention can be avoided by having all systems utilize the same shared pool of resources.
[0045] To illustrate the detailed structure and operation of computer device 110, Figure 3 is presented. While the elements of Figure 3 may apply to computer device 110 in Figure 1, it should be understood that computer device 110 may utilize other structures and operations. Figure 3 includes computer device 300 as an example of a computer device, server, computer system, blade server, etc. Computer device 300 may include one or more network links, each having its own associated network address. For example, as described herein, a first network link 380 may connect computer device 300 to a dynamic server rebalancing system or network switch for coupling remote peripheral devices. Another network link 381 may communicate with other computer devices, further networks, or the Internet, among various other networks and endpoints, for transmissions unrelated to remote peripheral data traffic. The example described with respect to Figure 3 also utilizes TCP / IP-style networking for communication with computer device 300 and PCIe fabric for communication with local and emulated or networked peripheral devices.
[0046] Computer device 300 includes user space 302 and kernel space 301. Kernel space 301 may be a software system comprising core operating system (OS) elements, such as the OS kernel, device drivers, hardware interface subsystems, network stack, memory management subsystem, mechanical clock / time module, and other low-level elements used to manage the resources of computer device 300 between user-level and kernel-level software, in addition to acting as an interface between hardware components and user applications. User space 302 may include user applications, tools, games, graphical or command-line user interface elements, and other similar elements. Typically, user space elements interface with device driver elements in kernel space 301 through application programming interfaces (APIs) or other software-defined interfaces to share access to low-level hardware elements among all user software, such as network controllers, graphics cards, audio devices, video devices, user interface hardware, and various communication interfaces. These device driver elements receive user-level traffic and interact with hardware elements that ultimately drive link-layer communication, data transfer, data processing, logic, or other low-level functions.
[0047] Within kernel space 301, computer device 300 may include a network stack 330, a peripheral overfabricue (PoF) unit 320, a PCI / PCIe module 340, and device drivers 350 for connected peripheral devices, etc. Other kernel space elements are omitted for clarity and to focus on kernel-level elements relevant to the operation described herein. User space 302 includes user commands 360 and user applications 361. The network stack 330 comprises a TCP / IP stack and includes various layers or modules typical of a network stack, although some elements are omitted for clarity. The Ethernet driver 334 includes functions for link layer, media access controller (MAC) addressing, Ethernet frame processing, and interfacing with a network interface controller (not shown) that handles physical layer operation and structure. The IP module 333 performs packet processing, IP addressing, and internetworking operations. The TCP / UDP module 332 interfaces between the user application's data structure and the IP module 333, and also packets user data, handles error correction and retransmission, forwarding acknowledgments, etc. The socket layer 331 interfaces with the user application and other components of the computer device 300 and acts as an endpoint for packetized communication. Individual sockets can be established, each handling a specific communication purpose, type, protocol, or other communication isolation. Several sockets can be established by the network stack 330, each of which can act as an endpoint for a different type of communication. In the case of TCP / UDP, sockets are typically identified by their IP address and port number, and a host device may have a single IP address as well as many such port numbers for multiple IP addresses, each having its own set of port numbers. Thus, many sockets may be established, each with a specific purpose.User-level applications, user processes, or even kernel-level processes, modules, and elements may interface with the network stack 330 through specific sockets.
[0048] During operation, in response to attach / detach commands forwarded by the server rebalancing / control entity (and directed to the socket described above), the PoF unit 320 can establish the functionality of a remote peripheral device as if it were a local peripheral device of computer device 300 by invoking the hot-plug / unplug functionality of the PCI / PCIe module 340 and emulating the hardware behavior for these functions. The PoF unit 320 interfaces with socket layer 331 to forward and receive packets carrying traffic related to peripheral devices that may be located remotely from computer device 300. Instead of directly interfaced with socket layer 331, the PoF unit 320 may use a TCP Offload Engine (TOE) stack and Remote Direct Memory Access (RDMA) for a particular network interface controller vendor type. Socket layer 331, or the equivalent described above, is identified by its IP address and port number and is typically dedicated to traffic related to a specific peripheral device for computer device 300 or all remote peripheral devices. Usernames / passwords or other security credentials may be passed along with packets received by socket layer 331. The PoF unit 320 has "hooks" or software interface functionality for communicating with the socket layer 331. Packets arrive from peripheral devices through the network stack 330, are interpreted by the PoF unit 320, and then converted into a format suitable for the PCI / PCIe module 340. Packets received by the interaction unit 320 may contain PCIe device state information of the peripheral device. The PCI / PCIe module 340 receives these communications from the PoF unit 320 as if they originated from a local peripheral device of the computer device 300. Thus, the PoF unit 320 emulates the behavior of a local peripheral device to the PCI / PCIe module 340, and such a peripheral device appears local to the PCI / PCIe module 340.Device drivers, device tools or toolsets, and device-centric libraries function similarly for locally connected PCIe devices or remote PCIe devices mapped through the PoF unit 320. To achieve this emulation of local devices, the PoF unit 320 may establish several functions or libraries that present the PCI / PCIe module 340 with goals for communication for I / O transactions, configuration transactions, read / write, and various other communications.
[0049] Advantageously, user applications can interact with peripheral devices located remotely from the computer device 300 using a standard device driver 350 that interfaces with the PCI / PCIe module 340. Communications issued by the PCI / PCIe module 340, which are normally intended for local hardware devices, are intercepted by the PoF unit 320 and interpreted for transmission over the network stack 330 and network link 380. When GPU peripheral devices are utilized, for example, even if the GPU is located remotely with respect to the computer device 300, graphics drivers can be utilized without modification by the user application, such as machine learning, deep learning, artificial intelligence, or game applications.
[0050] To discuss the detailed structure and operation of the dynamic server rebalancing system, Figure 4 is presented. While the elements of Figure 4 may apply to the dynamic server rebalancing system 130 in Figure 1, it should be understood that the dynamic server rebalancing system 130 may utilize other structures and operations. Figure 4 includes a dynamic server rebalancing system 400 as an example, comprising a management node, computer devices, servers, computer systems, and blade servers. The dynamic server rebalancing system 400 can establish and manage a communication fabric for sharing resources, such as peripheral devices, between connected computer nodes, such as the computer device 110 and peripheral device pool 153 in Figure 1. The dynamic server rebalancing system 400 may include one or more network links, each having an associated network address. For example, as described herein, the first network link 480 may be coupled to a network switch and then to a computer device or server device for coupling peripheral devices. The examples described with respect to Figure 4 also utilize TCP / IP-style networking for communicating with the dynamic server rebalancing system 400, and in some examples, a PCIe interface for communicating with the peripheral device pool.
[0051] The dynamic server rebalancing system 400 may include a user space 402 and a kernel space 401. The kernel space 401 may be a software system that includes core operating system (OS) elements, such as the OS kernel, device drivers, hardware interface subsystems, network stack, memory management subsystem, mechanical clock / time module, and other low-level elements used to manage the resources of the dynamic server rebalancing system 400 between user-level and kernel-level software, in addition to acting as an interface between hardware components and user applications. The user space 402 may include user applications, tools, telemetry, event handlers, user interface elements, and other similar elements. Typically, elements of the user space 402 interface with device driver elements of the kernel space 401 through application programming interfaces (APIs) or other software-defined interfaces to share access to low-level hardware elements among all user software, such as network controllers, fabric interfaces, sideband communication / control interfaces, maintenance interfaces, user interface hardware, and various communication interfaces. These device driver elements receive user-level traffic and interact with hardware elements that ultimately drive link-layer communication, data transfer, data processing, logic, or other low-level functions.
[0052] Within kernel space 401, the dynamic server rebalancing system 400 may include a network stack 430, a fabric module 440, and a PCI / PCIe interface 460. Other kernel space elements are omitted for clarity and to focus on kernel-level elements relevant to the operation described herein. The network stack 430 comprises a TCP / IP stack and includes various layers or modules typical of a network stack, although some elements are omitted for clarity. The Ethernet driver 434 includes functions for link layer, media access controller (MAC) addressing, Ethernet frame processing, and interface with a network interface controller (not shown) that handles physical layer operation and structure. The IP module 433 performs packet processing, IP addressing, and internetwork operations. The TCP / UDP module 432 interfaces between the data structures of the user application and the IP module 433, and also packets user data, handles error correction and retransmission, forwarding acknowledgments, etc. The socket layer 431 interfaces with user applications and other components of the dynamic server rebalancing system 400 and acts as an endpoint for packetized communication. Individual sockets can be established, each handling a specific communication purpose, type, protocol, connected device, or other communication isolation. Several sockets can be established by the network stack, each of which may act as an endpoint for a different communication type or communication link. In the case of TCP / UDP, sockets are typically identified by their IP address and port number, and a device may have many such port numbers for multiple IP addresses, each having its own set of port numbers, in addition to a single IP address. Thus, many sockets may be established, each with a specific purpose. User-level applications, user processes, or even kernel-level processes, modules, and elements may interface with the network stack through specific sockets.
[0053] The fabric module 440 may include drivers and other elements for managing resource sharing or rebalancing between computer devices connected to the dynamic server rebalancing system 400. The fabric module 440 provides pathways for commands and controls of the fabric itself, such as for logical partitioning / isolation or attachment / detachment of peripheral devices. Traffic related to reading, writing, configuring, and I / O of peripheral devices may also be handled by the fabric module 440. In some examples, the fabric module 440 may include a PCI / PCIe subsystem, including a protocol stack equivalent for PCI / PCIe links. The fabric module 440 can interface with physical layer elements, such as the PCI / PCIe interface 460, and also provides a software / programming interface for the configuration handler 412. Furthermore, the fabric module 440 may interface user-space elements (e.g., a command processor 414) with the PCI / PCIe interface 460. The PCI / PCIe interface 460 may include a fabric chip or fabric switch circuit, which may provide one or more physical fabric links to the fabric module 440, the local devices of the dynamic server rebalancing system 400, and a pool of peripheral devices coupled via associated PCIe links (e.g., peripheral device 153 connected via link 152 in Figure 1), or to further fabric chips or fabric switch circuits that provide part of the fabric and further fabric links.
[0054] User space 402 includes a server rebalancing control element 410, which may further comprise a monitor 411, a configuration handler 412, an event handler 413, a command processor 414, and a user interface 415. The command processor 414 communicates with the fabric module 440 to control a communication fabric used to establish logical partitioning or allocation between peripheral devices coupled to the fabric, and may provide routing for communication / traffic to and from selected peripheral devices. When a selected peripheral device is removed, the peripheral device may be placed into a pool of unused devices. The user interface 415 may receive operator commands to manage the fabric or to control the addition or removal of peripheral devices to and from computer devices. The user interface 415 may display or indicate a list of computer devices and peripheral devices, along with their associated status or telemetry. The user interface 415 may display or indicate which peripheral devices are associated with which computer devices. The user interface 415 may display or indicate traffic histograms, logs, faults, alerts, and various other telemetry and statuses. The user interface 415 may include terminal interfaces, application programming interfaces (APIs), REST (representational state transfer) interfaces, or RestAPI, web interfaces, or WebSocket interfaces, among other types of user interfaces, including those transmitted via software, hardware, virtualization, or various intermediate links. The event handler 413 may initiate attach / detach and device discovery operations on computer devices. The configuration handler 412 facilitates traffic interaction between network-connected computer devices and optionally between sets of PCI / PCIe-connected peripheral devices from a peripheral device pool. The configuration handler 412 interfaces with the fabric module 440 for fabric communication and the network stack 430 for network communication.The configuration handler 412 may interact with the format, size, and type of frames or packets to transmit communications over network links and PCI / PCIe links. The configuration handler 412 interfaces with the network stack 430 through the socket layer 431 via a specific socket indicated by at least an IP address and port number. The monitor 411 may monitor various telemetry, operations, logs, and status for the dynamic server rebalancing system 400. In addition to indicating computer devices and associated sockets (IP addresses and ports), the monitor 411 may maintain data structures that indicate indicators, addresses, or identities of peripheral devices. The monitor 414 may maintain logs and data structures in computer-readable media, such as data structures 435 and 436 in a memory device 465 locally connected to the PCI / PCIe interface 460 of the dynamic server rebalancing system 400.
[0055] During operation, the event handler 413 may discover compatible computer devices coupled via the network interface and initiate operations to discover peripheral devices coupled to the communication fabric. The event handler 413 may instruct the configuration handler 412 to discover computer devices through the network stack 430. Socket information of compatible computer devices may be determined and stored, for example, in a data structure 435, for later use. In some examples, the event handler 413 may instruct the command processor 414 to discover peripheral devices via the fabric module 440 and the PCI / PCIe interface 460. The PCI / PCIe interface 460 scans the communication fabric to determine which peripheral devices and resources are available. The command processor 414 forms a pool of free peripheral devices and instructions for assigned peripheral devices, stores the device / vendor IDs of the peripheral devices, and may store instructions for the PCIe addressing and buffer structure or characteristics of each peripheral device, for example, in a data structure 436.
[0056] A request to add a peripheral device to a computer device is received, such as via the user interface 415 or via the network stack 430, and the command processor 414 may isolate or attach the selected peripheral device to a logical partition of the communication fabric via the fabric module 440. This may trigger an event handler 413 to begin notifying the host device where the selected peripheral device is physically located that the selected peripheral device has been assigned to the remote client computer device. The event handler 413 may also issue an attach command to the client device along with peripheral device information (e.g., vendor ID and device ID), which the client device then attaches as described herein. From here, communication between the client device and the peripheral device on the host device interacts using a configuration handler 412 that interprets and exchanges traffic between socket layers 431 via the fabric module 440. In some examples, the event handler 413 may connect the requesting client device with a peripheral device not associated with a host computer device via the fabric module 440 and the PCI / PCIe interface 460. At some point, it may be desirable for a mounted peripheral device to be removed or detached from a particular client device. The event handler 413 may detect these detachment events, such as those received from the computer device via the user interface 415 or the network stack 430. The event handler 413 then issues a detach command through the configuration handler 412 to detach the affected peripheral device, as described herein. The command processor 414 may remove the logical partition or allocation of the detached peripheral device and return the detached peripheral device to an inactive state or to a free pool of peripheral devices for later use.
[0057] Figure 5 is a system diagram showing computer system 500. Computer system 500 may comprise, but may also comprise, elements of computer system 100 in Figure 1, computer device 300 in Figure 3, or dynamic server rebalancing system 400 in Figure 4. Computer system 500 comprises a rack-mount arrangement of multiple modular chassis. One or more physical enclosures, such as modular chassis, may be further included in shelves or rack units. Chassis 510, 520, 530, 540, and 550 are included in computer system 500 and may be mounted in a common rack-mount arrangement within one or more data centers, or across multiple rack-mount arrangements. Within each chassis, modules or peripheral devices may be mounted on common circuit boards and switch elements, along with various power systems, structural supports, and connector elements. The enclosed modular system may include physical support structures and enclosures containing circuits, printed circuit boards, semiconductor systems, and structural elements. Modules comprising the components of the computer system 500 are insertable into and removable from rack-mount style enclosures. In some examples, the components in Figure 5 are contained in a "U" style chassis for mounting within a larger rack-mount environment. It should be understood that the components in Figure 5 can be included in any physical mounting environment and do not necessarily require associated enclosures or rack-mount elements.
[0058] The chassis 510 may include a management module or a top-of-rack (ToR) switch chassis, such as the dynamic server rebalancing systems 130, 300 in Figure 1 and 400 in Figure 4, and may include a management processor 511, an Ethernet switch 516, and a PCIe switch 560. The management processor 511 may include a management operating system (OS) 512, a user interface 513, and an interaction unit 514. The management processor 511 may be coupled to the Ethernet switch 516 via one or more network links through a network interface controller. The management processor 511 may be coupled to the PCIe switch 560 via one or more PCIe links having one or more PCIe lanes.
[0059] The Ethernet switch 516 may include network ports that provide switched network connectivity to the attached devices, as shown for network link 566. In an exemplary embodiment, network link 566 may connect the dynamic server rebalancing system chassis 510 to the blade server motherboards 561-563 of chassis 520, 530, and 540. The Ethernet switch 516 includes various network switching circuits for linking individual ports to other ports in a communicative manner based on traffic patterns, addressing, or other traffic properties. For example, the Ethernet switch 516 comprises an Ethernet or Wi-Fi (802.11xx) switch that hosts wired or wireless connections, which may refer to any of the various available network communication protocol standards and bandwidths, such as 10BASE-T, 100BASE-TX, 1000BASE-T, 10GBASE-T (10GB Ethernet), 40GBASE-T (40GB Ethernet), Gigabit (GbE), Terabit (TbE), 200GbE, 400GbE, 800GbE, or various other wired and wireless formats and speeds. The PCIe switch 560 may be coupled to a PCIe switch 564 in the chassis 550 via one or more PCIe links. These one or more PCIe links are represented by PCIe module-to-module connections 565.
[0060] The components of network link 566, PCIe link 565, and the dynamic server rebalancing system chassis 510 form a fabric that communicatively connects all of the various physical computer elements in Figure 5. In some examples, a management processor 511 may communicate with the elements of the fabric via a special management PCIe link or sideband signaling (not shown), such as an inter-integrated circuit (I2C) interface, to control the operation, partitioning, and mounting / removal of the elements of the fabric. These control operations may include configuring and disassembling peripheral devices, changing logical partitions within the fabric, monitoring fabric telemetry, controlling power-up / down operations of modules on the fabric, updating firmware for various circuits comprising the fabric, and other operations.
[0061] Chassis 520, 530, and 540 may contain a blade server computer system (such as computer system 110 in Figure 1). Chassis 520 may contain a blade server motherboard and Ethernet switch 561 having various components and resources, including a CPU 521, GPU 522, TPU 523, NIC 524, and SSD 525, but may contain different numbers and arrangements of system components. Similarly, chassis 530 may contain a blade server motherboard 562 and locally connected components 531-535, and chassis 540 may contain a blade server motherboard 563 and locally connected components 541-545. Power systems, monitoring elements, internal / external ports, mount / remove hardware, and other related functions may be included in each chassis. An exemplary component or module of the blade server 520 (e.g., PoF unit 115 in Figure 1 or PoF unit 320 in Figure 3) may generate a list of components 521-525 within the chassis 520 that are available for remote sharing with other computer devices in the system 500, and may provide a list of some or all of these components to a management device in the chassis 510 via link 566. For example, the list may include all components 521-525, along with indications of which are available for sharing and which are not (e.g., because they are being used by the blade server 520), or the list may include only the devices available for sharing.
[0062] The chassis 550 may comprise a deaggregated collection of peripheral devices (such as peripheral device 153 in Figure 1). The chassis 550 may comprise a number of GPUs 551-555, each coupled to a PCIe fabric via a backplane and PCIe switch 564 and associated PCIe links (not shown). In exemplary embodiments, the chassis 550 may comprise a JBOG (Just a Box for GPUs) system, but other arrangements of deaggregated peripheral devices may also be provided. For example, a chassis for a CPU, NIC, or SSD may be presented, or various combinations of deaggregated peripheral devices may be provided within a single chassis. For example, the chassis 550 may include modular bays for mounting modules comprising corresponding elements of each CPU, GPU, SSD, NIC, or other module type. In addition to the component types described above, other component types such as FPGAs, CoPUs, RAM, memory devices, or other components may also be included. Power systems, monitoring elements, internal / external ports, mount / remove hardware, and other related functions may be included in each chassis 550. Further descriptions of the individual elements of chassis 520, 530, 540, and 550 are included below.
[0063] When various CPU, GPU, TPU, SSD, or NIC components of computer system 500 are installed in the associated chassis or enclosure and reported or discovered by the management device 510, the components may be logically assigned or organized into any number of separate, arbitrarily defined arrangements and mounted on computer devices. These arrangements may consist of a selected number of CPUs, GPUs, SSDs, and NICs, including zero of any type of module. In the case of the exemplary computer device 540 shown in Figure 5, a remote peripheral GPU 532 (which may be physically located on the host computer device 530) is mounted. The GPU 532 may be mounted on the client computer device 540 via a network link 566 using a logical partition or assignment within the network fabric indicated by the logical domain 570. A logical arrangement of computer components, including locally connected physical components and remote or logically connected peripheral components, may be referred to as a “computation unit.” The network fabric may be configured by the management processor 511 to selectively route traffic between selected client devices and host devices, and peripheral devices locally attached to host devices, while maintaining logical isolation between components not physically or virtually included in the selected client devices. In this way, a deaggregated and flexible "bare-metal" configuration can be established among the components of the computer system 500. Individual peripheral devices may be positioned according to specific user identification information, computer device identification information, execution jobs, or usage policies.
[0064] In some examples, the management processor 511 may provide mounting or detaching of peripheral and host devices via one or more user interfaces or job interfaces. For example, the management processor 511 may provide a user interface 513 that can present instructions for available peripheral components to be mounted, instructions for available computer devices, and software and configuration information. In some examples, the creation user interface 513 may provide templates for mounting predetermined arrangements of peripheral devices to computer devices based on use cases or use categories. For example, the user interface 513 may provide proposed templates or configurations for a game server unit, an artificial intelligence learning and computing unit, a data analysis unit, and a storage server unit. For example, the game server unit or artificial intelligence processing template may specify additional graphics processing resources compared to the storage server unit template. Furthermore, the user interface 513 may provide customization of templates or placement configurations and options for the user to create placement templates from component types arbitrarily selected from a list or category of components.
[0065] In additional examples, the management processor 511 may provide policy-based dynamic adjustments to operational deployments. In some examples, the user interface 513 allows the user to define policies for adjustments to the operational configuration information of a computer device, in addition to adjustments to peripheral devices assigned to the computer device. In one example, during operation, the management processor 511 may analyze telemetry data to determine the current resource utilization by the computer device. Based on current utilization, the dynamic adjustment policy may specify allocating or removing general processing resources, graphics processing resources, storage resources, networking resources, memory resources, etc., from the host device. For example, if telemetry data indicates that the current usage level of allocated storage resources on a computer device is approaching a threshold level, additional storage devices may be allocated to the host device.
[0066] The management processor 511 may provide control and management of multiple protocol communication fabrics, including combining different communication protocols such as PCIe and Ethernet. For example, the management processor 511, and the devices connected via links 566 and 565, may provide communication coupling of physical components using multiple different implementations or versions of Ethernet, PCIe, and similar protocols. Furthermore, next-generation interfaces, such as Gen-Z, CCIX, CXL, OpenCAPI, or wireless interfaces including Wi-Fi or cellular wireless interfaces, may be utilized. Also, while Ethernet and PCIe are used in Figure 5, it should be understood that other interconnects, networks, and link interfaces, including different or additional communication links or buses such as NVMe, SAS, FibreChannel, Thunderbolt, and SATA Express, may be used instead.
[0067] Referring here to the description of the components of the computer system 500, the management processor 511 may comprise one or more microprocessors and other processing circuits that retrieve and execute software for managing the operating system 512, user interface 513, and interaction unit 514, or components or modules, or any combination thereof, from the associated storage system. The management processor 511 may be implemented within a single processing device, or it may be distributed across multiple processing devices or subsystems that cooperate in executing program instructions. Examples of the management processor 511 include general-purpose central processing units, application-specific processors, and logic devices, as well as any other types of processing devices, combinations thereof, or variations thereof. In some examples, the management processor 511 comprises Intel®, AMD®, Apple®, ARM®, field-programmable gate arrays (FPGAs), application-specific integrated circuits (ASICs), application-specific processors, or other microprocessors or processing elements.
[0068] The management operating system (OS) 512 is executed by the management processor 511 and provides the management of the resources of the computer system 500, as well as the execution of the user interface 513 and the interaction unit 514. The management OS 512 provides the functions and operations described herein for the management processor 511, specifically for the user interface 513 and the interaction unit 514.
[0069] The user interface 513 may present one or more users with a graphical user interface (GUI), an application programming interface (API), a REST interface, a RestAPI, a command-line interface (CLI), a WebSocket interface, or other interfaces. The user interface 513 may be used by end users or administrators to configure computer devices and peripheral devices, establish the placement of computer devices and peripheral devices, receive status, telemetry, and other notifications regarding the operation of computer devices and peripheral devices, and perform other actions. The user interface 513 may be used to manage, select, and modify templates. The user interface 513 may be used to manage, select, and modify policies. The user interface 513 may also provide users with telemetry information for the operation of the computer system 500 in one or more status interfaces or status views. The status of various components or elements of the computer system 500, including, among others, CPU status, GPU status, NIC status, SSD status, Ethernet status, PCIe switch / fabric status, may be monitored via the user interface 513. Various performance metrics and error statuses can be monitored using the user interface 513.
[0070] The interaction unit 514 provides various fabric interaction functions and operations described herein, along with the discovery, attachment, and detachment of peripheral devices related to computer devices. Specifically, the interaction unit 514 may discover computer devices and peripheral devices in a computer system 500 coupled via communication links (e.g., 565-566), receive instructions for available peripheral devices located on computer devices, and catalog these devices in one or more data structures 590. The data structures 590 are represented by exemplary server device data structure 591 and peripheral device data structure 592. Cataloging may include taking note of device identifiers, vendor identifiers, addresses, ports, sockets, assignments or attachments between devices, or other properties. The interaction unit 514 may receive operator instructions via a user interface 513 from computer devices to establish arrangements between computer devices and one or more peripheral devices (e.g., attaching peripheral devices from a host device to be emulated in a client device). Instructions may identify which computer devices and which peripheral devices, or what types of peripheral devices, should be coupled. In response to the command, the interaction unit 514 initiates the attachment of one or more peripheral devices from the available device pool to the client computer device, for example, by issuing one or more attach commands to the client device's PoF unit. This attach command alerts the client device's PoF unit to initiate the instantiation of the peripheral device to the client device's local peripheral device interconnect interface by at least emulating the behavior of a peripheral device connected via a network interface as a local peripheral device connected to the client system's peripheral device interconnect interface. The interaction unit 514 may also notify the host device of the linked peripheral device that the peripheral device has been assigned to or attached to the client device.The interaction unit 514 can then interact with traffic between the client system and the host system along the network link 566, and in some examples, between the network link 566 and the PCIe link 565. The interaction unit 514 can receive commands to detach a peripheral device from a client device and issue a detach command to the client device's PoF unit. Once detached, the peripheral device can be returned to the pool of available devices for later use by other computer devices.
[0071] An example of the installation operation is shown in Figure 5. The interaction unit 514 can receive commands to attach the GPU 532 (physically located on computer device 530) to computer device 540, and can provide the attach command and identifier of the GPU 532 to computer device 540 via the corresponding network link 566. Once attached, computer device 540 may become a client device for the GPU 532, and computer device 530 may be a host device. Table 591 shows the target computer device identifier 001236 (corresponding to client device 540) with the IP address 10.10.10.3, and the corresponding socket (not shown) with the port number. Table 592 shows the GPU 532 with the vendor ID and device ID provided to the target client device 540 for installation. To enable proper exchange of communication between the GPU 532 and the target client device 540 via the communication fabric without interference from other peripheral and computer devices, logical isolation of the GPU 532 can be established between the GPU 532 and the management processor 511 within the communication fabric. This arrangement can be referred to as a logical domain 570, which includes elements of a communication link (e.g., Ethernet link 566) between the host 530, the client 540, the fabric switch circuit 510, the fabric switch circuit itself, and the GPU 532. Communication can be directed from the client device 540 to the GPU 532 in the host device 530 via the network link 566, using the device ID and vendor ID of the GPU 532, which are intercepted and interpreted by the interaction unit 514 using the IP address 10.10.10.2 and associated port number of the socket corresponding to the host device 530. Communication from GPU 532 to client device 540 can be directed via network link 566 using the IP address of client device 540, i.e., 10.10.10.3, so that it may be intercepted and interpreted by interaction unit 514.
[0072] Multiple instances of elements 511-514 (e.g., two or more dynamic server rebalancing systems) may be included in the computer system 500. User commands, such as those received via a GUI, may be received by one of the management instances and forwarded by the receiving management instance to a handling management instance. Each management instance may have a unique or pre-assigned identifier that can assist in the delivery of user commands to the appropriate management instance. Furthermore, the management processors of each management instance may communicate with each other, for example, by using a mailbox process or other data exchange technology. This communication may take place via a dedicated sideband interface such as an I2C interface, or via a PCIe or Ethernet interface connecting each management processor.
[0073] Multiple CPUs 521, 531, and 541 are included in system 500. Each CPU may comprise a CPU module including one or more CPUs or microprocessors and other processing circuits that retrieve and execute software, such as operating systems, device drivers, and applications, from the associated storage system. Each CPU may be implemented within a single processing device, but may be distributed across multiple processing devices or subsystems that cooperate in executing program instructions. Examples of each CPU include general-purpose central processing units, application-specific processors, and logic devices, as well as any other types of processing devices, combinations thereof, or variations thereof. In some examples, each CPU comprises an Intel® microprocessor, an Apple® microprocessor, an AMD® microprocessor, an ARM® microprocessor, a graphics processor, a compute core, a graphics core, an ASIC, an FPGA, or other microprocessor or processing element. Each CPU may also communicate with other CPUs, such as CPUs in the same storage assembly / enclosure or a different storage assembly / enclosure, via one or more PCIe interfaces and PCIe fabrics.
[0074] System 500 includes multiple GPUs 522, 532, 542, 551-555 and TPUs 523, 533, 543, which may represent any type of CoCPU. Each GPU may comprise a GPU module containing one or more GPUs. Each GPU includes graphics processing resources that can be allocated to one or more host devices. A GPU may comprise a graphics processor, shaders, pixel rendering elements, frame buffers, texture mappers, graphics cores, graphics pipelines, graphics memory, or other graphics processing and processing elements. In some examples, each GPU comprises a graphics "card" comprising circuitry supporting the GPU chip. Exemplary GPU cards include nVIDIA® or AMD® graphics cards that include graphics processing elements along with various support circuits, connectors, and other elements. Similarly, each Tensor Processing Unit (TPU) may comprise a TPU module containing one or more TPUs. Each TPU may include circuitry and resources for AI acceleration and processing configured for neural network machine learning applications. Further examples may include other styles of graphics processing units, graphics processing assemblies, or coprocessing elements, such as machine learning processing units, AI accelerators, FPGAs, ASICs, or other specialized processors that may include specialized processing elements for concentrating processing and memory resources for processing specific datasets.
[0075] Multiple NICs 524, 534, and 544 are included within the system 500, each having an associated MAC address or Ethernet address. Each NIC may comprise a NIC module containing one or more NICs. Each NIC may include a network interface controller card for communication over a TCP / IP (Transmission Control Protocol (TCP) / Internet Protocol) network or for carrying user traffic, such as iSCSI (Internet Small Computer System Interface) or NVMe (NVM Express) traffic for elements of the associated host device. The NICs are equipped with Ethernet interface equipment and can communicate over wired, optical, or wireless links. External access to components of the computer system 500 may be provided via packet network links provided by the NICs. The NICs may communicate with other components of the associated host device via associated PCIe links in the PCIe fabric. In some examples, the NICs are provided for communication with a management processor 511 via an Ethernet link. In additional examples, the NICs are provided for communication with one or more other chassis, rack-mount systems, data centers, computer platforms, communication fabrics, or other elements via Ethernet links.
[0076] Multiple SSDs 525, 535, and 545 are included in System 500. Each SSD may comprise an SSD module containing one or more SSDs. Each SSD includes one or more storage drives, such as solid-state storage drives having a PCIe interface. Each SSD also includes a PCIe interface, a control processor, and power system elements. Each SSD may also include a processor or control system for traffic statistics and status monitoring, among other operations. In yet another example, each SSD may instead comprise a different data storage medium, such as a magnetic hard disk drive (HDD), crosspoint memory (e.g., Optane® device), static random access memory (SRAM) device, programmable read-only memory (PROM) device, or other magnetic, optical, or semiconductor-based storage medium, along with the associated enclosure, control system, power system, and interface circuitry.
[0077] In addition to the CPU, GPU, TPU, SSD, and NIC, the computer platform may utilize other dedicated devices. These other dedicated devices may include, among other circuits, dedicated coprocessing circuits, fabric-coupled RAM devices, ASIC circuits, or FPGA circuits, as well as coprocessing modules with various memory components, storage components, and interface components. Each of these other dedicated devices may include a PCIe interface or an Ethernet interface, thereby enabling integration into the network fabric of System 500 for remote access, either directly or by mounting it to a computer device. These other dedicated devices may include PCIe endpoint devices or other computer devices, which may or may not have a root complex.
[0078] FPGA devices can be used as an example of other dedicated devices. FPGA devices can receive processing tasks from other peripheral devices, such as a CPU or GPU, and offload those processing tasks to FPGA programmable logic circuits. FPGAs are typically initialized to a programmed state using configuration data, which includes various logic configurations, memory circuits, registers, processing cores, special circuits, and other functions that provide special or application-specific circuits. FPGA devices can be reprogrammed to perform different sets of processing tasks at different times, in addition to modifying the circuits implemented within them. FPGA devices can be used to perform machine learning tasks, implement artificial neural network circuits, implement custom interfaces or glue logic, perform encryption / decryption tasks, blockchain computation and processing tasks, or other tasks. In some examples, a CPU would provide data to be processed locally or remotely by the FPGA via a PCIe interface. The FPGA could process this data, generate results, and provide these results to the CPU via the PCIe interface. Two or more CPUs and / or FPGAs may be involved to parallelize tasks through two or more devices or to serialize data through two or more devices. In some examples, the FPGA placement may include locally stored configuration data that can be supplemented, replaced, or overwritten using configuration data stored in configuration data storage. This configuration data may include firmware, programmable logic programs, bitstreams, or objects, PCIe device initial configuration data, among other configuration data described herein. The FPGA placement may also include SRAM or PROM devices used to perform boot programming, power-on configuration, or other functions for establishing the initial configuration of the FPGA device. In some examples, the SRAM or PROM devices may be integrated into the FPGA circuit or package.
[0079] The blade server motherboards 561-563 may include printed circuit boards or backplanes on which computer components can be mounted or connected. For example, peripheral devices 521-525, 531-535, and 541-545 may be connected to PCIe ports or other slots on the blade server motherboards 561-563. Each of the blade server motherboards 561-563 may include one or more network switches or ports (not shown) for connecting to a network link 566, such as an Ethernet switch. The blade server motherboards 561-563 communicate with other components of the system 500 via the network link 566, thereby enabling access to remote peripheral devices and allowing external devices to access and utilize resources of local peripheral devices, such as devices 521-525 for chassis 520, devices 531-535 for chassis 530, and devices 541-545 for chassis 540. Blade server motherboards 561-563 can logically interconnect devices of system 500 so as to be managed by management processor 511. Attach or detach commands for remote peripheral devices can be sent or received through the blade server motherboards 561-563 via network link 566, and the blade server motherboards 561-563 can receive a list of available resources from management processor 511 or issue requests to management processor 511 to access remote resources.
[0080] PCIe switch 564 can communicate with other components of system 500 via associated PCIe link 565. In the example in Figure 5, PCIe switch 564 can be used to carry traffic between PCIe devices in chassis 564 and switching units in chassis 510, from which traffic can be directed to other chassis via network link 566. PCIe switch 564 may include PCIe cross-connect switches for establishing switching connections between any PCIe interfaces handled by PCIe switches in system 500. The PCIe switches described herein can logically interconnect various PCIe links among associated PCIe links based on at least the traffic carried by each PCIe link. These examples may include domain-based PCIe signaling distribution that can isolate PCIe ports of PCIe switches according to user-defined groups. User-defined groups may be managed by a management processor 511 that logically integrates components into associated logical units and logically isolates components and logical units from one another. In addition to or alternative to domain-based isolation, each PCIe switch port can be either a non-transparent (NT) or transparent port. NT ports allow for some form of logical isolation between endpoints, similar to a bridge, while transparent ports do not allow for logical isolation and have the effect of connecting endpoints in a purely switched configuration. Access via one or more NT ports may involve an additional handshake between the PCIe switch and the initiating endpoint to select a specific NT port or to enable visibility through the NT port.
[0081] In a further example, memory-mapped direct memory access (DMA) conduits may be formed between individual CPU / PCIe device pairs. This memory mapping can be performed on the PCIe fabric address space, among other configurations. The logical partitions described herein may be used to provide these DMA conduits on a shared PCIe fabric with many CPUs and GPUs. Specifically, NT port or domain-based partitioning on a PCIe switch can isolate individual DMA conduits between related CPUs / GPUs. The PCIe fabric has a 64-bit address space, which allows for 264 bytes of addressable space and provides at least 16 exbibytes of byte-addressable memory. The 64-bit PCIe address space can be shared by all compute units or separated between various compute units to form an arrangement for proper memory mapping to resources.
[0082] The PCIe interface supports multiple bus widths, such as x1, x2, x4, x8, x16, and x32, with each multiple of the bus width potentially having additional "lanes" for data transfer. PCIe also supports the transfer of sideband signaling, including associated clock, power, and bootstrap signals, in addition to other signaling interfaces such as the System Management Bus (SMBus) interface and the Joint Test Action Group (JTAG) interface. PCIe may also have different implementations or versions than those used herein. For example, PCIe version 3.0 or later (e.g., 4.0, 5.0, or later) may be used. Furthermore, next-generation interfaces such as Gen-Z, Cache Coherent CCIX, CXL, or OpenCAPI may be used. Also, while PCIe is used in Figure 5, it should be understood that other communication links or buses, such as NVMe, Ethernet, SAS, FibreChannel, Thunderbolt, and SATA Express, may be used instead among other interconnect, network, and link interfaces. NVMe is an interface standard for mass storage devices, such as hard disk drives and solid-state memory devices. NVMe can replace the SATA interface for interfacing with mass storage devices in personal computer and server environments. However, these NVMe interfaces, like SATA devices, are limited to one-to-one host-drive relationships. In the examples described herein, a PCIe interface may be used to transmit NVMe traffic and present a multi-drive system with many storage drives as one or more NVMe virtual logical unit numbers (VLUNs) via the PCIe interface.
[0083] Each of the links in Figure 5 may use a variety of communication media, including air, space, metal, optical fiber, or any other signal propagation path including a combination thereof. Each of the PCIe links in Figure 5 may include any number of composite PCIe links or single / multi-lane configurations. Each of the links in Figure 5 may be a direct link or may include various devices, intermediate components, systems, and networks. Each of the links in Figure 5 may be a common link, a shared link, an aggregated link, or may consist of individual separate links.
[0084] Here, a simple example of forming and mounting compute units of peripheral components from a host device to a remote client device is described. In Figure 5, configurable logical visibility may be provided to a computer device for any / all available peripheral devices or resources coupled to the network fabric of the computer system 500, as enumerated in data structure 592. The computer device may request access to available resources and may logically connect remote peripheral devices as if they were local resources. Logically linked components may form compute units. For example, the CPU 521, GPU 522, and SSD 525 of chassis 520 may request and be granted access to a remote GPU 532 from chassis 530, NIC 544 from chassis 540, and GPU 551 from chassis 550, forming a single compute unit. Thus, in some of the following examples, "m" SSDs or GPUs can be coupled with "n" CPUs to enable large, scalable architectures with high levels of performance, redundancy, and density. This division allows a GPU to interact with one or more CPUs as desired, and two or more GPUs, such as eight GPUs, may be associated with a particular computing unit. The management processor 511 may be configured to update one or more data structures maintained by the management processor 511 when resources or peripheral devices are attached to or detached from various computer devices or computer units.
[0085] Figure 6 is a block diagram showing an implementation of the management processor 600. The management processor 600 is an example of any of the control elements, control systems, interaction units, dynamic server rebalancing systems, PoF units, fabric control elements, or management processors described herein, such as the PoF system 115 in Figure 1, the dynamic server rebalancing system 130 in Figure 1, the PoF unit 320 in Figure 3, the server rebalancing control element 410 in Figure 4, or the management processor 511 in Figure 5. The management processor 600 includes a communication interface 601, a user interface 603, and a processing system 610. The processing system 610 includes a processing circuit 611 and a data storage system 612 which may include a random access memory (RAM) 613, but may include additional or different configurations of elements.
[0086] The processing system 610 is generally intended to represent a computer system on which at least the software 620 is deployed and executed in order to render or otherwise perform the operations described herein. However, the processing system 610 may also represent any computer system on which at least the software 620 and data 630 are staged, and from there the software 620 and data 630 may be delivered, transmitted, downloaded, or provided to another computer system for deployment, execution, or further distribution. The processing circuit 611 may be implemented within a single processing device, but may be distributed across multiple processing devices or subsystems that cooperate in executing program instructions. Examples of the processing circuit 611 include general-purpose central processing units, microprocessors, application-specific processors, and logic devices, as well as any other type of processing device. In some examples, the processing circuit 611 includes physically distributed processing devices, such as cloud computing systems.
[0087] Communication interface 601 includes one or more communication fabrics and / or network interfaces for communication over Ethernet links, PCIe links, packet networks, the Internet, and other networks. Communication interface 601 may include one or more local or wide-area network communication interfaces capable of communicating over Ethernet interfaces, PCIe interfaces, serial interfaces, serial peripheral interface (SPI) links, inter-integrated circuit (I2C) interfaces, universal serial bus (USB) interfaces, UART interfaces, wireless interfaces, or Ethernet or Internet Protocol (IP) links. Communication interface 601 may include network interfaces configured to communicate using one or more network addresses that may be associated with different network links. Examples of communication interface 601 include network interface controller equipment, transceivers, modems, and other communication circuits. Communication interface 601 may communicate with control elements of a network or other communication fabric to establish logical partitioning or remote resource allocation within the fabric, such as through management interfaces or control interfaces of one or more communication switches in the communication fabric. The communication interface 601 can communicate via the PCIe fabric to exchange traffic / communication with peripheral devices.
[0088] The user interface 603 may include a software-based interface or a hardware-based interface. A hardware-based interface may include a touchscreen, keyboard, mouse, voice input device, audio input device, or other touch input device for receiving user input. Output devices, such as displays, speakers, web interfaces, terminal interfaces, and other types of output devices, may also be included in the user interface 603. The user interface 603 may provide output and receive input via a network interface, such as a communication interface 601. In the network example, the user interface 603 may packetize display or graphics data for a remote display by a display system or computer system coupled via one or more network interfaces. The physical or logical elements of the user interface 603 may provide attention or visual output to the user or other operators. The user interface 603 may also include relevant user interface software executable by the processing system 610 that supports the various user input and output devices described above. User interface software and user interface devices may, individually or in conjunction with each other, and together with other hardware and software elements, support a graphical user interface, a natural user interface, or any other type of user interface.
[0089] The user interface 603 may present one or more users with a command-line interface (CLI), application programming interface (API), graphical user interface (GUI), REST interface, RestAPI, WebSocket interface, or other interface. The user interface may be used by operators or administrators to assign assets (compute units / resources / peripheral devices) to each host device. In some examples, the user interface provides an interface that allows end users to determine one or more templates and dynamic tuning policy sets to use or customize for use in creating compute units. The user interface 603 may be used to manage, select, modify machine templates and change policies. The user interface 603 may also provide telemetry information in one or more status interfaces or status views, etc. The status of various components or elements, among others, such as processor status, network status, storage unit status, and PCIe element status, may be monitored via the user interface 603. Various performance metrics and error statuses may be monitored using the user interface 603.
[0090] The storage system 612 and RAM 613 may both comprise a memory device or a non-temporary data storage system, but these can be modified. The storage system 612 and RAM 613 may each comprise any storage medium that is readable by the processing circuit 611 and capable of storing software and OS images. RAM 613 may include volatile and non-volatile, removable and non-removable media implemented in any method or technique for storing information, such as computer-readable instructions, data structures, program modules, or other data. The storage system 612 may include non-volatile storage media, such as solid-state storage media, flash memory, NAND flash or NOR flash, phase-change memory, magnetic memory, or other non-temporary storage media including combinations thereof. The storage system 612 and RAM 613 may each be implemented as a single storage device, but they may also be implemented across multiple storage devices or subsystems. The storage system 612 and RAM 613 may each comprise additional elements, such as a controller capable of communicating with the processing circuit 611.
[0091] The software 620 or data 630 may be stored on or within the storage system 612 or RAM 613, which may include computer program instructions, firmware, data structures, or any other form of machine-readable processing instructions having processes, which instruct the processor 600 to operate as described herein when the processing system is executed. The software 620 may reside in RAM 613 during the execution and operation of the processor 600, and may reside in the non-volatile portion of the storage system 612 during the power-off state, among other locations and states. The software 620 may be loaded into RAM 613 during the startup or boot procedure, as described for computer operating systems and applications. The software 620 may receive user input via the user interface 603. This user input may include user commands as well as other inputs, including combinations thereof.
[0092] Software 620 includes application 621 and operating system (OS) 622. Software 620 may receive user or computer device commands and drive processor 600 for mounting or detaching peripheral devices from computer devices. Software 620 may also drive processor 600 to receive and monitor telemetry data, statistics, operational data, and other data, to provide telemetry to users, and to modify operations according to telemetry data, policies, or other data and criteria. Software 620 may also drive processor 600 to manage peripheral device resources and computer device resources, to establish domain partitioning or NT partitioning between communication fabric elements, and to interface with individual communication switches and, among other operations, to control the operation of such communication switches. Software 620 may also include user software applications, application programming interfaces (APIs), or user interfaces. Software 620 may be implemented as a single application or as multiple applications. Generally, when the software 620 is loaded onto and executed in the processing system 610, it can transform the processing system 610 from a general-purpose device into a dedicated device customized as described herein.
[0093] The software application 621 may take different forms depending on the operation and devices implemented by the management processor 600, and may include a set of applications 640 or 650. For example, when the management processor 600 operates a dynamic server rebalancing system, an application set 640 may be deployed, which includes a discovery application 641, an event application 642, a fabric interaction application 643, and a fabric user interface application 644. Alternatively, when the management processor 600 operates a computer device such as a blade server, an application set 650 may be deployed, which includes a server interaction application 651 and a server user interface application 652. Each of the software applications 641-644 and 651-652 includes executable instructions that can be executed by the processor 600 to operate a computer system or processing circuit in accordance with the operation described herein.
[0094] Application set 640 includes a discovery application 641, an event application 642, a fabric interaction application 643, and a fabric user interface application 644. The discovery application 641 may obtain instructions for computer devices and associated local peripheral devices, or de-aggregated peripheral devices, that are available for remote association with client computer devices. Instructions for computer devices or peripheral devices may include addressing information, device identifiers, vendor identifiers, device specifications or requirements, associations between devices, or other information. The discovery application 641 may obtain instructions for computer devices and peripheral devices via a network interface, a PCIe interface, or other connection links between devices. The discovery application 641 may store these instructions in data 630. Based on the instructions, the event application 642 initiates the instantiation and de-instancing of peripheral devices from the host device or de-aggregated peripheral device to the client device's local peripheral device interconnect interface. The fabric interaction application 643 intercepts client-originating traffic for a remote peripheral device received via a network interface, interprets the client-originating traffic, performs any necessary format conversions to deliver the traffic to the target peripheral device, and routes the client-originating traffic in the appropriate format via the network or PCIe interface for delivery to the peripheral device. Similarly, the fabric interaction application 643 intercepts peripheral device-originating traffic from either a host device or a deaggregated peripheral device that is directed to a client device, interprets the peripheral device-originating traffic, performs any necessary format conversions to forward the traffic to the client device, and routes the peripheral device-originating traffic in the appropriate format via the network interface for delivery to the client device.The fabric user interface application 644 can receive operator commands from a computer device to attach or detach peripheral devices and can present various information, status, telemetry, logs, etc., to the device through various types of user interfaces. Commands or requests to attach or detach peripheral devices received from a network computer system may be received via the event application 642, the fabric interaction application 643, or the fabric user interface application 644, depending on the implementation.
[0095] Application set 650 includes a server interaction application 651 and a server host user interface application 652. The server interaction application 561 may interface with the network stack of a computer device to allow peripheral device traffic to interact with the local peripheral device interconnect interface. As a local peripheral device coupled to the peripheral device interconnect interface of a client system, the server interaction application 561 may emulate the behavior of peripheral devices coupled via a network interface. As a local peripheral device coupled to a client system, the server interaction application 561 emulates the hardware plugging process by providing at least an address space description to the local peripheral device interconnect interface in order to access the peripheral device. The server interaction application 561 removes the instantiation of peripheral devices from the local peripheral device interconnect interface by at least emulating the hardware removal process within the local peripheral device interconnect interface.
[0096] When instantiated on the client device's local peripheral device interconnect interface, the client device's device driver can interface with peripheral devices through the local peripheral device interconnect interface. The server interaction application 561 emulates the behavior of peripheral devices by at least intercepting client-derived traffic for peripheral devices from the local peripheral device interconnect interface, interpreting the client-derived traffic to convert it from a native peripheral device format (such as a PCIe frame or a memory-mapped format) to a network format suitable for transmission over the network interface (e.g., a frame or packet with relevant encapsulation and addressing / header / footer), and routing the client-derived traffic in packet format over the network interface for delivery to peripheral devices. The server interaction application 561 emulates the behavior of peripheral devices by at least receiving peripheral device-derived traffic in packet format from the network interface and interpreting the peripheral device-derived traffic in packet format to convert it to a native peripheral device format suitable for the local peripheral device interconnect interface. The server interaction application 561 initiates the instantiation of peripheral devices to the local peripheral device interconnect interface by triggering at least a start point process within the local peripheral device interconnect interface, in order to emulate the hardware plugging process for peripheral devices having a local peripheral device interconnect interface. In the case of a host server device, the server interaction application 651 may route traffic from remote client devices to local peripheral devices connected to the local peripheral device interconnect interface.For example, client-initiated traffic may be received by a host server interaction application 651 via a network connection and converted into a format for use by a host peripheral device interconnect interface, such as a PCIe frame. Thus, client-initiated traffic can be routed from the server interaction application 651 to the peripheral device interconnect interface, and from there to local peripheral devices for processing. In some examples, in addition to executing commands received from a remote client device, the server interaction application 651 may notify the host server system that a local peripheral device assigned to a remote client is unavailable for use by that host, and may send instructions, commands, or requests over the network indicating the availability status of the local peripheral device, requesting the removal of the local peripheral device from the remote client, or requesting access to the remote peripheral device from an available peripheral device pool.
[0097] The server user interface application 652 provides the computer device operator with local instructions for attaching and detaching peripheral devices, and may receive operator commands to attach or detach peripheral devices during other operations.
[0098] In addition to the software 620, other data 630, including various data structures, may be stored by the storage system 612 and RAM 613. Data 630 may include templates, policies, telemetry data, event logs, or fabric status. Data 630 may include instructions and identification of peripheral devices and computer devices. Data 630 may include the current assignment of peripheral devices to client devices. Fabric status includes information and properties of various communication fabrics, including pools of resources or components, such as fabric type, protocol version, technical descriptors, header requirements, addressing information, and other data. Fabric data may include relationships between components and the specific fabrics to which the components are connected.
[0099] This specification describes various peripheral devices, including data processing elements or other computer components, coupled via one or more communication fabrics or communication networks. Various communication fabric types or communication network types can be utilized. For example, a Peripheral Component Interconnect Express (PCIe) fabric can be used to couple to a CoPU, which may include various versions such as 3.0, 4.0, or 5.0. Instead of a PCIe fabric, other point-to-point communication fabrics or communication buses with associated physical layers, electrical signaling, protocols, and layered communication stacks can be used. These may include, in particular, Gen-Z, Ethernet, InfiniBand, NVMe, Internet Protocol (IP), Serial Attached SCSI (SAS), Fibre Channel, Thunderbolt, Serial Attached ATA Express (SATA Express), NVLink, Cache Coherent Interconnect for Accelerators (CCIX), Compute Express Link (CXL), Open Coherent Accelerator Processor Interface (OpenCAPI), Wi-Fi (802.11x), or cellular wireless technologies. The communication network is coupled to a host system and includes Ethernet or Wi-Fi (802.11x), which may refer to any of the various available network communication protocol standards and bandwidths, such as 10BASE-T, 100BASE-TX, 1000BASE-T, 10GBASE-T (10GB Ethernet), 40GBASE-T (40GB Ethernet), Gigabit (GbE), Terabit (TbE), 200GbE, 400GbE, 800GbE, or various other wired and wireless Ethernet formats and speeds. Cellular radio technology may also include various radio protocols and networks built around 3GPP standards, including, among others, 4G Long-Term Evolution (LTE), 5G NR (New Radio), and related 5G standards.
[0100] Some of the aforementioned signaling or protocol types are built on PCIe and therefore add additional functionality to the PCIe interface. Parallel, serial, or combined parallel / serial type interfaces may also be applicable to the examples herein. While many of the examples herein utilize PCIe as an exemplary fabric type for coupling peripheral devices, it should be understood that others may be used instead. PCIe is a high-speed serial computer expansion bus standard that typically has multi-lane point-to-point connections between host and component devices or between peer devices. PCIe typically has multi-lane serial links connecting individual devices to a root complex. PCIe communication fabrics can be established using various switching circuits and control architectures described herein.
[0101] The various computer system components described herein may be contained within one or more physical enclosures, such as rack-mountable modules that may be further included in shelves or rack units. Numerous components may be inserted into or installed within a physical enclosure, such as a modular framework into which modules can be inserted and removed, depending on the specific end-user requirements. Enclosed modular systems may include physical support structures and enclosures containing circuits, printed circuit boards, semiconductor systems, and structural elements. Modules containing components may be insertable into or removable from rack-mount style or rack unit (U) type enclosures. It should be understood that the components described herein can be included in any physical mounting environment and do not necessarily require the inclusion of associated enclosures or rack-mount elements.
[0102] The functional block diagrams, operating scenarios and sequences, and flowcharts provided in the figures represent exemplary systems, environments, and methodologies for carrying out novel embodiments of this disclosure. For the sake of simplicity, the methods included herein may be in the form of functional diagrams, operating scenarios or sequences, or flowcharts, or may be described as a series of operations, and it should be understood and recognized that the methods are not limited by the order of operations, as some operations may be performed in a different order and / or simultaneously with others, accordingly. For example, those skilled in the art will understand and recognize that the methods may be alternatively represented as a series of interrelated states or events, such as a state diagram. Furthermore, not all operations exemplified in the methodology are required for novel implementations.
[0103] The descriptions and figures included herein illustrate specific implementations to teach those skilled in the art how to create and use the best options. Some conventional embodiments have been simplified or omitted for the purpose of teaching the principles of the present invention. Those skilled in the art will understand variations from these implementations that fall within the scope of this disclosure. Those skilled in the art will also understand that the above-described features can be combined in various ways to form multiple implementations. Consequently, the present invention is not limited to the specific implementations described above, but is limited only by the claims and their equivalents.
Claims
1. A method performed by a server rebalancing system, The server rebalancing system includes the step of receiving identification information of peripheral devices that are available for data processing and are physically connected to a first computer device, The server rebalancing system receives a request from a second computer device to access the peripheral device, The server rebalancing system instructs the second computer device to emulate the peripheral device as a local device installed on the second computer device based on the request, The server rebalancing system includes the steps of routing data traffic from the second computer device to the first computer device for processing by the peripheral device, It includes, The aforementioned method, The server rebalancing system includes the steps of maintaining a data structure that identifies peripheral devices available for processing, including the peripheral devices, The server rebalancing system provides a list of peripheral devices available for processing to a networked computer device from the server rebalancing system. Based on the step of instructing the second computer device to emulate the peripheral device as a local device by the server rebalancing system, the steps include updating the data structure to indicate that the peripheral device is assigned to the second computer device, The server rebalancing system performs the steps of routing the data traffic based on the data structure, Methods that further include the above.
2. The server rebalancing system receives a notification regarding the removal of the peripheral device from the second computer device, The server rebalancing system, based on the notification, separates the peripheral device from the second computer device and updates the data structure to indicate that the peripheral device is available for processing; The method according to claim 1, further comprising:
3. A method performed by a server rebalancing system, The server rebalancing system includes the step of receiving identification information of peripheral devices that are available for data processing and are physically connected to a first computer device, The server rebalancing system receives a request from a second computer device to access the peripheral device, The server rebalancing system instructs the second computer device to emulate the peripheral device as a local device installed on the second computer device based on the request, The server rebalancing system includes the steps of routing data traffic from the second computer device to the first computer device for processing by the peripheral device, It includes, The aforementioned method, The server rebalancing system includes the step of discovering a second peripheral device available for processing via a peripheral interface, The server rebalancing system receives a second request from the second computer device to access the second peripheral device, The server rebalancing system instructs the second computer device to emulate a local installation of the second peripheral device on the second computer device based on the second request, The server rebalancing system performs the steps of converting a message received from the second computer device from a first data format to a second data format, The server rebalancing system includes the steps of routing messaging in the second data format to the second peripheral device via the peripheral interface, Methods that further include the above.
4. The method according to claim 1, wherein the identification information and the request are received via a network interface.
5. It is a system, A first computer device including a network interface, The first computer device, The process involves obtaining identification information of peripheral devices available via the network interface from the server rebalancing system via the network interface, wherein the peripheral devices are physically connected to a second computer device. To issue a request to access the aforementioned peripheral device to the server rebalancing system via the network interface, Based on the response from the server rebalancing system, the peripheral device is emulated as a local device installed on the first computer device, To issue data traffic for processing by the peripheral device to the server rebalancing system via the network interface, It is configured to implement the following: The server rebalancing system further includes a second network interface and a processor. The aforementioned processor, Receiving identification information of the peripheral devices that are physically connected to the second computer device and available for processing via the second network interface, The request to access the peripheral device is received from the first computer device via the second network interface, Based on the above request, the first computer device is instructed via the second network interface to emulate the peripheral device as a local device installed on the first computer device, For processing by the aforementioned peripheral devices, data traffic is routed from the first computer device to the second computer device, A system configured to perform the following actions.
6. The first computer device issues identification information to the server rebalancing system via a network interface, indicating that a peripheral device physically connected to the first computer device is available for processing. The first computer device receives second identification information from the server rebalancing system via the network interface, indicating that the peripheral device is assigned to a second computer device. The steps include receiving data traffic from the second computer device to the first computer device via the network interface from the server rebalancing system for processing by the peripheral device, The steps include providing the results of the processing of the data traffic by the peripheral device from the first computer device via the network interface, Methods that include...
7. The first computer device receives the data traffic in a packet format suitable for transfer via the network interface. The first computer device interprets the data traffic in the packet format in order to convert it into a native peripheral device format suitable for the peripheral device, The steps include processing the data traffic in the peripheral device, The first computer device interprets the result of the processing from the native peripheral device format to the packet format, The first computer device provides the results in the packet format to the server rebalancing system via the network interface, The method according to claim 6, further comprising:
8. The steps of receiving identification information of the peripheral device available for processing on the first computer device via the network interface in the server rebalancing system, The steps include receiving a request to access the peripheral device from the second computer device via the network interface in the server rebalancing system, The steps include: instructing the second computer device via the network interface to emulate the peripheral device as a local device installed on the second computer device based on the aforementioned request; The server rebalancing system performs the steps of routing data traffic from the second computer device to the first computer device via the server rebalancing system for processing by the peripheral devices, The method according to claim 6, further comprising:
9. The steps include obtaining identification information of the peripheral devices available for processing from the server rebalancing system to the second computer device, The second computer device issues a request to access the peripheral device to the server rebalancing system via the network interface, The second computer device emulates the peripheral device as a local device installed on the second computer device, based on the response from the server rebalancing system on the second computer device. The steps include issuing the data traffic for processing by the peripheral device from the second computer device to the server rebalancing system via the network interface, The method according to claim 6, further comprising: