Thread-level boot management control of a compute node for context switching using a boot controller

The remote boot control and network reconfiguration of compute nodes in data centers address the challenge of static resource configurations, enabling flexible and efficient utilization of computing resources to meet varying demands and support multiple services.

JP7760075B2Active Publication Date: 2025-10-24SONY INTERACTIVE ENTERTAINMENT LLC
View PDF 7 Cites 0 Cited by

Patent Information

Application Number
JP2024554891
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Priority Date
2022-03-14
Filing Date
2023-02-23
Publication Date
2025-10-24
Estimated Expiration
2043-02-23

AI Technical Summary

Technical Problem

Data centers dedicated to online gaming face challenges in dynamically reconfiguring computing resources to handle varying demands and provide other services due to static configurations, which limits flexibility and efficiency.

Method used

Implementing remote boot control of compute nodes and network reconfiguration of rack assemblies using a board management controller (BMC) and cloud management controller to facilitate dynamic reconfiguration based on demand, allowing flexible software selection and efficient resource utilization.

Benefits of technology

Enables dynamic reconfiguration of computing resources to adapt to varying demands, optimizing resource utilization and supporting multiple services efficiently.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007760075000001
    Figure 0007760075000001
  • Figure 0007760075000002
    Figure 0007760075000002
  • Figure 0007760075000003
    Figure 0007760075000003
Patent Text Reader

Abstract

A method for performing a system boot includes receiving boot configuration instructions at a board management controller (BMC) for booting a compute node with an operating system, the compute node being placed on a sled that includes a plurality of compute nodes, the BMC configured to manage a plurality of communication interfaces that provide communication to the plurality of compute nodes; transmitting the boot instructions from the BMC via the communication interface to a boot controller of the compute node to execute basic input / output system (BIOS) firmware stored external to the compute node; and executing the BIOS firmware on the compute node to initiate loading of an operating system to be executed by the compute node.
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The present disclosure relates to remote activation of computing resources, including remote boot control of compute nodes, such as compute nodes of a streaming array of compute threads in a rack assembly, and network reconfiguration of a rack assembly in a data center. [Background technology]

[0002] In recent years, there has been a continuous push for online services that enable streaming online or cloud gaming between cloud gaming servers and clients connected over a network. Streaming has become increasingly popular due to the availability of game titles on demand, the ability to run more complex games, the ability to network players in multiplayer games, the ability to share assets between players, the ability to share an instantaneous experience between players and / or spectators, the ability for friends to watch others play video games, and the ability for friends to join in on-going gameplay.

[0003] A data center may be configured with multiple computing resources to support online or cloud gaming. For example, each computing resource may be configured to run a gaming application for gameplay that may then be streamed to users. Demand for computing resources may vary depending on one or more parameters, including the duration of demand, the geographic area of ​​demand, the type of game being sought, etc. Limited demand for computing resources for online gaming may result in periods of idle computing resources.

[0004] A data center dedicated to online gaming may be limited in its ability to make short-term changes to computing resources to handle various other computing services distinct from gaming. That is, because the computing resources are configured for gaming, they are not configured to provide other types of services that require applications that require a different computing resource platform. When computing resources are statically configured, it may not be possible to change the configuration of these computing resources to support these other services. Changing the configuration of these computing resources may also prove difficult if configuration parameters must be changed locally at the computing resources before any configuration changes are implemented.

[0005] It is against this background that the embodiments of the present disclosure have been made. Summary of the Invention

[0006] Embodiments of the present disclosure relate to providing remote booting of computing resources, including remote boot control of compute nodes, such as compute nodes of a streaming array of compute threads in a rack assembly, and network reconfiguration of rack assemblies in a data center.

[0007] An embodiment of the present disclosure discloses a method for performing a system boot. The method includes receiving boot configuration instructions at a board management controller (BMC) to boot a compute node using an operating system, the compute node being arranged on a sled including multiple compute nodes, the BMC configured to manage multiple communication interfaces providing communication to the multiple compute nodes. The method includes sending the boot instructions from the BMC via the communication interface to a boot controller of the compute node to execute basic input / output system (BIOS) firmware stored external to the compute node. The method includes executing the BIOS firmware on the compute node to initiate loading of the operating system for execution by the compute node.

[0008] Another embodiment of the present disclosure discloses a method. The method includes detecting, at a cloud management controller, a decrease in demand for a first priority service supported by a data center including a plurality of rack assemblies, each of the plurality of rack assemblies being configured in a first configuration that facilitates the first priority service, the cloud management controller managing the configuration of the plurality of rack assemblies, the first priority service being implemented by a first plurality of applications. The method includes sending a reconfiguration message from the cloud management controller to a rack controller of the rack assembly to reconfigure the rack assembly from the first configuration to a second configuration, the second configuration facilitating a second priority service, the second priority service having a lower priority than the first priority service, and the second plurality of services being implemented by a second plurality of applications. The method includes configuring the rack assembly in the second configuration. Each rack assembly of the plurality of rack assemblies includes one or more network storage and one or more streaming arrays, each streaming array including one or more compute threads, each compute thread including one or more compute nodes.

[0009] Another embodiment of the present disclosure discloses a non-transitory computer-readable medium storing a computer program for performing a system startup. The non-transitory computer-readable medium includes program instructions for receiving, at a board management controller (BMC), startup configuration instructions for booting a compute node using an operating system, the compute node being arranged on a sled including multiple compute nodes, the BMC being configured to manage multiple communication interfaces providing communication to the multiple compute nodes. The non-transitory computer-readable medium includes program instructions for sending the boot instructions from the BMC via the communication interface to a boot controller of the compute node to execute basic input / output system (BIOS) firmware stored external to the compute node. The non-transitory computer-readable medium includes program instructions for executing the BIOS firmware on the compute node and initiating loading of an operating system to be executed by the compute node.

[0010] Other aspects of the present disclosure will become apparent from the following detailed description, taken in conjunction with the accompanying drawings, illustrated by way of example of the principles of the disclosure.

[0011] The present disclosure is best understood by reference to the following detailed description taken in conjunction with the accompanying drawings. [Brief explanation of the drawings]

[0012] [Figure 1] FIG. 1 is a diagram of a game cloud system for providing games over a network among one or more computing nodes located in one or more data centers, according to one embodiment of the present disclosure. [Figure 2A] FIG. 1 illustrates a diagram of multiple rack assemblies including multiple compute nodes in a representative data center of a gaming cloud system, according to one embodiment of the present disclosure. [Figure 2B]FIG. 1 is a diagram of multiple rack assemblies including multiple computing nodes in a representative data center of a gaming cloud system, each network storage accessible by a corresponding array of computing nodes, according to one embodiment of the present disclosure. [Figure 3] 1 is a diagram of a computing sled including multiple computing nodes arranged in a rack assembly configured for remote boot control of the computing nodes using a board management controller, according to one embodiment of the present disclosure. [Figure 4] FIG. 2 is a flow diagram illustrating steps in a method for performing a remote boot of a compute node using a board management controller, according to one embodiment of the present disclosure. [Figure 5] FIG. 1 is a flow diagram illustrating steps in a method for reconfiguring a rack assembly of a data center according to one embodiment of the present disclosure. [Figure 6] 1 illustrates components of an exemplary device that can be used to implement aspects of various embodiments of the present disclosure. DETAILED DESCRIPTION OF THE INVENTION

[0013] Although the following detailed description includes many specific details for purposes of explanation, those skilled in the art will recognize that many variations and modifications to the following details are within the scope of the present disclosure. Accordingly, the aspects of the present disclosure described below are set forth without any loss of generality to, and without imposing limitations on, the claims that follow this description.

[0014] Generally speaking, embodiments of the present disclosure provide remote boot control of compute nodes, such as compute nodes of a streaming array of compute threads in a rack assembly. Notably, external hardware is used in the boot process of the compute nodes, so no storage is required, allowing flexibility as to which software is used for the external boot. Embodiments of the present disclosure also provide network reconfiguration of a rack assembly in a data center. In particular, dark time utilization of data center computing resources is achieved through reconfiguration of the rack assembly's networking, such as reconfiguring the internal networking to provide communication between compute nodes within the rack assembly via network interfaces.

[0015] Based on the foregoing general understanding of the various embodiments, illustrative details of the embodiments will now be described with reference to the various figures.

[0016] Throughout the specification, references to an "application" or "game" or "video game" or "game application" or "game title" are meant to refer to any type of interactive application that is directed through the execution of input commands. By way of example only, interactive applications include applications for games, word processing, video processing, video game processing, etc. Furthermore, the terms introduced above are interchangeable.

[0017] 1 is a diagram of a system 100 for providing games over a network 150 between one or more computing nodes located in one or more data centers, according to one embodiment of the present disclosure. The system is configured to provide games over a network between one or more cloud gaming servers, and more specifically, for remote boot control of computing nodes, such as computing nodes of a streaming array of computing threads in a rack assembly, and for network reconfiguration of rack assemblies in a data center to provide better utilization of computing resources (e.g., for multiple purposes) in the data center. Cloud gaming involves running a video game on a server to generate game-rendered video frames, which are then sent to and displayed by clients.

[0018] It is also understood that cloud gaming and / or other services, in various embodiments (e.g., within a cloud gaming environment or a standalone system), can be executed using physical machines (e.g., central processing units (CPUs) and graphics processing units (GPUs)), virtual machines, or a combination of both. For example, virtual machines (e.g., instances) can be created using a hypervisor on host hardware (e.g., located in a data center) that utilizes one or more components of a hardware layer, such as multiple CPUs, memory modules, GPUs, network interfaces, communication components, etc. These physical resources can be arranged in racks, such as a rack of CPUs, a rack of GPUs, a rack of memory, etc., and the physical resources in the racks can be accessed using a top-of-rack switch that facilitates a fabric for assembling and accessing the components used in the instances (e.g., when building the virtualized components of the instances). Generally, the hypervisor can present multiple guest operating systems of multiple instances configured with virtual resources. That is, each operating system may be configured with a corresponding set of virtualized resources supported by one or more hardware resources (e.g., located in a corresponding data center). For example, each operating system may be supported by a virtual CPU, multiple virtual GPUs, virtual memory, virtualized communication components, etc. Furthermore, to reduce latency, the configuration of an instance may be transferred from one data center to another. Instant-uses defined for a user or game can be used when saving a user's game session. Instant-uses may include any number of configurations described herein to optimize fast rendering of video frames for a game session. In one embodiment, instant-uses defined for a game or user can be transferred between data centers as configurable settings. The ability to transfer instant-use settings enables efficient migration of game play from data center to data center when users connect to play games from different geographic locations.

[0019] System 100 includes a gaming cloud system 190 implemented across one or more data centers (e.g., data centers 1 through N). As shown, an instance of gaming cloud system 190 may be located in data center N that provides management functions, and the management functions of gaming cloud system 190 may be distributed across multiple instances of gaming cloud system 190 at each data center. In some implementations, the gaming cloud system management functions may be located outside of any of the data centers.

[0020] The gaming cloud system 190 includes an assigner 191 configured to assign each of the client devices (e.g., 1-N) to corresponding resources in a corresponding data center. In particular, when the client device 110 logs into the gaming cloud system 190, the client device 110 may be connected to an instance of the gaming cloud system 109 at data center N, which may be geographically closest to the client device 110. The assigner 191 may perform diagnostic tests to determine the available transmit and receive bandwidth to the client device 110. The diagnostic tests may also include determining the latency and / or round-trip time between the corresponding data center and the client device 110. Based on the tests, the assigner 191 may assign resources very specifically to the client device 110. For example, the assigner 191 may assign a particular data center to the client device 110. Furthermore, the assigner 191 may assign a particular compute thread, a particular streaming array, a particular rack assembly, or a particular compute node to the client device 110. The allocation may be performed based on knowledge of assets (e.g., games) available at the compute nodes. Previously, client devices were typically assigned to data centers without further assignment to rack assemblies. In this manner, the assigner 191 may assign a client device requesting execution of a particular computationally intensive gaming application to a compute node that may not be running the computationally intensive application. Additionally, load management of the allocation of compute-intensive gaming applications requested by clients may be performed by the assigner 191. For example, the same compute-intensive gaming application requested for a short period of time may be distributed to different compute nodes in different compute threads within one rack assembly or different rack assemblies to reduce the load on a particular compute node, compute thread, and / or rack assembly.

[0021] In some embodiments, allocation may be performed based on machine learning. In particular, resource demand may be predicted for a particular data center and its corresponding resources. For example, if a data center can predict that it will soon be handling many clients running computationally intensive gaming applications, assigner 191 can use that information to allocate client devices 110 and allocate resources that may not currently be utilizing all of their resource capabilities. In another case, assigner 191 may switch client device 110 from gaming cloud system 190 in data center N to resources available in data center 3 in anticipation of increased load at data center N. Additionally, future clients may be assigned resources in a distributed manner such that resource load and demand may be distributed throughout the gaming cloud system, across multiple data centers, across multiple rack assemblies, across multiple computational threads, and / or across multiple computational nodes. For example, client device 110 may be assigned resources from the gaming cloud systems in both data center N (e.g., via path 1) and data center 3 (e.g., via path 2).

[0022] Once a client device 110 is assigned to a particular computational node of a corresponding computational thread of a corresponding streaming array, the client device 110 connects to the corresponding data center over a network, i.e., the client device 110 may communicate with a data center different from the data center that performed the assignment, such as data center 3.

[0023] System 100 provides games via game cloud system 190, which, according to one embodiment of the present disclosure, are executed remotely from client devices (e.g., thin clients) of corresponding users playing the games. System 100 may provide game control to one or more users playing one or more games through cloud gaming network or game cloud system 190 via network 150 in either single-player or multiplayer mode. In some embodiments, cloud gaming network or game cloud system 190 may include multiple virtual machines (VMs) executing on a hypervisor of a host machine, where one or more virtual machines are configured to execute a game processor module that utilizes hardware resources available to the host's hypervisor. Network 150 may include one or more communication technologies. In some embodiments, network 150 may include fifth generation (5G) network technology with advanced wireless communication systems.

[0024] In some embodiments, communication may be facilitated using wireless technology. Such technology may include, for example, 5G wireless communication technology. 5G is the fifth generation of cellular network technology. 5G networks are digital cellular networks in which a provider's coverage area is divided into small geographic areas called cells. Analog signals representing sound and images are digitized within the phone, converted by an analog-to-digital converter, and transmitted as a bit stream. All 5G wireless devices within a cell communicate over electromagnetic waves with a local antenna array and low-power automatic transceiver (transmitter and receiver) within the cell via frequency channels assigned by the transceiver from a frequency pool reused by other cells. The local antennas are connected to the telephone network and the Internet by high-bandwidth optical fiber or wireless backhaul connections. As with other cellular networks, mobile devices moving from one cell to another are automatically transferred to the new cell. It should be understood that a 5G network is merely one example type of communication network, and embodiments of the present disclosure may utilize previous generations of wireless or wired communications, as well as later generations of wired or wireless technologies following 5G.

[0025] As shown, system 100, including game cloud system 190, can provide access to multiple game applications. In particular, each client device may request access to a different game application from the cloud gaming network. For example, game cloud system 190 may provide one or more game servers, which may be configured as one or more virtual machines running on one or more hosts to execute corresponding game applications. For example, a game server may manage virtual machines supporting game processors that instantiate instances of users' game applications. Thus, multiple game processors of one or more game servers associated with multiple virtual machines are configured to execute multiple instances of one or more game applications associated with gameplay for multiple users. In this manner, the backend server support provides streaming of gameplay media (e.g., video, audio, etc.) for multiple game applications to multiple corresponding users. That is, the game servers of the game cloud system 190 are configured to stream data (e.g., rendered images and / or frames of corresponding gameplay) back to corresponding client devices over the network 150. In this manner, computationally complex game applications can continue to execute on the backend servers in response to controller inputs received and forwarded by the client devices. Each server can render images and / or frames, which are then encoded (e.g., compressed) and streamed to corresponding client devices for display.

[0026] In embodiments, each virtual machine defines a resource environment capable of supporting an operating system on which game applications can be executed. In one embodiment, the virtual machine can be configured to emulate the hardware resource environment of a game console, and an operating system associated with the game console runs on the virtual machine to support the execution of game titles developed for that game console. In another embodiment, the operating system can be configured to emulate the game console's native operating system environment, but the underlying virtual machine may or may not be configured to emulate the game console's hardware. In another embodiment, an emulator application runs on the virtual machine's operating system, and the emulator is configured to emulate the game console's native operating system environment and support game applications and / or video games designed for that game console. It should be appreciated that a variety of current and traditional game consoles can be emulated by the cloud-based gaming system. In this manner, users can access game titles from a variety of game consoles via the cloud gaming system.

[0027] In one embodiment, the cloud gaming network or game cloud system 190 is a distributed game server system and / or architecture. In particular, a distributed game engine that executes game logic is configured as a corresponding instance of a corresponding game application. Generally, a distributed game engine takes each function of the game engine and distributes those functions to be executed by multiple processing entities. Individual functions may be further distributed across one or more processing entities. The processing entities may be configured in various configurations, such as physical hardware and / or virtual components or virtual machines and / or virtual containers; containers differ from virtual machines because they virtualize instances of game applications running on virtualized operating systems. The processing entities may utilize and / or rely on servers and their underlying hardware on one or more servers (computing nodes) of the cloud gaming network or game cloud system 190, which may be arranged on one or more racks. Coordination, allocation, and management of the execution of their functions across the various processing entities is performed by a distributed synchronization layer. In this manner, the execution of their functions is controlled by the distributed synchronization layer to generate media (e.g., video frames, audio, etc.) for the game application in response to controller inputs by the player. The distributed synchronization layer can efficiently execute their functions across the distributed processing entities (e.g., via load balancing) so that critical game engine components / functions are distributed and restructured for more efficient processing.

[0028] 2A is a diagram of multiple rack assemblies 210 including multiple computing nodes in a representative data center 200A of a gaming cloud system, according to one embodiment of the present disclosure. For example, multiple data centers may be distributed around the world, such as in North America, Europe, and Japan.

[0029] Data center 200 includes multiple rack assemblies 220 (e.g., rack assemblies 220A through 220N). Each rack assembly includes corresponding network storage and multiple compute threads. For example, representative rack assembly 220N includes network storage 211 and multiple compute threads 230 (e.g., threads 230A through 230N), as well as a rack controller 250 configured for internal and external network configuration of the components of rack assembly 220N. Other rack assemblies may be similarly configured, with or without modifications. In particular, each compute thread includes one or more compute nodes that provide hardware resources (e.g., processors, CPUs, GPUs, etc.). For example, compute thread 230N in the multiple compute threads 230 of rack assembly 220N is shown as including four compute nodes, although it is understood that a rack assembly may include one or more compute nodes. Each rack assembly is coupled to a cluster switch configured to provide communication with a management server configured for management of the corresponding data center. For example, rack assembly 220N is coupled to cluster switch 240N. The cluster switch also provides communication to external communication networks (eg, the Internet, etc.).

[0030] In particular, the cluster fabric (e.g., cluster switches, etc.) provides communication between one or more cluster rack assemblies, distributed storage 270, and a communications network. Additionally, the cluster fabric / switch also provides data center support services such as management, logging, monitoring, event generation, and boot management information tracking. The cluster fabric / switch may provide communication to an external communications network via a router system and a communications network (e.g., the Internet). The cluster fabric / switch also provides communication to storage 270. A cluster of rack assemblies in a typical data center for a gaming cloud system may include one or more rack assemblies, depending on design choice. In one embodiment, the cluster includes 50 rack assemblies. In other embodiments, the cluster may include more or less than 50 rack assemblies.

[0031] In one network configuration, each rack assembly provides high-speed access to corresponding network storage, such as within the rack assembly. In one embodiment, this high-speed access is provided via a PCIe fabric, which provides direct access between the compute nodes and the corresponding network storage. In other embodiments, the high-speed access is provided via other network fabric topologies and / or networking protocols, including Ethernet, InfiniBand, Remote Direct Memory Access (RDMA) over Converged Ethernet (RoCE), and the like. For example, in rack assembly 220N, high-speed access is configured to provide data path 201 between specific compute nodes of corresponding compute threads to corresponding network storage (e.g., storage 211). In particular, the network fabric (e.g., PCIe fabric) can provide network storage bandwidth (e.g., access) of over 4 gigabytes per second (GB / s) per compute node (e.g., rack assembly) at Non-Volatile Memory Express (NVMe) latency. Additionally, control path 202 is configured to communicate control and / or management information between network storage 210 and each compute node.

[0032] In another network configuration (not shown), each rack assembly provides high-speed access between compute nodes. For example, reconfiguration of the PCIe fabric within the rack assembly is performed to enable high-speed communication between compute nodes within the rack assembly, such as when forming a supercomputer or an artificial intelligence (AI) computer.

[0033] As shown, cloud management controller 210 and / or a corresponding rack controller (e.g., rack controller 220B) of data center 200 communicate with assigner 191 (shown in FIG. 1 ) to allocate resources to client devices 110. In particular, cloud management controller 210 and / or a corresponding rack controller (e.g., rack controller 220B) of data center 200 cooperate with an instance of gaming cloud system 190′ and, together with the initial instance of gaming cloud system 190 (e.g., of FIG. 1 ), can allocate resources to client devices 110. That is, assigner 191 cooperates with cloud management controller 210 and / or a corresponding rack controller to assign users to resources. In one embodiment, the corresponding rack controller is primarily responsible for operationalizing the rack assembly, including configuring internal and external networks, powering on the computing threads and / or computing nodes located within the corresponding rack controller, etc. The corresponding rack controller communicates status information, such as the number of work threads available for allocation and the number of compute nodes available for allocation, to the cloud management controller 210. This information may be communicated by the cloud management controller 210 and / or the corresponding rack controller to the assignor 191 so that the assignor can use the information to assign users to resources (e.g., compute nodes, etc.). In some embodiments, the corresponding rack controller and network storage pair are configured on the same server (e.g., network storage), while in other embodiments, the corresponding rack controller and network storage are configured on separate servers. In embodiments, the allocation is performed based on asset awareness, such as knowing the resources and bandwidth needed and that they exist in the data center. Thus, for illustrative purposes, embodiments of the present disclosure are configured to assign client devices 110 to specific compute nodes 232 of corresponding compute threads 231 of corresponding rack assembly 220B.

[0034] In one embodiment, the cloud management controller 210 is configured to manage remote boot control of computing resources, such as compute nodes of corresponding compute threads of corresponding streaking arrays in corresponding rack assemblies, in coordination with corresponding rack controllers. In another embodiment, the cloud management controller 210 is configured to directly manage remote boot control of computing resources, such as compute nodes of thread servers arranged in streaming arrays of rack assemblies. That is, the cloud management controller 210 and / or the corresponding rack controllers are configured to manage remote boot control of computing resources. For example, cloud management controller 210 and / or corresponding rack controllers may manage boot files (e.g., BIOS firmware) used for start-up and / or boot operations of compute nodes, which may be located at storage addresses 265A-N in remote storage 260. Additionally, cloud management controller 210 and / or corresponding rack controllers may control boot management information 270 for each computing resource in the data center, including boot images (e.g., operating system images) for compute nodes in the data center. In this manner, the data center's computing resources (e.g., rack assemblies, computing threads, computing nodes, etc.) may be dynamically reconfigured depending on the intended use of those resources (e.g., gaming during peak user demand, secondary use during off-hours (e.g., performing system maintenance, performing AI modeling, running supercomputer applications, etc.)).

[0035] The streaming rack assembly is configured around compute nodes, which run gaming applications, video games, and / or stream audio / video of game sessions to one or more clients. Additionally, within each rack assembly, game content can be stored on storage servers that provide network storage. The network storage provides large amounts of storage and a high-speed network to serve many compute nodes. In particular, storage protocols (e.g., network file systems) are implemented over a network fabric topology used to access the network storage. Data can be stored in the network storage using file storage, block storage, or object storage technologies, depending in part on the underlying storage protocol implemented. For example, a PCIe fabric storage protocol can access block storage data from the network storage.

[0036] 2B is a diagram of multiple rack assemblies 221 including multiple compute nodes in an exemplary data center 200B of a gaming cloud system, each with network storage accessible by a corresponding array of compute nodes, according to one embodiment of the present disclosure. Data center 200B is similar to data center 200A, with like-numbered components having similar functions. However, data center 200B has rack assemblies configured differently than the rack assemblies of data center 200A, with network storage accessed by a single streaming array of compute nodes, as described below.

[0037] Data center 200B includes multiple rack assemblies 221 (e.g., rack assemblies 221A through 221N). Each rack assembly includes one or more streaming arrays, each including corresponding network storage and multiple computational threads. For example, representative rack assembly 221N includes streaming arrays 225A through 225N. In one embodiment, rack assembly 221N includes two streaming arrays, each including network storage, a corresponding rack controller, and multiple computational threads. For example, streaming array 225N includes multiple computational threads 235 that access network storage 211N. In particular, each streaming array includes a corresponding rack controller and network storage. For example, in a representative rack assembly 221N, streaming array 225A includes rack controller 250A and network storage 211A, and streaming array 225N includes rack controller 250N and network storage 211N. In some embodiments, corresponding rack controller and network storage pairs for corresponding streaming arrays are configured on the same server (e.g., network storage), while in other embodiments, corresponding rack controllers and network storage are configured on separate servers. In yet other embodiments, a rack assembly includes multiple streaming arrays, each with corresponding network storage, and a single rack controller for the rack assembly manages boot control (e.g., remote boot control, control boot management information, etc.) for computing resources across the streaming arrays. In yet other embodiments, a rack assembly includes multiple streaming arrays accessing a single network storage, and a single rack controller for the rack assembly manages boot control (e.g., remote boot control, control boot management information, etc.) for computing resources across the streaming arrays. As shown, each computational thread includes one or more computational nodes that provide hardware resources (e.g., processors, CPUs, GPUs, etc.). For example, while computational thread 235X of streaming array 225N is shown to include four computational nodes, it is understood that a rack assembly can include one or more computational nodes.

[0038] Each rack assembly is coupled to a cluster switch configured to provide communication, via a corresponding rack controller, with a cloud management controller 210 configured for management of a corresponding data center, as described above. The corresponding rack controller provides internal management of the resources of the corresponding rack assembly, which includes handling the internal details of the resources in the corresponding rack assembly and ensuring that those resources start up successfully and remain operational by performing management of the compute nodes, compute threads, streaming arrays, etc. For example, rack assembly 221N is coupled to cluster switch 240N. The cluster switch also provides communication to other rack assemblies (e.g., via the corresponding cluster switch) and to an external communication network (e.g., the Internet, etc.).

[0039] In one network configuration, each streaming array in a corresponding rack assembly provides high-speed access to corresponding network storage, as described above. In one embodiment, this high-speed access is provided via a PCIe fabric, which provides direct access between the compute nodes and the corresponding network storage. In other embodiments, high-speed access is provided via other network fabric topologies and / or networking protocols, including Ethernet, InfiniBand, remote RDMA over RoCE, etc. The compute nodes are capable of executing gaming applications and streaming audio / video of gaming sessions to one or more clients, and corresponding network storage (e.g., a storage server) holds the gaming applications, game data, and user data. For example, in streaming array 225A of rack assembly 221N, the high-speed access is configured to provide data and control path 201A between a particular compute node of a corresponding compute thread and corresponding network storage (e.g., storage 211A). Path 201N is also configured to communicate control and / or management information between network storage 211N and each compute node in streaming array 225N.

[0040] In another network configuration (not shown), each rack assembly provides high-speed access between compute nodes. For example, reconfiguration of the PCIe fabric within the rack assembly is performed to enable high-speed communication between compute nodes within the rack assembly, such as when forming a supercomputer or an artificial intelligence (AI) computer.

[0041] As previously described, cloud management controller 210 of data center 200B, in conjunction with rack controllers of corresponding rack assemblies, communicates with assigner 191 to allocate resources to client devices 110 to support gaming cloud systems 190′ and / or 190. In embodiments, allocation is performed based on asset awareness, such as knowing the resources and bandwidth needed and that are present in the data center. Thus, for illustrative purposes, embodiments of the present disclosure are configured to assign client devices 110 to particular compute nodes 232B of corresponding compute threads 231B of corresponding streaming arrays in rack assembly 221B.

[0042] FIG. 3 is a diagram of a compute sled 300 including multiple compute nodes arranged in a rack assembly configured for remote boot control of the compute nodes using a board management controller (e.g., via management control from a corresponding rack controller) and / or for network reconfiguration of the corresponding rack assembly at the compute node level, in accordance with one embodiment of the present disclosure.

[0043] Each computational thread 300 includes one or more computational nodes (e.g., nodes 1-4) arranged within a corresponding rack assembly. While FIG. 3 illustrates a computational thread including four computational nodes, it is understood that any number of computational nodes may be provided for a computational thread including one or more computational nodes. Computational thread 300 may provide a hardware platform (e.g., a circuit board) that provides computational resources (e.g., via the computational nodes). For example, computational thread 300 illustrates multiple computational nodes (e.g., nodes 1-4) and supporting hardware that supports the operation of the computational nodes. Computational thread 300 may be implemented within any of the streaming arrays and / or rack assemblies described above in FIGS. 2A-2B, or any other rack assembly, streaming array, and / or data center configuration.

[0044] Thread switch / management panel 315A integrates switching functionality provided by PCIe switchboard 320A and thread management functionality provided in part through board management controller (BMC) 350. For example, PCIe switchboard 320A provides communication with compute nodes via NT ports or the like. BMC 350 may also be directly connected to PCIe switchboard 320A (e.g., via a 1x channel or the like). As shown, panel 315A integrates multiple functions in one embodiment, although in other embodiments, those functions may be implemented separately in independent controllers.

[0045] In one network configuration, compute sled 300 provides compute nodes with high-speed access to network storage using PCIe (e.g., Gen4) communications, according to one embodiment of the present disclosure. In another network configuration, compute sled 300 is configured to provide high-speed communications between compute nodes in a compute sled within a rack assembly and between compute nodes in different rack assemblies. In one embodiment, a compute node includes multiple I / O interfaces. For example, a compute node may include an M.2 port and multiple lanes for PCIe Gen4 (bidirectional) communications via channels and / or connections 335. In particular, compute sled 300 includes PCIe switch board 320A, which provides eight PCIe lanes to an array-level PCIe fabric, such as via PCIe cable 305 (e.g., quad form factor, double density (QSFP-DD)). The PCIe (e.g., Gen4) interface (e.g., four lanes) can be used to expand the system with additional devices. In particular, the PCIe interface is used to connect to a PCIe fabric, including PCI Express switch 320A, for high-speed storage.

[0046] Additionally, each compute node is configured for Ethernet connectivity 311 (e.g., Gigabit Ethernet) via Ethernet patch panel 310A, which is configured to connect Ethernet cables between the compute nodes (e.g., nodes 1-4) and a rack-level network switch (not shown).

[0047] Additionally, board management controller (BMC) 350 is configured, in part, to manage one or more communication interfaces (e.g., PCIe, USB, UART, GPIO, I2C, I3C, etc.), each of which may be used to communicate with a compute node on compute thread 300 via a corresponding communication channel (e.g., 330). That is, BMC 350 is configured to manage multiple communication interfaces that provide communication to multiple compute nodes (e.g., compute nodes 1-4) on compute thread 300. For example, a corresponding compute node includes one or more universal asynchronous transmitter-receiver (UART) connections configured for transmitting and / or receiving serial data. In particular, one or more UART ports may be present, which are used for management purposes (e.g., connecting compute nodes to BMC 350). Ports may also be used for remote control operations such as "power on," "power off," and diagnostics. Additionally, another UART port may provide serial console functionality.

[0048] Thread management of resources within compute thread 300 may be facilitated via one or more communication channels that implement a management interface utilized by a corresponding rack controller of a corresponding rack assembly. The management interface enables transmission of data packets to various different devices within compute thread 300, such as from a corresponding rack controller. The management interface may implement rack management (e.g., through the execution of software) as provided by the rack controller, and the rack management software may run on the rack controller or may be implemented via network storage. The rack management software executes to manage the operation of the rack via the management interface, including powering on / off systems and resources of the compute sleds 300 (e.g., BMC 350, PCIe switch 320A, fan 390, power interposer board 340A, compute nodes 1-4, boot controllers on the compute nodes, etc.), performing firmware updates, obtaining thread status (e.g., temperature, voltage, fan speed, etc.), and providing access to the UART interfaces of individual compute nodes and / or BMC 350 (each compute node and / or BMC may have a UART or other interface / port, including I2C, I3C, etc.).

[0049] In particular, management control may be provided from a corresponding rack controller via a management interface via Ethernet connection 335 and / or via PCIe cable 305, as a primary embodiment. For example, management interface 327 may be implemented as a separate wire (e.g., two separate wires for transmit and receive) within PCIe cable 305 and may be used or further transmitted via I2C, I3C, UART, or some other interface. In particular, in addition to an Ethernet connection, the management interface may be provided via a lower-speed interface (e.g., implemented by BMC 350) implemented using either I2C, I3C, UART, or other communications interface to transmit control packets to devices within compute sled 300. In this manner, management control over the management interface may be provided via Ethernet (i.e., a first option) and / or via PCIe (i.e., a second option) via management interface 327 and / or I2C, I3C, UART, or other communications interface.

[0050] In one embodiment, the management interface can be provided directly by or through BMC 350, such as when receiving board management control signals from PCIe cable 305 via management interface 327, and delivered to the compute node through communication interface 330 (e.g., I2C, I3C, UART, etc.). For example, one PCIe lane can be reserved for communication with BMC 350 for board management control. As previously mentioned, BMC 350 is configured to manage the communication interface for the purpose of providing remote booting of the compute node. In particular, BMC 350 is configured to communicate with corresponding compute nodes on compute sleds 300 for the purpose of performing remote boot control of the compute nodes. For example, because external hardware is used in the compute node boot process, which is handled by communication between BMC 350 and the corresponding compute nodes, no storage is required on the compute nodes to store boot configuration files. This allows for flexibility as to which software is used for external booting of the compute nodes. That is, a compute node may be configured with one operating system that provides services corresponding to a first priority (e.g., primary services) and then later reconfigured with a second operating system that provides services corresponding to a second priority (e.g., secondary services).

[0051] In another embodiment, the management interface may be provided by or through another helper chip, such as a complex programmable logic device (CPLD) 325, which manages communication over the management interface (i.e., manages the transmission of data packets). Transmission through the CPLD 325 provides a robust backup mechanism for delivering board management control signals across the compute threads 300. For example, a backup mechanism may be implemented if Ethernet communication fails for any reason, or if the PCIe switch 320A fails, or if the BMC 350 crashes (e.g., its firmware becomes corrupted). Under such circumstances, the management interface 327 is configured to implement a backup mechanism for delivering board management control signals. In particular, management interface 327 connects first CPLD 325, which in turn is connected to BMC 350 or directly to an I2C, I3, UART, or other communication interface (e.g., via channel 330). That is, the backup mechanism allows BMC 350 to be bypassed, so that board management control signals from the rack controller are routed through CPLD 325 and delivered via channel 330. While the backup mechanism is highly robust and reliable (i.e., CPLD 325 is not susceptible to failures because it is implemented in dedicated hardware), associated performance may be limited in range and / or speed. However, the backup mechanism provides implementation of at least low-level operations (i.e., operations programmed via CPLD 325), including power control, reset operations, and some limited diagnostic collection.

[0052] BMC 350 may provide board or thread management functions. Furthermore, in embodiments, board management provided by BMC 350 includes control, monitoring, and management of the compute nodes, which is performed using signals such as universal asynchronous receiver / transmitter (UART), I2C, or I3C that deliver serial data (e.g., power on / off, diagnostic, and logging information) for each compute node over connection 330. In other embodiments, board management may be provided over other interfaces, such as Ethernet (e.g., for direct communication with a rack controller). Furthermore, BMC 350 can be configured to provide electromagnetic compatibility (EMC) control for controlling electromagnetic energy and debugging delays using UART signals. BMC 350 can be configured to provide control status information to management panel 330A, such as via control status light-emitting diodes (LEDs), which are further configured to display status using LEDs and buttons. BMC 350 is configured to monitor temperatures and voltages. BMC 350 is also configured to manage fans configured for cooling. BMC 350 is also configured to manage Ethernet connections for board management.

[0053] Compute sleds 300 include power interposer boards 340A configured to provide power to corresponding compute sleds via one or more bus bar connections. That is, BMC 350 may be configured to enable power control / delivery to the compute nodes via rack management bus 360 using general-purpose input / output (GPIO) and / or I2C and / or some other communication interface to power interposer 340A (e.g., from a 12 volt or higher bus bar to each compute node). In particular, power management functionality is used to manage and / or monitor power supply to each compute node via power interposer board 340A via connections. For example, rack management bus 360 may be configured to provide sled management control signals. Power monitoring (e.g., sensors connected via I2C channels to measure current, voltage, and other conditions through the power connections to the compute nodes) may be performed and communicated via rack management bus 360 to BMC 350, etc. The BMC 350, in coordination with a corresponding rack controller, may decide to power off individual components (e.g., compute nodes) or entire compute threads under certain conditions, including, for example, detecting too high a voltage, too high a temperature, inoperable fans 390 on compute threads 300, and other adverse conditions. In particular, the BMC 350 provides management information (e.g., power conditions) to the corresponding rack controller and provides thread management instructions. While the BMC 350 can operate independently of the rack controller under certain extreme conditions, it generally works in coordination with the rack controller for thread management. For example, thread management is performed externally (e.g., via the rack controller) and communicated internally via the rack management bus 360. This includes disabling threads under various conditions, such as during scheduled maintenance on a rack assembly, during a power outage in the corresponding data center due to a fire in the data center, or when a backup power source is implemented. Each compute node also includes a power input connector (eg, 12 volts for designed power consumption) connected to the power interposer 340A via a corresponding bus bar 370 connection.

[0054] 4 is a flow diagram 400 illustrating steps in a method for performing a remote boot and / or remote start-up of a computing resource, according to one embodiment of the disclosure. For example, flow diagram 400 may be implemented to perform a remote boot of a compute node of a compute thread including one or more compute nodes using the compute thread's board management controller (BMC) and the compute node's corresponding boot controller (e.g., boot controller 320-1 of compute node 1, boot controller 320-2 of compute node 2, boot controller 320-3 of compute node 3, or boot controller 320-4 of compute node 4 in compute thread 300 of FIG. 3). Notably, in one embodiment, external hardware is used in the compute node boot process so that no storage is required to boot the OS, allowing flexibility as to which software is used for external booting. In other embodiments, there may be storage on the partially configured compute node for booting the OS. While flow diagram 400 is described in the context of remote booting compute nodes of compute threads configured in a streaming array of rack assemblies located within a cluster of rack assemblies in a data center, such as the data centers of FIGS. 2A-2B , it will be understood that the operations described in flow diagram 400 may be generally applicable to remote booting of any computing resource using external hardware for storing boot files (e.g., BIOS, operating system configuration files, etc.).

[0055] Traditionally, system startup may involve performing one or more processes to initialize hardware and load an operating system used by the computer system to run applications. Generally, computer systems typically load and execute basic input / output system (BIOS) firmware from read-only, non-volatile memory (e.g., electrically erasable programmable read-only memory (EEPROM), flash memory, etc.) located on the computer system to initialize the computer system's hardware during system startup. Once the hardware is initialized, the BIOS firmware is used to boot an operating system from a local storage device, such as a hard drive or solid-state drive (SSD), by executing one or more boot loader programs directed by the BIOS firmware. Because the BIOS firmware used to load the operating system is located in read-only memory, changing the operating system requires changing the BIOS firmware on the local system and / or modifying settings that may be stored in local volatile memory (which may be battery-powered), which can be difficult to accomplish in a time sufficient to meet the real-time demands of clients and / or users.

[0056] Embodiments of the present disclosure provide for remote booting of a computing node of a computing sled (e.g., a game console) that has no internal storage and no BIOS used for booting on the platform (computing node). In some embodiments, there may optionally be minimal storage for a boot controller (located inside or outside the boot controller), while in other embodiments, there is no storage for a boot controller. That is, remote booting includes swapping the BIOS implementation and correspondingly swapping the OS implementation. For example, the computing node may be configured to play games using a backend streaming server of a cloud gaming system, and the computing node is arranged on a computing sled of a streaming array including multiple computing sleds. The streaming array is arranged in a rack assembly of a data center including multiple clustered rack assemblies, each rack assembly including one or more network storages and one or more streaming arrays. The use of external hardware in the boot process eliminates the need for storage, such as non-volatile memory, for storing BIOS firmware and / or other boot configuration files (e.g., less hardware reduces the cost of each compute node). A compute node may be configured with volatile memory, such as random access memory (RAM), that is used during the operation of the compute node (e.g., to run an operating system, applications, etc.). In this manner, remote booting of a compute node provides flexibility as to which software (e.g., selecting from among multiple BIOS firmwares) to boot externally, thereby providing flexibility in loading a desired operating system to use to run an application on the compute node. For example, if a compute node is configured for gaming, an operating system appropriate for the game is loaded to run the game application on the compute node (e.g., it may not use a BIOS). On the other hand, a compute node may be configured for non-gaming services that require a different operating system. In that case, the compute node may be configured to load a new operating system using a remote boot process to run other applications that provide those non-gaming services on the compute node. The required BIOS firmware required for gaming and non-gaming use cases differs significantly. Non-gaming use cases transform the compute node and / or system into something more akin to a personal computer (PC) running a typical or standard OS. In this case, the BIOS may be stored locally in non-volatile memory (e.g., SSD). On the other hand, gaming use cases are more specialized and do not necessarily require traditional BIOS firmware. For example, some functions associated with BIOS firmware (e.g., functions typically performed by the BIOS to implement an OS on a PC) are not necessary for gaming use cases, and BIOS functions for implementing a game console may be performed by the game OS itself during load without using dedicated BIOS firmware.

[0057] In particular, at 410, the method includes receiving, at a board management controller of the compute sled (e.g., BMC 350 of compute sled 300 of FIG. 3 ), a boot configuration instruction (from the cloud management controller) for booting a compute node with an operating system. The instruction may be delivered to the BMC by the cloud management controller via a corresponding rack controller. The cloud management controller may be configured to manage remote booting of compute nodes in the data center in coordination with the corresponding rack controller and to collect information related to booting the compute nodes, such as storing boot images for each compute node. More specifically, the BMC of the corresponding compute sled is configured to manage multiple communication interfaces providing communication with the multiple compute nodes. For example, the communication interfaces may include an inter-integrated circuit (I2C) bus interface (e.g., providing communication using PCI, PCIe, etc.), I3C, PCI, PCIe, UART, USB, etc.

[0058] In particular, the BMC communicates with a boot controller within the compute node using one or more of these communication interfaces. In one embodiment, the BMC may be configured on a system-on-chip (SoC), which includes a central processing unit (CPU) and other components. The BMC executes a corresponding OS and / or firmware that may be loaded from non-volatile memory (e.g., flash, ROM, etc.) located on the integrated circuit of the SoC. The BMC's OS may require management and / or updating. While the BMC may provide management and update functionality, management of the BMC may be provided via a side channel, such as using the management interface described above in connection with FIG. 3. Thus, the management interface may be used to manage and update the OS and / or firmware, and may also perform other operations on the BMC described above, including performing power on / off, power resets, etc. Additionally, the boot controller on a compute node may require firmware to function, which may be loaded via various methods depending on the hardware design. For example, in one embodiment, the boot controller may have its firmware integrated into a small ROM located on the motherboard of the corresponding compute node. In another embodiment, the boot controller firmware may be loaded remotely on the corresponding compute thread via the BMC, which pushes the boot controller firmware through a low-level interface via channel 330 (e.g., I2C, UART, serial peripheral interface (SPI), etc.).

[0059] The boot controller may be configured to load BIOS firmware onto the compute node's main CPU (e.g., RAM or system memory used by the CPU) according to the boot instructions. More specifically, in embodiments of the present disclosure, the BMC is configured to provide instructions to the boot controller to load the BIOS and operating system from one or more remote locations, where the instructions may be provided via PCIe, USB, UART, or another interface, and the BIOS firmware may be accessed via the PCIe fabric. In addition to triggering the boot process, the BMC may perform other operations that enable it to return information about the compute node's motherboard and / or boot controller. For example, the board information may include the board type, board version, and serial number. The board information may also include the compute node's CPU information, including the model number, serial number, hardware encryption key, and other information. This board information can be used to verify compatibility between BIOSes or operating systems booting on corresponding compute nodes. In another embodiment, the BIOS or operating system may be locked (e.g., by encryption) to a specific compute node system.

[0060] In particular, at 420, the method includes sending boot instructions from the BMC to a boot controller of the compute node via the selected communication interface to execute BIOS firmware stored remotely from (or external to) the compute node. That is, the BIOS firmware is stored on external hardware (i.e., external to) the compute node. For example, the BIOS firmware may be stored in memory of the BMC (memory on a BMC chip), in remote storage of a corresponding rack assembly, or in storage external to the corresponding rack assembly (e.g., remote storage accessible by one or more rack assemblies of a data center, such as storage addresses 260A-N in FIGS. 2A-2B).

[0061] In one embodiment, a pull method is implemented to access BIOS firmware located outside the compute node. Specifically, the pull method includes, at the BMC, receiving a request from the compute node to access a storage address storing the BIOS firmware. That is, the storage address of the BIOS firmware is known to the corresponding rack controller, and the storage address is included in the startup configuration instructions provided by the rack controller to the BMC, and the storage address is included in the boot instructions delivered from the BMC to the compute node. The pull method further includes the compute node facilitating access to the storage address to obtain the BIOS firmware. That is, the BMC forwards the request over an appropriate network to provide access and delivery of the BIOS firmware to the compute node. While the corresponding rack controller knows the internal details necessary for remote boot control of the corresponding compute node, a higher-level cloud management controller may provide high-level instructions, such as moving sleds and / or rack assemblies from one operating mode to another. In this manner, certain compute node management may be performed at the rack assembly level.

[0062] In another embodiment, a push method is implemented to access BIOS firmware located outside the compute node. In particular, the push method includes the BMC accessing the BIOS firmware at a storage address outside the compute node. That is, the storage address of the BIOS firmware is known to the corresponding rack controller, and the storage address is included in a boot configuration instruction provided to the BMC by the rack controller. In this manner, the BIOS firmware is accessible by the BMC from the storage address. The push method further includes transmitting the BIOS firmware from the BMC to the compute node in association with a boot instruction for executing the BIOS firmware by the compute node. For example, the BIOS firmware may be included in the boot instruction. In this manner, the compute node does not require an additional step to access the BIOS firmware.

[0063] In one embodiment, the communication interface is selected based on the operating systems initialized and loaded through execution of the BIOS firmware. For example, each operating system may be associated with a corresponding communication interface. Control signals for each operating system are delivered to the compute node via the corresponding communication interface. For example, control signals for a first operating system may be delivered via a first communication interface (e.g., I2C), control signals for a second operating system may be delivered via a second communication interface (e.g., UART), control signals for a third operating system may be delivered via a third communication interface (e.g., USB), etc. In this manner, boot instructions delivered from the BMC to the boot controller of the compute node based on the operating system are delivered via a corresponding communication interface managed by the BMC for communicating control signals for that operating system. Thus, the BMC may be configured as a switch that delivers control signals for each of multiple operating systems loadable by the compute node via a corresponding communication interface selected from the multiple communication interfaces managed by the BMC. That is, the BMC sled switch board provides a communication interface to the compute node, such as PCI, PCIe, I2C, UART, USB, etc.

[0064] At 430, the method includes executing BIOS firmware on the compute node (CPU) to initiate loading of an operating system to be executed by the compute node. As previously mentioned, execution of the BIOS firmware (e.g., a first-stage boot loader) may include execution of other associated firmware and / or software or programs (e.g., second-stage boot loader(s)) to complete loading of the operating system for use by the compute node. That is, in some cases, the BIOS firmware may be complete and may not require any additional firmware and / or software or programs to complete loading of the operating system, while in other cases, the BIOS firmware may be used to call and / or execute other firmware and / or software or programs.

[0065] For purposes of explanation only, execution of the BIOS firmware may include initializing system hardware and loading an operating system. In particular, the BIOS firmware may perform a power-on self-test (POST) process to initialize system hardware (e.g., video cards and other hardware devices), may perform memory tests to configure memory and drive parameters, configure plug-and-play devices, and may identify any boot devices for loading a subsequent boot loader (e.g., second stage, etc.) that is executed to load an operating system (e.g., an operating system configuration file).

[0066] The method further includes loading at least a portion of the operating system into system memory (e.g., RAM) of the compute node from a storage address, the storage address including one or more operating system configuration files. In one embodiment, the entire operating system may be loaded into system memory. However, some operating systems may be very large, and it may be more efficient to load a portion of the operating system into system memory so that the remaining portion can be used and / or loaded as needed. For example, a first portion of the operating system may be loaded and executed in system memory of the compute node. Additionally, a second portion of the operating system may be stored in corresponding network storage such that it can be accessed by the compute node (e.g., via a PCIe fabric) while the compute node is executing the operating system. For example, data and / or files in the second portion may be accessed and loaded into system memory as needed, or may be accessed remotely while the operating system is running.

[0067] In another embodiment, the cloud management controller, in coordination with a corresponding rack controller, is configured to manage one or more images of operating systems running on the compute nodes. That is, the cloud management controller and / or corresponding rack controller of the data center is configured to manage software images (e.g., operating system images) for each compute node. In this manner, the cloud management controller and / or corresponding rack controller manages the operating systems running on the various compute nodes and rack assemblies available in the data center. For example, the management information may be stored in storage 260 of FIG. 2A or 2B.

[0068] In one embodiment, during a switch of a compute node's operating system, one or more jobs executing on the compute node prior to the switch may be paused (or paused) and / or transferred for continued execution. For example, a job may be executed by a compute node executing an application using a first operating system that is different from the new operating system loaded via remote boot. To illustrate, the job may relate to a gaming application, and the gaming application may be paused to be started at a later time or transferred to another compute node for seamless transfer and execution. For example, the state (e.g., configuration) of a compute node during the execution of a job is captured and stored. The state is transferred to another compute node, such as on the same compute thread, or on the same streaming array, or on the same rack assembly, or on a compute node in a different rack assembly in the same data center, or on a compute node in a different data center. The state is initialized on the other compute node, and then the application is executed on the new compute node using the transferred state, and the job is resumed. In one embodiment, a paused job is paused at a selected pause or suspend point that is not a predetermined pause point, such as the end of a level in a game application.

[0069] Alternative jobs (e.g., enabled via remote boot) running on a compute node during the dark time utilization of the compute node may be paused and / or transferred to another data center to continue running, such as to a less utilized data center or to a less utilized portion of the same data center. For example, these alternative jobs may include machine learning workloads, or some other type of CPU / GPU workload, or video encoding workload, etc. In various implementations, the decision to pause and / or transfer these alternative jobs may depend on what business model is implemented or selected for these alternative workloads. Some cost-effective jobs purchased by a customer may be performed by purchasing a fixed amount of computing time (i.e., not including any transfers) and may run during different dark time periods. Other customers may pay for a fixed number of jobs to be completed during a period (e.g., a dark time period, a period exceeding a dark time period) and require pausing and transferring from one resource to another. In another scenario, jobs may be paused without being transferred when more expensive jobs require processing. In this manner, embodiments of the present disclosure support pausing and transferring jobs under various scenarios.

[0070] A key difference between gaming use cases and dark time utilization cases (i.e., running replacement jobs) is scale. In particular, gaming workloads typically involve a single user occupying a single compute node, whereas multiple compute nodes support multiple users in a one-to-one relationship. In a dark time utilization use case, the single user may be, for example, an organization, potentially requiring the use of many server racks or rack assemblies, if not an entire data center. For security reasons, resources for organizational users in a dark time utilization use case are allocated on a rack assembly-by-rack assembly basis (e.g., rack controllers and network storage) to ensure that resources allocated for dark time utilization for that organization are isolated from other customers. This may involve placing those resources on a different network with different firewall rules, etc., requiring reconfiguration of those resources. It may also involve providing resources with additional network access, such as access to a different computer network of the customer, requesting organization, or another organization, to access data.

[0071] Purely for purposes of example, the paused job may be associated with a game, or a video game or game application. In particular, an ongoing game may be paused at any point within the game, and the game may be resumed at the same point in the game at some future point (e.g., immediately after transfer or later). The game state of the game and the configuration parameters of the compute nodes on which the game is running are captured and saved. That is, the paused game state is saved with enough data to reconstruct the game state when the game is resumed. While the game is so paused, the game state is collected and saved in storage, thereby eliminating the need for the cloud gaming system to store the state in active memory or registers of the hardware. This frees up the system for other gameplay or allows the compute nodes to provide other services (e.g., dark time utilization of the compute nodes), and allows gameplay to be resumed at any time, forming any remote compute node (e.g., client). When a game is resumed, the game state is loaded onto a new compute node (e.g., the same or a different compute node) tasked with resuming the game. Loading the game state involves generating the game state from multiple saved files and data structures, and the reconstructed game state can leave the compute node in the same, or nearly the same, state as when the game was paused.

[0072] For example, a cloud gaming system may include a cloud management controller, a rack controller, storage, and multiple computing nodes managed by the cloud gaming system and coupled via a network. Each computing node configured as a game console may include a hardware layer, an operating system layer, and an application layer. The operating system layer is configured to interact with the hardware layer, and the operating system layer includes a state manager. The application layer is configured to instruct at least a portion of the operating system and a portion of the hardware layer. The application layer includes a game, and the state manager is configured to capture game state data of the computing node and store the captured game state data when the game is paused. The state manager is also configured to apply the game state data to the same or a different computing node to resume the game at the point where the game was paused.

[0073] 5 is a flow diagram 500 illustrating steps in a method for reconfiguring a rack assembly of a data center, according to one embodiment of the present disclosure. Flow diagram 500 may be implemented within the data storage system shown in FIGS. 2A-2B. In particular, the flow diagram may be implemented to provide for reconfiguration of a rack assembly at any appropriate time.

[0074] For example, reconfiguration of rack assemblies may be performed to enable utilization of computing resources during dark times in a cloud gaming system configured primarily for gaming. That is, in certain geographic areas served by a regional data center, there are periods of low or inactive computing resources (i.e., dark times) during which computing resources are underutilized, such as during work or school hours, large social gatherings (e.g., local football games), or late nights or early mornings (e.g., sleep time) when users are not playing games on the cloud gaming system. Embodiments of the present disclosure provide for utilization of computing resources during those dark times or periods. As such, monetization of computing resources during those dark times is important for cost reasons (i.e., to increase the net profit of each computing resource within a 24-hour period).

[0075] Thus, one or more rack assemblies in a data center may be repurposed to process different workloads other than gaming during dark times. Generally, it is more efficient to convert all of the computing nodes in the corresponding rack assembly to perform similar tasks, e.g., tasks performed during dark times, and a network reconfiguration of the rack assembly may be required to accomplish those tasks. Because all of the computing nodes may be newly configured to process new tasks performed during dark times, the reconfiguration of the rack assembly does not adversely affect any of the computing nodes that may be performing tasks (e.g., gaming) under the previous network configuration of that rack assembly. That is, the computing nodes processing games may be transferred to other rack assemblies in the same data center or a different data center, as described above.

[0076] At 510, the method includes sending a reconfiguration message from the cloud management controller to a rack controller of the rack assembly to reconfigure the rack assembly from a first configuration to a second configuration. Other resources within the data center may also be reconfigured, including resources external to the cloud gaming rack assembly. In particular, some use cases may require network / Internet connectivity to other networks of third-party cloud hosting companies (e.g., AWS Cloud Computing) or other networks. This connectivity may be established as part of the network configuration. In other cases, reconfiguration of resources within the data center may be performed for security purposes to isolate harmful workloads (e.g., from hackers, or external companies requesting services, or untrusted customers) and / or prevent them from in any way adversely affecting the overall functionality and integrity of the data center (e.g., primarily providing cloud gaming services). As an example, reconfiguration may be utilized to combat a distributed denial of service (DDOS) or hacking of portions of one or more data centers by an external party or untrusted customer. In a first configuration, the rack assembly is configured to facilitate first priority services implemented by a first plurality of applications (e.g., games). In a second configuration, the rack assembly is configured to facilitate second priority services implemented by a second plurality of applications (e.g., running artificial intelligence or AI services). The second priority services (e.g., AI) may have a lower priority than the first priority services (gaming), such that the rack assembly primarily supports gaming, but when demand for gaming decreases (i.e., dark times), the rack assembly may be configured to make more computing resources available throughout a 24-hour period by handling the second priority services.

[0077] In one embodiment, the rack assembly includes one or more network storages and one or more streaming arrays, as described above in Figures 2A-2B, each of which includes one or more computational threads, each of which includes one or more computational nodes.

[0078] A rack assembly may be sent a reconfiguration message when it is deemed appropriate to switch to a different configuration. For example, reconfiguration of one or more rack assemblies may occur at a specific, predetermined time (e.g., after 2:00 AM, indicating the start of dark time) or may occur in anticipation of when dark time will occur based on projected usage of computing resources in the data center. In one embodiment, dark time is measured through metrics, and the cloud management controller is configured to detect a decrease in demand for a first priority service supported by a data center including multiple rack assemblies. For example, each of the multiple rack assemblies is configured in a first configuration that facilitates a first priority service (e.g., gaming).

[0079] Generally, switching to a different configuration of a rack assembly may occur during dark times. However, there may be periods outside of dark times when it may be beneficial to perform a reconfiguration of the rack assembly. For example, if a second-priority service (e.g., non-gaming workload) generates more profit than a first-priority service (e.g., gaming), a priority inversion may occur, in which case the second-priority service is prioritized over the first-priority service (e.g., gaming). That is, the data center may lower the priority of the first-priority service (e.g., gaming).

[0080] In one embodiment, a cloud management controller manages the network configuration of multiple rack assemblies. For example, the cloud management controller may maintain an inventory of the number of rack assemblies and the number of compute nodes configured for or available for gaming (e.g., first priority service), such as within a particular geographic region. As such, when demand for gaming decreases over a particular period of time, indicating the onset of dark time, the cloud management controller may reconfigure one or more rack assemblies to support second priority service over gaming, while maintaining a sufficient number of rack assemblies to support current and forecasted gaming demand by users of the data center.

[0081] At 520, the method includes configuring the rack assembly in a second configuration so that the rack assembly can run applications that support second priority services (e.g., AI), thereby utilizing compute nodes that would otherwise be idle.

[0082] In one embodiment, reconfiguring a rack assembly may include rebooting the compute nodes of the rack assembly with a different operating system. Generally, it may be beneficial for the compute nodes of the rack assembly to run the same operating system, and even the same version of the same operating system. For example, the different operating system may be configured to support secondary priority services (e.g., running applications during dark times). Rebooting the compute nodes may include the operations previously described in FIG. 4 and may include receiving a startup configuration instruction at the BMC to boot the compute nodes with a new operating system. The compute nodes are arranged on a sled that includes one or more compute nodes, and the BMC is configured to manage multiple communication interfaces that provide communication to the compute nodes. The rebooting process may include sending a boot instruction from the BMC via the communication interface to a boot controller of the compute node to execute basic input / output system (BIOS) firmware. The rebooting process may include executing BIOS firmware on the compute node to load the operating system to be executed by the compute node.

[0083] In one embodiment, reconfiguring a rack assembly, which includes rebooting the compute nodes of the rack assembly to a different operating system, may include loading new operating system software or other software components into the rack assembly's corresponding network storage, or into the cloud management controller or other storage accessible by the cloud management controller.

[0084] In one embodiment, reconfiguring a rack assembly may include changing the network configuration of the rack assembly. That is, the network configuration of the rack assembly may need to be reconfigured to provide communication between the rack assembly and a remote rack assembly, where the network configuration defines the communication path between the rack assembly and another rack assembly. For example, external access of a compute node using a virtual local area network may need to be reconfigured, and correspondingly, a firewall may need to be reconfigured to provide appropriate access to other remote rack assemblies (e.g., compute nodes located on the remote rack assemblies). As previously mentioned, resources within the data center may be reconfigured during network configuration to include access to resources outside of the cloud gaming rack assembly, such as providing network / internet connectivity to other networks of third-party cloud hosting companies (e.g., AWS Cloud Computing) or other networks. Resource reconfiguration may also be performed for security purposes, such as to isolate harmful workloads (e.g., DDOS, hacking workloads, or untrusted external companies or untrusted customers requesting services) and / or prevent them from adversely affecting in any way the overall functionality and integrity of the data center (e.g., primarily providing cloud gaming services).

[0085] In one embodiment, reconfiguring the rack assembly may include modifying the boot storage architecture, for example, a network reconfiguration may provide faster access to boot or operating system files for use during boot and / or execution of the corresponding operating system.

[0086] In one embodiment, reconfiguring the rack assembly may include reconfiguring a PCIe (or equivalent) fabric to enable high-speed communication between compute nodes to form a "supercomputer" or AI computer where high-speed communication between the compute nodes is required. For example, the rack assembly may be initially configured for direct access between multiple compute nodes and at least one network storage of the rack assembly, such as for in-game implementation. However, during dark times, the rack assembly may be reconfigured in a second configuration. For example, the second configuration may configure a PCIe fabric for direct communication between multiple compute nodes in a rack assembly instead of providing direct access to network storage (e.g., in the first configuration). Thus, in the second configuration, the PCIe fabric facilitates direct communication between a first compute node and a second compute node in the rack assembly. Additionally, rack reconfiguration may be performed to establish a special network configuration between the rack assemblies. Typically, cloud gaming rack assemblies operate independently of each other, but in some use cases, two or more rack assemblies may be interconnected and operate in conjunction through reconfiguration. Typically, without reconfiguration, network settings such as firewalls and / or cluster-level switches prevent such cooperation.

[0087] FIG. 6 illustrates components of an exemplary device 600 that can be used to implement aspects of various embodiments of the present disclosure. For example, FIG. 6 illustrates an exemplary hardware system suitable for providing remote boot control of compute nodes, such as compute nodes of a streaming array of compute threads in a rack assembly, and network reconfiguration of a rack assembly in a data center, in accordance with embodiments of the present disclosure. The block diagram illustrates device 600, which may incorporate or be any of a personal computer, a server computer, a gaming console, a mobile device, or other digital device, each of which is suitable for practicing embodiments of the present invention. Device 600 includes a central processing unit (CPU) 602 for executing software applications and optionally an operating system. CPU 602 may be comprised of one or more homogeneous or heterogeneous processing cores.

[0088] According to various embodiments, CPU 602 is one or more general-purpose microprocessors having one or more processing cores. Further embodiments can be implemented using one or more CPUs with a microprocessor architecture specifically adapted for highly parallel, computationally intensive applications (e.g., media and interactive entertainment applications) configured to perform graphics processing during game execution.

[0089] Memory 604 stores applications and data used by CPU 602 and GPU 616. Storage 606 provides non-volatile storage and other computer-readable media for applications and data and may include fixed disk drives, removable disk drives, flash memory devices, and CD-ROM, DVD-ROM, Blu-ray, HD-DVD, or other optical storage devices, as well as signal transmission and storage media. User input device 608 communicates user input from one or more users to device 600 and may include, for example, a keyboard, mouse, joystick, touchpad, touchscreen, still or video recorder / camera, and / or microphone. The network interface 609 enables the device 600 to communicate with other computer systems over electronic communications networks, which may include wired or wireless communications over local area networks and wide area networks such as the Internet. The audio processor 612 is adapted to generate analog or digital audio output from instructions and / or data provided by the CPU 602, memory 604, and / or storage 606. The components of the device 600, including the CPU 602, graphics subsystem including the GPU 616, memory 604, data storage 606, user input devices 608, network interface 609, and audio processor 612, are connected via one or more data buses 622.

[0090] Graphics subsystem 614 is further connected to data bus 622 and the components of device 600. Graphics subsystem 614 includes at least one graphics processing unit (GPU) 616 and graphics memory 618. Graphics memory 618 includes display memory (e.g., a frame buffer) used to store pixel data for each pixel of an output image. Graphics memory 618 may be integrated into the same device as GPU 616, connected as a separate device from GPU 616, and / or incorporated within memory 604. Pixel data may be provided to graphics memory 618 directly from CPU 602. Alternatively, CPU 602 may provide data and / or instructions defining desired output images to GPU 616, from which GPU 616 generates pixel data for one or more output images. Data and / or instructions defining desired output images may be stored in memory 604 and / or graphics memory 618. In an embodiment, GPU 616 includes 3D rendering functionality for generating pixel data for output images from instructions and data defining scene geometry, lighting, shading, texturing, motion, and / or camera parameters. GPU 616 may further include one or more programmable execution units capable of executing shader programs.

[0091] Graphics subsystem 614 periodically outputs pixel data for images from graphics memory 618 to be displayed on display device 610 or projected by a projection system (not shown). Display device 610 can be any device capable of displaying visual information in response to signals from device 600, including CRT, LCD, plasma, and OLED displays. Device 600 can provide analog or digital signals to display device 610, for example.

[0092] In other embodiments, graphics subsystem 614 includes multiple GPU devices coupled to perform graphics processing for a single application running on a corresponding CPU. For example, multiple GPUs may perform multi-GPU rendering of geometry for an application by pre-testing geometry against a region of the screen (which may be interleaved) before rendering objects for an image frame. In another example, multiple GPUs may perform interleaved frame rendering, where GPU1 renders a first frame, GPU2 renders a second frame, and so on, in successive frame cycles, until the last GPU is reached. The first GPU then renders the next video frame (e.g., if only two GPUs are present, GPU1 renders the third frame). That is, the GPUs perform rotation operations as they render frames. Rendering operations may overlap, and GPU2 may begin rendering the second frame before GPU1 finishes rendering the first frame. In another implementation, different shader operations in the rendering and / or graphics pipeline may be assigned to multiple GPU devices. A master GPU performs primary rendering and compositing. For example, in a group including three GPUs, master GPU1 may perform main rendering (e.g., a first shader operation) and compositing the output from slave GPU2 and slave GPU3, slave GPU2 may perform a second shader (e.g., a fluid effect such as a river) operation, slave GPU3 may perform a third shader (e.g., particle smoke) operation, and master GPU1 may composite the results from each of GPU1, GPU2, and GPU3. In this manner, various GPUs may be assigned to perform various shader operations (e.g., flag flapping, wind, smoke generation, fire, etc.) to render a video frame. In yet another embodiment, each of the three GPUs may be assigned to a different object and / or portion of a scene corresponding to a video frame. In the above-described embodiments and implementations, these operations may occur in the same frame cycle (concurrently in parallel) or in different frame cycles (sequentially in parallel).

[0093] Accordingly, this disclosure describes methods and systems configured to provide remote boot control of computing nodes, such as computing nodes of a streaming array of computing threads in a rack assembly, and network reconfiguration of a rack assembly in a data center.

[0094] It should be understood that the various embodiments defined herein may be combined or assembled into specific implementations that use various features disclosed herein. Thus, the examples provided are only some of the possible examples and are not intended to limit the various implementations that are possible by combining various elements to define more implementations. In some instances, some implementations may include fewer elements without departing from the spirit of the disclosed or equivalent implementations.

[0095] Embodiments of the present disclosure may be practiced with a variety of computer system configurations, including handheld devices, microprocessor systems, microprocessor-based or programmable consumer electronics, minicomputers, mainframe computers, etc. Embodiments of the present disclosure may also be practiced in distributed computing environments where tasks are performed by remote processing devices that are linked through a wire-based or wireless network.

[0096] With the foregoing embodiments in mind, it should be understood that embodiments of the present disclosure can employ various computer-implemented operations involving data stored in computer systems. These operations are operations requiring physical manipulation of physical quantities. Any of the operations described herein that form part of embodiments of the present disclosure are useful machine operations. Embodiments of the present disclosure also relate to devices or apparatus for performing these operations. An apparatus can be specially constructed for the required purpose, or the apparatus can be a general-purpose computer selectively activated or configured by a computer program stored in the computer. In particular, various general-purpose machines can be used with computer programs written in accordance with the teachings herein. Alternatively, it may be more convenient to construct a more specialized apparatus to perform the required operations.

[0097] The present disclosure can also be embodied as computer-readable code on a computer-readable medium. A computer-readable medium is any data storage device that can store data, which can thereafter be read by a computer system. Examples of computer-readable media include hard drives, network-attached storage (NAS), read-only memory, random-access memory, CD-ROMs, CD-Rs, CD-RWs, magnetic tape, and other optical and non-optical data storage devices. The computer-readable medium can also include computer-readable tangible media distributed over network-connected computer systems so that the computer-readable code is stored and executed in a distributed fashion.

[0098] Although the method operations have been described in a particular order, it should be understood that other housekeeping operations may be performed between operations, or operations may be adjusted to occur at slightly different times, or may be distributed within a system that allows processing operations to occur at various intervals relative to processing, so long as the processing of the overlay operation is performed in the desired manner.

[0099] Although the foregoing disclosure has been described in some detail for clarity of understanding, it will be apparent that certain changes and modifications may be practiced within the scope of the appended claims. Accordingly, the present embodiments are to be considered as illustrative and not restrictive, and embodiments of the present disclosure are not limited to the details provided herein, but may be modified within the scope of the appended claims and their equivalents.

Claims

1. 1. A method for performing a system boot, comprising: receiving, at a board management controller (BMC), a boot configuration instruction for booting the compute node with an operating system; The computing nodes are arranged in a thread including a plurality of computing nodes; the BMC is configured to manage a plurality of communication interfaces that provide communications to the plurality of compute nodes; transmitting a boot command from the BMC to a boot controller of the compute node via a first communication interface of the plurality of communication interfaces to execute basic input / output system (BIOS) firmware stored externally to the compute node; executing the BIOS firmware on the compute node to initiate loading of the operating system to be executed by the compute node; selecting the first communication interface based on the operating system; the BMC is configured as a switch that delivers a control signal to each of a plurality of operating systems via a corresponding communication interface selected from the plurality of communication interfaces; the operating system control signals are delivered to the compute node via the first communication interface; pausing a job being executed by the compute node while the compute node is executing an application using a first operating system that is different from the operating system loaded for execution; capturing a state of the compute node and storing the state; transferring the state of the compute node to another compute node; and program instructions for executing the application on the other compute node to resume the job; the job is paused at a selected pause point that is not a predetermined pause point; method.

2. receiving, from the computing node, a request to access a storage address storing the BIOS firmware at the BMC; facilitating access by the compute node to the storage address to retrieve the BIOS firmware; the storage address is included in the boot command from the BMC; The method of claim 1.

3. The sending of the boot command comprises: accessing the BIOS firmware at a storage address by the BMC; the storage address is included in the boot configuration command from the BMC; transmitting the BIOS firmware from the BMC to the computing node in association with the boot instructions for the execution of the BIOS firmware by the computing node; The method of claim 1.

4. The execution of the BIOS firmware includes: loading at least a portion of the operating system from a storage address into a system memory of the compute node; The method of claim 1.

5. loading the at least a portion of the operating system from a storage address into a system memory of the compute node, loading a first portion of the operating system into the system memory of the compute node for execution; loading a second portion of the operating system into network storage for access by the compute node over a PCIe fabric during said execution of the operating system by the compute node; The method of claim 4.

6. the thread is one of a plurality of threads of a streaming array; each rack assembly of the plurality of rack assemblies including one or more network storages and one or more streaming arrays; each of the one or more streaming arrays includes one or more threads; each computational thread in the one or more threads includes one or more computational nodes; The method of claim 1.

7. and managing, at a cloud management controller at a data center, one or more images of the operating system on the compute nodes. The method of claim 1.

8. sending a reconfiguration message from the cloud management controller to a rack controller of the rack assembly to reconfigure the rack assembly from a first configuration to a second configuration; the rack assemblies in the first configuration are configured to facilitate a first priority service implemented by a first plurality of applications; the second configuration is configured to facilitate a second priority service implemented by a second plurality of applications; the second priority service has a lower priority than the first priority service; configuring the rack assembly in the second configuration; the rack assembly includes one or more network storage arrays and one or more streaming arrays; each of the one or more streaming arrays includes one or more computational threads; each of the one or more computational threads includes one or more computational nodes; In the configuration of the rack assembly in the second configuration, reconfiguring the PCIe fabric in the second configuration for direct communication between a plurality of compute nodes in the rack assembly, such that the PCIe fabric facilitates direct communication between a first compute node and a second compute node in the rack assembly; The method, wherein in the first configuration, the PCIe fabric is configured for direct access between the plurality of compute nodes and at least one network storage in the rack assembly.

9. detecting, at the cloud management controller, a decrease in demand for the first priority service supported by a data center including a plurality of rack assemblies; each of the plurality of rack assemblies configured in the first configuration that facilitates the first priority of service; the cloud management controller manages the configuration of the plurality of rack assemblies; The method of claim 8.

10. The configuration of the rack assembly in the second configuration includes: rebooting each of the plurality of compute nodes of the rack assembly to execute an operating system configured to support execution of the second priority service; Rebooting the compute node receiving, at a board management controller (BMC), a boot configuration instruction for booting the compute node with the operating system; the computing nodes are arranged in a computing thread that includes a second plurality of computing nodes; the BMC is configured to manage a plurality of communication interfaces that provide communication to the second plurality of compute nodes; sending a boot command from the BMC to a boot controller of the compute node via a first communication interface of the plurality of communication interfaces to execute basic input / output system (BIOS) firmware; executing the BIOS firmware on the computing node to load the operating system to be executed by the computing node; The method of claim 8.

11. The configuration of the rack assembly in the second configuration includes: changing a network configuration between the rack assembly and another rack assembly; the network configuration defines a communication path between the rack assembly and the other rack assembly; The method of claim 8.

12. A computer-readable medium storing a computer program for executing system startup, program instructions for receiving, at a board management controller (BMC), startup configuration instructions for booting a compute node with an operating system, the instructions comprising: The computing nodes are arranged in a thread including a plurality of computing nodes; the BMC is configured to manage a plurality of communication interfaces that provide communications to the plurality of compute nodes; transmitting a boot command from the BMC to a boot controller of the compute node via a first communication interface of the plurality of communication interfaces to execute basic input / output system (BIOS) firmware stored externally to the compute node; program instructions for executing the BIOS firmware on the computing node to initiate loading of the operating system for execution by the computing node; further comprising program instructions for selecting the first communication interface based on the operating system; the BMC is configured as a switch that delivers a control signal to each of a plurality of operating systems via a corresponding communication interface selected from the plurality of communication interfaces; the operating system control signals are delivered to the compute node via the communication interface; program instructions for pausing a job being executed by the compute node during execution of an application by the compute node using a first operating system that is different from the operating system loaded for execution; program instructions for capturing a state of the compute node and storing the state; transferring the state of the compute node to another compute node; and program instructions for executing the application on the other compute node to resume the job; the job is paused at a selected pause point that is not a predetermined pause point; Computer-readable medium.

13. program instructions for receiving, at the BMC, a request from the computing node to access a storage address storing the BIOS firmware; and program instructions for facilitating access to the storage address by the computing node to obtain the BIOS firmware; the storage address is included in the boot command from the BMC; The computer-readable medium of claim 12.

14. The program instructions for transmitting the boot command include: accessing, by the BMC, the BIOS firmware at a storage address; the storage address is included in the boot configuration command from the BMC; and program instructions for transmitting the BIOS firmware from the BMC to the computing node in association with the boot instructions for the execution of the BIOS firmware by the computing node. The computer-readable medium of claim 12.

15. The program instructions for executing the BIOS firmware include: program instructions for loading at least a portion of the operating system from a storage address into a system memory of the compute node; The computer-readable medium of claim 12.

16. 1. A method for performing a system boot, comprising: receiving, at a board management controller (BMC), a boot configuration instruction for booting the compute node with an operating system; The computing nodes are arranged in a thread including a plurality of computing nodes; the BMC is configured to manage a plurality of communication interfaces that provide communications to the plurality of compute nodes; transmitting a boot command from the BMC to a boot controller of the compute node via a first communication interface of the plurality of communication interfaces to execute basic input / output system (BIOS) firmware stored externally to the compute node; executing the BIOS firmware on the compute node to initiate loading of the operating system to be executed by the compute node; pausing a job being executed by the compute node while the compute node is executing an application using a first operating system that is different from the operating system loaded for execution; capturing a state of the compute node and storing the state; transferring the state of the compute node to another compute node; and program instructions for executing the application on the other compute node to resume the job; A method wherein the job is paused at a selected pause point that is not a predetermined pause point.

Citation Information

Patent Citations

  • Information processor having crossbar switch and method for controlling crossbar switch

    JP1999232237A

  • Method and system for verifying the proper operation of computing devices after system changes

    JP2015512535A

  • D / a converter

    US20150263743A1

  • Mechanism for hardware configuration and software deployment

    US20190020540A1

  • Compression techniques for distributed data

    US20190196907A1