Reconfigurable Computing Pods Using Optical Networks
Optical networks dynamically reconfigure compute nodes into workload clusters, addressing scalability and failure issues in supercomputers by enhancing availability and performance through fault isolation and reduced delay.
Patent Information
- Application Number
- JP2023035587
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Priority Date
- 2019-04-11
- Filing Date
- 2023-03-08
- Publication Date
- 2025-10-02
- Estimated Expiration
- 2039-12-18
AI Technical Summary
Fixed configurations of supercomputers with interconnection networks limit scalability, availability, and performance due to failure of processing nodes, and vary in performance based on configuration.
Utilize optical networks to reconfigure compute nodes into workload clusters by selecting building blocks that match a specified target configuration and dynamically route data through optical circuit switches.
Enhances availability and performance by allowing easy replacement of failed nodes, provides fault isolation, and reduces data transmission delay, while enabling flexible and secure workload execution.
Smart Images

Figure 0007748409000001 
Figure 0007748409000002 
Figure 0007748409000003
Abstract
Description
[Background technology]
[0001] background Some computational workloads, such as machine learning training, require many processing nodes to efficiently process the workload. The processing nodes can communicate with each other through an interconnection network. For example, in the case of machine learning training, the processing nodes can converge on an optimal deep learning model by communicating with each other. The interconnection network is important for the speed and efficiency with which the processing units achieve convergence.
[0002] As machine learning and other workloads vary in size and complexity, a fixed configuration of a supercomputer including multiple processing nodes may limit the availability, scalability, and performance of the supercomputer. For example, if some processing nodes of a supercomputer having a fixed interconnection network connecting a particular configuration of multiple processing nodes fail, the supercomputer will be unable to replace the failed processing nodes, resulting in reduced availability and performance. Also, the performance of some particular configurations may be higher than the performance of other configurations regardless of the failed nodes. Summary of the Invention [Means for solving the problem]
[0003] overview This document describes techniques that allow optical networks to be used to reconfigure superports of compute nodes to create workload clusters.
[0004] Generally, one inventive aspect of the subject matter described herein can be embodied in a method including receiving request data specifying compute nodes required to execute a computational workload. The request data specifies an n-dimensional target configuration of compute nodes, where n is two or more. From a superpod including a set of building blocks, each including an m-dimensional configuration of compute nodes, a subset of building blocks that, when combined, matches the n-dimensional target configuration specified by the request data is selected. The set of building blocks is connected to an optical network including one or more optical circuit switches for each of the n dimensions. A workload cluster is generated of compute nodes that include the subset of building blocks. A workload cluster is a cluster of compute nodes dedicated to computing or executing a particular computational workload. The generating includes, for each dimension of the workload cluster, configuring respective routing data for one or more optical circuit switches for that dimension. The routing data, respectively corresponding to each dimension of the workload cluster, specifies how data of the computational workload is to be routed between the compute nodes along each dimension of the workload cluster. The compute nodes in the workload cluster execute the computational workload.
[0005] These and other implementations may include one or more of the following features, as desired: In some aspects, the request data specifies different types of compute nodes, and selecting the subset of building blocks may include, for each type of compute node specified by the request data, selecting a building block that includes one or more compute nodes of the specified type.
[0006] In some embodiments, the respective routing data for each dimension of the superpod is stored in an optical circuit switch routing table for one of the one or more optical circuit switches. In some aspects, the optical network includes, for each dimension of the n dimensions, one or more optical circuit switches of the optical network that route data between compute nodes along that dimension. Each building block may include multiple segments of compute nodes along each dimension of the building block. The optical network may include, for each segment of each dimension, an optical circuit switch of the optical network that routes data between compute node segments corresponding to each building block in a workload cluster.
[0007] In some embodiments, each building block comprises one of a three-dimensional torus-shaped computing node or a mesh-shaped computing node. In some embodiments, the superpod comprises multiple workload clusters, each of which comprises a different subset of the building blocks and can execute a different workload than the other workload clusters.
[0008] Some aspects include receiving data indicating that a particular building block of a workload cluster has failed and replacing the particular building block with an available building block. Replacing the particular building block with the available building block can include updating data routing of one or more optical circuit switches of the optical network to stop data routing between the particular building block and one or more other building blocks of the workload cluster and updating data routing of the one or more optical circuit switches of the optical network to route data between the available building block and the one or more other building blocks of the workload cluster.
[0009] In some aspects, selecting a subset of building blocks that, when combined, matches the n-dimensional target configuration specified by the request data includes determining that the n-dimensional configuration specified by the request data requires a first quantity of building blocks that exceeds a second quantity of building blocks that are available and healthy in the superpod, and, in response to determining that the n-dimensional configuration specified by the request data requires the first quantity of building blocks that exceeds the second quantity of building blocks that are available and healthy in the superpod, identifying one or more second computational workloads that have a lower priority than the computational workload and are being executed by other building blocks of the superpod, and reallocating one or more building blocks of the one or more second computational workloads to a workload cluster of the computational workload. Generating a workload cluster of compute nodes including the subset of building blocks can include including one or more building blocks of the one or more second computational workloads in the subset of building blocks.
[0010] In some aspects, generating a workload cluster of compute nodes including a subset of the building blocks includes, for each dimension of the workload cluster, reconfiguring routing data of each of the one or more optical circuit switches for that dimension such that each building block of the one or more building blocks of the one or more second computational workloads communicates with other building blocks of the workload cluster other than the building blocks of the one or more second computational workloads.
[0011] The subject matter described herein may be implemented in particular embodiments to realize one or more of the following advantages: By using an optical network to dynamically configure a cluster of compute nodes to run workloads, compute nodes are more available because other compute nodes can easily replace failed or offline compute nodes. Flexible configuration of compute nodes allows for higher performance of compute nodes to run each workload. The optical network can instead enable workload clusters of various shapes, in which compute nodes operate adjacent to one another, even though they may be located in any physical location.
[0012] Using optical networks to configure pods also provides fault isolation and better workload security. For example, some conventional supercomputers route traffic between the various computers that make up the supercomputer. If one computer fails, the communication path is interrupted. Using optical networks, data can be quickly rerouted and / or available compute nodes can be used to replace the failed compute node. Additionally, the physical separation between workloads provided by optical circuit switching (OCS) switches, e.g., physical separation of different optical paths, provides better security between various workloads running on the same supercomputer compared to managing the separation using fragile software.
[0013] Additionally, connecting building blocks using optical networks reduces the delay of transmitting data between building blocks compared to packet switching networks. For example, packet switching increases the delay because the switch must receive the packet, buffer it, and then transmit it again on another port. Connecting building blocks using OCS switches provides a true end-to-end optical path without packet switching or buffering along the way.
[0014] Various features and advantages of the aforementioned subject matter are described below with reference to the drawings. Further features and advantages will be apparent from the subject matter described herein and in the claims. [Brief explanation of the drawings]
[0015] [Figure 1] 1 is a block diagram illustrating an environment in which an exemplary processing system creates a workload cluster of compute nodes and executes a computational workload using the workload cluster. [Figure 2] FIG. 2 illustrates an example logical superpod and an example workload cluster created from some of the building blocks in the superpod. [Figure 3] FIG. 1 illustrates an exemplary building block and an exemplary workload cluster created using the building block. [Figure 4] FIG. 2 illustrates an exemplary optical link from a compute node to an optical circuit switching (OCS) switch. [Figure 5] FIG. 1 illustrates a logical computation tray for forming building blocks. [Figure 6] FIG. 1 illustrates sub-blocks of an exemplary building block with one dimension omitted. [Figure 7] FIG. 1 illustrates an exemplary building block. [Figure 8] FIG. 1 illustrates a superpod OCS fabric topology. [Figure 9] FIG. 1 illustrates components of an exemplary superpod. [Figure 10] 1 is a flow diagram illustrating an example process for creating a workload cluster and executing a computational workload using the workload cluster. [Figure 11] 1 is a flow chart illustrating an example process for reconfiguring an optical network to replace a failed building block. DETAILED DESCRIPTION OF THE INVENTION
[0016] Detailed Description Like reference numbers and designations in the various drawings indicate like elements.
[0017] Generally, the systems and techniques described herein can create workload clusters of compute nodes from superpods by configuring an optical network fabric. A superpod is a group of multiple building blocks of compute nodes connected to each other via an optical network. For example, a superpod can include a set of interconnected building blocks. Each building block can include multiple compute nodes in an m-dimensional configuration, such as a two-dimensional or three-dimensional configuration.
[0018] A user can specify a target configuration of compute nodes for executing a particular workload. For example, a user can provide a machine learning workload and specify a target configuration of compute nodes for executing machine learning operations. The target configuration can define the number of compute nodes along each dimension of n dimensions (n is, for example, 2 or greater). That is, the target configuration can define the size and shape of the workload cluster. For example, some machine learning models and calculations perform better in non-square topologies.
[0019] The bandwidth cross-section can also limit computation across compute nodes that are waiting to transfer data or that are going off of idle computation cycles, for example. The shape of a workload cluster can affect the performance of the compute nodes in the workload cluster, depending on how work is allocated across the compute nodes and how much data needs to be transferred across the network in various dimensions.
[0020] For workloads that use all compute nodes to compute all compute node data traffic, a cubic workload cluster can minimize the number of hops between compute nodes. If a workload has many local communications, transferring data to a set of adjacent compute nodes along a particular dimension and chaining many of these adjacent communications together, a configuration with more compute nodes along a particular dimension than other dimensions is advantageous. Thus, by allowing a user to specify the configuration of compute nodes in a workload cluster, the user can specify a configuration that will result in better performance for executing the workload.
[0021] If different types of compute nodes are included in the superpod, the request can also specify the number of each type of compute node to include in the workload cluster, allowing the user to specify a configuration of compute nodes that will perform better for running a particular workload.
[0022] The workload scheduler can select building blocks for a workload cluster based on, for example, the availability of the building blocks, the health of the building blocks (e.g., operational or faulty), and / or the priority of the workloads within the superpod (e.g., the priority of the workloads to be executed by the compute nodes of the superpod). The workload scheduler can provide data identifying the selected building blocks and a target configuration of the building blocks to an optical circuit switching (OCS) manager. The OCS manager can create a workload cluster by configuring one or more OCS switches of the optical network to connect the building blocks to each other. The workload scheduler can then execute the computational workloads on the compute nodes in the workload cluster.
[0023] If one of the building blocks in a workload cluster fails, another building block can be quickly replaced by simply reconfiguring the OCS switch. For example, the workload scheduler can select an available building block from a superpod to replace the failed building block. The workload scheduler can instruct the OCS manager to replace the failed building block with the selected building block. The OCS manager can then reconfigure the OCS switch to connect the selected building block to other building blocks in the workload cluster and not connect the failed building block to building blocks in the workload cluster.
[0024] 1 is a block diagram illustrating an environment 100 in which an example processing system 130 generates a workload cluster of compute nodes and executes a computational workload using the workload cluster. The processing system 130 can receive a computational workload 112 from a user device 110 over a data communications network 120, such as a local area network (LAN), a wide area network (WAN), the Internet, a mobile network, or a combination thereof. Illustratively, the workload 112 includes software applications, machine learning models, such as training and / or using machine learning models, video encoding and decoding, and digital signal processing workloads.
[0025] A user may specify a cluster 114 of compute nodes required to execute a workload 112. For example, the user may specify a target shape and target size of the required cluster of compute nodes. That is, the user may specify the quantity of compute nodes along multiple dimensions and the shape of the compute nodes. For example, if the compute nodes are arranged along three dimensions x, y, and z, the user may specify the number of compute nodes in each dimension. The user may also specify one or more types of compute nodes to be included in the cluster. As described below, the processing system 130 may include different types of compute nodes.
[0026] As described below, the processing system 130 can use the building blocks to generate workload clusters that match a cluster of a target shape and a target size. Each building block can include multiple compute nodes arranged in m dimensions, e.g., three dimensions. Thus, a user can specify the target shape and size by specifying the number of building blocks in each of the multiple dimensions. For example, the processing system 130 can provide a user interface to the user device 110. The user interface allows the user to select the maximum number of building blocks along each dimension.
[0027] The user device 110 may provide the workload 112 and data specifying the requested clusters 114 to the processing system 130. For example, the user device 110 may provide the processing system 130, via the network 120, with request data including the workload 112 and data specifying the requested clusters 114.
[0028] The processing system 130 includes a cell scheduler 140 and one or more cells 150. A cell 150 is a group of one or more superpods. For example, the illustrated cell 150 includes four superpods 152-158. Each superpod 152-158 includes a set of building blocks 160, also referred to herein as a building block pool. In this example, each superpod 152-158 includes 64 building blocks. Each superpod 152-158 includes 64 building blocks 160. However, the superpods 152-158 may include other numbers of building blocks 160, such as 20, 50, 100, or another suitable number of building blocks 160. The superpods 152-158 may also include different numbers of building blocks 160. For example, the superpod 152 may include 64 building blocks, while the superpod 152 may include 100 building blocks.
[0029] As described in more detail below, each building block 160 can include multiple computational nodes arranged in two or more dimensions. For example, a building block 160 can include 64 computational nodes arranged along three dimensions, specifically, four computational nodes arranged along each dimension. Such a configuration of computational nodes is referred to herein as a 4x4x4 building block, including four computational nodes arranged along the x dimension, four computational nodes arranged along the y dimension, and four computational nodes arranged along the z dimension. Other numbers of dimensions, e.g., two dimensions, and other numbers of computational nodes along each dimension, e.g., 3x1, 2x2x2, 6x2, 2x3x4, are also possible.
[0030] Also, a building block may include only one compute node. However, as described below, optical links between the building blocks are configured to connect the building blocks to each other to create a workload cluster. Therefore, smaller building blocks, such as those including only one compute node, may provide greater flexibility in creating workload clusters, but require more OCS switch configurations and more optical network elements (e.g., cables and switches). The number of compute nodes in a building block may be selected based on a trade-off between the flexibility of the desired workload cluster, the building blocks that need to be connected to each other to create a workload cluster, and the number of OCS switches required.
[0031] Each computing node of building block 160 may include an application-specific integrated circuit (ASIC), such as a tensor processing unit (TPU) for machine learning workloads, a graphics processing unit (GPU), or other types of processing units. For example, each computing node may be a single processor chip that includes a processing unit.
[0032] In some implementations, all building blocks 160 in a superpod have similar compute nodes. For example, superpod 152 may include 64 building blocks, each with 64 TPUs in a 4x4x4 configuration, for running machine learning workloads. A superpod may also include different types of compute nodes. For example, superpod 154 may include 60 building blocks with TPUs and four building blocks with dedicated processing units that perform tasks other than machine learning workloads. In this manner, a workload cluster for running workloads may include different types of compute nodes. A superpod may include multiple building blocks of each type of compute node for redundancy and / or to enable the execution of multiple workloads within the superpod.
[0033] In some implementations, all building blocks 160 in a superpod have a similar configuration, e.g., a similar size and shape. For example, each building block 160 in superpod 152 can have a 4x4x4 configuration. A superpod can also include building blocks with different configurations. For example, superpod 154 can include 32 building blocks with a 4x4x4 configuration and 32 building blocks with a 16x8x16 configuration. Different building block configurations A configuration may include similar or different compute nodes. For example, a building block that includes a TPU may have a different configuration than a building block that includes a GPU.
[0034] A superpod can include building blocks of different hierarchies. For example, superpod 152 can include basic building blocks having a 4x4x4 configuration. Superpod 152 can also include intermediate building blocks consisting of more compute nodes. For example, an intermediate building block can have an 8x8x8 configuration formed from, say, eight basic building blocks. In this way, larger workload clusters can be created using intermediate building blocks with fewer links than would be possible by connecting basic building blocks. Including basic building blocks in a superpod also allows for the flexibility of creating smaller workload clusters that do not require the number of compute nodes in the intermediate building blocks.
[0035] The superpods 152-158 within the cell 150 can include building blocks of similar or different types of compute nodes. For example, the cell 150 can include one or more superpods of TPU building blocks and one or more superpods of GPU building blocks. The size and shape of the building blocks of the different superpods 152-158 within the cell 150 can be similar or different.
[0036] Each cell 150 also includes shared data storage 162 and shared auxiliary compute elements 164. Each superpod 152-158 within the cell 150 can use the shared data storage 162 to store data, for example, generated by workloads running on the superpods 152-158. The shared data storage 162 can include hard drives, solid-state drives, flash memory, and / or other suitable data storage devices. The shared auxiliary compute elements can include CPUs (e.g., general-purpose CPU devices), GPUs, and / or other accelerators (e.g., video decoding accelerators or image decoding accelerators) shared within the cell 150. The auxiliary compute elements 164 can also include storage devices, memory devices, and / or other compute elements that can be shared by compute nodes over a network.
[0037] The cell scheduler 140 may select a cell 150 and / or a superpod 152-158 of the cell 150 to execute each workload received from the user device 110. The cell scheduler 140 may select a superpod based on the target configuration specified for executing the workload, the availability of building blocks 160 in the superpod 152-158, and the health of the building blocks in the superpod 152-158. For example, the cell scheduler 140 may select a superpod to execute the workload that includes at least a sufficient number of available and healthy building blocks to generate a workload cluster of the target configuration. If the request data specifies a type of compute node, the cell scheduler 140 may select a superpod that includes at least a sufficient number of available and healthy building blocks that include the specified type of compute node.
[0038] As described below, each superpod 152-158 may include a workload scheduler and an OCS manager. When cell scheduler 140 selects a superpod for a cell 150, cell scheduler 140 may provide the workload scheduler for that superpod 150 with data specifying the workload and the requested cluster. As described in more detail below, the workload scheduler may determine the availability and health of the building blocks, the workload scheduler schedules the workloads within the superpod, if necessary, and the OCS manager schedules the workloads within the superpod. Based on the priority of the building blocks, the workload scheduler may select a set of building blocks from the building blocks of the superpod to be connected to generate a workload cluster. For example, as described below, if the workload scheduler receives a request for a workload cluster that includes more building blocks than the number of available and healthy building blocks in the superpod, the workload scheduler may reallocate building blocks for executing lower priority workloads to the requested workload cluster. The workload scheduler may provide data identifying the selected building blocks to the OCS manager. The OCS manager may create a workload cluster by configuring one or more OCS switches to connect the building blocks to each other. The workload scheduler may then execute the workloads on the compute nodes in the workload cluster.
[0039] In some implementations, the cell scheduler 140 balances the load among various cells 150 and superpods 152-158, for example, when selecting a superpod 152-158 to execute a workload. For example, when selecting a superpod among two or more superpods containing building blocks capable of processing a workload, the cell scheduler 140 may select the superpod with the highest capacity, such as the most available and healthy building block, or the superpod of the cell with the highest overall capacity. An available and healthy building block is a building block that is not running another workload or part of a starting workload cluster and is not faulty. The cell scheduler may store building block indices. Each building block index may include data indicating whether the building block is healthy (e.g., not faulty) and / or available (e.g., not running another workload or part of a starting workload cluster).
[0040] In some implementations, cell scheduler 140 can determine a target configuration for executing a workload. For example, cell scheduler 140 can determine a target configuration of building blocks based on the estimated computational demand of the workload and the throughput of one or more types of available compute nodes. In this example, cell scheduler 140 can provide the determined target configuration to a workload scheduler of a superpod.
[0041] 2 illustrates an example logical superpod 210 and example workload clusters 220, 230, and 240 created from some of the building blocks in superpod 210. In this example, superpod 210 includes 64 building blocks, each having a 4x4x4 configuration. While many examples herein describe building blocks with a 4x4x4 configuration, the techniques may be applied to building blocks with other configurations.
[0042] As described below, the shaded building blocks in superpod 210 are building blocks that are assigned to workloads, the white building blocks are available and healthy building blocks, and the black building blocks are unhealthy building blocks that cannot be used to create workload clusters, for example due to a failure.
[0043] The workload cluster 220 is an 8x8x4 pod that includes four 4x4x4 building blocks from the superpod 210. That is, the workload cluster 220 has eight compute nodes along the x-dimension and eight compute nodes along the y-dimension. Each building block has four compute nodes along the x dimension, and four compute nodes along the z dimension. Because each building block has four compute nodes along each dimension, workload cluster 220 includes two building blocks along the x dimension, two building blocks along the y dimension, and one building block along the z dimension.
[0044] The four building blocks of workload cluster 220 are shown with diagonal lines to indicate their location within superpod 210. As shown, the building blocks of workload cluster 220 are not adjacent to one another. As described in more detail below, by using an optical network, a workload cluster can be created from any combination of building blocks within superpod 210, regardless of the relative locations of the building blocks within superpod 210.
[0045] Workload cluster 230 is an 8x8x8 pod that includes eight of the building blocks of superpod 210. Specifically, workload cluster 230 includes two building blocks along each dimension, thereby resulting in workload cluster 230 including eight compute nodes along each dimension. Building blocks within workload cluster 230 are indicated by vertical lines to indicate their location within superpod 210.
[0046] Workload cluster 240 is a 16x8x16 pod that includes 32 of the building blocks of superpod 210. Specifically, workload cluster 240 includes four building blocks along the x dimension, two building blocks along the y dimension, and four building blocks along the z dimension. This results in the workload cluster including 16 compute nodes along the x dimension, eight compute nodes along the y dimension, and 16 compute nodes along the z dimension. The building blocks in workload cluster 240 are shown with a crosshatch to indicate their location within superpod 210.
[0047] Workload clusters 220, 230, and 240 are simply some examples of clusters of superpods 210 that may be created to execute workloads. Workload clusters may have many other configurations. The example workload clusters 220, 230, and 240 have rectangular shapes, but may have other shapes.
[0048] The geometry of the workload clusters, including workload clusters 220, 230, and 240, is a logical geometry, not a physical geometry. The optical network is configured so that the building blocks communicate with each other along each dimension, such that the workload clusters are physically connected in a logical configuration. However, the physical building blocks and corresponding compute nodes may be physically arranged within a data center in a variety of ways. The building blocks of workloads 220, 230, and 240 may be selected from any available and healthy building blocks, regardless of the physical relationship between the building blocks within superpod 210, except that all building blocks within superpod 210 are connected to the optical network. For example, as described above and illustrated in FIG. 2, workload clusters 220, 230, and 240 include building blocks that are not physically adjacent.
[0049] Furthermore, the logical configuration of a workload cluster is not limited by the physical configuration of the building blocks within the superpod. For example, one may arrange eight rows and eight columns of building blocks, with only one building block along the z-dimension. However, a workload cluster can be configured by configuring the optical network to create a logical configuration that includes multiple building blocks along the z-dimension. do.
[0050] FIG. 3 illustrates an example building block 310 and example workload clusters 320, 330, and 340 created using the building block 310. The building block 310 is a 4x4x4 building block that includes four compute nodes along each dimension. In this example, each dimension of the building block 310 includes 16 segments, each of which includes four compute nodes. For example, there are 16 compute nodes at the top of the building block 310. Of the 16 compute nodes, a segment along the y dimension includes one compute node and three other compute nodes, including the corresponding final compute node at the bottom of the building block 310. For example, one segment along the y dimension includes compute nodes 301-304.
[0051] The computing nodes within building block 310 can be connected to each other via internal links 318 made of conductive material, such as copper cables. The computing nodes within each segment of each dimension can be connected via internal links 318. For example, one internal link 318 connects computing node 301 to computing node 302. Another internal link 318 connects computing node 302 to computing node 303. Another internal link 318 connects computing node 303 to computing node 304. Similarly, internal data communication between the computing nodes of building block 310 can be provided by connecting computing nodes in other segments.
[0052] Building block 310 also includes external links 311-316 for connecting building block 310 to an optical network. The optical network connects building block 310 to other building blocks. In this example, building block 310 includes 16 external input links 311 in the x dimension. That is, building block 310 includes an external input link 311 for each of the 16 segments along the x dimension. Similarly, building block 310 includes an external output link 312 for each segment along the x dimension, an external input link 313 for each segment along the y dimension, an external output link 314 for each segment along the y dimension, an external input link 315 for each segment along the z dimension, and an external output link 316 for each segment along the z dimension. Because some building blocks can have a four- or more-dimensional configuration, such as a torus, building block 310 can include similar external links for each dimension of building block 310.
[0053] Each external link 311-316 may be an optical fiber link for connecting a compute node on the corresponding compute node segment to an optical network. For example, each external link 311-316 may connect the corresponding compute node to an OCS switch in the optical network. As described below, the optical network may include one or more OCS switches for each dimension of the segment of the building block 310. That is, the external links 311 and 312 in the x-dimension are connected to a different OCS switch than the external links 313 and 314. As described in more detail below, the OCS switches may be configured to create workload clusters by connecting building blocks to other building blocks.
[0054] The building block 310 has a 4x4x4 mesh configuration. A 4x4x4 (or other size) building block may have other configurations. For example, the building block 310 may have a three-dimensional torus configuration including wrap-around torus links, similar to the workload cluster 320. The workload cluster 320 may be generated from a single mesh building block 310 by configuring an optical network to form wrap-around torus links 321-323.
[0055] Torus links 321-323 provide wraparound data communication between one end of each segment and the other end of that segment. For example, torus link 321 connects a computation node located at each end of each segment along the x-dimension to a computation node located at the other end of that segment. Torus link 321 may include a link connecting computation node 325 to computation node 326. Similarly, torus link 322 may include a link connecting computation node 325 to computation node 327.
[0056] Torus links 321-323 may be conductive cables, such as copper cables, or optical links. For example, the optical links of torus links 321-323 may connect corresponding compute nodes to one or more OCS switches. The OCS switches may be configured to route data from one end of each segment to the other end of each segment. Building block 310 may include an OCS switch for each dimension. For example, torus link 321 may be connected to a first OCS switch, which may route data between one end of each segment along the x dimension and the other end of each segment along the x dimension. Similarly, torus link 322 may be connected to a second OCS switch, which may route data between one end of each segment along the y dimension and the other end of each segment along the y dimension. Torus link 322 may be connected to a third OCS switch, which may route data between one end of each segment along the z dimension and the other end of each segment along the z dimension.
[0057] Workload cluster 330 includes two building blocks 338 and 339 that create a 4x8x4 pod. Each of building blocks 338 and 339 may be similar to building block 310 or workload cluster 320. The two building blocks are connected along the y-dimension via external link 337. For example, one or more OCS switches may be configured to route data between the y-dimension segment of building block 338 and the y-dimension segment of building block 339.
[0058] One or more OCS switches may also be configured to form wraparound links 331-333 between one end of each segment and the other end of each segment along all three dimensions. In this example, wraparound link 333 connects one end of the y-dimension segment of building block 338 to one end of the y-dimension segment of building block 339, thereby providing full wraparound communication for the y-dimension segment created by combining two building blocks 338 and 339.
[0059] Workload cluster 340 includes eight building blocks (one not shown) that create an 8x8x8 cluster. Each building block 348 may be similar to building block 310. Building blocks connected along the x dimension are connected via external links 345A-345C. Similarly, building blocks connected along the y dimension are connected via external links 344A-344C, and building blocks connected along the z dimension are connected via external links 346A-346C. For example, one or more OCS switches may be configured to route data between x-dimension segments, one or more OCS switches may be configured to route data between y-dimension segments, and one or more OCS switches may be configured to route data between z-dimension segments. Along each dimension, building blocks not shown in FIG. 3 connect to adjacent building blocks. One or more OCS switches may also be configured to form wrap-around links 341-343 between one end of each segment and the other end of each segment along all three dimensions.
[0060] 4 illustrates an exemplary optical link 400 from a compute node to an OCS switch. The compute nodes of a superpod may be installed in trays in a data center rack. Each compute node may include six high-speed electrical links. Two of the electrical links may be connected to the compute node's circuit board, and four of the electrical links may be connected to external electrical connectors connected to ports 410, e.g., Octal Small Form Factor Pluggable (OSFP) ports. The port 410 may be routed to a connector, such as an OSFP connector. In this example, the port 410 is connected to an optical module 420 via an electrical link 412. The optical module 420 can convert the electrical link to an optical link that extends the length of the external link, e.g., greater than 1 kilometer (km), as needed to provide data communication between compute nodes located in a large data center. The type of optical module may be varied based on the required length and the desired speed and bandwidth of the link between the building block and the OCS switch.
[0061] The optical module 420 is connected to the circulator 430 via fiber optic cables 422 and 424. The fiber optic cable 422 can include one or more fiber optic cables for transmitting data from the optical module 420 to the circulator 430. The fiber optic cable 424 can include one or more fiber optic cables for receiving data from the circulator 430. For example, the fiber optic cables 422 and 424 can include bidirectional optical fibers or unidirectional TX / RX optical fiber pairs. The circulator 430 can reduce the number of optical fibers (e.g., by converting the unidirectional optical fibers to bidirectional optical fibers, two pairs of fiber optic cables 432 can be reduced to one pair). This typically corresponds to a single OCS port 445 of the OCS switch 440, which accommodates the converted pair of optical paths (two fibers). In some implementations, the circulator 430 can be integrated into the optical module 420 or omitted from the optical link 400.
[0062] Figures 5-7 show how multiple computational trays can be used to form a 4x4x4 building block. Similar techniques can be used to form building blocks of other sizes and shapes.
[0063] FIG. 5 illustrates a logical compute tray 500 for forming a 4×4×4 building block. The basic hardware block of a 4×4×4 building block is a single compute tray 500 with a 2×2×1 topology. In this example, compute tray 500 includes two compute nodes along the x-dimension, two compute nodes along the y-dimension, and one compute node along the z-dimension. For example, compute nodes 501 and 502 form an x-dimension segment, and compute nodes 503 and 504 form an x-dimension segment. Similarly, compute nodes 501 and 503 form a y-dimension segment, and compute nodes 502 and 504 form a y-dimension segment.
[0064] Each compute node 501-504 is connected to two other compute nodes via internal links 510, e.g., copper cables or traces on a printed circuit board. Each compute node is also connected to four external ports. Computing node 501 is connected to external port 521. Similarly, computing node 502 is connected to external port 522, computing node 503 is connected to external port 523, and computing node 504 is connected to external port 524. As mentioned above, external ports 521-524 may be OSFP ports or other ports that connect the compute nodes to OCS switches. These ports may be connected to copper cables or It can accommodate fiber optic modules attached to fiber optic cables.
[0065] Each of the external ports 521-524 of the computing nodes 501-504 includes one x-dimension port, one y-dimension port, and two z-dimension ports because each computing node 501-504 is already connected to other computing nodes along the x and y dimensions via internal links 510. By including two z-dimension external ports, each computing node 501-504 can connect to two computing nodes along the z dimension.
[0066] 6 illustrates a sub-block 600 of an exemplary building block with one dimension (z) omitted. Specifically, the sub-block 600 is a 4×4×1 block formed by 2×2 compute trays, such as the 2×2 compute trays 500 of FIG. 1. The sub-block 600 includes four 2×2 compute trays 620A-620D. Each compute tray 620A-620D may be similar to the compute tray 500 of FIG. 5, which includes four 2×2×1 compute nodes 622.
[0067] The compute nodes 622 in compute trays 620A-620D may be connected via internal links 631-634, e.g., copper cables. For example, two compute nodes 622 in compute tray 620A are connected along the y-dimension to two compute nodes 622 in compute tray 620B via internal link 632.
[0068] Furthermore, two computing nodes 622 in each of the computing trays 620A-620D are connected to an external link 640 along the x dimension. Similarly, two computing nodes in each of the computing trays 620A-620D are connected to an external link 641 along the y dimension. Specifically, the computing nodes located at the end of each x-dimensional segment and the computing nodes located at the end of each y-dimensional segment are connected to the external link 640. These external links 640 may be optical fiber cables connecting the computing nodes, i.e., the building blocks including the computing nodes, to the OCS switch, for example, via the optical link 400 in FIG. 4.
[0069] A 4x4x4 building block may be formed by connecting four of the sub-blocks 600 together along the z-dimension. For example, the compute nodes 622 of each compute tray 620A-620A may be connected via internal links to one or two compute nodes of corresponding compute trays on other sub-blocks 600 located in the z-dimension. The compute nodes located at the end of each z-dimension segment, as well as the compute nodes located at the ends of the x-dimension and y-dimension segments, may include external links 640 connected to the OCS switch.
[0070] Figure 7 illustrates an exemplary building block 700. Building block 700 includes four sub-blocks 710A-710D connected along the z-dimension. Each of sub-blocks 710A-710D may be similar to sub-block 600 of Figure 6. Figure 7 illustrates some of the connections between sub-blocks 710A-710D along the z-dimension.
[0071] Specifically, building block 700 includes internal links 730-733 along the z dimension between corresponding compute nodes 716 on compute trays 715 of sub-blocks 710A-710D. For example, internal link 730 connects the segment of compute node 0 along the z dimension. Similarly, internal link 731 connects the segment of compute node 1 along the z dimension, internal link 732 connects the segment of compute node 8 along the z dimension, and internal link 733 connects the segment of compute node 9 along the z dimension. Although not shown, similar internal links connect the segments of compute nodes 2-7 and A-F.
[0072] Building block 700 also includes external links 720 located at the ends of each segment along the z dimension. While the illustration only shows external links 720 for the segments of compute nodes 0, 1, 8, and 9, each of the other segments of compute nodes 2-7 and A-F also includes external links 720. The external links, similar to the external links located at the ends of the x- and y-dimensional segments, can connect the segments to OCS switches.
[0073] FIG. 8 illustrates a superpod OCS fabric topology 800. In this example, the OCS fabric topology includes a separate OCS switch for each segment along each dimension of a 4x4x4 building block of a superpod, including 64 building blocks 805, i.e., building blocks 0 through 63. The 4x4x4 building block 805 includes 16 segments along the x dimension, 16 segments along the y dimension, and 16 segments along the z dimension. In this example, the OCS fabric topology includes 16 x-dimension OCS switches, 16 y-dimension OCS switches, and 16 z-dimension OCS switches, for a total of 48 OCS switches that can be configured to form various workload clusters.
[0074] For the x dimension, OCS fabric topology 800 includes 16 OCS switches, including OCS switch 810. Each building block 805 includes external input links 811 and external output links 812 connected to OCS switches 810 in each segment along the x dimension. These external links 811 and 812 may be the same as or similar to optical links 400 of FIG. 4.
[0075] In the y-dimension, OCS fabric topology 800 includes 16 OCS switches, including OCS switch 820. Each building block 805 includes external input links 821 and external output links 822 connected to OCS switches 810 in each segment along the y-dimension. These external links 821 and 822 may be the same as or similar to optical links 400 in FIG. 4 .
[0076] In the z dimension, OCS fabric topology 800 includes 16 OCS switches, including OCS switch 830. Each building block 805 includes external input links 821 and external output links 822 connected to OCS switches 810 in each segment along the y dimension. These external links 821 and 822 may be the same as or similar to optical links 400 in FIG. 4.
[0077] In another example, multiple segments can share the same OCS switch, depending on, for example, the number of OCSs and / or the number of building blocks in a superpod. For example, if one OCS switch has a sufficient number of ports for all x-dimensional segments of all building blocks in a superpod, all x-dimensional segments can be connected to that OCS switch. In another example, if one OCS switch has a sufficient number of ports, two segments in each dimension can share that OCS switch. However, by connecting all building block segments in a superpod to the same OCS switch, a single routing table can be used for data communication between the compute nodes in these segments. Using separate OCS switches for each segment or dimension can also simplify fault response and diagnosis. For example, if there is a problem with data communication in a particular segment or dimension, it may be easier to identify the potentially failed OCS than if multiple OCSs were used for that segment or dimension.
[0078] 9 is a diagram illustrating components of an exemplary superpod 900. For example, superpod 900 may be one of the superpods of processing system 130 of FIG. An exemplary superpod 900 may include 64 4x4x4 building blocks 960, which may be used to form workload clusters for running computational workloads, such as machine learning workloads. As described above, each 4x4x4 building block 960 includes 32 compute nodes, four compute nodes arranged along each of three dimensions. For example, building block 960 may be the same as or similar to building block 310, workload cluster 320, or building block 700 described above.
[0079] The exemplary superpod 900 includes an optical network 970 that includes 48 OCS switches 930, 940, and 950 connected to the building blocks via 96 external links 931, 932, and 933 in each building block 960. Each external link may be a fiber optic link the same as or similar to optical link 400 in FIG.
[0080] Optical network 970 includes an OCS switch for each segment in each dimension of each building block, similar to OCS fabric topology 800 of FIG. 8. In the x dimension, optical network 970 includes 16 OCS switches 950, one for each segment along the x dimension. Optical network 970 also includes, for each building block 960, input and output external links corresponding to each segment of building block 960 along the x dimension. These external links connect the compute nodes on the segment to the OCS switch 950 for that segment. Because each building block 960 includes 16 segments along the x dimension, optical network 970 includes 32 external links 933 (i.e., 16 input links and 16 output links) for connecting the x-dimensional segments of each building block 960 to the OCS switch 950 corresponding to that segment.
[0081] For the y-dimension, optical network 970 includes 16 OCS switches 930, one for each segment along the y-dimension. Optical network 970 also includes, for each building block 960, input and output external links for each segment of building block 960 along the y-dimension. These external links connect the compute nodes on the segment to the OCS switches 930 for that segment. Because each building block 960 includes 16 segments along the y-dimension, optical network 970 includes 32 external links 931 (i.e., 16 input links and 16 output links) for connecting the y-dimension segments of each building block 960 to the OCS switches 930 corresponding to that segment.
[0082] For the z dimension, optical network 970 includes 16 OCS switches 932, one for each segment along the y dimension. Optical network 970 also includes, for each building block 960, input and output external links corresponding to each segment of building block 960 along the z dimension. These external links connect the compute nodes on a segment to the OCS switches 940 for that segment. Because each building block 960 includes 16 segments along the z dimension, optical network 970 includes 32 external links 932 (i.e., 16 input links and 16 output links) for connecting the z-dimension segments of each building block 960 to the OCS switches 940 corresponding to that segment.
[0083] The workload scheduler 910 may receive request data including a workload and data specifying a cluster of building blocks 960 required to execute the workload. The request data may include a priority for the workload. The priority may be expressed as a level of high, medium, or low, and may range, for example, from 1 to 100, or another number. It may be expressed numerically within an appropriate range. For example, the workload scheduler 910 may receive request data from a user device or a cell scheduler, such as the user device 110 or the cell scheduler 140 of FIG. 1. As described above, the request data may specify an n-dimensional target configuration of compute nodes, such as a target configuration of building blocks that include compute nodes.
[0084] The workload scheduler 910 can select a set of building blocks 960 to create a workload cluster that matches the target configuration specified by the request data. For example, the workload scheduler 910 can identify a set of building blocks that are available and healthy in the superpod 900. As described above, available and healthy building blocks are building blocks that are not running another workload or part of a starting workload cluster and are not faulty.
[0085] For example, the workload scheduler 910 may store and update, e.g., in the form of a database, state data indicating the state of each building block 960 in the superpod. The availability state of a building block 960 may indicate whether the building block 960 is assigned to a workload cluster. The health state of a building block 960 may indicate whether the building block is operational or faulty. The workload scheduler 910 may identify building blocks 960 that have an availability state indicating that they are not assigned to a workload and a health state indicating that they are operational. When a building block 960 is assigned to a workload, e.g., when it is used to create a workload cluster for executing a workload, or when its health state changes from operational to faulty or vice versa, the workload scheduler may update the state data of the building block 960 accordingly.
[0086] The workload scheduler 910 may select a number of building blocks 960 from the identified plurality of building blocks 960 that matches the number defined by the target configuration. If the request data specifies one or more types of compute nodes, the workload scheduler 910 may select building blocks that include the requested types of compute nodes from the identified building blocks 960. For example, if the request data specifies a 2x2 configuration of building blocks consisting of two TPU building blocks and two GPU building blocks, the workload scheduler 910 may select two available and healthy TPU building blocks and two available and healthy GPU building blocks.
[0087] Additionally, workload scheduler 910 may select building blocks 960 based on the priority of each workload running in the superpod and the priority of the workload included in the request data. If superpod 900 does not have enough available, healthy building blocks to create a workload cluster for executing the requested workload, workload scheduler 910 may determine whether a workload having a lower priority than the requested workload is running in superpod 900. If a workload having a lower priority than the requested workload is running, workload scheduler 910 may reallocate building blocks of a workload cluster executing one or more lower priority workloads to a workload cluster for executing the requested workload. For example, workload scheduler 910 may allocate building blocks of a workload cluster executing one or more lower priority workloads to a workload cluster for executing the requested workload by terminating the lower priority workload, delaying the lower priority workload, or reducing the size of the workload cluster for the lower priority workload. The building blocks for loading can be released.
[0088] The workload scheduler 910 can reassign a building block from one workload cluster to another simply by reconfiguring the optical network (e.g., by configuring the OCS switches, as described below), so that the building block is connected to a building block for running a higher priority workload rather than to a building block for running a lower priority workload. Similarly, if a building block for running a higher priority workload fails, the workload scheduler 910 can reconfigure the optical network to reassign a building block in a workload cluster for running a lower priority workload to a workload cluster for running a higher priority workload.
[0089] The workload scheduler 910 can generate and provide per-job configuration data 912 to the OCS manager 920 of the superpod 900. The per-job configuration data 912 can specify the building blocks 960 selected to execute the workload and the configuration of the building blocks. For example, if the configuration is a 2x2 configuration, it includes four spots for placing the building blocks. The per-job configuration data can specify that the selected building blocks 960 are to be placed in each of the four spots.
[0090] The per-job configuration data 912 can identify the selected building blocks 960 using a logical identifier for each building block. For example, each building block 960 can include a unique logical identifier. In a particular embodiment, the 64 building blocks 960 can be assigned numbers from 0 to 63, which can be unique logical identifiers.
[0091] OCS manager 920 uses per-job configuration data 912 to configure OCS switches 930, 940, and / or 950 to generate workload clusters that match the configuration specified by the per-job configuration data. Each OCS switch 930, 940, and 950 includes a routing table used when routing data between the physical ports of the OCS switch. For example, assume that an outgoing external link of an x-dimensional segment of a first building block is connected to an incoming external link of a corresponding x-dimensional segment of a second building block. In this case, the routing table of OCS switch 950 for the x-dimensional segment indicates that data between the physical ports of the OCS switch to which these segments are connected is routed between these physical ports.
[0092] OCS manager 920 can store port data that maps each port of each OCS switch 920, 930, and 940 to each logical port of each building block. The port data for each x-dimensional segment of a building block can specify which physical port of OCS switch 950 an external input link is connected to and which physical port of OCS switch 950 an external output link is connected to. The port data for each dimension of each building block 960 of superpod 900 can include similar data.
[0093] The OCS manager 920 can use the port data to configure the routing tables of the OCS switches 930, 940, and / or 950 to create workload clusters for running workloads. For example, in a 2x1 configuration where a first building block is located to the left of a second building block along the x dimension, Suppose a first building block attempts to connect to a second building block. The OCS manager 920 updates the routing table of the x-dimensional OCS switch 950 to route data between the x-dimensional segment of the first building block and the x-dimensional segment of the second building block. As each x-dimensional segment of a building block needs to be connected, the OCS manager 920 can update the routing table of each OCS switch 950.
[0094] The OCS manager 920 can update the routing table of the OCS switch 950 for each x-dimensional segment. Specifically, by updating the routing table, the OCS manager 920 can map the physical port of the OCS switch 950 to which the segment of the first building block is connected to the physical port of the OCS switch to which the segment of the second building block is connected. Because each x-dimensional segment includes an input link and an output link, the OCS manager 920 can update the routing table so that the input link of the first building block is connected to the output link of the second building block and the output link of the first building block is connected to the input link of the second building block.
[0095] The OCS manager 920 can update the routing tables by obtaining current routing tables from each OCS switch. In another example, the OCS manager 920 can update the appropriate routing tables and send the updated routing tables to the appropriate OCS switches. In another example, the OCS manager 920 can send update data specifying the updates to the OCS switches, and the OCS switches can update their routing tables according to the update data.
[0096] After configuring the OCS switch with the updated routing table, a workload cluster is created. The workload scheduler 910 can then cause the compute nodes in the workload cluster to execute the workload. For example, the workload scheduler 910 can provide the workload to the compute nodes in the workload cluster for execution.
[0097] After the workload finishes executing, the workload scheduler 910 can update the state of each building block used to generate the workload cluster to return the state of each building block to an available state. The workload scheduler 910 can also instruct the OCS manager 920 to release the connections between the building blocks used to generate the workload cluster. Accordingly, the OCS manager 920 can update the routing table to release the mappings between the physical ports of the OCS switch that were used to route data between the building blocks.
[0098] In this way, by configuring an optical fabric topology using OCS switches to generate workload clusters for executing workloads, a superpod can dynamically and securely host multiple workloads. The workload scheduler 920 can instantly generate workload clusters when a new workload is received and instantly release the workload clusters once the workload is processed. The inter-segment routing provided by OCS switches provides better security between different workloads running on the same superpod than traditional supercomputers. For example, OCS switches physically separate workloads using air gaps between them. Traditional supercomputers use software to separate workloads, which makes them more susceptible to information leakage.
[0099] 10 is a flow diagram illustrating an example process 1000 for generating a workload cluster and executing a computational workload using the workload cluster. The operations of process 1000 may be performed by a system including one or more data processing devices. For example, the operations of process 1000 may be performed by processing system 130 of FIG. 1.
[0100] The system receives 1010 request data specifying a cluster of required compute nodes. The request data may be received from a user device. The request data may include a computational workload and data specifying an n-dimensional target configuration of the compute nodes. For example, the request data may specify an n-dimensional target configuration of building blocks that include the compute nodes.
[0101] In some implementations, the request data can specify the types of compute nodes to generate the building blocks. A superpod can include building blocks of different types of compute nodes. For example, a superpod can include 90 building blocks, each including a 4x4x4 TPU, and 10 dedicated building blocks including a 2x1 dedicated compute node. The request data can specify the number of building blocks of each type of compute node and the configuration of these building blocks.
[0102] The system selects 1020 a subset of building blocks from a superpod containing a set of building blocks to generate the requested cluster. As described above, a superpod may contain a set of building blocks containing a three-dimensional configuration of computing nodes, such as a 4x4x4 configuration of computing nodes. The system may select a quantity of building blocks that matches the quantity defined by the target configuration. As described above, the system may select building blocks that are healthy and available to generate the requested cluster.
[0103] The subset of building blocks may be a proper subset of the building blocks. The proper subset is a subset that does not include all members of a set of building blocks. For example, fewer than all building blocks may be required to generate a workload cluster that matches the target configuration of compute nodes.
[0104] The system generates 1030 a workload cluster that includes the selected subset of compute nodes. The workload cluster can include building blocks with a configuration that matches the target configuration specified by the request data. For example, if the request data specifies a 4x8x4 configuration of compute nodes, the workload cluster can include two building blocks arranged as in workload cluster 330 of FIG. 3.
[0105] To generate a workload cluster, the system can configure routing data for each dimension of the workload cluster. For example, as described above, a superpod can include an optical network including one or more OCS switches for each dimension of building blocks. The routing data for a dimension can include routing tables for one or more OCS switches. As described above with reference to FIG. 9, the routing tables of the OCS switches can be configured to route data between appropriate segments of compute nodes along each dimension.
[0106] The system causes the compute nodes in the workload cluster to execute the computational workload (1040). For example, the system may provide the computational workload to the compute nodes in the workload cluster. While the computational workload is executing, the configured O The CS switch can route data between the building blocks of the workload cluster. The configured OCS switch can route data between the compute nodes of the building blocks as if they were physically connected, even though the compute nodes are not physically connected in the target configuration.
[0107] For example, compute nodes in each segment of a dimension can communicate data to other compute nodes in segments in different building blocks through OCS switches so that they are physically connected to other compute nodes in segments in different building blocks on a single physical segment. This configuration of workload clusters differs from packet switching networks because it provides a true end-to-end optical path without packet switching or buffering along the way. Packet switching introduces longer latency because the switch must receive the packet, buffer it, and then transmit it again on a different port.
[0108] After the computational workload has finished executing, the system can free up the building blocks to execute other workloads, for example, by updating the building block's state to an available state and updating data routing to not route data between building blocks in the workload cluster.
[0109] 11 is a flow diagram illustrating an example process 1100 for reconfiguring an optical network to replace a failed building block. The operations of process 1100 may be performed by a system including one or more data processing devices. For example, the operations of process 1100 may be performed by processing system 130 of FIG. 1.
[0110] The system causes the compute nodes in the workload cluster to execute the computational workload 1110. For example, the system may create a workload cluster and cause the compute nodes to execute the computational workload according to process 1000 of FIG.
[0111] The system receives data indicating that a building block of a workload cluster has failed 1120. For example, if one or more compute nodes of the building block have failed, another element, such as a monitoring element, may determine that the building block has failed and send data to the system indicating that the building block has failed.
[0112] The system identifies 1130 available building blocks. For example, the system can identify available and healthy building blocks in a superpod similar to other building blocks in the workload cluster. For example, the system can identify available and healthy building blocks based on building block state data stored by the system.
[0113] The system replaces the failed building block with the identified available building block (1140). The system can update data routing in one or more OCS switches in the optical network connecting the building blocks to replace the failed building block with the identified available building block. For example, the system can update routing tables in the one or more OCS switches to remove connections between the failed building block and other building blocks in the workload cluster. The system can also update routing tables in the one or more OCS switches to connect the identified building block to other building blocks in the workload cluster.
[0114] The system assigns the identified building blocks to the failed building block spots. The OCS switch may logically place the building block in a logical spot in the network. As described above, the OCS switch's routing table may map the physical port of the OCS switch connected to a segment of one building block to the physical port of the OCS switch connected to the corresponding segment of another building block. In this case, the system can perform the replacement by updating the mapping with the corresponding segment of the identified available building block instead of the failed building block.
[0115] For example, assume that an incoming external link of a particular x-dimensional segment of a failed building block is connected to a first port of an OCS switch, and that an incoming external link of a corresponding x-dimensional segment of an identified available building block is connected to a second port of the OCS switch. Further, assume that a routing table maps the first port to a third port of the OCS switch that is connected to a corresponding x-dimensional segment of another building block. To effect the replacement, the system can update the mapping in the routing table by mapping the second port to the third port, rather than mapping the first port to the third port. The system can do the same for each segment of the failed building block.
[0116] Embodiments of the subject matter and operations described herein may be implemented in digital electronic circuitry, computer software, firmware, hardware, or one or more combinations thereof, including the structures and equivalents disclosed herein. Embodiments of the subject matter described herein may be implemented as one or more computer programs, i.e., one or more modules of computer program instructions encoded on a computer storage medium and executed by or controlling the operation of a data processing device. Alternatively or additionally, the program instructions may be encoded in an artificially generated propagated signal, e.g., a machine-generated electrical, optical, or electromagnetic signal, encoding information that is transmitted to an appropriate receiving device for execution by the data processing device. A computer storage medium may be, or may be included in, a computer-readable storage device, a computer-readable storage substrate, a random access memory array or device, or a serial access memory array or device, or one or more combinations thereof. Furthermore, a computer storage medium may not be a propagated signal, but may be the source or destination of computer program instructions encoded in an artificially generated propagated signal. A computer storage medium may also be, or be included in, one or more separate physical elements or media (e.g., multiple CDs, disks, or other storage devices).
[0117] The operations described herein may be implemented as operations performed by a data processing apparatus on data stored in one or more computer-readable storage devices or on data received from other sources.
[0118] The term "data processing device" includes all types of equipment, devices, and machines for processing data, including, for example, programmable processors, computers, systems-on-chips, or combinations thereof. A device may include special-purpose logic circuitry, such as a field-programmable gate array (FPGA) or an application-specific integrated circuit (ASIC). In addition to hardware, a device may also include code that creates an execution environment for the computer program, such as code constituting processor firmware, a protocol stack, a database management system, an operating system, a cross-platform runtime environment, a virtual machine, or a combination thereof. The device and execution environment may implement a variety of different computing model infrastructures, such as web services, distributed computing infrastructures, and grid computing infrastructures.
[0119] A computer program (also known as a program, software, software application, script, or code) may be written in any programming language, including compiled or interpreted languages, declarative or procedural languages, and may be deployed in any form, including a stand-alone program, module, component, subroutine, object, or other unit suitable for use in a computing environment. A computer program may, but need not, correspond to a file in a file system. A program may be stored as part of a file that holds other programs or data (e.g., one or more scripts stored in a markup language document), in a single file dedicated to the program, or in multiple associated files (e.g., files storing one or more modules, subprograms, or portions of code). A computer program may be deployed and executed on one computer or on multiple computers located at one site or distributed across multiple sites and interconnected by a communications network.
[0120] The processes and logic flows described herein may be performed by one or more programmable processors that execute one or more computer programs by processing input data and generating output to perform operations. The processes and logic flows may also be performed by, and apparatus may be implemented as, special purpose logic circuitry, such as an FPGA (field programmable gate array) or an ASIC (application-specific integrated circuit).
[0121] Processors suitable for executing a computer program include, by way of example, general-purpose microprocessors, special-purpose microprocessors, and one or more processors of any kind of digital computer. Typically, a processor receives instructions and data from a read-only memory, a random-access memory, or both. The essential elements of a computer include a processor for performing operations in accordance with the instructions and one or more memory devices for storing instructions and data. Typically, a computer also includes one or more mass storage devices, such as magnetic, magneto-optical, or optical disks, for storing data, or is operatively coupled to one or more mass storage devices to receive data from, transfer data to, or both. However, a computer need not have such devices. Furthermore, a computer can be incorporated into another device, such as a mobile phone, a personal digital assistant (PDA), a mobile audio or video player, a game console, a global positioning system (GPS) receiver, or a portable storage device (e.g., a USB flash drive). Suitable devices for the storage of computer program instructions and data include all non-volatile memories, media, and memory devices, including, for example, semiconductor memory devices such as EPROM, EEPROM, and flash memory devices, magnetic disks, for example internal hard disks or removable disks, magneto-optical disks, CD-ROM and DVD-ROM disks. The processor and memory may be supplemented by, or incorporated in, special purpose logic circuitry.
[0122] To provide for user interaction, embodiments of the subject matter described herein may be implemented on a computer that includes a display device, such as a CRT (cathode ray tube) or LCD (liquid crystal display) monitor, for displaying information to the user, and a keyboard and pointing device, such as a mouse or trackball, by which the user can provide input to the computer. Other types of devices may also be used to provide user interaction. Feedback provided to the user may be, for example, any sensory feedback, such as visual feedback, auditory feedback, or tactile feedback. Input received from the user may include acoustic input, speech input, or tactile input. Additionally, the computer can interact with the user by sending documents to and receiving documents from a device being used by the user, for example, by sending a web page to a web browser on the user's client device in response to a request received from the web browser.
[0123] Embodiments of the subject matter described herein may be implemented in a computing system including a back-end component, e.g., a data server, or a computing system including a middleware component, e.g., an application server, or a computing system including a front-end component, e.g., a client computer with a graphical user interface or web browser through which a user can interact with an implementation of the subject matter described herein, or any combination of the back-end, middleware, or front-end components described above. The components of the system may be interconnected by any digital data communication medium, e.g., a communications network. Examples of communications networks include local area networks (LANs) and wide area networks (WANs), internetworks (e.g., the Internet), and peer-to-peer networks (e.g., ad hoc peer-to-peer networks).
[0124] A computing system may include clients and servers. Clients and servers are generally remote from each other and typically interact through a communication network. The relationship of client and server is formed by computer programs running on the respective computers and having a client-server relationship to each other. In some embodiments, a server sends data (e.g., HTML pages) to a client device (e.g., to display the data to a user interacting with the client device and receive user input). The server can receive data generated by the client device (e.g., the result of user interaction) from the client device.
[0125] While this specification contains many specific implementation details, these implementation details should not be construed as limitations on the scope of the invention or the claims, but rather as descriptions of features specific to particular embodiments of a particular invention. Specific features described herein with respect to separate embodiments may also be implemented in a single embodiment in any combination. Conversely, various features described with respect to a single embodiment may also be implemented in multiple embodiments separately or in any suitable subcombination. Furthermore, even if features are operative in a certain combination and are claimed to operate in that combination, one or more features can be removed from the claimed combination, and the claimed combination can be divided into subcombinations, as appropriate.
[0126] Similarly, while the figures depict operations in a particular order, this should not be understood as requiring that these operations be performed in the particular order or sequentially depicted, or that all of the operations depicted be performed, to achieve desirable results. Depending on the circumstances, multitasking and parallel processing may be advantageous. Furthermore, the separation of various system elements in the above-described embodiments should not be understood as requiring such separation in all embodiments, and it should be understood that the described program elements and systems may generally be integrated into a single software product or packaged into multiple software products.
[0127] Thus, specific embodiments of the present invention have been described. Other embodiments are within the scope of the following claims. In some cases, the actions recited in the claims can be performed in a different order and still achieve desirable results. Furthermore, the processes illustrated in the accompanying drawings The steps do not necessarily have to be performed in the particular order shown, or sequentially, to achieve desirable results. In some implementations, multitasking and parallel processing may be advantageous.
Claims
1. 1. A method executed by one or more data processing devices, said method comprising: Identifying a target configuration of compute nodes for running a computational workload; generating a workload cluster of compute nodes having a configuration that matches the target configuration of the compute nodes; The generating step comprises: and configuring, for each dimension of n dimensions, each of the one or more optical circuit switches of the optical network of that dimension such that computational nodes of a plurality of building blocks arranged in that dimension communicate with each other via the one or more optical circuit switches of that dimension, the plurality of building blocks being communicatively connected to the optical network, each building block including an n-dimensional configuration of computational nodes, the plurality of computational nodes arranged in each dimension of the n dimensions (n being 2 or greater); causing the compute nodes of the workload cluster to execute the computational workload.
2. configuring each of the one or more optical circuit switches in each dimension includes configuring routing data for each of the one or more optical switches in each dimension; The method of claim 1 , wherein the respective routing data specifies how to route data of the computational workload among the compute nodes along the dimension.
3. The method of claim 1 or 2, further comprising selecting a subset of building blocks from a set of building blocks of the computational workload as the plurality of building blocks.
4. each building block including a plurality of segments of computational nodes along each of the n dimensions; The method of claim 1 , wherein the optical network includes an optical circuit switch corresponding to each segment in each dimension of the plurality of building blocks.
5. The method of claim 1 , wherein each building block comprises one of a three-dimensional torus-shaped computing node or a mesh-shaped computing node.
6. receiving data indicating that a particular building block of the workload cluster has failed; and replacing the particular building block with an available building block from the plurality of building blocks; 6. The method of claim 1, wherein the replacing comprises reconfiguring each of the one or more optical circuit switches corresponding to the one or more dimensions so that the compute nodes of the available building block communicate with other compute nodes of the workload cluster corresponding to the one or more dimensions.
7. 7. The method of claim 1, wherein identifying a target configuration of compute nodes for executing the computational workload comprises receiving request data specifying the target configuration of compute nodes and different types of compute nodes to be included in the target configuration of compute nodes.
8. The method of claim 7 , further comprising: for each type of computing node specified by the requirement data, determining a building block that includes one or more computing nodes of the specified type.
9. 1. A system comprising: a data processing device; a computer storage medium storing a computer program; A system, wherein the computer program, when executed by the data processing apparatus, causes the data processing apparatus to perform the method of any one of claims 1 to 8.
10. A program that causes a computer to execute the method according to any one of claims 1 to 8.
Citation Information
Patent Citations
Scheduling in high-performance computing (HPC) system
JP2006146864A
Job management program, job management method, and job management apparatus
JP2016091069A
Processing System with Distributed Processors
JP2016504668A
Configuration of cluster server using cellular automaton
JP2017527031A