Resilient interconnection networks in large-scale computing
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- GOOGLE LLC
- Filing Date
- 2025-11-27
- Publication Date
- 2026-08-07
AI Technical Summary
随着计算组中处理节点数量的增加,互连组件故障的风险随着互连组件数量的增加而增加
Smart Images

Figure CN122534007A_ABST
Abstract
Description
Background Technology
[0001] Machine learning (ML) techniques based on large language models (LLM) are revolutionizing multiple industries. However, LLM training requires a large number of processing nodes deployed in groups such as computing pods. Some pods can include thousands or tens of thousands of processing units, such as tensor processing units (TPUs). Processing nodes communicate with each other through an interconnect network that provides communication paths between them. For ML training requirements, the processing power of the processing nodes and the communication bandwidth of the interconnect network must be sufficient to provide speed and efficiency for the training process. Furthermore, LLM workloads require synchronized operation of all processing nodes, so that any failure of a processing node or interconnect network component can interrupt the processing job. As the number of processing nodes in a computing pod increases, the risk of interconnect component failure also increases. To achieve robustness and reliability in ML training networks, the following capabilities are desired: providing system-level redundancy in the interconnect network to prevent job interruptions due to component failures. Summary of the Invention
[0002] This technology generally involves providing redundant interconnect communication paths in a computing network, such as reconfigurable superpods. A protected interconnect network for a computing network can include multiple processing nodes connected by the protected interconnect network. At least one building block can be defined as having a predetermined number of processing nodes. For example, a building block can be defined as the number of processing nodes housed within a computer rack. A circuit switch (ECS) can be associated with each building block. The ECS can route a portion of the data traffic for each building block to a protected communication path that operates in parallel with the regular communication path. The computing network can be a reconfigurable superpod. A reconfigurable superpod can include multiple dimensions, with each processing node communicating with each dimension.
[0003] Multiple Optical Circuit Switches (OCS) comprise an OCS corresponding to each dimension of the supergroup, with each dimension's OCS communicating with the ECS of each building block. Protected communication paths are created from the processing nodes via associated ECSs and protection OCSs connected to them. In one example, a building block could be defined by 32 processing nodes housed within a single computer rack. Depending on some aspects of the technology, the processing nodes could be Tensor Processing Units (TPUs). The building block could be contained within the computer rack, with the ECS associated with it implemented as a top-of-rack (ToR) switch within the rack.
[0004] A supergroup may include a first number of inter-chip interconnect (ICI) links connecting processing nodes to a regular communication path, and a second number of ICI links connecting processing nodes to a protection communication path via an ECS. According to some examples, the ECS may include 32 building-block-oriented ports for receiving protection ICI links from each processing node in the building block. The ECS may further include 20 ports facing a protection OCS in the protection communication path.
[0005] In other examples, a building block can be divided into several groups of processing nodes. Each group of processing nodes can connect to an ECS, with ports facing the building block connected to the divided groups of processing nodes. For example, a building block with 32 processing nodes can be divided into two groups of 16 processing nodes each. A pair of ECSs can connect to the building block, with each ECS having 16 ports facing the building block and 10 ports facing the protection path.
[0006] In one aspect of this technology, the protection path may include a single protection OCS that communicates with all dimensions of the supergroup. Each building block of the supergroup communicates with a single protection OCS. In a configuration using a single protection OCS, the ECS associated with each building block may include 32 ports facing the building block and 4 ports facing the protection communication path.
[0007] To provide redundant interconnect communication paths, each processing node can have N regular ICI links for general communication and one additional protective ICI link for secure communication. In one example, each dimension may include two outward-facing hyperplanes to facilitate communication with the dimension in both inbound and outbound directions. For instance, the N+1 redundant interconnect networks for a 5-dimensional supergroup may include two regular ICI links for each dimension and one additional protective ICI link for communication with the ECS connected to the secure communication path. Therefore, for 5 dimensions, the processing node comprises 10 regular ICI links and 1 protective ICI link (10+1).
[0008] In another aspect of the described technology, a method for protecting an interconnected network of a computing network includes: establishing a first number of communication paths connected to regular processing paths in a building block of a computing network having multiple processing nodes, and providing an additional protective communication path connected to the protective processing path, wherein data routing from the building block to the protective processing path is achieved via a circuit switch (ECS) associated with the building block. For a building block comprising 32 processing nodes, each protective ICI link of each processing node is connected to an input port of a 32×20-port ECS. 20 output ports are connected to a protective OCS communicating with each dimension in the dimension. Alternatively, the protective path may include multiple OCSs, each OCS connected to a corresponding dimension of the computing network. Attached Figure Description
[0009] Figure 1 It is a block diagram of a reconfigurable supergroup based on various aspects of the described technology.
[0010] Figure 2 It is a block diagram of faults in the interconnection network of supergroups based on various aspects of the described technology.
[0011] Figure 3 It is a block diagram of a redundant interconnected network at the node level, based on the various aspects of the described technology.
[0012] Figure 4 It is a block diagram of a reconfigurable supergroup based on various aspects of the described technology.
[0013] Figure 5 It is a block diagram of a redundant protection interconnection network based on various aspects of the described technology.
[0014] Figure 6 It is a block diagram of an example system based on various aspects of this disclosure.
[0015] Figure 7 It is a flowchart of the process of establishing a redundant protection communication path based on the various aspects of the described technology. Detailed Implementation
[0016] This technology generally involves an N+1 protected interconnect technique for computing supergroups, where N represents the number of regular ICI links per processing node. Therefore, redundant protected ICI links are introduced on a per-processing-node (e.g., TPU) basis. A single protected ICI link can be used to protect all working ICI links of that processing node from all dimensions. Protected ICI links are provided through rack-level circuit switches (ECS) and supergroup-level protected optical circuit switches (OCS). This technology can protect all optical and electrical interconnects, including the OCS, without performance degradation and with negligible latency increases.
[0017] For a 5D torus supergroup with 10 ICI links per TPU, only additional protection ICI links need to be added to achieve 10 + 1 ICI link protection. In the following description, for illustrative purposes, an example of a 5D torus-based supergroup architecture is used. It should be understood that the features and techniques discussed will be applicable to other network architectures and topologies.
[0018] Figure 1 A supergroup concept 100 of the described technology is illustrated. Redundant protective ICI bidirectional links 110 are provided based on processing nodes (TPUs) 130. In addition to 10 regular ICI links 140 connecting the processing nodes to each dimension 150, there are protective ICI links 130. Each building block 130 provides a low-latency cross-switch protective ECS 111. According to one aspect of the technology, the protective ECS 111 has 32 duplex ports facing the processing nodes, which connect to the 32 protective ICI links 130 originating from the 32 processing nodes in the building block. The 32 processing nodes of the building block typically reside within a single hardware rack. The protective ECS may include 20 externally facing duplex ports used as optical ports to connect to five redundant protective OCSs 112, representing one protective OCS for each dimension 112x, 112y, 112z, 112a, and 112b. Figure 1 The system allows any failure in any dimension 150 of the conventional optical ICI link path 140, including optical modules, optical links, or OCS, to be protected by the established dedicated protection path 115 for that dimension. Furthermore, the protection device ECS 111 can be configured to allow full-pair full connectivity between any pair of ports among the 32 duplex ports facing the processing node, thus also protecting against any electrical ICI link failures within the building block.
[0019] exist Figure 1In the example shown, a single 32×20 protection ECS 415 is provided for each 32 TPU building block. However, these 32 TPUs can be divided into two sub-building blocks within the same building block, each containing 16 TPUs. In this case, the protection ICI links corresponding to the two sub-building blocks can be connected to two smaller ECSs (not shown), for example, two 16×10 protection ECSs could be used. Therefore, the actual ECS base requirement, depending on the compute group topology and the number of ICI links bundled in the optical domain, can be adapted at the building block level through partitioning.
[0020] Figure 2 This demonstrates that when supergroups scale to approximately 4,000 nodes or larger, OCS and optical ICI link failures pose a greater risk to providing improved compute group availability. Figure 2 As shown, a single OCS failure 210 will cause the loss of four ICI links 203, 205, and 207 in each building block 220, essentially paralyzing the entire supergroup. While the effective radius of a single optical ICI link failure is smaller than the impact of a failed OCS 210, it still leads to the failure of the entire building block 220. Considering that the number of optical ICI links is several orders of magnitude higher than the number of OCSs, the probability of optical link failures could significantly reduce the availability of the computing group.
[0021] Figure 3 This is a block diagram of a redundant interconnected network at the processing node level, based on various aspects of the described technology. TPU1 processing node 310 and TPU2 processing node 311 provide computing services to the supergroup system. Processing nodes 310 and 311 can be housed within the same computer rack. Processing nodes 310 and 311 can communicate with each other via a cross-board rack connection 315. This cross-board rack connection 315 can correspond to the dimensions configured in the supergroup. For example, for a 5D supergroup, its dimensions can be represented as dimension x, dimension y, dimension z, dimension a, and dimension b. Each dimension can contain two hyperplanes, represented as + and –, which indicate the communication direction.
[0022] In normal network communication, processing node 310 communicates via optical transceiver module 320, which converts data into optical signals and communicates via conventional operating optical path 312. Similarly, processing node 311 communicates via optical transceiver module 321, which converts data into optical signals and communicates via conventional operating optical path 313.
[0023] In addition to the conventional communication paths, each processing node 310, 311 is also provided with additional protection links 322, 323. Processing node 310 includes protection ICI link 322, while processing mode 311 includes protection ICI link 323. Protection ICI links 322, 323 are connected to the input ports of protection cross switch ECS 330. Protection cross switch ECS 330 communicates with protection optical path 335 via protection ICI links 322, 323. Protection optical path 335 may include optical transceiver modules, protection OCS for managing protection communication paths, and other components that provide redundant communication paths to supplement these optical paths in the event of a failure of the conventional operating optical paths 312, 313.
[0024] Figure 4 Additional aspects of the described technique using a single protective OCS 415 to communicate with all dimensions 440 of the computing group are illustrated. In this case, the protective ECS 411 for each building block is configured as a 32×4 protective ECS 411. 32 ports facing the processing nodes are connected to 32 protective ICI links 412 originating from the same building block, while 4 externally facing ports are connected to a single protective OCS 415. The single protective OCS 415 is used to protect all dimensions from optical ICI link failures. Therefore, failures occurring in different dimensions 440 require reconfiguration of the building blocks, OCSs, and protective ICI links. Reconfiguration can be achieved through messaging between components protecting the communication path. For example, a failure at one link in the link to dimension 440 can be reported to the OCS 415, which relays the information to the building block 410. The building block 410 can then reconfigure traffic to avoid link breaks and redirect affected data traffic through the working channel. Similarly, the protection OCS 415 can receive messages from the building blocks indicating a failure. The protection OCS 415 can be reconfigured to reroute other traffic via the path used by the failure in building block 410.
[0025] Figure 4 A schematic diagram of a reconfigurable supergroup 200 is shown, where a 5D toroidal topology uses 2×2×2×2×2 (32 TPUs within a rack) as building blocks 210 of the 5D supergroup. It is assumed that each TPU has 10 1.6 Tbps ICI links 421, one ICI link per dimension in each direction. For this supergroup example, each building block 410 has 10 outward-facing hyperplanes 431 represented as X+, X-, Y+, Y-, Z+, Z-, a+, a-, b+, and b-. Here, X, Y, Z, a, and b represent each of the five dimensions 440, and + and – indicate the direction of each dimension 440.
[0026] For each building block 410, a total of 32 optical ICI links in each dimension 421 are connected to the 8 OCSs associated with that dimension 440 via two hyperplanes 431 in each dimension. Referring to dimension X 440x, the connection strategy can be understood as follows: Optical ICI links 1 and 2 from X+ and optical ICI links 1 and 2 from X- are connected to OCS 1 in dimension X. Optical ICI links 3 and 4 from X+ and optical ICI links 3 and 4 from X- are connected to OCS 2. The remaining 24 optical ICI links in this dimension are directed to the remaining 6 OCSs to complete the connection of this dimension.
[0027] Since all building blocks 410 are connected to the OCS, a healthy compute group can be built using the OCS even if some processing nodes within certain building blocks 410 are corrupted. In this way, reconfigurable supergroups can increase the availability of the group if the interconnection network is reliable enough to support reconfiguration.
[0028] Figure 5 A high-level conceptual diagram of a reconfigurable supergroup with a redundant interconnect network according to various aspects of the described technology is shown. The supergroup can be arranged from multiple building blocks. Each building block includes multiple processing nodes, which are grouped to form the building block. The processing nodes can be TPUs used to perform computational tasks. TPU group 503 is associated with a building block. For example, TPU 5031 is associated with building block 1, TPU 5032 with building block 2, and TPU 503n with building block n. Figure 5 In the example, a building block may include 32 processing nodes or TPUs 503. For instance, a building block containing 32 TPUs 503 can be physically housed within an associated computer rack 501. The building blocks can be arranged according to their corresponding computer racks 5011, 5012, and 501n.
[0029] In normal operation, each building block communicates with each dimension defined in the supergroup topology via regular communication link 509. Figure 5 In the example shown, there are five dimensions 510, represented as dimensions x, y, z, a, and b. Each dimension 510 includes two hyperplanes to allow communication along the incoming and outgoing directions. Each dimension 510 includes a set of OCSs that allow information routing throughout the supergroup.
[0030] Based on various aspects of the technology, Figure 5The supergroup further introduces a protection path 520 to provide redundancy for network traffic in the event of a failure of the normal communication path 530. The protection path 520 includes one or more protection OCSs 540 that communicate with dimension 510. For example, a protection OCS 540 may be provided for each dimension. Figure 5 In the example, there are five protection OCS 540s, each communicating with one of the corresponding dimensions 510x, 510y, 510z, 510a, and 510b. Figure 4 An alternative to the technology shown could be a single protective OCS 540 that communicates with each of the multiple dimensions 510.
[0031] The protection OCS 540 communicates with the processing node 503 via multiple protection ICI links 507 located between the building blocks and the protection OCS 540. Each building block's associated computer rack 501 includes a low-latency ECS 505, which connects the processing node 503 to the protection path 520 via the protection OCS 540. According to an illustrative example, the ECS 505 includes 32 ports facing the building block. In an example where the building block includes 32 processing nodes 503, the number of ports facing the building block in the ECS 505 can be 32, thus providing one port for each processing node 503. The ECS 505 further includes network-facing ports that connect the ECS 505 to the protection path 520 via the protection OCS(s) 540. By a non-limiting example, the ECS 505 may include 20 network-facing ports communicating with the protection OCS 540. In this example, the ECS 505 is configured as a 32×20-port ECS. Other base topologies for the ECS 505 can be used. In some implementations, the set of processing nodes 503 can be partitioned to correspond to the number of building block-oriented ports available in the ECS 505. The ECS 505 can be implemented as a top-of-rack (ToR) switch or a moR switch, which can be located within the same computer rack 501 as the processing nodes that make up the building block.
[0032] Supergroups define two parallel communication paths between the building blocks and the supergroup dimension. This technique provides redundant communication paths to prevent failures in conventional interconnect networks. Any failures in components such as ICI links, optical modules, or OCSs in the interconnect network can be overcome by implementing protective communication paths. Using low-latency ECSs to establish protective communication paths provides an easy-to-implement solution that is less complex and less costly than alternative solutions, while supporting synchronous operation in the system by preventing the introduction of significant latency.
[0033] Figure 6 An example system 600 in which the features described above can be implemented is shown. It should not be construed as limiting the scope of this disclosure or the usefulness of the features described herein. In this example, system 600 may include apparatus 606, server computing apparatus 630, storage system 640, and network 660.
[0034] Each device 606 may be a personal computing device intended for use by a corresponding user. Device 606 may include one or more processors 636, memory 646, data 666, and instructions 656. Each device 606 may also include outputs 676, user inputs 686, and a position sensor 696. By way of example only, device 606 may be a mobile phone or device such as a wireless-enabled PDA, smartphone, tablet PC, wearable computing device (e.g., smartwatch, AR / VR headset, smart helmet, etc.), netbook capable of accessing information via the Internet or other networks, or smart home device (e.g., home assistant, smart thermostat, smart doorbell, smart light, etc.).
[0035] The memory 646 of device 606 may store information accessible by processor 636. Memory 646 may also include data that can be retrieved, manipulated, or stored by processor 636. Memory 646 may be any non-transitory type capable of storing information accessible by processor 636, including non-transitory computer-readable media or other media that store data that can be read with the aid of electronic devices, such as hard disk drives, memory cards, read-only memory (“ROM”), random access memory (“RAM”), optical discs, and other writable and read-only memories. Memory 646 may store information accessible by processor 636, including instructions 656 and data 666 that can be executed by processor 636.
[0036] Data 666 can be retrieved, stored, or modified by processor 636 according to instruction 656. For example, although this disclosure is not limited to any particular data structure, data 666 can be stored in a computer register, stored as a table with multiple different fields and records in a relational database, stored in an XML document, or stored in a flat file. Data 666 can also be formatted in a computer-readable format such as, but not limited to, binary values, ASCII, or Unicode. Further, by way of example only, data 666 may include information sufficient to identify relevant information, such as numbers, descriptive text, proprietary codes, pointers, references to data stored in other memory (including other network locations), or information used by functions to calculate relevant data.
[0037] Instruction 656 can be any set of instructions (such as machine code) to be executed directly by processor 636, or any set of instructions (such as scripts) to be executed indirectly by the processor. In this regard, the terms "instruction," "application," "step," and "program" are used interchangeably herein. Instructions can be stored in an object code format for direct processing by the processor, or in any other computing device language, including sets of scripts or stand-alone source code modules that are interpreted on demand or compiled in advance. The function, methods, and routines of the instructions are explained in more detail below.
[0038] One or more processors 636 may include any conventional processor, such as a commercially available CPU or microprocessor. Alternatively, the processor may be a special-purpose component, such as an ASIC or other hardware-based processor. Although not required, computing device 606 may include specialized hardware components to perform specific computing functions faster or more efficiently.
[0039] although Figure 6 While the processor, memory, and other elements of device 606 are shown functionally within the same corresponding box, those skilled in the art will understand that a processor or memory may actually include multiple processors or memories that may or may not be stored in the same physical housing. Similarly, memory may be a hard disk drive or other storage medium located in a housing different from that of device 606. Therefore, references to processors or devices will be understood to include references to a collection of processors, devices, or memories that may or may not operate in parallel.
[0040] Output 676 may be a display, such as a monitor with a screen, a touch screen, a projector, or a television. The display 676 of one or more computing devices 606 may electronically display information to a user via a graphical user interface (“GUI”) or other type of user interface. For example, as will be discussed below, the display 676 may electronically display query results.
[0041] The user can input 686 using a mouse, keyboard, touchscreen, microphone, or any other type of input.
[0042] Device 606 can be located at various nodes of network 660 and can communicate directly and indirectly with other nodes of network 660. Although in Figure 6A device is described herein, but it should be understood that a typical system may include one or more devices, each located at a different node of network 660. Network 660 and the intermediate nodes described herein can be interconnected using various protocols and systems, such that the network may be part of the Internet, the World Wide Web, a specific intranet, a wide area network, or a local area network. Network 660 may utilize one or more proprietary standard communication protocols, such as WiFi, Bluetooth, 4G, 5G, etc. While certain advantages are gained when transmitting or receiving information as described above, other aspects of the subject matter described herein are not limited to any particular mode of transmission.
[0043] In one example, system 600 may include one or more server computing units 630 having multiple computing devices, such as a load-balanced server farm. These server computing units exchange information with different nodes in a network to receive data from other computing units, process data, and transfer data to other computing units. For example, the one or more server computing units 630 may be a web server capable of communicating with one or more client computing units 606 via network 660. Additionally, server computing units 630 may use network 660 to transmit and present information to a user of one of the other computing units 606.
[0044] Server computing device 630 may include one or more processors, memory, instructions, data, etc. These components operate in the same or similar manner as those described above with reference to computing device 606.
[0045] According to some examples, server computing device 630 can be connected via a network to data center 610, which houses any number of hardware accelerators. Data center 610 can be one of multiple data centers or other facilities housing various types of computing devices such as hardware accelerators. The computing resources housed in the data center can be designated for duplicate result monitoring, including identifying duplicate query results, etc.
[0046] Server computing device 630 can be configured to receive queries from client computing device 606 regarding computing resources in data center 610. For example, the environment may be part of a computing platform configured to provide various services to users through various user interfaces and / or application programming interfaces (APIs) that expose platform services. These services may include identifying the content in response to the query, determining whether the query result is a duplicate query result, etc. Client computing device 606 may transmit input data associated with the query. Server computing device 630 may receive the input data and, in response, identify and provide the query result as output. When identifying the query result, server computing device 630 may generate a signature for the query result. The generated signature may be compared with other signatures associated with the query result and / or historical query signatures. Based on the comparison, server computing device 630 may determine whether the query result is a duplicate query result. In an example where the query result is a duplicate query result, server computing device 630 may enable one or more preventative measures.
[0047] As another example of the potential services provided by the platform of the implementation environment, server computing devices can maintain multiple models depending on the different constraints available at the data center. For example, server computing devices can maintain different series of models for deployment on various types of TPUs and / or GPUs housed in the data center or otherwise available for processing.
[0048] Figure 7 This is a flowchart illustrating the process of establishing redundant protective communication paths according to various aspects of the described technology. For computing systems such as reconfigurable supergroups, processing nodes are arranged in groups represented as building blocks. For example, a building block may include a number of processing nodes housed within a physical computer rack. For nodes in a building block, multiple conventional communication paths (701) are established between the processing nodes and the interconnect network of the supergroup. For each processing node, an additional protective communication path (703) is established. A low-latency circuit switch is associated with each building block and receives data from the protective communication path of each processing node in building block (705). In the event of a failure in a component of a conventional communication path, data is transmitted to the network via the protective communication path (707). The protective communication paths from the processing nodes in the building block can prevent failures including, but not limited to, ICI links, optical modules, or optical circuit switches.
[0049] The aspects of this disclosure can be implemented in digital electronic circuit systems, in tangibly embodied computer software or firmware, and / or in computer hardware—such as the structures disclosed herein, their structural equivalents, or combinations thereof. The aspects of this disclosure can further be implemented as one or more computer programs, such as one or more modules of computer program instructions encoded on a tangible, non-transitory computer storage medium, for execution by or control of the operation of one or more data processing devices. The computer storage medium can be a machine-readable storage device, a machine-readable storage substrate, a random or serial access memory device, or a combination thereof. The computer program instructions can be encoded on artificially generated propagated signals (such as machine-generated electrical, optical, or electromagnetic signals), which are generated to encode information for transmission to a suitable receiver device for execution by the data processing device.
[0050] The term "configured" is used herein in conjunction with both system and computer program components. For a system of one or more computers configured to perform a specific operation or action, it means that the system has software, firmware, hardware, or a combination thereof installed thereon to cause the system to perform that operation or action. For one or more computer programs configured to perform a specific operation or action, it means that the one or more programs include instructions that, when executed by one or more data processing devices, cause those devices to perform those operations or actions.
[0051] The term "data processing device" refers to data processing hardware and encompasses a variety of devices, apparatuses, and machines used for processing data, including programmable processors, computers, or combinations thereof. Data processing devices may include dedicated logic circuit systems, such as field-programmable gate arrays (FPGAs) or application-specific integrated circuits (ASICs). Data processing devices may include code that creates an execution environment for computer programs, such as code that constitutes processor firmware, protocol stacks, database management systems, operating systems, or combinations thereof.
[0052] Data processing equipment may include dedicated hardware accelerator units for implementing machine learning models to handle the general and computationally intensive portions of machine learning training or production, such as inference or workloads. Machine learning models may be implemented and deployed using one or more machine learning frameworks.
[0053] The term "computer program" refers to a program, software, software application, app, module, software module, script, or code. Computer programs can be written in any type of programming language, including compiled, interpreted, declarative, or procedural languages, or combinations thereof. Computer programs can be deployed in any form, including as standalone programs or as modules, components, subroutines, or other units suitable for a computing environment. Computer programs can correspond to files in a file system and can be stored as part of a file that holds other programs or data (such as one or more scripts stored in a markup language document), in a single file dedicated to the program in question, or in multiple coordinating files (such as files storing one or more modules, subroutines, or code portions). Computer programs can execute on a single computer or on multiple computers located at a single site or distributed across multiple sites and interconnected through a data communication network.
[0054] The term "database" refers to any collection of data. The data can be unstructured or structured in any way. The data can be stored on one or more storage devices in one or more locations. For example, an indexed database may include multiple collections of data, each of which can be organized and accessed differently.
[0055] The term "engine" refers to a software-based system, subsystem, or process programmed to perform one or more specific functions. An engine can be implemented as one or more software modules or components or can be installed on one or more computers at one or more locations. A particular engine may have its own dedicated one or more computers, or multiple engines may be installed and run on the same one or more computers.
[0056] The processes and logical flows described herein can be executed by one or more computers, which execute one or more computer programs to perform functions by manipulating input data and generating output data. The processes and logical flows can also be executed by a dedicated logic circuit system or a combination of a dedicated logic circuit system and one or more computers.
[0057] A computer or special-purpose logic circuit system that executes one or more computer programs may include a central processing unit (including a general-purpose or special-purpose microprocessor) for running or executing instructions, and one or more memory devices for storing instructions and data. The central processing unit may receive instructions and data from one or more memory devices (such as read-only memory, random access memory, or combinations thereof), and may run or execute the instructions. The computer or special-purpose logic circuit system may also include or be operatively coupled to one or more storage devices such as a magnetic disk, magneto-optical disk, or optical disk for storing data, to receive data from or transfer data to the storage device. The computer or special-purpose logic circuit system may be embedded in another device, such as, for example, a mobile phone, a personal digital assistant (PDA), a mobile audio or video player, a game console, a global positioning system (GPS), or a portable storage device, such as a universal serial bus (USB) flash drive.
[0058] Computer-readable media suitable for storing one or more computer programs may include any form of volatile or non-volatile memory, medium, or memory device. Examples include: semiconductor memory devices, such as EPROM, EEPROM, or flash memory devices; magnetic disks, such as internal hard disks or removable disks; magneto-optical disks; CD-ROM disks; DVD-ROM disks; or combinations thereof.
[0059] The aspects of this disclosure can be implemented in a computing system comprising: a back-end component, such as a data server; a middleware component, such as an application server; or a front-end component, such as a client computer having a graphical user interface, a web browser, or an app, or any combination thereof. The components of the system can be interconnected via any form or medium of digital data communication—such as a communication network. Examples of communication networks include local area networks (LANs) and wide area networks (WANs), such as the Internet.
[0060] A computing system may include clients and servers. Clients and servers may be geographically separated and interact via a communication network. The client-server relationship arises from computer programs running on respective computers and having a client-server relationship with each other. For example, a server may transmit data, such as an HTML page, to a client device, for example, to display data to a user interacting with the client device and to receive user input from that user. Data generated at the client device—for example, the result of user interaction—may be received from the client device at the server.
[0061] Unless otherwise stated, the foregoing alternative examples are not mutually exclusive, but can be implemented in various combinations to achieve unique advantages. Since these and other variations and combinations of the features discussed above can be utilized without departing from the subject matter defined by the claims, the foregoing description of the examples should be illustrative rather than restrictive of the subject matter defined by the claims. Furthermore, the examples described herein and the provision of terms expressed as "such as," "including," etc., should not be construed as limiting the subject matter of the claims to the specific examples; rather, these examples are intended to illustrate only one of many possible implementations. Moreover, the same reference numerals in different figures may identify the same or similar elements.
Claims
1. A protected interconnection network for computing networks, characterized in that, include: Multiple processing nodes are connected via the protected interconnect network; At least one building block, the building block comprising a predetermined number of processing nodes from the plurality of processing nodes; A circuit switch (ECS) associated with each of the at least one building blocks, the ECS routing a portion of the data traffic of each building block to a protection communication path that operates in parallel with the regular communication path.
2. The protected interconnection network as described in claim 1, characterized in that, The computing network described therein is a reconfigurable supergroup.
3. The protected interconnection network as described in claim 2, characterized in that, The reconfigurable supergroup defines multiple dimensions, with each processing node communicating with each dimension.
4. The protected interconnection network as described in claim 3, characterized in that, include: An optical circuit switch (OCS) communicates with each ECS corresponding to each of the at least one building block, and with each of the plurality of dimensions.
5. The protected interconnection network as described in claim 3, characterized in that, include: An optical circuit switch (OCS) corresponds to each of the multiple dimensions, and the OCS in each dimension communicates with the ECS of each building block and its corresponding dimension in the at least one building block.
6. The protected interconnection network as claimed in claim 1, characterized in that, The building blocks in at least one of the building blocks include 32 processing nodes.
7. The protected interconnection network as claimed in claim 1, characterized in that, The processing node is a Tensor Processing Unit (TPU).
8. The protected interconnection network as claimed in claim 1, characterized in that, Further includes: A computer rack for accommodating one of the at least one building blocks.
9. The protected interconnection network as claimed in claim 8, characterized in that, Each ECS in each building block is configured as a top-of-rack ToR switch or a mid-rack MoR switch on the computer rack.
10. The protected interconnection network as claimed in claim 4, characterized in that, include: A first number of conventional inter-chip interconnect (ICI) links, the first number of ICI links connecting each processing node to a conventional OCS of the computing network; as well as A second number of protection chip interconnect (ICI) links connect each ECS to a protection OCS.
11. The protected interconnection network as claimed in claim 1, characterized in that, The ECS for each building block includes: A first number of building block-oriented ports, the first number of building block-oriented ports being connected to each processing node of the building block; and A second number of externally facing ports communicate with the protection communication path of the building block; The first number of ports and the second number of ports are determined by the topology of the computing network.
12. The protected interconnection network as claimed in claim 1, characterized in that, Each building block consists of two ECSs, and each ECS includes: A first number of building block-oriented ports, the first number of building block-oriented ports being connected to a subset of the processing nodes of the building block; and A second number of externally facing ports communicate with the protection communication path of the building block; The first number of ports and the second number of ports are determined by the topology of the computing network.
13. The protected interconnection network as claimed in claim 1, characterized in that, The ECS for each building block includes: 32 building block-oriented ports, each connected to a processing node of the building block; and Twenty externally facing ports, which communicate with a single OCS in the protection communication path of the building block.
14. The protected interconnection network as claimed in claim 1, characterized in that, Each processing node communicates with the following: A first number of conventional inter-chip interconnect (ICI) links, the first number of ICI links communicating with the conventional communication path of the computing network; as well as A protected ICI link communicates with the protected communication path of the computing network.
15. The protected interconnection network as claimed in claim 14, characterized in that, Each dimension includes two outward-facing hyperplanes to facilitate communication with the dimension in the inbound and outbound directions.
16. The protected interconnection network as claimed in claim 15, characterized in that, Each processing node includes: For each of the dimensions, there are two inter-chip interconnect (ICI) links; and An additional ICI link communicates with the ECS containing the building block of the processing node.
17. A method for protecting an interconnected network of computing networks, characterized in that, include: In the building blocks of the computing network, which includes multiple processing nodes, a first number of communication paths are established to connect to the regular processing paths; The building block provides an additional protection communication path that connects to the protection processing path; as well as Data is routed from the building block to the protection processing path via a circuit switch (ECS) associated with the building block of the computing network.
18. The method as described in claim 17, characterized in that, Further includes: In a building block comprising 32 processing nodes, each protection ICI link of each processing node of the building block is connected to the input port of a 32×20 port ECS.
19. The method as described in claim 18, characterized in that, Further includes: From the 32×20-port ECS, the 20 output ports of the 32×20-port ECS are connected to the protection optical circuit switch (OCS), which communicates with each of the multiple dimensions of the computer network.
20. The method as described in claim 18, characterized in that, Further includes: From the 32×20-port ECS, the 20 output ports of the 32×20-port ECS are connected to multiple protection optical circuit switches (OCS), each of the multiple OCSs being connected to the corresponding dimension of the computer network.