Method and device for adjusting network topology of computing cluster and related equipment
By adjusting the network topology of the computing cluster and optimizing the connection links between the optical switches and sub-clusters based on communication and physical topology information, the problem of low bandwidth utilization in the optoelectronic hybrid network was solved, and the overall performance of the computing cluster was improved.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- HUAWEI TECH CO LTD
- Filing Date
- 2024-11-21
- Publication Date
- 2026-05-22
AI Technical Summary
In hybrid optical-electrical networks, the continuous changes in traffic load lead to significant differences in traffic load on different forwarding paths, resulting in low bandwidth utilization of optical and electrical switches and affecting the overall performance of the computing cluster.
By adjusting the network topology of the computing cluster, the communication and physical topology information of the tasks are obtained, the link communication traffic requirements between sub-clusters are determined, and the connection links between the optical switches and sub-clusters are adjusted according to the load balancing strategy to avoid traffic congestion and overload and improve bandwidth utilization.
This effectively avoids traffic congestion between optical and electrical switches, improves bandwidth utilization between them, and thus enhances the overall performance of the computing cluster.
Smart Images

Figure CN122073570A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of computing technology, and in particular to a method, apparatus, and related equipment for adjusting the network topology of a computing cluster. Background Technology
[0002] Hybrid optoelectronic networking is a networking method that combines optical and electrical communication technologies. It typically includes electrical switches and optical switches, and different electrical switches can exchange data via optical switches. Optical switches not only offer greater bandwidth capacity and lower transmission latency, but are also generally better suited to meet the latency requirements of artificial intelligence (AI) tasks (such as training AI models). Furthermore, optical switches can dynamically adjust the topology between themselves and different electrical switches based on optical circuit switching (OCS) technology. Figure 1 As shown, optical switches can use micro-electro-mechanical system (MEMS) lenses to connect input and output fiber optic ports; furthermore, optical switches can control the connection state of input fiber optic ports with different output fiber optic ports by adjusting the rotation angle of the lenses.
[0003] Typically, the traffic load in a hybrid optical-electrical network may fluctuate continuously, which can lead to significant differences in traffic load on different forwarding paths. This can result in low bandwidth utilization of both optical and electrical switches, affecting the overall performance of the computing cluster. Summary of the Invention
[0004] This application provides a method for adjusting the network topology of a computing cluster to improve the bandwidth utilization of optical and electrical switches and enhance the overall performance of the computing cluster. Furthermore, this application also provides corresponding apparatus, computing devices, computer-readable storage media, and computer program products for adjusting the network topology of a computing cluster.
[0005] Firstly, this application provides a method for adjusting the network topology of a computing cluster. The computing cluster includes multiple sub-computing clusters, each sub-computing cluster includes multiple computing nodes, and the computing nodes in each sub-computing cluster are connected via electrical switches, such as single-layer or multi-layer electrical switches. Furthermore, the multiple sub-computing clusters are connected via optical switches. In practical applications, each sub-cluster may include multiple computing nodes within a computer room. At least one task runs within the computing cluster; this task may be an AI task or other types of tasks. During the adjustment of the network topology of the computing cluster, the adjustment device can acquire communication information between the multiple sub-tasks included in the at least one task running in the computing cluster. This communication information may include, for example, the communication domain to which the computing node belongs, the communication operator to be executed, and the amount of data to be communicated. Furthermore, the adjustment device can also acquire physical topology information of the computing cluster, which indicates the connection relationships between computing nodes, electrical switches, and optical switches. Then, the adjustment device determines the communication traffic requirements of each link in the multiple links connecting the first and second sub-clusters via optical switches, based on the communication information between the multiple subtasks and the physical topology information of the computing cluster. It also determines the remaining traffic of each link in the multiple links between the first and second sub-clusters based on the physical topology information of the computing cluster. Therefore, the adjustment device adjusts the connection links between the optical switches and the first and second sub-clusters based on the remaining traffic of each link in the multiple links between the first and second sub-clusters and the communication traffic requirements of each link in the multiple links between the first and second sub-clusters.
[0006] In this way, the adjustment device adjusts the connection links between the optical switch and the first and second sub-clusters based on the communication traffic requirements of multiple links between different sub-clusters participating in the running tasks in the computing cluster. This ensures that the adjusted connection links can meet the communication traffic requirements between the first and second sub-clusters, effectively preventing the traffic generated when the computing cluster is running tasks from being blocked on the physical link between the optical switch and the electrical switch. This improves the bandwidth utilization between the optical switch and the electrical switch, thereby improving the overall performance of the computing cluster.
[0007] In one possible implementation, when the adjustment device adjusts the connection links between the optical switch and the first and second sub-clusters based on the remaining traffic and communication traffic requirements of each link in the multiple links between the first and second sub-clusters, specifically, it may adjust the connection links between the optical switch and the first and second sub-clusters when the remaining traffic of the first link in the multiple links between the first and second sub-clusters cannot meet the communication traffic requirements of the first link in the multiple links between the first and second sub-clusters. Similarly, for the second link between the first and second sub-clusters, the adjustment device may also adjust the connection links between the optical switch and the first and second sub-clusters based on the remaining traffic and communication traffic requirements of the second link. Thus, when the remaining traffic on multiple links between two sub-clusters cannot meet the communication traffic demand, the connection links between the optical switch and these two sub-clusters and the second sub-cluster can be adjusted to ensure that the connection links between the optical switch and these two sub-clusters can meet the communication traffic demand between the two sub-clusters. This can prevent traffic congestion caused by the computing nodes in these two sub-clusters and prevent traffic overload on the links between the optical switch and these two clusters, thereby improving the communication efficiency between different sub-clusters and thus improving the overall performance of the computing cluster.
[0008] In one possible implementation, when the adjusting device adjusts the connection links between the optical switch and the first and second sub-clusters based on the remaining traffic and communication traffic requirements of each link among the multiple links between the first and second sub-clusters, specifically, it may adjust the connection links between the optical switch and the first and second sub-clusters based on a load balancing strategy. In this way, by adjusting the connection links between the optical switch and the first and second sub-clusters through a load balancing strategy, the traffic between the first and second sub-clusters can be relatively evenly distributed across different links of the optical switch. This avoids traffic overload on some links of the optical switch, which could affect the communication efficiency between different sub-clusters, thereby improving the overall performance of the computing cluster.
[0009] In one possible implementation, the communication information between multiple subtasks includes the communication relationships and communication data volume between multiple computing units executing the multiple subtasks. These computing units are distributed across computing nodes in a first sub-cluster and a second sub-cluster. In practical applications, each computing node may include two or more computing units. The physical topology information of the computing cluster includes the links between each computing node in the first and second sub-clusters. Therefore, when the adjusting device determines the communication traffic requirements of each link in the multiple links based on the communication information between the multiple subtasks and the physical topology information of the computing cluster, it can specifically determine the communication traffic requirements of each link in the multiple links based on the communication relationships, communication data volume, and physical topology information of the multiple computing units executing the multiple subtasks. Thus, by determining the communication traffic requirements of each link in the multiple links between the two sub-clusters, it is possible to subsequently adjust the connection links between the optical switch and the first and second sub-clusters based on these communication traffic requirements, thereby improving the bandwidth utilization between the optical switch and the electrical switch.
[0010] In one possible implementation, when the adjusting device determines the communication traffic requirements of each link in a plurality of links based on the communication relationships, communication data volume, and physical topology information of the computing cluster among the multiple computing units executing multiple sub-tasks, specifically, it may determine a traffic matrix based on the communication relationships and communication data volume among the multiple computing units executing multiple sub-tasks. This traffic matrix indicates the communication relationships and communication data volume among the multiple computing units; that is, it can indicate which computing units need to communicate and the amount of data communicated, and it can also indicate which computing units do not need to communicate. Then, based on the traffic matrix and the physical topology information of the computing cluster, the communication traffic requirements of each link in the plurality of links are determined. In this way, by converting the communication information between multiple sub-tasks into a traffic matrix, the communication traffic requirements between sub-clusters can be determined, so that the connection links between the optical switch and the first and second sub-clusters can be adjusted according to the communication traffic requirements between the sub-clusters, thereby improving the overall performance of the computing cluster.
[0011] In one possible implementation, when the adjustment device determines the communication traffic requirements of each link in multiple links based on the traffic matrix and the physical topology information of the computing cluster, it can specifically determine the logical topology of the links between different sub-clusters in multiple sub-clusters based on the traffic matrix and the physical topology information of the computing cluster. The logical topology is used to indicate the communication traffic requirements of each link in multiple links connecting multiple sub-clusters via optical switches. In this way, the adjustment device can determine the communication traffic requirements of each link in multiple links connecting different sub-clusters via optical switches based on the traffic matrix and physical topology information, so that the connection links between the optical switches and the first and second sub-clusters can be adjusted according to the communication traffic requirements between the sub-clusters, thereby improving the overall performance of the computing cluster.
[0012] In one possible implementation, when the adjustment device determines the logical topology of links between different sub-clusters in multiple sub-clusters based on the traffic matrix and the physical topology information of the computing cluster, specifically, it can determine the links used for data communication between multiple computing units based on the traffic matrix and the physical topology information of the computing cluster. Then, the adjustment device aggregates the communication data volume between multiple computing units using the second link to obtain the communication traffic requirement of the second link among the multiple links connecting multiple sub-clusters via optical switches. Thus, by aggregating the traffic between multiple computing nodes using the same link, the adjustment device can determine the traffic volume required for data forwarding between different sub-clusters via that link, i.e., determine the communication traffic requirement of that link. Subsequently, by adjusting the connection links between the optical switch and the first and second sub-clusters based on this communication traffic requirement, the overall performance of the computing cluster can be improved.
[0013] In one possible implementation, when adjusting the connection links between the optical switch and the first and second sub-clusters based on the remaining traffic and communication traffic requirements of each link in the multiple links connecting the first and second sub-clusters via the optical switch, the adjusting device may specifically determine the physical topology between the optical switch and the first and second sub-clusters based on the remaining traffic and logical topology of each link. Based on the physical topology, the device calculates the order for dismantling and recreating physical links. The physical links are the links between the sub-clusters and the optical switch. Thus, the adjusting device can adjust the connection links between the optical switch and the first and second sub-clusters according to the correct order. By determining the order of dismantling and recreating the physical links between the optical switch and the first and second sub-clusters, traffic loss during the connection link adjustment process can be avoided. This allows for physical topology reconstruction without traffic loss, preventing the physical topology reconstruction from affecting the computing cluster's operational tasks.
[0014] In one possible implementation, the adjustment device can also adjust the logical topology between different switches in the first sub-cluster, or the logical topology between different switches in the second sub-cluster, based on the communication traffic requirements of each link in the multiple links between the first and second sub-clusters. In this way, by optimizing the logical topology between different switches in the sub-clusters, the adjustment device can avoid traffic congestion or overload on the connection links between switches within the sub-clusters, thereby further improving the bandwidth utilization between switches and / or between switches and computing nodes, thus improving the overall performance of the computing cluster.
[0015] In one possible implementation, when the adjustment device acquires communication information between multiple subtasks included in at least one task, it may specifically acquire initial communication information reported by computing units in computing nodes of a first sub-cluster and a second sub-cluster for each of the multiple subtasks included in at least one task, and aggregate the initial communication information reported by computing units in computing nodes of the first and second sub-clusters to obtain the communication information between the multiple subtasks included in at least one task. In this way, the adjustment device obtains global communication information between multiple subtasks by clustering the information reported by each computing node, so that subsequent optimization of the network topology of the computing cluster can be performed based on this communication information.
[0016] In one possible implementation, the adjustment device can also acquire resource indication information, which indicates the resources corresponding to each task in at least one task. The resources corresponding to each task include electrical switches and optical switches that the computing cluster can use when running the task, and the resources corresponding to different tasks in at least one task do not overlap. Therefore, when the adjustment device determines the communication traffic requirements of each link in the multiple links connecting the first and second sub-clusters via optical switches based on the communication information between multiple sub-tasks and the physical topology information of the computing cluster, specifically, it can determine the communication traffic requirements of each link in the multiple links connecting the first and second sub-clusters via optical switches based on the resource indication information, the communication information between multiple sub-tasks, and the physical topology information of the computing cluster. In this way, for different tasks, the adjustment device can optimize the network topology of the computing cluster for each task within different resource ranges, thereby meeting the resource isolation requirements of different tasks in practical application scenarios and improving the overall network bandwidth utilization of the computing cluster.
[0017] In one possible implementation, the communication relationship between multiple computing units includes the communication domains to which the multiple computing units belong and the communication operators executed by the multiple computing units.
[0018] In one possible implementation, the task running in the computing cluster is an AI (Artificial Intelligence) task. In this case, because the computing cluster typically generates similar or identical traffic periodically during the execution of AI tasks, and because it executes known communication operators, the traffic generated during AI task execution usually exhibits a certain regularity and predictability. The adjustment device then adjusts the connection links between the optical switch and multiple sub-clusters based on the communication traffic requirements of the multiple links between the different sub-clusters participating in the task. This allows the adjusted connection links to better adapt to the traffic distribution requirements generated by the AI task, thereby maintaining a consistently high level of bandwidth utilization between the optical and electrical switches, and consequently, maintaining a consistently high level of overall computing cluster performance. Alternatively, the task running in the computing cluster may be an HPC (High-Performance Computing) task, or other types of tasks.
[0019] Secondly, this application provides an apparatus for adjusting the network topology of a computing cluster, characterized in that the computing cluster includes multiple sub-computing clusters, computing nodes in each sub-computing cluster are connected via electrical switches, the multiple sub-computing clusters are connected via optical switches, and at least one task runs in the computing cluster; the apparatus includes: an acquisition module, used to acquire communication information between multiple subtasks included in at least one task; acquire physical topology information of the computing cluster; a determination module, used to determine the communication traffic requirements of each link in multiple links connecting a first sub-cluster and a second sub-cluster in the multiple sub-clusters via optical switches based on the communication information between the multiple subtasks and the physical topology information of the computing cluster; determine the remaining traffic of each link in the multiple links based on the physical topology information of the computing cluster; and an adjustment module, used to adjust the connection links between the optical switches and the first and second sub-clusters based on the remaining traffic of each link in the multiple links and the communication traffic requirements of each link in the multiple links.
[0020] In one possible implementation, when adjusting the connection link between the optical switch and the first sub-cluster and the second sub-cluster, the adjustment module is specifically used to: adjust the connection link between the optical switch and the first sub-cluster and the second sub-cluster when the remaining traffic of the first link among multiple links cannot meet the communication traffic requirements of the first link.
[0021] In one possible implementation, when adjusting the connection links between the optical switch and the first sub-cluster and the second sub-cluster, the adjustment module is specifically used to: adjust the connection links between the optical switch and the first sub-cluster and the second sub-cluster based on the remaining traffic of each link in the multiple links and the communication traffic requirements of the multiple links, according to a load balancing strategy.
[0022] In one possible implementation, the communication information between multiple subtasks includes the communication relationship and communication data volume between multiple computing units executing multiple subtasks. The multiple computing units are distributed among computing nodes in a first sub-cluster and a second sub-cluster. The physical topology information of the computing clusters includes the links between each computing node in the first sub-cluster and the second sub-cluster. When determining the communication traffic requirements of each link in the multiple links, the determining module is specifically used to: determine the communication traffic requirements of each link in the multiple links based on the communication relationship, communication data volume, and physical topology information of the computing clusters.
[0023] In one possible implementation, when the determining module determines the communication traffic requirements of each link in multiple links based on the communication relationships, communication data volume, and physical topology information of the computing cluster among multiple computing units executing multiple sub-tasks, it is specifically used to: determine a traffic matrix based on the communication relationships and communication data volume among multiple computing units executing multiple sub-tasks, wherein the traffic matrix is used to indicate the communication relationships and communication data volume among multiple computing units; and determine the communication traffic requirements of each link in multiple links based on the traffic matrix and the physical topology information of the computing cluster.
[0024] In one possible implementation, when the determining module determines the communication traffic requirements of each link in multiple links based on the traffic matrix and the physical topology information of the computing cluster, it is specifically used to: determine the logical topology of the links between different sub-clusters in multiple sub-clusters based on the traffic matrix and the physical topology information of the computing cluster. The logical topology is used to indicate the communication traffic requirements of each link in multiple links connected between multiple sub-clusters through optical switches.
[0025] In one possible implementation, when the determining module determines the logical topology of the links between different sub-clusters in multiple sub-clusters based on the traffic matrix and the physical topology information of the computing cluster, it is specifically used to: determine the links used for data communication between multiple computing units based on the traffic matrix and the physical topology information of the computing cluster; aggregate the communication data volume between multiple computing units using the second link to obtain the communication traffic requirements of the second link among the multiple links connected between multiple sub-clusters through optical switches.
[0026] In one possible implementation, when adjusting the connection links between the optical switch and the first sub-cluster and the second sub-cluster, the adjustment module is specifically used to: determine the physical topology between the optical switch and the first sub-cluster and the second sub-cluster based on the remaining traffic and logical topology of each link among multiple links; calculate the order of dismantling physical links and creating new physical links based on the physical topology, wherein the physical links are the links between the sub-cluster and the optical switch; and adjust the connection links between the optical switch and the first sub-cluster and the second sub-cluster according to the order.
[0027] In one possible implementation, the adjustment module is further configured to: adjust the logical topology between different switches in the first sub-cluster, or adjust the logical topology between different switches in the second sub-cluster, based on the communication traffic requirements of each link in the multiple links between the first sub-cluster and the second sub-cluster.
[0028] In one possible implementation, when the acquisition module acquires communication information between multiple subtasks included in at least one task, it is specifically used to: acquire initial communication information reported by computing units in computing nodes in the first sub-cluster and the second sub-cluster for the multiple subtasks included in at least one task; and aggregate the initial communication information reported by computing units in computing nodes in the first sub-cluster and the second sub-cluster to obtain communication information between multiple subtasks included in at least one task.
[0029] In one possible implementation, the acquisition module is further configured to acquire resource indication information, which indicates the resources corresponding to each task in at least one task. The resources corresponding to each task include electrical switches and optical switches that can be used when the computing cluster runs the task, and the resources corresponding to different tasks in at least one task do not overlap. When determining the communication traffic requirements of each link in the multiple links connected between the first sub-cluster and the second sub-cluster in multiple sub-clusters via optical switches, the determination module is specifically configured to: determine the communication traffic requirements of each link in the multiple links connected between the first sub-cluster and the second sub-cluster in multiple sub-clusters via optical switches based on the resource indication information, the communication information between multiple sub-tasks, and the physical topology information of the computing cluster.
[0030] In one possible implementation, the communication relationship between multiple computing units includes the communication domains to which the multiple computing units belong and the communication operators executed by the multiple computing units.
[0031] In one possible implementation, the tasks running in the computing cluster are either artificial intelligence (AI) tasks or high-performance computing (HPC) tasks.
[0032] The apparatus for adjusting the network topology of a computing cluster provided in the second aspect corresponds to the method for adjusting the network topology of a computing cluster provided in the first aspect. Therefore, the technical effects of the second aspect and any implementation thereof can be referred to the technical effects of the corresponding implementation thereof in the first aspect, and will not be elaborated here.
[0033] Thirdly, this application provides a computing device including a processor and a memory. The processor and the memory communicate with each other. The processor executes instructions stored in the memory to cause the computing device to perform a method for adjusting the network topology of a computing cluster, as described in the first aspect or any implementation thereof. It should be noted that the memory may be integrated into the processor or may be independent of the processor. The computing device may also include a bus. The processor is connected to the memory via the bus. The memory may include readable storage and random access memory.
[0034] Fourthly, this application provides a computer-readable storage medium storing instructions that, when executed on a computing device, cause the computing device to perform the operational steps of the method for adjusting the network topology of a computing cluster as described in the first aspect or any implementation thereof.
[0035] Fifthly, this application provides a computer program product containing instructions that, when run on a computing device, cause the computing device to perform the operational steps of the method for adjusting the network topology of a computing cluster as described in the first aspect or any implementation thereof.
[0036] Based on the implementation methods provided in the above aspects, this application can be further combined to provide more implementation methods. Attached Figure Description
[0037] Figure 1 This is a schematic diagram illustrating how the connection status between the input fiber optic port and the output fiber optic port can be controlled by adjusting the rotation angle of a lens within an optical switch.
[0038] Figure 2 A schematic diagram of an exemplary computing cluster provided in this application;
[0039] Figure 3 A flowchart illustrating a method for adjusting the network topology of a computing cluster provided in this application;
[0040] Figure 4 This application provides a flowchart illustrating the process of adjusting the logical and physical topology for AI tasks.
[0041] Figure 5 This diagram illustrates how the initial communication information reported by each computing node is aggregated to obtain the communication information for the AI task.
[0042] Figure 6 A schematic diagram illustrating the generation of an N*N traffic matrix based on communication information between multiple subtasks;
[0043] Figure 7a To adjust the network topology between sub-cluster 101 and sub-cluster 102;
[0044] Figure 7b To adjust the network topology between sub-cluster 101 and sub-cluster 102;
[0045] Figure 8 A schematic diagram of a device for adjusting the network topology of a computing cluster provided in this application;
[0046] Figure 9 This is a schematic diagram of the hardware structure of a computing device provided in this application. Detailed Implementation
[0047] The terms "first," "second," etc., used in the specification, claims, and accompanying drawings of this application are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such terms can be used interchangeably where appropriate; this is merely a method of distinction used in describing objects with the same attributes in the embodiments of this application.
[0048] The technical solutions in this application will now be described with reference to the accompanying drawings.
[0049] See Figure 2 This is a schematic diagram of the structure of an exemplary computing cluster 10. Figure 2 As shown, the computing cluster 10 includes multiple computing nodes, multiple electrical switches, multiple optical switches, and a management node. Figure 2 The following example illustrates the use of a computing cluster 10, which includes 4 computing nodes (compute nodes 1 to 4), 8 electrical switches (electrical switches 1 to 8), 2 optical switches (optical switch 1 and optical switch 2), and a management node 200.
[0050] In computing cluster 10, a computing node refers to a node with data computing capabilities, such as a server or other computing device. Each computing node may include one or more computing units. Figure 2 This example uses a single compute node comprising four technical units, such as compute node 1 including compute units 1 through 4. In practical applications, different compute units deployed within the same compute node can communicate via a bus, such as a compute express link (CXL) bus.
[0051] The computing units in a computing node can be implemented using a processor. For example, a processor can be any type of processor or any combination thereof, such as a central processing unit (CPU), an accelerator, an application-specific integrated circuit (ASIC), a programmable logic device (PLD), a complex programmable logical device (CPLD), a field-programmable gate array (FPGA), a generic array logic (GAL), a system-on-chip (SoC), a software-defined infrastructure (SDI) chip, an artificial intelligence (AI) chip, or a data processing unit (DPU). An accelerator can be, for example, a graphics processing unit (GPU), a neural network processing unit (NPU), or a tensor processing unit (TPU).
[0052] The computing cluster 10 can run at least one task, such as an AI task or a high-performance computing (HPC) task. For example, multiple computing units in the computing cluster 10 can be used to train an AI model, or multiple computing units in the computing cluster 10 can be used to run an AI model and support the AI model inference, such as supporting large language model (LLM) inference to output human-computer dialogue text. For example, the AI model can be, in addition to large language model (LLM), a large language model meta AI (LLaMA) model, a bidirectional encoder representations from transformers (BERT) model, or a generative pre-trained Transformer 3 (GPT-3) model, or other types of models, such as GPT-4, etc., without limitation.
[0053] exist Figure 2 The computing cluster 10 shown may include multiple sub-clusters, such as Figure 2 The sub-clusters 101 and 102 shown can each include one or more computing nodes. Different computing nodes within each sub-cluster can be connected via electrical switches, and different sub-clusters can be connected via optical switches. An electrical switch is a switch that uses electrical signals for data exchange, which can be transmitted via twisted-pair or coaxial cables. An optical switch is a switch that uses optical signals for data exchange, which can be transmitted via optical fibers. For example, assuming computing node 1 in sub-cluster 101 communicates with computing node 3 in sub-cluster 102 (specifically, computing unit 1 in computing node 1 needs to communicate with computing unit 9 in computing node 3), computing node 1 can forward the communication data to electrical switch 1 in sub-cluster 101, and electrical switch 1 forwards the communication data to optical switch 1 via electrical switch 5. Then, optical switch 1 forwards the communication data to electrical switch 7 in sub-cluster 102, and electrical switch 7 forwards the communication data to computing node 3 via electrical switch 3.
[0054] In practical applications, multiple switches within each sub-cluster can be deployed in multiple tiers. For example, taking a two-tier deployment as an example... Figure 2As shown, some electrical switches can act as leaf nodes, while others can act as spine nodes. Different leaf nodes can interact with each other through spine nodes. Furthermore, electrical switches acting as spine nodes in different sub-clusters (such as spine nodes) can interact with each other through optical switches, thereby establishing communication connections between different sub-clusters.
[0055] The management node 200 can be implemented using a processor, or using a computing device that includes a processor. For example... Figure 2 As shown, the management node 200 may include an electrical management device 201 and an optical management device 202. The electrical management device 201 is used to monitor and manage the electrical switches in the hybrid optical-electrical network, such as configuring the routing tables in the electrical switches. The optical management device 202 is used to manage the optical switches in the hybrid optical-electrical network, such as configuring the connection between the input and output fiber optic ports of the optical switches to be open or closed, thereby controlling the connection or disconnection of physical links between different electrical switches (such as spine nodes).
[0056] Furthermore, the management node 200 may also include an adjustment device 203, which can be used to adjust the network topology used in the optoelectronic hybrid network when the computing cluster 10 is running tasks. This includes adjusting the logical topology in the optoelectronic hybrid network (such as adjusting the routing table in the optical switch), or adjusting the physical topology in the optoelectronic hybrid network (such as adjusting the connection status between the input fiber optic port and the output fiber optic port in the optical switch), or adjusting both the logical topology and the physical topology simultaneously.
[0057] The electrical management device 201, optical management device 202, and adjustment device 203 can all be implemented in software or hardware. Taking the adjustment device 203 as an example, when implemented in software, the adjustment device 203 can be, for example, a process running on the management node 200. When implemented in hardware, the adjustment device 203 can be implemented using a processor, such as any processor or any combination thereof, including CPU, ASIC, PLD, CPLD, FPGA, GAL, SoC, SDI chip, AI chip, and DPU (data processing unit).
[0058] During the execution of one or more tasks in the computing cluster 10, the traffic generated by these tasks dynamically changes. This causes the traffic load on multiple links between optical and electrical switches in the hybrid optical-electrical network to fluctuate continuously, resulting in significant differences in traffic load across these links. For example, traffic generated during service operation may experience congestion (excessive traffic load) on some links between optical and electrical switches, while other links may have lower traffic loads or be idle. Consequently, the bandwidth utilization between optical and electrical switches in the hybrid optical-electrical network is low, which affects the overall performance of the computing cluster 10.
[0059] Based on this, Figure 2 In the computing cluster 10 shown, the adjustment device 203 can adjust the physical topology between the optical switch and the electrical switch according to the communication traffic requirements between different sub-clusters during the operation of the computing cluster 10.
[0060] Specifically, the adjustment device 203 acquires communication information between multiple subtasks included in at least one task running by the computing cluster 10. When the task executed by the computing cluster 10 is specifically a distributed training task for an AI model, the communication information acquired by the adjustment device 203 may include, for example, the communication domain, communication operators, and communication data volume corresponding to the AI model. Each subtask may be a task to be executed by a computing node within the communication domain. Thus, the adjustment device 203 can determine the communication traffic requirements of each link in the multiple links connected by optical switches between different sub-clusters based on the acquired communication information between the multiple subtasks. The following determines... Figure 2Taking sub-clusters 101 and 102 as examples, the adjustment device 203 can also acquire the physical topology information of computing cluster 10. This physical topology information can be used to indicate the links between each computing node in sub-cluster 101 and sub-cluster 102. The links between two computing nodes can include links between a computing node and an electrical switch, links between different electrical switches, and links between an electrical switch and an optical switch. Therefore, the adjustment device 203 can determine the remaining traffic of each link among the multiple links between sub-cluster 101 and sub-cluster 102 based on this physical topology information. Then, the adjustment device 203 adjusts the connection links between the optical switches in the computing cluster 10 and the sub-clusters 101 and 102 based on the remaining traffic of each link in the multiple links between sub-clusters 101 and 102, and the communication traffic demand of each link in the multiple links between sub-clusters 101 and 102. For example, it may add new links between the optical switches and the sub-clusters 101 and 102, or remove and rebuild the connection links between the optical switches and the sub-clusters 101 and 102, so as to balance the number of links between the multiple optical switches and the sub-clusters 101 and 102 respectively.
[0061] It is understood that the adjustment device 203 adjusts the connection links between the optical switch and sub-cluster 101 and sub-cluster 102 according to the communication traffic requirements of multiple links between different sub-clusters participating in the running tasks in the computing cluster 10. This enables the adjusted connection links to meet the communication traffic requirements between sub-cluster 101 and sub-cluster 102. This can effectively prevent the traffic generated when the computing cluster 10 is running tasks from being blocked on the link between the optical switch and the electrical switch, thereby improving the bandwidth utilization between the optical switch and the electrical switch and thus improving the overall performance of the computing cluster 10.
[0062] Furthermore, when the task run by computing cluster 10 is specifically an AI task, the computing cluster 10 typically generates similar or identical traffic periodically during the AI task execution process. For example, the traffic generated in each iteration cycle during the iterative training of the AI model by computing cluster 10 is basically the same. Moreover, computing cluster 10 executes known communication operators during the AI task execution process, which makes the traffic generated by computing cluster 10 during AI task execution generally have a certain regularity and predictability. At this time, the adjustment device 203 adjusts the connection links between the optical switch and sub-cluster 101 and sub-cluster 102 based on the communication traffic requirements of multiple links between different sub-clusters participating in the task within computing cluster 10. This allows the adjusted connection links to better adapt to the traffic distribution requirements generated by the AI task, thereby ensuring that the bandwidth utilization between the optical switch and the electrical switch remains at a high level, and consequently, that the overall performance of computing cluster 10 remains at a high level.
[0063] In addition, compared to the adjustment device 203, which predicts the traffic distribution of the optoelectronic hybrid network in the future based on the traffic distribution data of the optoelectronic hybrid network in the past time period and adjusts the network topology of the optoelectronic hybrid network according to the prediction results, the adjustment device 203 adjusts the connection links between the optical switch and sub-cluster 101 and sub-cluster 102 based on the communication traffic requirements of multiple links between different sub-clusters participating in the running tasks in the computing cluster 10. This can adapt to the traffic changes in the optoelectronic hybrid network caused by the dynamic changes of the running tasks in the computing cluster 10 (such as the computing cluster 10 running new tasks), and avoid large differences between the actual traffic and the predicted traffic, which would lead to large differences in the traffic load of multiple links between the optical switch and the electrical switch. This can ensure that the bandwidth utilization between the optical switch and the electrical switch can be effectively improved, thereby ensuring that the overall performance of the computing cluster 10 can be maintained at a high level.
[0064] It is worth noting that the above Figure 2The computing cluster 10 shown is merely an illustrative example and is not intended to limit the scope of the system. For instance, in other possible implementations, the data processing system may include more types of nodes, such as computing power management nodes and task scheduling nodes. The computing power management node can be used to manage the computing power in the computing cluster 10, such as putting some computing nodes in the computing cluster 10 into hibernation to reduce power consumption. The task scheduling node can be used to schedule AI tasks to computing nodes in the computing cluster 10 (specifically, to one or more computing units within that computing node), and different AI tasks can be scheduled to different computing nodes. Furthermore, in other possible implementations, the power management device 201, the optical management device 202, and the adjustment device 203 can be deployed independently in the data processing system; for example, the power management device 201, the optical management device 202, and the adjustment device 203 can each be implemented through a separate management server. Moreover, in other possible implementations, the multiple computing nodes in the computing cluster 10 can be divided into a greater number of sub-clusters.
[0065] For ease of understanding, embodiments of the method for adjusting the network topology of a computing cluster provided in this application are described below with reference to the accompanying drawings.
[0066] See Figure 3 , Figure 3 This application provides a flowchart illustrating a method for adjusting the network topology of a computing cluster, which can be applied to... Figure 2 The aforementioned computing cluster 10 can also be applied to other suitable data processing systems. For ease of explanation, this embodiment uses an application... Figure 2 The computing cluster 10 shown is used as an example for illustration.
[0067] in, Figure 3 The method for adjusting the network topology of the computing cluster shown can specifically include:
[0068] S301: Adjustment device 203 acquires communication information between multiple subtasks included in at least one task running in computing cluster 10.
[0069] The tasks run by computing cluster 10 can be tasks that generate traffic that requires forwarding by electrical and optical switches, such as AI tasks, HPC tasks, or other types of tasks. AI tasks refer to tasks that require execution based on AI technology, such as distributed training of AI models or AI model inference tasks.
[0070] The tasks run by computing cluster 10 can include multiple subtasks, each of which can be executed by at least one computing unit within computing cluster 10. Furthermore, different computing units can communicate with each other while executing their assigned subtasks. When computing units in different computing nodes need to communicate with each other, traffic that needs to be forwarded through electrical switches and optical switches may be generated.
[0071] For example, when the task run by computing cluster 10 is specifically an AI task, the multiple sub-tasks included in the AI task can specifically be model training tasks or model inference tasks that the computing units need to execute. Furthermore, the communication information between the multiple sub-tasks included in the AI task can, for example, include the communication relationships and communication data volume between the multiple computing units participating in the execution of the multiple sub-tasks. The communication relationships between the multiple computing units can be used to indicate whether communication is required between different computing units. For example, the communication relationships between the multiple computing units can specifically include the communication domain to which the multiple computing units belong and the communication operators that the computing units need to execute.
[0072] In this context, a communication domain refers to a set of multiple computing units within the computing cluster 10 that participate in executing the AI task. This set defines the communication relationships between these computing units; for example, they can communicate data according to the rules corresponding to one or more communication operators. In practical applications, a communication domain may include identifiers for multiple computing units, indicating which units belong to that domain. Thus, by acquiring information such as the communication domain for the AI task, the adjustment device 203 can determine which computing units in the computing cluster 10 need to communicate. Furthermore, a communication domain may also include other types of information, such as its identifier. For instance, during the execution of the AI task, multiple computing units in the computing cluster 10 may be assigned to different communication domains, such as different computing units executing multiple different communication operators corresponding to the AI task. The adjustment device 203 can then distinguish between different communication domains based on their identifiers.
[0073] Communication operators refer to the operators executed to achieve data exchange (or data synchronization) between different computing units, such as broadcast and all-reduce operators. Thus, by acquiring the communication operators required by the computing units, the adjustment device 203 can determine the data interaction method between different computing units. For example, when multiple computing units in a communication domain execute a broadcast operator, data from one computing unit is sent to each of the other computing units in that communication domain. This generates unidirectional traffic from one computing unit to multiple computing units in the same communication domain in a hybrid optoelectronic network. As another example, when multiple computing units in a communication domain execute an all-reduce operator, one computing unit retrieves data from the other computing units in that communication domain, reduces the data from all computing units in that communication domain, and then sends the reduction result to each of the other computing units in the same communication domain. This generates bidirectional traffic between one computing unit and multiple computing units in a hybrid optoelectronic network.
[0074] Communication data volume refers to the amount of data that needs to be transmitted when different computing units exchange data. Thus, by obtaining the amount of communication data required between different computing units, the adjustment device 203 can calculate the amount of traffic that the optoelectronic hybrid network needs to forward when different computing units in the communication domain exchange data, which means determining the bandwidth requirements of the optoelectronic hybrid network between different computing units in the communication domain.
[0075] In practical applications, the communication information between multiple subtasks acquired by the adjustment device 203 can also be other information that can be used to indicate the communication traffic requirements between different computing units.
[0076] The following describes a non-limiting implementation method for the adjustment device 203 to obtain communication information between multiple subtasks. Specifically, the communication information between the multiple subtasks includes the communication domain, communication operators, and communication data volume.
[0077] In a first possible implementation, multiple computing units in the computing cluster 10 participating in the execution of one or more tasks can each report initial communication information for their respective sub-tasks. This initial communication information may include, for example, the communication domain to which each computing unit belongs, the communication operators to be executed by the computing units within that communication domain, and the amount of communication data exchanged between the computing units. Therefore, the adjustment device 203 can aggregate the initial communication information reported by each computing unit to obtain communication information between multiple sub-tasks included in at least one task running by the computing cluster 10.
[0078] For example, suppose computing cluster 10 is running an AI task, such as... Figure 4 As shown, a computing node can deploy multiple computing units. The computing node itself can be a server (such as an Atlas server), and the computing units within it can be accelerators (such as GPUs). One or more containers can run within the computing node. Taking running a single container as an example, this container can run an acceleration component, such as one named "ascend speed." This acceleration component can collect the communication domain of each computing unit, the communication operators to be executed, and the amount of communication data to be exchanged—that is, it collects the initial communication information of the subtasks executed by the computing unit. For example, each computing unit can report the initial communication information of its executed subtasks to the acceleration component. Then, the acceleration component can provide the initial communication information of the AI tasks corresponding to each computing unit to an agent component within the computing node (such as an agent component named "NetMindAgent"). In practical applications, such as... Figure 4 As shown, the acceleration component can be configured with runtime tools, such as a runtime tool named "NetMind RumTime". Furthermore, the acceleration component can use this runtime tool to provide the initial communication information of the AI tasks collected from each computing unit to the agent component in the computing node.
[0079] Next, the proxy component can aggregate the initial communication information of multiple subtasks collected by the acceleration component and forward the aggregated initial communication information to the adjustment device 203. For example, for the eight computing units in the computing node, the initial communication information of each computing unit collected by the acceleration component includes the identifier of the communication domain, the identifier of the computing unit, the communication operator to be executed (the name of the communication operator), and the amount of communication data to be exchanged, such as... Figure 5 As shown. The initial communication information may also include the communication direction between the computing unit and other computing units in the same communication domain. Then, the proxy component can aggregate the initial communication information reported by each computing unit to obtain aggregated initial communication information. Thus, after receiving the initial communication information provided by the proxy components in each computing node, the adjustment device 203 can further aggregate the received initial communication information to obtain the communication information between the multiple subtasks included in the AI task. For example, as... Figure 4 As shown, the adjustment device 203 may include a service module, which may be, for example, a service process supporting the NetMind platform, and this service module is responsible for performing aggregation operations on the initial communication information. The implementation method of the adjustment device 203 aggregating the initial communication information of multiple subtasks provided by various agent components is related to the aggregation... Figure 5 The implementation of the initial communication information shown is similar, so it will not be described in detail here.
[0080] In practical applications, for AI tasks involving AI models, when the AI model is large, the computing cluster 10 typically deploys the AI model in a distributed manner. For example, it may deploy different network layers of the AI model to different computing nodes (or different computing units within the same computing node) in a pipelined parallel manner. In this case, for the parallel approach of the AI model, the initial communication information reported by each computing unit for executing subtasks may also include parallel domain information, such as the identifier of the parallel domain. Typically, a parallel domain can include multiple communication domains, such as... Figure 5 The communication domains A and B are shown.
[0081] In a second possible implementation, the computing cluster 10 may further include a task scheduling node ( Figure 2 (Not shown in the image), during the process of scheduling a task to multiple computing units in the computing cluster 10, the task scheduling node can parse the task to determine the subtasks to be executed by different computing units participating in the task. It can also locally store the initial communication information of the subtasks executed by each computing unit, such as the communication domain to which each computing unit participating in the AI task belongs, the communication operators executed, and the amount of data required for communication. Therefore, when the adjustment device 203 needs to obtain the communication information between multiple subtasks included in at least one task running in the computing cluster 10, it can access it from the task scheduling node, such as by accessing the task scheduling node's local hard drive.
[0082] It is understood that the above-described method of adjusting device 203 acquiring communication information of multiple subtasks is only an example. In actual application, adjusting device 203 may also acquire communication information between multiple subtasks included in at least one task running in computing cluster 10 based on other methods, and there is no limitation on this.
[0083] S302: Adjustment device 203 acquires the physical topology information of computing cluster 10.
[0084] In this embodiment, the following two non-limiting implementation examples for obtaining the physical topology of computing cluster 10 are provided.
[0085] In the first implementation, the electrical management device 201 in the management node 200 is responsible for monitoring and managing the electrical switches in the optoelectronic hybrid network. Therefore, the electrical management device 201 can query the connection relationships between different electrical switches and the connection relationships between the electrical switches and computing nodes, and thereby generate the physical topology between multiple electrical switches and multiple computing nodes, such as... Figure 4As shown. For example, the electrical management device 201 can query the routing table in each electrical switch, and determine the connection between different electrical switches, the connection between electrical switches and computing nodes, and the physical link used to implement the connection based on the routing table between each electrical switch, and thereby generate the physical topology between electrical switches and between electrical switches and computing nodes.
[0086] Furthermore, the optical management device 202 in management node 200 is responsible for monitoring and managing the optical switches in the optoelectronic hybrid network. Therefore, the optical management device 202 can query the ports used by each optical switch to connect to the electrical switch and whether different ports are connected. Based on the queried ports used to connect to the electrical switch and the connectivity between different ports, the optical management device 202 can generate the physical topology corresponding to the optical switches, such as... Figure 4 As shown, this physical topology can be used to indicate the connection relationship between optical switches and electrical switches.
[0087] Then, the electrical management device 201 and the optical management device 202 respectively provide their generated physical topologies to the adjustment device 203. The adjustment device 203 can then aggregate the physical topologies provided by the electrical management device 201 and the optical management device 202 to obtain the global physical topology information of the computing cluster 10, such as... Figure 4 As shown. For example, this physical topology information can indicate links between different computing nodes. Specifically, these links may include physical links between the computing node and an electrical switch, and may also include physical links between the electrical switch and an optical switch. In practical applications, the physical topology information of the computing cluster 10 can also indicate available physical links, bandwidth, and other information between computing nodes, electrical switches, and optical switches.
[0088] In the second implementation, the electrical management device 201 can maintain the physical topology between multiple electrical switches and multiple computing nodes. For example, during the networking phase, the electrical management device 201 can configure a corresponding routing table for each electrical switch based on the connection relationships between the multiple electrical switches and the connection relationships between the electrical switches and the computing nodes. During this process, the electrical management device 201 can locally store the physical topology between the multiple electrical switches and the multiple computing nodes. Simultaneously, the optical management device 202 can also maintain the physical topology between multiple optical switches and multiple electrical switches. Therefore, when the adjustment device 203 needs to obtain the physical topology of the computing cluster 10, the adjustment device 203 can send query requests for the physical topology to both the electrical management device 201 and the optical management device 202, and obtain the global physical topology information of the computing cluster 10 by summarizing the physical topologies returned by the electrical management device 201 and the optical management device 202 respectively.
[0089] It is understood that the above implementation is only an example. In other embodiments, the adjustment device 203 can also obtain the topology of the optoelectronic hybrid network in other ways, such as saving the topology of the optoelectronic hybrid network locally during the networking process, so that when the adjustment device 203 needs to query the topology, it can obtain the topology from the local query.
[0090] S303: The adjustment device 203 determines the communication traffic requirements of each link in the multiple links between sub-clusters 101 and 102 connected by optical switches in the multiple sub-clusters based on the communication information between multiple sub-tasks and the physical topology information of computing cluster 10.
[0091] In this embodiment, the communication information between multiple subtasks acquired by the adjustment device 203 (specifically, the communication relationship and communication data volume between multiple computing units executing the multiple subtasks) can be used to indicate the traffic characteristics between the multiple computing units executing the multiple subtasks. Thus, the adjustment device 203 can determine the communication traffic requirements between different computing units based on the communication information between the multiple subtasks.
[0092] In one possible implementation, taking the task executed by the computing cluster 10 as an example, specifically an AI task, the adjustment device 203 can determine the traffic matrix corresponding to the AI task based on the communication relationship and the amount of communication data between multiple sub-tasks. The traffic matrix is used to indicate whether different computing units participating in the execution of the AI task need to communicate and the amount of data to be exchanged when communication is required. It can also indicate the bandwidth required by the electrical switch and optical switch when forwarding the traffic between the multiple computing units.
[0093] For example, suppose computing cluster 10 uses N computing units to execute the same AI task (e.g., using N computing units to perform a distributed training process for the same AI model), and the identifiers of these N computing units are 1, 2, ..., N. Then, the traffic matrix generated by adjustment device 203 can specifically be as follows: Figure 6 The diagram shows an N*N two-dimensional matrix, where N is a positive integer, each column corresponds to one of the N computational units, and each row corresponds to one of the N computational units. Furthermore, the elements of the two-dimensional matrix indicate the bandwidth requirements for communication between different computational units. For example, Figure 6The N*N two-dimensional matrix shown has an element in the first row and second column indicating whether communication is required between computing unit 1 and computing unit 2; and an element in the second row and third column indicating that communication is not required between computing unit 2 and computing unit 3, therefore this element has a value of 0. Furthermore, the larger the value of an element in the flow matrix, the greater the bandwidth required for communication between the two computing units; the smaller the value, the smaller the bandwidth required for communication between the two computing units. For example, Figure 6 The N*N two-dimensional matrix shown has an element value of 8 in the first row and second column, which is greater than the element value of 1 in the first row and third column. This indicates that the bandwidth required for communication between computing unit 1 and computing unit 2 is greater than the bandwidth required for communication between computing unit 1 and computing unit 3.
[0094] For example, the adjustment device 203 can determine whether different computing units communicate with each other and the bandwidth required when different computing units communicate with each other based on the communication operator and the amount of communication data in the communication information, and determine the value of the element in the traffic matrix based on the bandwidth.
[0095] Then, the adjustment device 203 can determine the logical topology 1 of the links between different sub-clusters in multiple sub-clusters participating in forwarding the AI task, based on the traffic matrix and the physical topology information of the computing cluster 10, such as... Figure 4 As shown. The determined logical topology 1 can be used to indicate the links through which traffic between computing units in multiple sub-clusters is located (i.e., which links or links forward traffic between different computing units), and can also be used to indicate the communication traffic requirements of each link in the multiple links connected by optical switches between multiple sub-clusters. That is, it is used to indicate the amount of traffic that needs to be forwarded by the electrical switches and the connection links between the optical switches on each link in the multiple links connected by the optical switches between sub-clusters 101 and 102, and the amount of traffic that needs to be forwarded by the connection links between the electrical switches.
[0096] As an example of generating logical topology 1, the adjustment device 203 can determine the link used for data communication between multiple computing units in sub-cluster 101 and sub-cluster 102 based on a heuristic search algorithm or other algorithms, according to the traffic matrix and the physical topology information of computing cluster 10. Typically, different computing units can forward traffic through the same link. For example, in... Figure 2In the computing cluster 10 shown, the four computing units in computing node 1 communicate with the four computing units in computing node 3. Therefore, the traffic generated during communication between computing units 1 to 4 and between computing units 9 to 12 can be forwarded via the link: electrical switch 1 → electrical switch 5 → optical switch 1 → electrical switch 7 → electrical switch 3. Then, for each link (hereinafter referred to as the second link), the adjustment device 203 can aggregate the communication data volume between multiple computing units using the same second link for traffic forwarding, obtaining the communication traffic demand of the second link among the multiple links connecting sub-cluster 101 and sub-cluster 102 via optical switches. For example, assuming that computing unit 1 sends 1TB of traffic to computing unit 9, computing unit 2 sends 1.5TB of traffic to computing unit 10, computing unit 3 sends 1.5TB of traffic to computing unit 11, and computing unit 4 sends 1TB of traffic to computing unit 12, then the communication traffic requirement of the second link is 5TB (i.e., 1TB + 1.5TB + 1.5TB + 1TB). Similarly, the adjustment device 203 can determine the communication traffic requirements on the other links based on the above method. In this way, the adjustment device 203 can generate the above logical topology 1 by planning the links where traffic between different computing units is located and the communication traffic requirements of each link.
[0097] S304: The adjustment device 203 determines the remaining traffic of each link among the multiple links connected between sub-cluster 101 and sub-cluster 102 via the optical switch based on the physical topology information of the computing cluster 10.
[0098] The remaining traffic, based on the current physical topology of computing cluster 10, refers to the available traffic that the links between electrical switches (used to forward traffic from sub-clusters) and optical switches can forward within a unit of time. These links can include one or more physical links; therefore, the available traffic of the links between optical and electrical switches is the sum of the available traffic of all physical links between them. For example, in... Figure 2In the computing cluster 10 shown, assuming that the bandwidth supported by a physical link between electrical switch 5 and optical switch 1 is 1TB (terabits), and there are a total of 10 physical links between optical switch 1 and electrical switch 5, then if none of these 10 physical links are allocated to other tasks in computing cluster 10, the maximum communication bandwidth that electrical switch 5 and optical switch 1 can provide for the tasks executed by sub-cluster 101 and sub-cluster 102 based on the current physical topology is 10TB. That is, the remaining traffic of the link connecting sub-cluster 101 and sub-cluster 102 through optical switch 1 (i.e., the link between electrical switch 5 and optical switch 1) is 10TB.
[0099] As an example, the adjustment device 203 can determine, based on the physical topology information of the computing cluster 10, all (available) physical links connecting sub-cluster 101 and sub-cluster 102 via the same optical switch, and the maximum forwarding traffic that each physical link can support per unit time. Thus, the adjustment device 203 can calculate the sum of the traffic bandwidth of multiple physical links and use the calculated sum as the remaining traffic of each link among the multiple links connecting sub-cluster 101 and sub-cluster 102 via the optical switch. When sub-cluster 101 and sub-cluster 102 are connected via multiple optical switches, the adjustment device 203 can determine the remaining traffic of the links connected via each optical switch in the above manner. Similarly, when the sub-clusters participating in the task in the computing cluster 10 also include other sub-clusters, the adjustment device 203 can determine the remaining traffic of each link among the multiple links connecting different sub-clusters via at least one optical switch in the above manner, which will not be elaborated further.
[0100] In other embodiments, the adjustment device 203 may also determine the remaining traffic of each link among the multiple links connected by the optical switch between sub-cluster 101 and sub-cluster 102 in other ways, without limitation.
[0101] S305: The adjustment device 203 adjusts the connection links between the optical switch and sub-cluster 101 and sub-cluster 102 based on the remaining traffic of each link in the multiple links connected between sub-cluster 101 and sub-cluster 102 via the optical switch, and the communication traffic requirements of each link in the multiple links connected between sub-cluster 101 and sub-cluster 102 via the optical switch.
[0102] In practical applications, the remaining traffic on the multiple links currently connected between sub-cluster 101 and sub-cluster 102 via optical switches may not be sufficient to meet the communication traffic requirements of each link within those links. In this case, the adjustment device 203 can adjust the connection links between sub-cluster 101 and sub-cluster 102 to meet the communication traffic requirements between different computing nodes within sub-cluster 101 and sub-cluster 102.
[0103] In this embodiment, the following implementation examples of adjusting the connection links between the optical switch and sub-cluster 101 and sub-cluster 102 are provided.
[0104] As a first implementation example, when a new task is running in computing cluster 10 (or in other scenarios), the electrical switch and optical switch forward all the tasks running in computing cluster 10 based on the current network topology, which may cause traffic overload on some links between the electrical switch and the optical switch, thereby affecting the overall performance of computing cluster 10 in executing tasks.
[0105] Based on this, during the execution of a new task in computing cluster 10 by the computing units in sub-cluster 101 and sub-cluster 102, the adjustment device 203 can determine the communication traffic requirements of each link among the multiple links connected to sub-cluster 101 and sub-cluster 102 via optical switches, based on the aforementioned method. Then, the adjustment device 203 can compare the communication traffic requirements of each link with the remaining traffic of each link. When the remaining traffic of a first link cannot meet its communication traffic requirements, the adjustment device 203 adjusts the connection links between the optical switch and sub-cluster 101 and sub-cluster 102. Specifically, this can be achieved by increasing the number of physical links used for forwarding task traffic between the optical switch and the electrical switch (e.g., increasing the number of physical links from 200 to 300), thereby increasing the remaining traffic of the first link so that its communication traffic requirements can be met. Similarly, for the second link, third link, etc., between sub-cluster 101 and sub-cluster 102 connected by an optical switch, the adjustment device 203 can adjust the connection link between the optical switch and sub-cluster 101 and sub-cluster 102 in the above manner so that the remaining traffic of multiple links between sub-cluster 101 and sub-cluster 102 can meet the communication traffic requirements of these multiple links.
[0106] For example, such as Figure 7aAs shown, assume that computing cluster 10 also includes sub-cluster 103 and sub-cluster 104, and that optical switch 1 is connected to electrical switch 9 in sub-cluster 103 and electrical switch 10 in sub-cluster 104. Sub-cluster 101 and sub-cluster 102 are connected by link 1 via optical switch 1 and link 2 via optical switch 2; sub-cluster 103 and sub-cluster 104 are connected by link 3 via optical switch 1. Assume that links 1 and 2 are implemented using 4 physical links, and link 3 is implemented using 2 physical links, with each physical link having a remaining capacity of 10TB. Based on the current physical topology of computing cluster 10, the remaining capacity of link 1 (connected by optical switch 1) between sub-cluster 101 and sub-cluster 102 is 40TB, the remaining capacity of link 2 (connected by optical switch 1) is 40TB, and the remaining capacity of link 3 (connected by optical switch 1) between sub-cluster 103 and sub-cluster 104 is 20TB. Taking the adjustment of link 1 as an example, if the communication traffic demand of link 1 between sub-cluster 101 and sub-cluster 102 calculated by the adjustment device 203 is 50T, then, for link 1, the adjustment device 203 can compare the remaining traffic of link 1 with the communication traffic demand. When the remaining traffic is less than the communication traffic demand, the adjustment device 203 can calculate the difference between the communication traffic demand and the remaining traffic of link 1, and based on this difference and the upper limit of bandwidth that a single physical link can provide, determine the number of physical links to be added between optical switch 1 and electrical switch 5 and electrical switch 7. The number of physical links to be added is 1 (i.e., (50-40) / 10). Therefore, the adjustment device 203 can add this number of physical links between optical switch 1 and electrical switch 5 in sub-cluster 101, and between optical switch 1 and electrical switch 7 in sub-cluster 102, so that the remaining traffic of link 1 increases from 40TB to 50TB, thereby ensuring that the remaining traffic of the adjusted link 1 meets the communication traffic demand of link 1.
[0107] In specific implementation, such as Figure 7bAs shown, the adjustment device 203 can dismantle the physical links in link 3 (specifically, it can dismantle one physical link between optical switch 1 and electrical switch 9, and one physical link between optical switch 1 and electrical switch 10), and use the released port resources to create a new physical link between optical switch 1 and electrical switch 5 and electrical switch 7 respectively. In this way, the remaining traffic of link 3 connected to sub-cluster 103 and sub-cluster 104 via optical switch 1 is reduced to 10TB, and the remaining traffic of link 1 connected to sub-cluster 101 and sub-cluster 102 via optical switch 1 can be increased to 50TB. This effectively avoids problems such as traffic overload or traffic congestion on some physical links due to insufficient physical links, thereby improving the overall performance of the computing cluster 10 in executing tasks.
[0108] As a second implementation example, when electrical switches and optical switches forward traffic for tasks running in computing cluster 10 based on the current network topology, it may be difficult to achieve high bandwidth utilization. For example, for multiple communication domains corresponding to AI tasks, the amount of communication data between computing units in different communication domains may vary significantly. This makes the bandwidth requirements high when computing units in some communication domains communicate with each other, while the bandwidth requirements are low when computing units in other communication domains communicate with each other. As a result, during the process of electrical switches and optical switches forwarding traffic for computing units in various communication domains, some links may experience traffic congestion (large amount of communication data between computing units in the communication domain), while other links may be idle for a long time (small amount of communication data between computing units in the communication domain).
[0109] Based on this, the adjustment device 203 can obtain a load balancing strategy, which instructs the relatively even distribution of traffic across different links. For example, the adjustment device 203 can obtain the load balancing strategy from a pre-configured file. Then, based on the load balancing strategy, the adjustment device 203 can adjust the connection links between the optical switch and sub-clusters 101 and 102 according to the remaining traffic on multiple links between sub-clusters 101 and 102, and the communication traffic requirements of multiple links between sub-clusters 101 and 102. In this way, the adjustment device 203 can evenly distribute the traffic that needs to communicate between sub-clusters 101 and 102 across multiple links, thereby effectively avoiding excessive traffic on some links leading to traffic overload or congestion, while simultaneously preventing insufficient traffic on other links leading to idleness. This effectively improves the bandwidth utilization of the links between the optical switch and the electrical switch.
[0110] In practical applications, the adjustment device 203 can also adjust the connection link between sub-cluster 101 and sub-cluster 102 based on other strategies, without limitation.
[0111] The adjustment device 203 can adjust the connection links between different sub-clusters according to changes in the physical topology.
[0112] In specific implementation, the adjustment device 203 can determine the logical topology 1 between multiple sub-clusters in the computing cluster 10 based on the aforementioned heuristic search and other methods. Then, based on this logical topology 1, the adjustment device 203 determines the physical topology 1 between sub-clusters 101 and 102, that is, determines the physical topology 1 between the electrical switches in sub-clusters 101 and 102 and the optical switches, respectively. Figure 4 As shown. In practical applications, some links between sub-cluster 101 and sub-cluster 102 may be forwarding some of the traffic generated by computing cluster 10 while it is running tasks. In this case, if the connection links between sub-cluster 101 and sub-cluster 102 are directly adjusted according to the determined physical topology 1, it may affect the normal transmission of this part of the traffic between sub-cluster 101 and sub-cluster 102, thereby affecting the normal operation of tasks in computing cluster 10. Therefore, the adjustment device 203 can first calculate the order of dismantling and establishing physical links in the optoelectronic hybrid network according to the physical topology 1. For example, the order of dismantling and establishing physical links can be calculated by a heuristic search algorithm. Specifically, the physical links to be dismantled and established are the links between the electrical switch in sub-cluster 101 and the optical switch in sub-cluster 102, respectively. For example, the adjustment device 203 can acquire physical topology 2, which indicates the current links between the electrical switches in sub-cluster 101 and sub-cluster 102 and the optical switches. Thus, the adjustment device 203 can determine the physical links that need to be newly established or dismantled between the electrical switches in sub-cluster 101, sub-cluster 102, and optical switches by comparing physical topology 1 and physical topology 2, and determine the order of establishing and dismantling physical links through heuristic search or other methods. Then, the adjustment device 203 adjusts the physical links between sub-cluster 101 and sub-cluster 102 according to the determined order of dismantling and establishing physical links, thereby achieving the adjustment of physical topology 1 between sub-cluster 101 and sub-cluster 102. Figure 4 As shown. In practical applications, the adjustment device 203 can control the connection status between different ports by adjusting the rotation angle of the lens in the optical switch. Specifically, the optical management device 202 can control the status of the ports between the optical switch and each electrical switch, thereby controlling the removal and creation of physical links between the optical switch and different electrical switches.
[0113] For example, when determining the physical topology 1 based on the logical topology 1, the adjustment device 203 can calculate the physical topology 1 that satisfies the logical topology 1 using algorithms such as heuristic search or mixed integer linear programming (MILP). When the adjustment device 203 determines the physical topology 1 based on the MILP algorithm, it can also determine the constraints on the physical topology 1 to solve for the physical topology 1 that simultaneously satisfies the constraints and the logical topology 1. The constraints can be implemented in several exemplary ways, as follows.
[0114] In the first implementation, the constraint could be, for example, a constraint on the cabling balance between optical switches and electrical switches. Thus, in the calculated physical topology 1, the difference in the number of physical links between the optical switches and each electrical switch involved in forwarding traffic is less than the upper limit indicated by this constraint, thereby achieving a balance in the number of physical links between the optical switches and multiple electrical switches.
[0115] In the second implementation, the constraints could be, for example, robustness constraints for forwarding traffic in a hybrid optoelectronic network. Thus, in the calculated physical topology 1, redundant connections can exist between the optical switches and electrical switches involved in the forwarding task. This means that if some optical switches (or some electrical switches) fail, the traffic for that task can still be forwarded through other optical switches (or other electrical switches), thereby improving the robustness of the traffic forwarding task in the hybrid optoelectronic network.
[0116] In the third implementation, the constraints could be, for example, constraints on the fault propagation of traffic in the optoelectronic hybrid network forwarding task. Thus, in the calculated physical topology 1, multiple electrical switches participating in the forwarding task can be connected to multiple optical switches. This allows other optical switches to continue forwarding the traffic even if some optical switches fail, thereby preventing the failure of one optical switch from causing the optoelectronic hybrid network to fail in forwarding the traffic and reducing the fault propagation of traffic in the optoelectronic hybrid network forwarding task.
[0117] In the fourth implementation, the constraint condition could be, for example, the difference between the solved physical topology 1 and the currently used physical topology 2 between sub-clusters 101 and 102. Thus, after calculating physical topology 1, the cost of link breaking and link building in adjusting the physical links between sub-clusters 101 and 102 based on physical topology 1 is minimized, thereby reducing the cost of adjusting the physical topology.
[0118] In the fifth implementation, the constraint could be, for example, the efficiency or accuracy of solving the physical topology 1. Generally, the longer the adjustment device 203 takes to solve the physical topology 1 using heuristic search or MILP algorithms, the higher the accuracy of the solved physical topology 1 as the optimal solution; however, the lower the efficiency of solving the physical topology 1. Therefore, by constraining the efficiency or accuracy of solving the physical topology 1, it is possible to ensure that the adjustment device 203 achieves a high level of efficiency or accuracy in determining the physical topology 1.
[0119] In practical applications, the adjustment device 203 determines the constraints used in the physical topology 1. These constraints can be other conditions or combinations of the above-mentioned constraints, and there are no limitations on this.
[0120] For example, when adjusting the physical links between sub-cluster 101 and sub-cluster 102 according to the determined order of dismantling and rebuilding physical links, the adjustment device 203 can reconstruct the physical topology in a lossless manner, such as... Figure 4 As shown.
[0121] As a first implementation example, during the adjustment of the physical link between sub-cluster 101 and sub-cluster 102, some physical links between optical switches and electrical switches may need to forward traffic. Therefore, the adjustment device 203 can instruct the electrical management device 201 to redirect the traffic on these physical links between optical switches and electrical switches. Taking the redirection of traffic on the physical link between optical switch A and an electrical switch as an example, the electrical management device 201 can generate an access control list (ACL) for optical switch A and distribute the ACL to the corresponding electrical switch. The electrical switch can then forward traffic to other optical switches based on the ACL, thereby redirecting the traffic on the physical link between the electrical switch and optical switch A to the physical link between the electrical switch and other optical switches. After the electrical management device 201 completes the ACL configuration, the adjustment device 203 can instruct the optical management device 202 to dismantle the physical link between the electrical switch and optical switch A and establish a physical link between the electrical switch and other optical switches. Thus, the adjustment device 203 can reconstruct the physical topology in the optoelectronic hybrid network without the AI task traffic being noticed.
[0122] As a second implementation example, during the adjustment of the physical link between sub-cluster 101 and sub-cluster 102, the physical link between optical switch B (i.e., some optical switches in the optoelectronic hybrid network) and electrical switch may be forwarding traffic. Therefore, the adjustment device 203 can instruct the electrical management device 201 to drain the traffic on this physical link. For example, when the electrical switch communicates with optical switch B using Border Gateway Protocol (BGP), the adjustment device 203 can send a traffic drain command to the electrical management device 201. This traffic drain command can include indication information of the physical link between the electrical switch and optical switch B. Thus, the electrical management device 201 can drain the traffic on the physical link according to the traffic drain command and then delete the physical link between the electrical switch and optical switch B. Specifically, it can delete the peer relationship (also called BGP-Peer relationship) between the electrical switch and optical switch B based on the BGP protocol. Then, the electrical management device 201 can notify the adjustment device 203 that the traffic drain is complete. The adjustment device 203 can notify the optical management device 202 of port changes in the optical switch. The optical management device 202 can then adjust the ports on optical switch B and other optical switches based on this information, thereby dismantling the physical link between the electrical switch and optical switch B and establishing a physical link between the electrical switch and other optical switches. Furthermore, after completing the physical link adjustment, the optical management device 202 can notify the adjustment device 203 that the adjustment is complete. The adjustment device 203 can then notify the electrical management device 201 of the newly added physical link, and the electrical management device 201 will add a BGP-Peer relationship for the electrical switch. Thus, during the physical topology adjustment process, the adjustment device 203 does not affect the smooth forwarding of task traffic in the hybrid optical-electrical network, enabling the reconstruction of the physical topology between sub-cluster 101 and sub-cluster 102 without affecting task traffic flow.
[0123] It is understood that the above-described method of reconstructing physical topology based on lossless flow is only an example. In actual applications, the adjustment device 203 may also use other methods to reconstruct physical topology without loss of flow, and there is no limitation on this.
[0124] It is worth noting that the above description is based on adjusting the connection links between two sub-clusters. In actual applications, when there are three or more (including three) sub-clusters in computing cluster 10 participating in the execution of tasks, the adjustment device 203 can adjust the connection links between sub-clusters 101 / 102 and other sub-clusters, as well as the connection links between other sub-clusters, based on the above method. This will not be elaborated further.
[0125] In this embodiment, the adjustment device 203 can not only adjust the physical topology between multiple sub-clusters, but also adjust the logical topology for the electrical switches in each sub-cluster. For example, in Figure 2 In the computing cluster 10 shown, electrical switch 1 and its associated traffic may both forward to electrical switch 5, which in turn forwards the traffic to optical switch 1. During this process, the link between electrical switch 5 and optical switch 1 experiences high traffic load, which can easily lead to traffic congestion on that link. Conversely, the link between electrical switch 6 and optical switch 1 experiences low traffic load (or even no load), resulting in a prolonged period of idleness. Therefore, the adjustment device 203 can also adjust the logical topology between electrical switches in each sub-cluster. That is, the physical links between different electrical switches in each sub-cluster remain unchanged; however, the forwarding paths used by the electrical switches to forward task traffic change. For example, the electrical switches used to forward task traffic are adjusted from electrical switches 1, 2, and 3 to electrical switches 1, 2, 3, 4, 5, and 6, etc.
[0126] In specific implementation, taking the adjustment of the logical topology of the electrical switches in sub-cluster 101 as an example, the adjustment device 203 can determine the logical topology 2 of the links between multiple electrical switches participating in forwarding task traffic in sub-cluster 101 based on the traffic matrix and the physical topology information of cluster 10, and obtain the logical topology 3 currently used by the electrical switches in sub-cluster 101, such as by determining the logical topology 3 based on the routing tables of each electrical switch in sub-cluster 101. Then, the adjustment device 203 can compare the topology differences between logical topology 2 and logical topology 3, and generate corresponding path configuration information based on the topology differences. This path configuration information is used to configure the forwarding paths of the electrical switches in sub-cluster 101. Then, the adjustment device 203 can send the path configuration information to the electrical management device 201.
[0127] Accordingly, the power management device 201 adjusts the forwarding path used by the sub-cluster 101 when forwarding task traffic based on the path configuration information, so as to update the logical topology of multiple power switches in the sub-cluster 101 to logical topology 2. For example, the power management device 201 can modify the routing table in the power switch based on the path configuration information, so that when the power switch forwards data based on the modified routing table, it can adjust the forwarding path used by the power switch when forwarding traffic.
[0128] Alternatively, the electrical management device 201 can directly reconfigure the routing tables in each electrical switch in the sub-cluster 101 according to the logical topology 2, so as to update the logical topology of multiple electrical switches in the sub-cluster 101 to logical topology 2.
[0129] In practical applications, the adjustment device 203 can adjust only the physical topology between multiple sub-clusters (i.e., the connection links between multiple sub-clusters); or, the adjustment device 203 can simultaneously adjust the connection links between multiple sub-clusters and the logical topology between multiple electrical switches within each sub-cluster. In this case, the adjustment device 203 can not only adjust the status of the ports connecting the optical switches to the electrical switches through the optical management device 202 (to achieve the adjustment of the physical topology), but also adjust the routing tables in some electrical switches through the electrical management device 201 (to achieve the adjustment of the logical topology).
[0130] Furthermore, this embodiment may also include the following steps.
[0131] S306: After the network topology adjustment for computing cluster 10 is completed, multiple electrical switches and multiple optical switches in computing cluster 10 forward traffic generated when computing cluster 10 is running at least one task based on the adjusted network topology.
[0132] Because the adjusted network topology is better suited to the communication traffic requirements generated by the tasks running in computing cluster 10, it can effectively improve the utilization rate of network bandwidth in the optoelectronic hybrid network, improve the overall performance of the optoelectronic hybrid network, and also improve the efficiency of the tasks running in computing cluster 10.
[0133] It is worth noting that this embodiment uses the example of adjusting the network topology adopted by the adjustment device 203 within the global topology of the computing cluster 10 to adjust the forwarding traffic for one or more tasks. In practical applications, multiple tasks may already be running in the computing cluster 10, and the adjustment device 203 may have already optimized the network topology of the computing cluster 10 for these multiple tasks. For example, the adjustment device 203 can do so by... Figure 4The agent component shown is aware of the tasks being run by the computing unit. At this time, some optical switches and some electrical switches in the computing cluster 10 have been assigned to the multiple tasks so that the traffic generated during the execution of the multiple tasks by the computing cluster 10 can be forwarded using these optical switches and electrical switches. When the computing cluster 10 runs a new task, the current network topology in the computing cluster 10 may be insufficient to meet the communication traffic requirements of the multiple sub-clusters in the computing cluster 10 when executing the new task. At this time, the adjustment device 203 can determine the remaining physical links in the computing cluster 10 that are not used to forward traffic for other tasks. Thus, the adjustment device 203 can determine the physical topology corresponding to the physical link and optimize the physical topology between different sub-clusters for the new task based on the physical topology, or optimize the physical topology between different sub-clusters and the logical topology within each sub-cluster. However, the adjustment device 203 does not adjust the topology between the multiple electrical switches and multiple optical switches used in the computing cluster to forward traffic for other tasks.
[0134] Furthermore, when multiple tasks are running in the computing cluster 10, the adjustment device 203 can acquire communication information between multiple subtasks included in each of the multiple tasks when acquiring communication information. For example, each computing unit in the computing cluster 10 can report the initial communication information of the subtask executed by that computing unit. Each computing unit can participate in the execution of only one AI task, so each computing unit can report the initial communication information of only one subtask. Accordingly, the adjustment device 203 can aggregate the initial communication information reported by each computing unit to obtain the communication information between multiple subtasks included in each task. For example, the initial communication information reported by each computing unit may include the identifier of the task to which the subtask executed by that computing unit belongs. Thus, the adjustment device 203 can aggregate the initial communication information with the same task identifier to obtain the communication information between multiple subtasks included in each task. Or, for example, the initial communication information reported by each computing unit may include the number of computing units participating in the execution of the same task. Thus, the adjustment device 203 can cluster the initial communication information with the same number of computing units to obtain the communication information between multiple subtasks included in each task. Then, the adjustment device 203 adjusts the network topology used by the computing cluster 10 to forward traffic for each task based on the communication information between the multiple subtasks included in each task. For tasks whose network topology optimization has already been completed, the adjustment device 203 does not need to optimize the network topology of the computing cluster 10 again. During this process, the adjustment device 203 can determine a task-level traffic matrix based on the communication information between the multiple subtasks included in each task; different tasks can correspond to different traffic matrices. Therefore, the adjustment device 203 can optimize the network topology used by the computing cluster 10 to forward traffic for a task based on the task-level traffic matrix, thereby achieving task-level optimization of the computing cluster 10's network topology.
[0135] In real-world applications, tasks from different tenants may require resource isolation. Specifically, this means that the optical and electrical switches used to forward traffic generated by tasks from different tenants should not overlap. Alternatively, different tasks from the same tenant may also require the use of isolated resources within computing cluster 10 for traffic forwarding.
[0136] Therefore, in a further possible implementation, the adjustment device 203 can also acquire resource indication information, which is used to indicate the resources corresponding to each task in at least one task run by the computing cluster 10. These resources are the resources used in the computing cluster 10 for forwarding traffic. The resources corresponding to each task may include electrical switches and optical switches that the computing cluster 10 can use when running the task, and the resources corresponding to different tasks do not overlap; that is, there is no overlap between the optical switches and electrical switches used to forward traffic from different tasks.
[0137] Then, for the task whose network topology needs to be optimized (such as a new task running in computing cluster 10), when adjusting the network topology of computing cluster 10 for the task, the adjustment device 203 may specifically adjust the network topology of the resources in computing cluster 10 indicated by the resource indication information based on the acquired resource indication information, the communication information of the task (i.e., the communication information between multiple subtasks included in the task), and the physical topology information of computing cluster 10. For example, the adjustment device 203 may first determine the physical topology information corresponding to the resource indicated by the resource indication information based on the resource indication information and the physical topology information of computing cluster 10; then, the adjustment device 203 may determine the communication traffic requirements of multiple links between multiple sub-clusters participating in the execution of the task based on the communication information of the task and the physical topology information corresponding to the resource; and based on the communication traffic requirements, optimize the network topology structure of the resource indicated by the resource indication information, including optimizing the physical topology between multiple sub-clusters participating in the execution of the task, and may also include optimizing the logical topology in each sub-cluster participating in the execution of the task. In this way, for different tasks, the adjustment device 203 can optimize the network topology of the computing cluster 10 for each task within different resource ranges, thereby meeting the resource isolation requirements for different tasks in actual application scenarios and improving the overall bandwidth utilization of electrical switches and optical switches in the computing cluster 10.
[0138] It is worth noting that other reasonable combinations of steps that can be conceived by those skilled in the art based on the above description also fall within the scope of protection of this application. Secondly, those skilled in the art should also be aware that the embodiments described in the specification are preferred embodiments, and the actions involved are not necessarily essential to this application.
[0139] The above combination Figures 1 to 7b The method for adjusting the network topology of a computing cluster provided in the embodiments of this application will be introduced. Next, the structure of the apparatus and computing device for adjusting the network topology of a computing cluster provided in the embodiments of this application will be described with reference to the accompanying drawings.
[0140] See Figure 8The diagram illustrates a device for adjusting the network topology of a computing cluster. The computing cluster comprises multiple sub-computing clusters, with computing nodes in each sub-cluster connected via electrical switches. The sub-clusters are interconnected via optical switches. Each computing cluster runs at least one task, such as an AI task.
[0141] like Figure 8 As shown, the apparatus 800 for adjusting the network topology of the computing cluster includes:
[0142] The acquisition module 801 is used to acquire communication information between multiple subtasks included in at least one task; and to acquire physical topology information of the computing cluster.
[0143] The determination module 802 is used to determine the communication traffic requirements of each link in the multiple links connecting the first sub-cluster and the second sub-cluster via optical switches in multiple sub-clusters; and to determine the remaining traffic of each link in the multiple links based on the physical topology information of the computing cluster.
[0144] The adjustment module 803 is used to adjust the connection links between the optical switch and the first sub-cluster and the second sub-cluster based on the remaining traffic of each link in the multiple links and the communication traffic requirements of each link in the multiple links.
[0145] In one possible implementation, when adjusting the connection link between the optical switch and the first sub-cluster and the second sub-cluster, the adjustment module 803 is specifically used for:
[0146] When the remaining traffic of the first link in a multi-link network cannot meet the communication traffic requirements of the first link, adjust the connection link between the optical switch and the first sub-cluster and the second sub-cluster.
[0147] In one possible implementation, when adjusting the connection link between the optical switch and the first sub-cluster and the second sub-cluster, the adjustment module 803 is specifically used for:
[0148] Based on the remaining traffic of each link in the multiple links and the communication traffic requirements of the multiple links, the connection links between the optical switch and the first sub-cluster and the second sub-cluster are adjusted based on the load balancing strategy.
[0149] In one possible implementation, the communication information between multiple subtasks includes the communication relationship and communication data volume between multiple computing units executing multiple subtasks. The multiple computing units are distributed in computing nodes in the first sub-cluster and the second sub-cluster. The physical topology information of the computing clusters includes the links between each computing node in the first sub-cluster and the second sub-cluster.
[0150] When determining the communication traffic requirements of each link in a multi-link network, module 802 is specifically used for:
[0151] Based on the communication relationships between multiple computing units executing multiple subtasks, the amount of communication data, and the physical topology information of the computing cluster, the communication traffic requirements of each link in the multiple links are determined.
[0152] In one possible implementation, when determining the communication traffic requirements of each link in multiple links based on the communication relationships, communication data volume, and physical topology information of the computing cluster among multiple computing units executing multiple subtasks, the determining module 802 is specifically used for:
[0153] Based on the communication relationships and communication data volume between multiple computing units executing multiple subtasks, a flow matrix is determined. The flow matrix is used to indicate the communication relationships and communication data volume between multiple computing units.
[0154] Based on the traffic matrix and the physical topology information of the computing cluster, the communication traffic requirements of each link in the multiple links are determined.
[0155] In one possible implementation, when determining the communication traffic requirements of each link in multiple links based on the traffic matrix and the physical topology information of the computing cluster, the determining module 802 is specifically used for:
[0156] Based on the traffic matrix and the physical topology information of the computing cluster, the logical topology of the links between different sub-clusters in multiple sub-clusters is determined. The logical topology is used to indicate the communication traffic requirements of each link in the multiple links connected by optical switches between multiple sub-clusters.
[0157] In one possible implementation, when determining the logical topology of links between different sub-clusters in multiple sub-clusters based on the traffic matrix and the physical topology information of the computing cluster, the determining module 802 is specifically used for:
[0158] Based on the traffic matrix and the physical topology information of the computing cluster, determine the links used for data communication between multiple computing units;
[0159] The communication data volume between multiple computing units using the second link is aggregated to obtain the communication traffic requirements of the second link among multiple links connected by optical switches between multiple sub-clusters.
[0160] In one possible implementation, when adjusting the connection link between the optical switch and the first sub-cluster and the second sub-cluster, the adjustment module 803 is specifically used for:
[0161] Based on the remaining traffic and logical topology of each link in the multiple links, determine the physical topology between the optical switch and the first and second sub-clusters;
[0162] Based on the physical topology, calculate the order of dismantling and rebuilding physical links. The physical link is the link between the sub-cluster and the optical switch.
[0163] Adjust the connection links between the optical switch and the first and second sub-clusters in sequence.
[0164] In one possible implementation, the adjustment module 803 is further configured to:
[0165] Based on the communication traffic requirements of each link in the multiple links between the first sub-cluster and the second sub-cluster, adjust the logical topology between different switches in the first sub-cluster, or adjust the logical topology between different switches in the second sub-cluster.
[0166] In one possible implementation, when acquiring communication information between multiple subtasks included in at least one task, the acquisition module 801 is specifically used for:
[0167] Obtain the initial communication information reported by the computing units in the computing nodes of the first sub-cluster and the second sub-cluster for each of the multiple subtasks included in at least one task;
[0168] The initial communication information reported by the computing units in the computing nodes of the first sub-cluster and the second sub-cluster is aggregated to obtain the communication information between multiple subtasks included in at least one task.
[0169] In one possible implementation, the acquisition module 801 is further configured to acquire resource indication information, which is used to indicate the resources corresponding to each task in at least one task. The resources corresponding to each task include electrical switches and optical switches that can be used when the computing cluster runs the task. The resources corresponding to different tasks in at least one task do not overlap.
[0170] When determining the communication traffic requirements of each link in the multiple links connecting the first and second sub-clusters via optical switches in multiple sub-clusters, module 802 is specifically used for:
[0171] Based on resource indication information, communication information between multiple subtasks, and physical topology information of the computing cluster, the communication traffic requirements of each link in the multiple links connected by optical switches between the first and second sub-clusters in the multiple sub-clusters are determined.
[0172] In one possible implementation, the communication relationship between multiple computing units includes the communication domains to which the multiple computing units belong and the communication operators executed by the multiple computing units.
[0173] In one possible implementation, the tasks running in the computing cluster are artificial intelligence (AI) tasks or high-performance computing (HPC) tasks.
[0174] because Figure 8 The apparatus 800 shown is for adjusting the network topology of a computing cluster, corresponding to the above. Figure 3 The adjustment device 203 in the illustrated embodiment, therefore Figure 8 For a detailed description of the implementation and technical effects of the device 800 for adjusting the network topology of the computing cluster, please refer to the above. Figure 3 The relevant descriptions in the illustrated embodiments are not repeated here.
[0175] Figure 9 This application provides a schematic diagram of the hardware structure of a computing device 900, which, for example, can implement the above-described... Figure 3 Adjustment device 203, etc., in the illustrated embodiment.
[0176] like Figure 9 As shown, the computing device 900 includes a processor 901, a memory 902, and a communication interface 903. The processor 901, memory 902, and communication interface 903 communicate via a bus 904, or via other means such as wireless transmission. The memory 902 stores instructions, and the processor 901 executes the instructions stored in the memory 902. Further, the computing device 900 may also include a memory unit 905, which is connected to the processor 901, the storage medium 902, and the communication interface 903 via the bus 904. The memory 902 stores program code, and the processor 901 can read the program code stored in the memory 902 into the memory unit 905 and perform the following operations based on the read program code:
[0177] Acquire communication information between multiple subtasks included in at least one task, wherein at least one task runs on a computing cluster, the computing cluster includes multiple sub-computing clusters, the computing nodes in each sub-computing cluster are connected via electrical switches, and the multiple sub-computing clusters are connected via optical switches.
[0178] Obtain the physical topology information of the computing cluster;
[0179] Based on the communication information between multiple subtasks included in at least one task and the physical topology information of the computing cluster, determine the communication traffic requirements of multiple links between the first sub-cluster and the second sub-cluster in the multiple sub-clusters.
[0180] Obtain the physical topology information of the computing cluster, and based on the physical topology information of the computing cluster, determine the communication traffic requirements of each link in the multiple links between the first sub-cluster and the second sub-cluster connected by optical switches in multiple sub-clusters.
[0181] Based on the physical topology information of the computing cluster, determine the remaining traffic of each link in the multiple links;
[0182] Based on the remaining traffic of each link in the multiple links and the communication traffic requirements of each link in the multiple links, adjust the connection links between the optical switch and the first sub-cluster and the second sub-cluster.
[0183] It should be understood that in this embodiment, the processor 901 can be a CPU, but it can also be other general-purpose processors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete device assemblies, etc. A general-purpose processor can be a microprocessor or any conventional processor.
[0184] The memory 902 may include read-only memory and random access memory, and provides instructions and data to the processor 901. The memory 902 may also include non-volatile random access memory.
[0185] The memory 902 can be volatile memory or non-volatile memory, or it can include both. The non-volatile memory can be read-only memory (ROM), programmable read-only memory (PROM), erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), or flash memory. The volatile memory can be random access memory (RAM), which is used as an external cache. By way of example, but not limitation, many forms of RAM are available, such as static random access memory (SRAM), dynamic random access memory (DRAM), synchronous dynamic random access memory (SDRAM), double data rate synchronous dynamic random access memory (DDR SDRAM), enhanced synchronous dynamic random access memory (ESDRAM), synchronous link dynamic random access memory (SLDRAM), and direct rambus RAM (DR RAM).
[0186] The communication interface 903 is used to communicate with other devices connected to the computing device 900. The bus 904 may include a data bus, a power bus, a control bus, and a status signal bus, etc. However, for clarity, all buses are labeled as bus 904 in the figure.
[0187] It should be understood that the computing device 900 according to the embodiments of this application may correspond to the adjustment device 203 in the embodiments of this application, and may correspond to the execution of the device according to the embodiments of this application. Figure 3 The method performed by the adjustment device 203 in the illustrated method, and the above-mentioned and other operations and / or functions implemented by the computing device 900, are respectively for the purpose of realizing Figure 3 The process of the corresponding methods in [the document] will not be elaborated here for the sake of brevity.
[0188] This application also provides a computer-readable storage medium. The computer-readable storage medium can be any available medium that a computing device can store, or a data storage device such as a data center containing one or more available media. The available medium can be a magnetic medium (e.g., floppy disk, hard disk, magnetic tape), an optical medium (e.g., DVD), or a semiconductor medium (e.g., solid-state drive). The computer-readable storage medium includes instructions that instruct the computing device to perform the above-described method for adjusting the network topology of a computing cluster.
[0189] This application also provides a computer program product. The computer program product includes one or more computer instructions. When the computer instructions are loaded and executed on a computing device, all or part of the processes or functions described in this application are generated.
[0190] The computer instructions may be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another. For example, the computer instructions may be transmitted from one website, computer, or data center to another website, computer, or data center via wired (e.g., coaxial cable, fiber optic, digital subscriber line (DSL)) or wireless (e.g., infrared, wireless, microwave, etc.) means.
[0191] The computer program product can be a software installation package. In cases where any of the aforementioned methods for adjusting the network topology of the computing cluster is required, the computer program product can be downloaded and executed on the computing device.
[0192] The above embodiments can be implemented, in whole or in part, by software, hardware, firmware, or any other combination thereof. When implemented using software, the above embodiments can be implemented, in whole or in part, as a computer program product. The computer program product includes one or more computer instructions. When the computer program instructions are loaded or executed on a computer, all or part of the processes or functions described in the embodiments of this application are generated. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device. The computer instructions can be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another. For example, the computer instructions can be transmitted from one website, computer, server, or data center to another website, computer, server, or data center via wired (e.g., coaxial cable, fiber optic, digital subscriber line (DSL)) or wireless (e.g., infrared, wireless, microwave, etc.) means. The computer-readable storage medium can be any available medium that a computer can access or a data storage device such as a server or data center that includes one or more sets of available media. The available medium can be a magnetic medium (e.g., floppy disk, hard disk, magnetic tape), an optical medium (e.g., DVD), or a semiconductor medium. A semiconductor medium can be a solid-state drive.
[0193] The terminology used in the above embodiments is for the purpose of describing specific embodiments only and is not intended to be limiting of this application. As used in the specification and appended claims of this application, the singular expressions “a,” “an,” “the,” “the,” “the,” and “this” are intended to also include expressions such as “one or more,” unless the context clearly indicates otherwise. It should also be understood that in the embodiments of this application, “one or more” refers to one, two, or more; the character “ / ” generally indicates that the preceding and following objects are in an “or” relationship. In the embodiments of this application, “simultaneously” means within the same time period, including situations where they are at the same moment.
[0194] References to "one embodiment" or "some embodiments" as described in this specification mean that one or more embodiments of this application include a specific feature, structure, or characteristic described in connection with that embodiment. Therefore, the phrases "in one embodiment," "in some embodiments," "in other embodiments," "in still other embodiments," etc., appearing in different parts of this specification do not necessarily refer to the same embodiment, but rather mean "one or more, but not all, embodiments," unless otherwise specifically emphasized. The terms "comprising," "including," "having," and variations thereof mean "including but not limited to," unless otherwise specifically emphasized.
[0195] The above description is merely a specific embodiment of this application, but the scope of protection of this application is not limited thereto. Any person skilled in the art can easily conceive of various equivalent modifications or substitutions within the technical scope disclosed in this application, and these modifications or substitutions should all be covered within the scope of protection of this application. Therefore, the scope of protection of this application should be determined by the scope of the claims.
Claims
1. A method for adjusting the network topology of a computing cluster, characterized in that, The computing cluster includes multiple sub-computing clusters, with computing nodes in each sub-computing cluster connected via electrical switches, and the multiple sub-computing clusters connected to each other via optical switches. At least one task runs within each computing cluster. The method includes: Obtain communication information between multiple subtasks included in the at least one task; Obtain the physical topology information of the computing cluster; Based on the communication information between the multiple subtasks and the physical topology information of the computing cluster, the communication traffic requirements of each link in the multiple links between the first sub-cluster and the second sub-cluster connected by the optical switch are determined. Based on the physical topology information of the computing cluster, determine the remaining traffic of each of the multiple links; Based on the remaining traffic of each link in the multiple links and the communication traffic requirements of each link in the multiple links, the connection links between the optical switch and the first sub-cluster and the second sub-cluster are adjusted.
2. The method according to claim 1, characterized in that, The step of adjusting the connection links between the optical switch and the first sub-cluster and the second sub-cluster based on the remaining traffic of each link in the multiple links and the communication traffic requirements of each link in the multiple links includes: When the remaining traffic of the first link in the multiple links cannot meet the communication traffic requirements of the first link, the connection links between the optical switch and the first sub-cluster and the second sub-cluster are adjusted.
3. The method according to claim 1, characterized in that, The step of adjusting the connection links between the optical switch and the first sub-cluster and the second sub-cluster based on the remaining traffic of each of the multiple links and the communication traffic requirements of the multiple links includes: Based on the remaining traffic of each of the multiple links and the communication traffic requirements of the multiple links, the connection links between the optical switch and the first sub-cluster and the second sub-cluster are adjusted according to the load balancing strategy.
4. The method according to any one of claims 1 to 3, characterized in that, The communication information between the multiple subtasks includes the communication relationship and communication data volume between the multiple computing units executing the multiple subtasks. The multiple computing units are distributed in the computing nodes in the first sub-cluster and the second sub-cluster. The physical topology information of the computing cluster includes the links between each computing node in the first sub-cluster and the second sub-cluster. The step of determining the communication traffic requirements of each link in the multiple links based on the communication information between the multiple subtasks and the physical topology information of the computing cluster includes: Based on the communication relationships between the multiple computing units executing the multiple subtasks, the amount of communication data, and the physical topology information of the computing cluster, the communication traffic requirements of each link in the multiple links are determined.
5. The method according to claim 4, characterized in that, The step of determining the communication traffic requirements of each link in the multiple links based on the communication relationships between the multiple computing units executing the multiple sub-tasks, the amount of communication data, and the physical topology information of the computing cluster includes: A flow matrix is determined based on the communication relationships and communication data volume between the multiple computing units executing the multiple subtasks. The flow matrix is used to indicate the communication relationships and communication data volume between the multiple computing units. Based on the traffic matrix and the physical topology information of the computing cluster, the communication traffic requirements of each link in the multiple links are determined.
6. The method according to claim 5, characterized in that, The step of determining the communication traffic requirements of each link in the multiple links based on the traffic matrix and the physical topology information of the computing cluster includes: Based on the traffic matrix and the physical topology information of the computing cluster, the logical topology of the links between different sub-clusters in the plurality of sub-clusters is determined. The logical topology is used to indicate the communication traffic requirements of each link in the plurality of links connected by the optical switch between the plurality of sub-clusters.
7. The method according to claim 6, characterized in that, Determining the logical topology of links between different sub-clusters in the plurality of sub-clusters based on the traffic matrix and the physical topology information of the computing cluster includes: Based on the traffic matrix and the physical topology information of the computing cluster, the link used for data communication between the multiple computing units is determined; The communication data volume between multiple computing units using the second link is aggregated to obtain the communication traffic demand of the second link among the multiple links connected by the optical switch between the multiple sub-clusters.
8. The method according to any one of claims 1 to 7, characterized in that, The method further includes: Based on the communication traffic requirements of each link in the multiple links between the first sub-cluster and the second sub-cluster, adjust the logical topology between different switches in the first sub-cluster, or adjust the logical topology between different switches in the second sub-cluster.
9. The method according to any one of claims 1 to 8, characterized in that, The tasks running in the computing cluster are either artificial intelligence (AI) tasks or high-performance computing (HPC) tasks.
10. An apparatus for adjusting the network topology of a computing cluster, characterized in that, The computing cluster includes multiple sub-computing clusters. The computing nodes in each sub-computing cluster are connected via electrical switches. The multiple sub-computing clusters are connected to each other via optical switches. At least one task is running in the computing cluster. The device includes: The acquisition module is used to acquire communication information between multiple subtasks included in the at least one task; and to acquire physical topology information of the computing cluster. The determination module is used to determine the communication traffic requirements of each link in the multiple links between the first sub-cluster and the second sub-cluster connected by the optical switch, based on the communication information between the multiple subtasks and the physical topology information of the computing cluster; and to determine the remaining traffic of each link in the multiple links based on the physical topology information of the computing cluster. The adjustment module is used to adjust the connection links between the optical switch and the first sub-cluster and the second sub-cluster based on the remaining traffic of each link in the multiple links and the communication traffic requirements of each link in the multiple links.
11. The apparatus according to claim 10, characterized in that, When adjusting the connection links between the optical switch and the first sub-cluster and the second sub-cluster, the adjustment module is specifically used for: When the remaining traffic of the first link in the multiple links cannot meet the communication traffic requirements of the first link, the connection links between the optical switch and the first sub-cluster and the second sub-cluster are adjusted.
12. The apparatus according to claim 10, characterized in that, When adjusting the connection links between the optical switch and the first sub-cluster and the second sub-cluster, the adjustment module is specifically used for: Based on the remaining traffic of each of the multiple links and the communication traffic requirements of the multiple links, the connection links between the optical switch and the first sub-cluster and the second sub-cluster are adjusted according to the load balancing strategy.
13. The apparatus according to any one of claims 10 to 12, characterized in that, The communication information between the multiple subtasks includes the communication relationship and communication data volume between the multiple computing units executing the multiple subtasks. The multiple computing units are distributed in the computing nodes in the first sub-cluster and the second sub-cluster. The physical topology information of the computing cluster includes the links between each computing node in the first sub-cluster and the second sub-cluster. When determining the communication traffic requirements of each of the multiple links, the determining module is specifically used for: Based on the communication relationships between the multiple computing units executing the multiple subtasks, the amount of communication data, and the physical topology information of the computing cluster, the communication traffic requirements of each link in the multiple links are determined.
14. The apparatus according to claim 13, characterized in that, When determining the communication traffic requirements of each link in the multiple links based on the communication relationships, communication data volume, and physical topology information of the computing cluster among the multiple computing units executing the multiple sub-tasks, the determining module is specifically used for: A flow matrix is determined based on the communication relationships and communication data volume between the multiple computing units executing the multiple subtasks. The flow matrix is used to indicate the communication relationships and communication data volume between the multiple computing units. Based on the traffic matrix and the physical topology information of the computing cluster, the communication traffic requirements of each link in the multiple links are determined.
15. The apparatus according to claim 14, characterized in that, When determining the communication traffic requirements of each link in the multiple links based on the traffic matrix and the physical topology information of the computing cluster, the determining module is specifically used for: Based on the traffic matrix and the physical topology information of the computing cluster, the logical topology of the links between different sub-clusters in the plurality of sub-clusters is determined. The logical topology is used to indicate the communication traffic requirements of each link in the plurality of links connected by the optical switch between the plurality of sub-clusters.
16. The apparatus according to claim 15, characterized in that, When determining the logical topology of links between different sub-clusters in the plurality of sub-clusters based on the traffic matrix and the physical topology information of the computing cluster, the determining module is specifically used for: Based on the traffic matrix and the physical topology information of the computing cluster, the link used for data communication between the multiple computing units is determined; The communication data volume between multiple computing units using the second link is aggregated to obtain the communication traffic demand of the second link among the multiple links connected by the optical switch between the multiple sub-clusters.
17. The apparatus according to any one of claims 10 to 16, characterized in that, The adjustment module is also used for: Based on the communication traffic requirements of each link in the multiple links between the first sub-cluster and the second sub-cluster, adjust the logical topology between different switches in the first sub-cluster, or adjust the logical topology between different switches in the second sub-cluster.
18. The apparatus according to any one of claims 10 to 17, characterized in that, The tasks running in the computing cluster are either artificial intelligence (AI) tasks or high-performance computing (HPC) tasks.
19. A computer-readable storage medium, characterized in that, Includes instructions that, when executed on at least one computing device, cause the at least one computing device to perform the steps of the method as described in any one of claims 1 to 9.
20. A computer program product containing instructions, characterized in that, When it is run on at least one computing device, it causes the at least one computing device to perform the method as described in any one of claims 1 to 9.