Port adjustment method and device for Leaf switch in intelligent computing cluster network
By adjusting the number of uplink and downlink ports on the Leaf switch in the intelligent computing cluster network, the problem of underutilization of uplink bandwidth resources was solved, improving overall throughput and reducing costs.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-12-12
- Publication Date
- 2026-04-14
AI Technical Summary
In intelligent computing cluster networks, the same number of ports configured for the uplink and downlink of Leaf switches leads to underutilization of uplink network bandwidth resources, affecting overall throughput.
After iterative training, the total transmission time of the Spine and Leaf switches is determined, and the number of uplink and downlink ports of the Leaf switch is dynamically adjusted so that the number of uplink ports is less than the number of downlink ports to match traffic demand.
It improves the overall throughput of the intelligent computing cluster network, makes full use of network bandwidth resources, and reduces construction and operation costs.
Smart Images

Figure CN121864586A_ABST
Abstract
Description
Technical Field
[0001] This disclosure relates to the field of computer technology, and in particular to a method and apparatus for adjusting ports of a Leaf switch in an intelligent computing cluster network. Background Technology
[0002] Currently, in the distributed training of large models using intelligent computing cluster networks, related technologies typically configure the uplink and downlink ports of Leaf switches within the network to have the same number of ports. However, this configuration often results in underutilization of uplink network bandwidth resources. Therefore, adjusting the number of uplink and downlink ports on Leaf switches is crucial for improving the overall throughput of intelligent computing cluster networks. Summary of the Invention
[0003] This disclosure provides a method and apparatus for adjusting ports on a Leaf switch in an intelligent computing cluster network.
[0004] In a first aspect, this disclosure provides a method for adjusting the ports of Leaf switches in an intelligent computing cluster network. The method includes: after completing one iteration of training on a large model through the intelligent computing cluster network, determining a first total transmission time required by the Spine switch in the intelligent computing cluster network to forward all gradient data of the large model in this iteration; determining a second total transmission time required by the Leaf switch in the intelligent computing cluster network to forward all intermediate representation data between network layers of the large model in this iteration; and adjusting the number of uplink ports and the number of downlink ports of the Leaf switches in the intelligent computing cluster network according to the first total transmission time and the second total transmission time, wherein the adjusted number of uplink ports of the Leaf switches is less than the adjusted number of downlink ports of the Leaf switches.
[0005] Secondly, this disclosure provides a port adjustment device for Leaf switches in an intelligent computing cluster network. The device includes: a first determining module, used to determine, after completing one iteration of training on a large model through the intelligent computing cluster network, the first total transmission time required by the Spine switch in the intelligent computing cluster network to forward all gradient data of the large model in this iteration; a second determining module, used to determine, in this iteration, the second total transmission time required by the Leaf switch in the intelligent computing cluster network to forward all intermediate representation data between network layers of the large model; and an adjusting module, used to adjust the number of uplink ports and the number of downlink ports of the Leaf switch in the intelligent computing cluster network according to the first total transmission time and the second total transmission time, wherein the adjusted number of uplink ports of the Leaf switch is less than the adjusted number of downlink ports of the Leaf switch.
[0006] Thirdly, this disclosure provides an electronic device, including: a processor, and a memory communicatively connected to the processor; the memory stores computer-executable instructions; the processor executes the computer-executable instructions stored in the memory to implement the port adjustment method of Leaf switches in a smart computing cluster network disclosed in this disclosure embodiment.
[0007] Fourthly, this disclosure provides a computer-readable storage medium storing computer-executable instructions, which, when executed by a processor, are used to implement the port adjustment method of Leaf switches in a smart computing cluster network disclosed in embodiments of this disclosure.
[0008] Fifthly, this disclosure provides a computer program product, including a computer program that, when executed by a processor, implements the port adjustment method for Leaf switches in a smart computing cluster network disclosed in embodiments of this disclosure.
[0009] The technical solutions provided by the embodiments of this disclosure have at least the following beneficial effects: After completing one iteration of training on a large model using an intelligent computing cluster network, the first total transmission time required for the Spine switch in the network to forward all gradient data of the large model in this iteration was determined, and the second total transmission time required for the Leaf switch to forward all intermediate representation data between network layers of the large model in this iteration was also determined. Based on the first and second total transmission times, the number of uplink and downlink ports on the Leaf switches in the network was adjusted. This dynamically adjusted the uplink and downlink port ratio of the Leaf switches to more efficiently match uplink and downlink traffic demands, thereby improving the overall throughput capacity of the intelligent computing cluster network. Attached Figure Description
[0010] The accompanying drawings, which are incorporated in and form a part of this specification, illustrate embodiments consistent with this disclosure and, together with the description, serve to explain the principles of this disclosure.
[0011] Figure 1 This is a flowchart illustrating a port adjustment method for a Leaf switch in an intelligent computing cluster network according to an exemplary embodiment; Figure 2 This is a detailed flowchart illustrating step 101 according to an exemplary embodiment; Figure 3 This is a detailed flowchart illustrating step 102 according to an exemplary embodiment; Figure 4 This is a flowchart illustrating another method for port adjustment of a Leaf switch in an intelligent computing cluster network according to an exemplary embodiment; Figure 5 This is an example diagram illustrating the interaction between GPU servers in a smart computing cluster network using data parallelism (DP), pipelined parallelism (PP), and tensor parallelism (TP). Figure 6 This is an example diagram of intelligent computing cluster networks in related technologies; Figure 7 This is an example diagram of the intelligent computing cluster network adjusted using the method of this embodiment; Figure 8 This is a schematic diagram illustrating the structure of a port adjustment device for a Leaf switch in an intelligent computing cluster network according to an exemplary embodiment; Figure 9 This is a structural block diagram of an electronic device according to an exemplary embodiment.
[0012] The accompanying drawings have illustrated specific embodiments of this disclosure, which will be described in more detail below. These drawings and descriptions are not intended to limit the scope of the concept in any way, but rather to illustrate the concepts of this disclosure to those skilled in the art through reference to particular embodiments. Detailed Implementation
[0013] Exemplary embodiments will now be described in detail, examples of which are illustrated in the accompanying drawings. When the following description relates to the drawings, unless otherwise indicated, the same numerals in different drawings denote the same or similar elements. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with this disclosure. Rather, they are merely examples of apparatuses and methods consistent with some aspects of this disclosure as detailed in the appended claims.
[0014] The technical solutions of this disclosure and how they solve the aforementioned technical problems will be described in detail below with specific embodiments. These specific embodiments can be combined with each other, and the same or similar concepts or processes may not be repeated in some embodiments. The embodiments of this disclosure will now be described with reference to the accompanying drawings.
[0015] First, combine Figure 1 The present disclosure provides an exemplary description of the port adjustment method for Leaf switches in an intelligent computing cluster network.
[0016] Figure 1 This is a flowchart illustrating a port adjustment method for a Leaf switch in an intelligent computing cluster network according to an exemplary embodiment.
[0017] It should be noted that the port adjustment method for Leaf switches in the intelligent computing cluster network provided in this embodiment can be executed by a port adjustment device for Leaf switches in the intelligent computing cluster network, which can be implemented in hardware. This port adjustment device can be an electronic device or can be configured within an electronic device.
[0018] It should be noted that the electronic device can be any device with computing capabilities, such as a terminal device or a server. This embodiment does not specifically limit the electronic device.
[0019] like Figure 1 As shown, the port adjustment method for the Leaf switch in this intelligent computing cluster network includes the following steps: Step 101: After completing one iteration of training on the large model through the intelligent computing cluster network, determine the first total transmission time required for the Spine switch in the intelligent computing cluster network to forward all gradient data of the large model in this iteration.
[0020] In some embodiments, when training a large model through an intelligent cluster network, the intelligent cluster network in this embodiment may include multiple groups of graphics processing unit (GPU) servers, wherein each group of GPU servers includes multiple GPU servers, and multiple network layers of the large language model are deployed on multiple GPU servers respectively, with each network layer corresponding to one of the multiple GPU servers. The training data subset used by each group of GPU servers when training the large language model is different, and the training data subset is obtained by partitioning the training data of the large language model.
[0021] In this embodiment, data communication between different groups of GPU servers is achieved via Spine switches. Specifically, for any first group of GPU servers, when the i-th GPU server in the first group needs to send gradient data of a large model to the i-th GPU server in the second group of GPU servers, the i-th GPU server in the first group sends the gradient data to the first Leaf switch connected to it. Correspondingly, the first Leaf switch transmits the gradient data to the first Spine switch connected to it. The first Spine switch then sends the gradient data to the second Leaf switch, which in turn sends the gradient data to the i-th GPU server in the second group of GPU servers connected to it. Here, i is an integer greater than or equal to 1 and less than or equal to N, where N represents the number of servers in the first group of GPU servers. The number of servers in different groups of GPU servers is the same.
[0022] In this setup, GPU servers within the same group communicate with each other via Leaf switches.
[0023] In some embodiments, the Spine switch in the intelligent computing cluster network can be one or more.
[0024] In some embodiments, when there are multiple Spine switches in the intelligent computing cluster network, the first total transmission time required for all Spine switches in the intelligent computing cluster network to forward all gradient data of the large model in this iteration can be determined.
[0025] In this embodiment, the large model can be of various types, such as a large language model. This embodiment does not specifically limit the large model.
[0026] Step 102: Determine the second total transmission time required by the Leaf switches of the intelligent computing cluster network to forward all intermediate representation data between the network layers of the large model in this iteration.
[0027] Step 103: Based on the first total transmission time and the second total transmission time, adjust the number of uplink ports and downlink ports of the Leaf switch in the intelligent computing cluster network, wherein the adjusted number of uplink ports of the Leaf switch is less than the adjusted number of downlink ports of the Leaf switch.
[0028] In some embodiments, after obtaining the first total transmission time and the second total transmission time, comparing the first total transmission time and the second total transmission time reveals that the second total transmission time is much larger than the first total transmission time. That is, it can be known that during the training of a large model through the intelligent cluster network, the total communication consumption of pipeline parallelism is much greater than the total communication consumption of data parallelism. Therefore, the traffic load (i.e., the amount of communication data) carried by the Leaf switch is much greater than the traffic load carried by the Spine switch. Therefore, to efficiently match uplink and downlink traffic demands and avoid underutilization of uplink network bandwidth resources on Leaf switches, the port allocation ratio can be dynamically adjusted based on the first total transmission time and the second total transmission time. Then, based on the total number of ports on the Leaf switches and the dynamically adjusted port allocation ratio, the first target port number for the uplink and the second target port number for the downlink of the Leaf switches can be determined. Furthermore, the number of uplink ports on the Leaf switches in the intelligent computing cluster network can be adjusted according to the first target port number, ensuring that the adjusted number of uplink ports on the Leaf switches equals the first target port number. Similarly, the number of downlink ports on the Leaf switches can be adjusted according to the second target port number, ensuring that the adjusted number of downlink ports on the Leaf switches equals the second target port number. In other words, the second target port number is the same as the adjusted number of uplink ports on the Leaf switches, and the second target port number is the same as the adjusted number of downlink ports on the Leaf switches.
[0029] It should be noted that in this embodiment, the network bandwidth of each port of the Leaf switch is the same, for example, it can be 200G (Gigabit).
[0030] In other embodiments, a target ratio of a first total transmission time and a second total transmission time is determined. Based on the target interval where the target ratio falls, a first target number of uplink ports and a second target number of downlink ports of the Leaf switches are determined. The number of uplink ports of the Leaf switches in the intelligent computing cluster network is adjusted according to the first target number of ports, so that the adjusted number of uplink ports of the Leaf switches equals the first target number of ports. Similarly, the number of downlink ports of the Leaf switches is adjusted according to the second target number of ports, so that the adjusted number of downlink ports of the Leaf switches equals the second target number of ports. That is, the second target number of ports is the same as the adjusted number of uplink ports of the Leaf switches, and the second target number of ports is the same as the adjusted number of downlink ports of the Leaf switches.
[0031] In some embodiments, the first total transmission time can be compared with the second total transmission time to obtain a target ratio, that is, the first total transmission time is divided by the second total transmission time to obtain the target ratio.
[0032] In this context, the uplink of a Leaf switch refers to the link that connects from the Leaf switch to the Spine switch.
[0033] The downlink of the Leaf switch is the link that connects the Leaf switch to the GPU server.
[0034] The port adjustment method for Leaf switches in an intelligent computing cluster network provided in this embodiment determines the first total transmission time required for the Spine switches in the intelligent computing cluster network to forward all gradient data of the large model in this iteration, and the second total transmission time required for the Leaf switches in the intelligent computing cluster network to forward all intermediate representation data between network layers of the large model in this iteration. The number of uplink and downlink ports of the Leaf switches in the intelligent computing cluster network is adjusted based on the first and second total transmission times. This dynamically adjusts the ratio of uplink to downlink ports of the Leaf switches to more efficiently match uplink and downlink traffic demands, thereby improving the overall throughput capacity of the intelligent computing cluster network.
[0035] In some embodiments, to clearly understand the process of determining the first total transmission time required by the Spine switch in the intelligent computing cluster network to forward all gradient data of the large model in this iteration, the following is combined with... Figure 2 An exemplary description is provided of one possible implementation for determining the first total transmission time required by the Spine switch in the intelligent computing cluster network to forward all gradient data of the large model in this iteration.
[0036] Figure 2 This is a detailed flowchart illustrating step 101 according to an exemplary embodiment.
[0037] like Figure 2 As shown, the method may include: Step 201: Determine the first data size of the gradient data of the large model transmitted by the Spine switch in the intelligent computing cluster network in a single forwarding operation during this iteration.
[0038] In some embodiments, when the large model is a transformer-based large language model, the first data size of the gradient data of the large model transmitted by the Spine switch in a single forwarding operation can be equal to the sum of the gradient data sizes of each encoder and decoder layer of the large model multiplied by the number of network layers n of the large model, plus the size of the word embeddings and output layers of the large model.
[0039] In some embodiments, each layer of normalization includes a scaling factor and a bias, the size of which is... s 1 is:
[0040] The attention layer contains an input projection matrix (Q, K, V) and an output projection matrix, with a parameter count of... s 2 is:
[0041] Generally ascending to a higher dimension The multi-layer perceptron (MLP) layer has a large number of parameters. s 3 is:
[0042] The word embedding layer and the output layer are the same size, where This refers to the vocabulary size; the two combined are: s 4:
[0043] In some embodiments, the first data size of the gradient data of a large model transmitted by the Spine switch in a single forwarding operation is... It can be represented as:
[0044] Where h represents the hidden layer dimension of the network layer of the large model; n represents the number of network layers n of the large model; and V represents the vocabulary size of the large model.
[0045] Step 202: Based on the data parallelism of the large model in the intelligent computing cluster network and the first data size, determine the first total transmission time required for the Spine switch in the intelligent computing cluster network to forward all gradient data in this iteration.
[0046] In some embodiments, when training a large model through an intelligent cluster network, the intelligent cluster network in this embodiment may include multiple groups of GPU servers, wherein each group of GPU servers includes multiple GPU servers, and multiple network layers of the large language model are deployed on multiple GPU servers respectively, with each network layer corresponding to one of the multiple GPU servers. The training data subset used by each group of GPU servers when training the large language model is different, and the training data subset is obtained by partitioning the training data of the large language model.
[0047] The data parallelism is the number of GPU servers in the set.
[0048] In some embodiments, a possible implementation of step 202 above may be: determining, based on the data parallelism, the first total number of data packet forwardings generated by the Spine switch in the intelligent computing cluster network transmitting gradient data in this iteration; determining, based on the first data size and the data parallel bus bandwidth of the large model in the intelligent computing cluster network, the first single transmission time required for the Spine switch in the intelligent computing cluster network to transmit the gradient data of the large model in this iteration; and determining, based on the first single transmission time and the first total number of data packet forwardings, the first total transmission time required for the Spine switch in the intelligent computing cluster network to forward all gradient data in this iteration.
[0049] It is understandable that, in different application scenarios, the method for determining the number of times the first total data packet is forwarded by the Spine switch in the intelligent computing cluster network to transmit gradient data in this iteration differs based on the degree of data parallelism. An example is illustrated below: As an example, based on the data parallelism, the first total number of packet forwardings generated by the Spine switch in the intelligent computing cluster network transmitting gradient data in this iteration can be determined from the pre-established correspondence between the data parallelism and the total number of packet forwardings.
[0050] As another example, the data parallelism can be input into a formula for calculating the first total number of packet forwardings, so as to obtain the first total number of packet forwardings through the formula.
[0051] In some embodiments, the first total number of data packet forwardings is calculated. The calculation formula can be expressed as: In the formula, d represents the degree of data parallelism.
[0052] In some embodiments, the first total transmission time is obtained. The calculation formula can be expressed as:
[0053] Among them, in the formula Indicates the time of the first single transmission, where, ,in, Indicates the size of the first data item. d represents the data parallel bus bandwidth, and d represents the data parallelism.
[0054] In some embodiments, to clearly understand the specific process of determining the second total transmission time required by the Leaf switches of the intelligent computing cluster network to forward all intermediate representation data between the network layers of the large model in this iteration, the following is combined with... Figure 3 An exemplary description is provided of one possible implementation of the second total transmission time required by the Leaf switch of the intelligent computing cluster network to forward all intermediate representation data between the network layers of the large model in this iteration.
[0055] Figure 3 This is a detailed flowchart illustrating step 102 according to an exemplary embodiment.
[0056] like Figure 3 As shown, it may include: Step 301: Determine the second data size of the intermediate representation data between network layers of the large model transmitted by the Leaf switch of the intelligent computing cluster network in a single forwarding operation during this iteration.
[0057] In some embodiments, the parameters can be determined based on the micro-batch size between network layers, the sequence length of the semantic unit sequence of each input sample in the micro-batch, and the dimension of the hidden layer in the network layer.
[0058] In some embodiments, the second data size It can be represented as:
[0059] Where b represents the micro-batch size, s represents the sequence length, and h represents the hidden layer dimension of the network layer. The sequence length refers to the sequence length of the semantic unit token sequence of each input sample in the micro-batch. The semantic unit sequence is obtained by the large language model from the input sample through word segmentation.
[0060] It should be noted that this embodiment uses the example of a large model where the hidden layer dimensions of each network layer are the same for illustration.
[0061] Step 302: Based on the pipeline parallelism of the large model in the intelligent computing cluster network and the second data size, determine the second total transmission time required for the Leaf switch of the intelligent computing cluster network to forward all intermediate representation data in this iteration.
[0062] In some embodiments, when training a large model using an intelligent cluster network, the intelligent cluster network in this embodiment may include multiple groups of GPU servers. Each group of GPU servers includes multiple GPU servers, and the multiple network layers of the large language model are deployed on the multiple GPU servers respectively. Each network layer corresponds one-to-one with the multiple GPU servers. The training data subset used by each group of GPU servers when training the large language model is different; the training data subset is obtained by partitioning the training data of the large language model. The number of servers included in each group of GPU servers is the pipeline parallelism.
[0063] In some embodiments, a possible implementation of step 302 above may be: determining the second total number of packet forwardings generated by the Leaf switch of the intelligent computing cluster network transmitting intermediate representation data in this iteration based on the pipeline parallelism; determining the second single transmission time required by the Spine switch of the intelligent computing cluster network to transmit the gradient data of the large model in this iteration based on the second data size and the total pipeline parallel bandwidth of the large model in the intelligent computing cluster network; and determining the second total transmission time required by the Leaf switch of the intelligent computing cluster network to forward all intermediate representation data in this iteration based on the second total number of packet forwardings and the second single transmission time.
[0064] In some embodiments, the implementation methods for determining the second total number of packet forwardings generated by the Leaf switch of the intelligent computing cluster network in transmitting intermediate representation data during this iteration differ in different application scenarios, based on the pipeline parallelism. Examples are illustrated below: As an example, based on the pipeline parallelism, the number of times the second total data packet is forwarded during the transmission of intermediate representation data by the Leaf switch of the intelligent computing cluster network in this iteration can be determined from the correspondence between the preset pipeline parallelism and the second total data packet forwarding count.
[0065] As another example, the pipeline parallelism can be input into the formula used to calculate the second total number of packet forwardings, so that the second total number of packet forwardings can be obtained through the formula.
[0066] In some embodiments, the second total number of data packet forwardings is obtained. The calculation formula can be expressed as: ,in, , Where B represents the number of training samples in a single transmission, b represents the micro-batch size, d represents the data parallelism, and p represents the pipeline parallelism. Here, b is the micro-batch size, which is obtained by further dividing the training samples in a single transmission. It should be noted that... This represents the total number of packet forwards by the Spine switch when only the forward propagation phase is performed versus when only the backward propagation phase is performed. This represents the number of forward and backward communications per microbatch during the steady-state phase in hybrid parallel training that simultaneously employs data parallelism and pipeline parallelism.
[0067] In some embodiments, the second total transmission time is obtained. The calculation formula can be expressed as:
[0068] Among them, in the formula Indicates the time of the second single transmission; Indicates the size of the second data; denoted by , where represents the total parallel bandwidth of the pipeline; B represents the number of samples used for training in a single pass; b represents the micro-batch size; d represents the data parallelism; and p represents the pipeline parallelism.
[0069] Based on any of the above embodiments, in order to fully utilize the downlink network bandwidth resources of the Spine switch, in some embodiments, the electronic device may also perform the following operations: grouping the adjusted uplink ports of the Leaf switch to obtain multiple first port groups; and grouping the downlink ports of the Spine switch according to the number of ports in the first port groups to obtain multiple second port groups, wherein the number of ports in the second port groups is the same as the number of ports in the first port groups.
[0070] In this context, the downlink of a Spine switch refers to the link that connects from the Spine switch down to the Leaf switch.
[0071] To facilitate a clear understanding of this disclosure, the following will be combined with... Figure 4 The method of this embodiment is described by way of example.
[0072] Figure 4 This is a flowchart illustrating another method for port adjustment of a Leaf switch in an intelligent computing cluster network according to an exemplary embodiment.
[0073] like Figure 4 As shown, it may include: Step 401: After completing one iteration of training on the large model through the intelligent computing cluster network, determine the first data size of the gradient data of the large model transmitted by the Spine switch in the intelligent computing cluster network in a single forwarding operation during this iteration.
[0074] Step 402: Based on the data parallelism of the large model in the intelligent computing cluster network and the first data size, determine the first total transmission time required for the Spine switch in the intelligent computing cluster network to forward all gradient data in this iteration.
[0075] Step 403: Determine the second data size of the intermediate representation data between network layers of the large model transmitted by the Leaf switch of the intelligent computing cluster network in a single forwarding operation during this iteration.
[0076] Step 404: Based on the pipeline parallelism of the large model in the intelligent computing cluster network and the second data size, determine the second total transmission time required for the Leaf switch of the intelligent computing cluster network to forward all intermediate representation data in this iteration.
[0077] It should be noted that the execution order of steps 402 and 403 is not important, and this embodiment does not impose any specific restrictions on this.
[0078] For a detailed description of steps 401-404, please refer to the relevant descriptions in other embodiments, which will not be repeated here.
[0079] Step 405: Based on the first total transmission time and the second total transmission time, adjust the number of uplink ports and downlink ports of the Leaf switch in the intelligent computing cluster network, wherein the adjusted number of uplink ports of the Leaf switch is less than the adjusted number of downlink ports of the Leaf switch.
[0080] It should be noted that for a detailed description of step 405, please refer to the relevant descriptions in other embodiments, which will not be repeated here.
[0081] Figure 5 This is an example diagram illustrating the interaction and transmission of data in a smart computing cluster network, encompassing data parallelism (DP), pipeline parallelism (PP), and tensor parallelism (TP). Figure 5As can be seen, during the training of large models via the intelligent computing cluster network, TP parallel data communicates between GPU cards on the GPU server, without needing to pass through Leaf and Spine switches in the cluster network. Correspondingly, PP parallel communication occurs between GPU cards with the same ID on different GPU servers within the same Leaf switch. DP communication also occurs between GPU cards with the same ID on different GPU servers, but not within the same Leaf switch; instead, it requires communication across Spine switches. Therefore, since TP parallel traffic does not pass through Leaf and Spine switches in the cluster network, this embodiment only calculates the communication overhead of DP and PP.
[0082] To facilitate a clear understanding of this disclosure, the following will be combined with... Figure 6 as well as Figure 7 The method of this embodiment is described by way of example.
[0083] For example, taking a computing cluster network with a capacity of 20 GPU servers as an example, during the training of a large model through this computing cluster network, there are 4 Spine switches and 8 Leaf switches in the network. Furthermore, all Leaf switches are 40*200G box-type wireless bandwidth (InfiniBand, IB) switches. An example diagram of the computing cluster network using these technologies is shown below. Figure 6 As shown. If a traditional symmetrical traffic load networking method is used, each Leaf switch's 40 200G ports are split in two, with uplink bandwidth of 20*200G and downlink bandwidth of 20*200G, resulting in a symmetrical uplink and downlink traffic load configuration. It should be noted that... Figure 6 Server-1 to Server-20 represent 20 GPU servers; Leaf Switch-1 to Leaf Switch-8 represent 8 Leaf switches; Spine Switch-1 to Spine Switch-4 represent 4 Spine switches.
[0084] In this intelligent computing cluster network, there are 4 Spine switches. Each Spine switch has 40 200G ports on the downlink, divided into 8 groups of 5 200G ports each, which are connected to 8 Leaf switches.
[0085] Each Leaf switch has 20 uplink 200G ports, divided into 4 groups of 5 200G ports each, connected to 4 Spine switches. Each Leaf switch also has 20 downlink 200G ports, connected to 20 GPU servers. In this configuration, the access layer below the Leaf switch and the aggregation layer above it have symmetrical traffic loads, both at 20*200G.
[0086] As can be seen from the method proposed in the embodiments of this disclosure, the traffic load of the aggregation layer for carrying DP parallel is much smaller than that of the access layer for carrying PP parallel, and the traffic load of the aggregation layer can be reduced. At the same time, the released traffic load capacity is provided to the access layer, which simplifies the aggregation layer network and improves the access layer capacity of the intelligent computing cluster network.
[0087] If, through calculation using the method of this embodiment, it is determined that the first target port number of the uplink of the Leaf switch is 16 and the second target port number of the downlink of the Leaf switch is 24, based on the first total transmission time and the second total transmission time, the ports of the uplink of the Leaf switch can be adjusted according to the first target port number, so that the adjusted uplink port number of the Leaf switch is 16; and the ports of the downlink of the Leaf switch can be adjusted according to the second target port number, so that the adjusted downlink port number of the Leaf switch is 24.
[0088] An example diagram of the intelligent computing cluster network adjusted using the method of this embodiment is shown below. Figure 7 As shown. It should be noted that, Figure 7 Server-1 to Server-24 represent 24 GPU servers; Leaf Switch-1 to Leaf Switch-10 represent 10 Leaf switches; Spine Switch-1 to Spine Switch-4 represent 4 Spine switches.
[0089] Understandably, the adjusted intelligent computing cluster network still contains 4 Spine switches. Each Spine switch has 10 downlink groups, each with 4 200G ports, connected to 10 Leaf switches. Each Leaf switch has 16 uplink ports, divided into 4 groups of 4, each connected to 4 Spine switches. Each Leaf switch has 24 downlink ports, connected to 24 GPU servers.
[0090] It is understandable that in the adjusted intelligent computing cluster network, the number of GPU servers that each Leaf switch can connect to has increased from 20 to 24, which is a 20% increase compared to the previous intelligent computing cluster network.
[0091] In addition, each Spine switch in the adjusted intelligent computing cluster network can connect to 10 Leaf switches. Compared with the previous intelligent computing cluster network, the number of Leaf switches connected to each Spine switch has also increased, which helps to reduce the aggregation layer network, thereby reducing construction costs and maintenance burden.
[0092] In this embodiment, by adjusting the traditional symmetrical traffic load network to asymmetrical based on different levels of network traffic load, the number of GPU servers that the Leaf switch can connect to is increased, thereby improving the cluster's computing power. Furthermore, by simplifying the aggregation layer network through asymmetrical traffic load, construction costs and operational burdens can be reduced.
[0093] Figure 8 This is a schematic diagram illustrating the structure of a port adjustment device for a Leaf switch in an intelligent computing cluster network according to an exemplary embodiment.
[0094] like Figure 8 As shown, the port adjustment device 800 of the Leaf switch in the intelligent computing cluster network includes: a first determining module 801, a second determining module 802, and an adjustment module 803, wherein: The first determining module 801 is used to determine, after completing one iteration of training of the large model through the intelligent computing cluster network, the first total transmission time required for the Spine switch in the intelligent computing cluster network to forward all gradient data of the large model in this iteration.
[0095] The second determining module 802 is used to determine, in this iteration, the second total transmission time required by the Leaf switch of the intelligent computing cluster network to forward all intermediate representation data between the network layers of the large model.
[0096] The adjustment module 803 is used to adjust the number of uplink ports and downlink ports of the Leaf switch in the intelligent computing cluster network according to the first total transmission time and the second total transmission time, wherein the adjusted number of uplink ports of the Leaf switch is less than the adjusted number of downlink ports of the Leaf switch.
[0097] In one embodiment of this disclosure, the first determining module 801 may include: The first determining unit is used to determine the first data size of the gradient data of the large model transmitted by the Spine switch in the intelligent computing cluster network in a single forwarding operation during this iteration. The second determining unit is used to determine, based on the data parallelism of the large model in the intelligent computing cluster network and the first data size, the first total transmission time required for the Spine switch in the intelligent computing cluster network to forward all gradient data in this iteration.
[0098] In one embodiment of this disclosure, the second determining unit is specifically configured to: determine, based on the data parallelism, the first total number of data packet forwardings generated by the Spine switch in the intelligent computing cluster network transmitting gradient data in the current iteration; determine, based on the first data size and the data parallel bus bandwidth of the large model in the intelligent computing cluster network, the first single transmission time required by the Spine switch in the intelligent computing cluster network to transmit the gradient data of the large model in the current iteration; and determine, based on the first single transmission time and the first total number of data packet forwardings, the first total transmission time required by the Spine switch in the intelligent computing cluster network to forward all gradient data in the current iteration.
[0099] In one embodiment of this disclosure, the second determining module 802 may include: The third determining unit is used to determine the second data size of the intermediate representation data between the network layers of the large model transmitted by the Leaf switch of the intelligent computing cluster network in a single forwarding operation during this iteration. The fourth determining unit is used to determine, based on the pipeline parallelism of the large model in the intelligent computing cluster network and the second data size, the second total transmission time required for the Leaf switch of the intelligent computing cluster network to forward all intermediate representation data in this iteration.
[0100] In one embodiment of this disclosure, the fourth determining unit is specifically configured to: determine, based on the pipeline parallelism, the second total number of data packet forwardings generated by the Leaf switch of the intelligent computing cluster network transmitting intermediate representation data in the current iteration; determine, based on the second data size and the total pipeline parallel bandwidth of the large model in the intelligent computing cluster network, the second single transmission time required by the Spine switch of the intelligent computing cluster network to transmit the gradient data of the large model in the current iteration; and determine, based on the second total number of data packet forwardings and the second single transmission time, the second total transmission time required by the Leaf switch of the intelligent computing cluster network to forward all intermediate representation data in the current iteration.
[0101] In one embodiment of this disclosure, the apparatus may further include: a processing module, configured to group the adjusted uplink ports of the Leaf switch to obtain multiple first port groups; and to group the downlink ports of the Spine switch according to the number of ports in the first port groups to obtain multiple second port groups, wherein the number of ports in the second port groups is the same as the number of ports in the first port groups.
[0102] It should be noted that the aforementioned description of the port adjustment method embodiment for Leaf switches in intelligent computing cluster networks also applies to the port adjustment device for Leaf switches in intelligent computing cluster networks in this embodiment, and will not be repeated here.
[0103] The port adjustment device for Leaf switches in the intelligent computing cluster network provided in this embodiment determines, after completing one iteration of training on a large model through the intelligent computing cluster network, the first total transmission time required for the Spine switches in the intelligent computing cluster network to forward all gradient data of the large model in this iteration, and the second total transmission time required for the Leaf switches in the intelligent computing cluster network to forward all intermediate representation data between network layers of the large model in this iteration. Based on the first and second total transmission times, the number of uplink ports and downlink ports of the Leaf switches in the intelligent computing cluster network are adjusted. This dynamically adjusts the ratio of uplink to downlink ports of the Leaf switches to more efficiently match uplink and downlink traffic demands, thereby improving the overall throughput capacity of the intelligent computing cluster network.
[0104] It should be noted that the acquisition, storage, use, and processing of data in this disclosed technical solution all comply with the relevant provisions of national laws and regulations.
[0105] According to embodiments of this disclosure, an electronic device is also provided, including: a processor and a memory communicatively connected to the processor; the memory stores computer-executable instructions; the processor executes the computer-executable instructions stored in the memory to implement the port adjustment method of Leaf switches in a smart computing cluster network disclosed in embodiments of this disclosure.
[0106] To implement the above embodiments, this disclosure also proposes a computer-readable storage medium storing computer-executable instructions, which, when executed by a processor, are used to implement the port adjustment method of Leaf switches in the intelligent computing cluster network disclosed in this disclosure.
[0107] To implement the above embodiments, this disclosure also provides a computer program product.
[0108] The computer program product includes a computer program that, when executed by a processor, implements the port adjustment method for Leaf switches in a smart computing cluster network disclosed in this embodiment.
[0109] Figure 9 This is a structural block diagram of an electronic device according to an exemplary embodiment. Figure 9 The electronic device shown is merely an example and should not be construed as limiting the functionality and scope of the embodiments disclosed herein.
[0110] like Figure 9 As shown, the electronic device 1000 includes a processor 111, which can perform various appropriate actions and processes according to a program stored in read-only memory (ROM) 112 or a program loaded from memory 116 into random access memory (RAM) 113. The RAM 113 also stores various programs and data required for the operation of the electronic device 1000. The processor 111, ROM 112, and RAM 113 are interconnected via a bus 114. An input / output (I / O) interface 115 is also connected to the bus 114.
[0111] The following components are connected to I / O interface 115: memory 116 including hard disks, etc.; and communication section 117 including network interface cards such as local area network (LAN) cards, modems, etc., communication section 117 performs communication processing via a network such as the Internet; and driver 118 is also connected to I / O interface 115 as needed.
[0112] In particular, according to embodiments of this disclosure, the processes described above with reference to the flowcharts can be implemented as computer software programs. For example, embodiments of this disclosure include a computer program carried on a computer-readable medium, the computer program containing program code for performing the methods shown in the flowcharts. In such embodiments, the computer program can be downloaded and installed from a network via communication section 117. When the computer program is executed by processor 111, it performs the functions defined in the methods of this disclosure.
[0113] In an exemplary embodiment, a storage medium including instructions is also provided, such as a memory including instructions, which can be executed by the processor 111 of the electronic device 1000 to perform the above-described method. Optionally, the storage medium may be a non-transitory computer-readable storage medium, such as a ROM, random access memory (RAM), CD-ROM, magnetic tape, floppy disk, and optical data storage device.
[0114] In this disclosure, a computer-readable storage medium can be any tangible medium containing or storing a program that can be used by or in connection with an instruction execution system, apparatus, or device. In this disclosure, a computer-readable signal medium can include a data signal propagated in baseband or as part of a carrier wave, carrying computer-readable program code. Such propagated data signals can take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination thereof. A computer-readable signal medium can also be any computer-readable medium other than a computer-readable storage medium, which can transmit, propagate, or transfer a program for use by or in connection with an instruction execution system, apparatus, or device. The program code contained on the computer-readable medium can be transmitted using any suitable medium, including but not limited to: wireless, wireline, optical fiber, RF, etc., or any suitable combination thereof.
[0115] Other embodiments of this disclosure will readily occur to those skilled in the art upon consideration of the specification and practice of the invention disclosed herein. This disclosure is intended to cover any variations, uses, or adaptations of this disclosure that follow the general principles of this disclosure and include common knowledge or customary techniques in the art not disclosed herein. The specification and examples are to be considered exemplary only, and the true scope and spirit of this disclosure are indicated by the following claims.
[0116] It should be understood that this disclosure is not limited to the precise structures described above and shown in the accompanying drawings, and various modifications and changes can be made without departing from its scope. The scope of this disclosure is limited only by the appended claims.
Claims
1. A method for adjusting ports of a Leaf switch in an intelligent computing cluster network, characterized in that, The method includes: After completing one iteration of training on a large model through an intelligent computing cluster network, it is determined that the Spine switch in the intelligent computing cluster network is the first total transmission time required to forward all gradient data of the large model in this iteration. In this iteration, the Leaf switch of the intelligent computing cluster network is determined to be the second total transmission time required to forward all intermediate representation data between the network layers of the large model; Based on the first total transmission time and the second total transmission time, the number of uplink ports and downlink ports of the Leaf switches in the intelligent computing cluster network are adjusted, wherein the adjusted number of uplink ports of the Leaf switches is less than the adjusted number of downlink ports of the Leaf switches.
2. The method as described in claim 1, characterized in that, The determination of the first total transmission time required for the Spine switch in the intelligent computing cluster network to forward all gradient data of the large model in this iteration includes: Determine the first data size of the gradient data of the large model transmitted by the Spine switch in the intelligent computing cluster network in a single forwarding operation during this iteration; Based on the data parallelism of the large model in the intelligent computing cluster network and the first data size, the first total transmission time required for the Spine switch in the intelligent computing cluster network to forward all gradient data is determined in this iteration.
3. The method as described in claim 2, characterized in that, The step of determining the first total transmission time required for the Spine switch in the intelligent computing cluster network to forward all gradient data in this iteration, based on the data parallelism of the large model in the intelligent computing cluster network and the first data size, includes: Based on the data parallelism, determine the first total number of data packet forwardings generated by the Spine switch in the intelligent computing cluster network transmitting the gradient data in this iteration; Based on the first data size and the data parallel bus bandwidth of the large model in the intelligent computing cluster network, the first single transmission time required for the Spine switch in the intelligent computing cluster network to transmit the gradient data of the large model in this iteration is determined. Based on the first single transmission time and the first total number of data packet forwardings, the first total transmission time required for the Spine switch in the intelligent computing cluster network to forward all gradient data is determined in this iteration.
4. The method as described in claim 1, characterized in that, The determination of the second total transmission time required for the Leaf switches of the intelligent computing cluster network to forward all intermediate representation data between the network layers of the large model in this iteration includes: In this iteration, the second data size of the intermediate representation data between the network layers of the large model transmitted by the Leaf switch of the intelligent computing cluster network in a single forwarding operation is determined. Based on the pipeline parallelism of the large model in the intelligent computing cluster network and the second data size, the second total transmission time required for the Leaf switch of the intelligent computing cluster network to forward all intermediate representation data is determined in this iteration.
5. The method as described in claim 4, characterized in that, Based on the pipeline parallelism of the large model in the intelligent computing cluster network and the second data size, the second total transmission time required for the Leaf switch of the intelligent computing cluster network to forward all intermediate representation data in this iteration is determined, including: Based on the pipeline parallelism, determine the second total number of data packet forwardings generated by the Leaf switch of the intelligent computing cluster network transmitting the intermediate representation data in this iteration; Based on the second data size and the total pipeline parallel bandwidth of the large model in the intelligent computing cluster network, the second single transmission time required for the Spine switch in the intelligent computing cluster network to transmit the gradient data of the large model in this iteration is determined. Based on the second total number of data packet forwardings and the second single transmission time, the second total transmission time required for the Leaf switch of the intelligent computing cluster network to forward all intermediate representation data is determined in this iteration.
6. The method according to any one of claims 1-5, characterized in that, The method further includes: The adjusted uplink ports of the Leaf switch are grouped to obtain multiple first port groups; Based on the number of ports in the first port group, the downlink ports of the Spine switch are grouped to obtain multiple second port groups, wherein the number of ports in the second port groups is the same as the number of ports in the first port group.
7. A port adjustment device for a Leaf switch in an intelligent computing cluster network, characterized in that, The device includes: The first determining module is used to determine, after completing one iteration of training on the large model through the intelligent computing cluster network, the first total transmission time required for the Spine switch in the intelligent computing cluster network to forward all gradient data of the large model in this iteration. The second determining module is used to determine the second total transmission time required for the Leaf switch of the intelligent computing cluster network to forward all intermediate representation data between the network layers of the large model in this iteration; An adjustment module is used to adjust the number of uplink ports and downlink ports of the Leaf switch in the intelligent computing cluster network according to the first total transmission time and the second total transmission time, wherein the adjusted number of uplink ports of the Leaf switch is less than the adjusted number of downlink ports of the Leaf switch.
8. An electronic device, characterized in that, include: A processor, and a memory communicatively connected to the processor; The memory stores computer-executed instructions; The processor executes computer execution instructions stored in the memory to implement the method as described in any one of claims 1-6.
9. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores computer-executable instructions, which, when executed by a processor, are used to implement the method as described in any one of claims 1-6.
10. A computer program product, characterized in that, Includes a computer program that, when executed by a processor, implements the method as described in any one of claims 1-6.