Communication scheduling method for distributed large model training, electronic equipment and medium

By dynamically dividing and building distributed large-model training communication topology, the problem that traditional static topology cannot adapt to dynamic communication needs is solved, and communication efficiency optimization and topological reconstruction cost are achieved.

CN120075122AActive Publication Date: 2025-05-30ZHEJIANG LAB
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202510528050.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-25
Publication Date
2025-05-30
Estimated Expiration
2045-04-25

AI Technical Summary

Technical Problem

In the training environment of distributed deep learning, traditional static network topology cannot adapt to dynamically changing communication needs, resulting in low network resource utilization, large communication overhead, and limited model training speed.

Method used

By obtaining the number of servers, the degree of each server, the transmission requirements of AllReduce traffic and MP traffic, it is divided into AllReduce traffic quantum topology and MP traffic quantum topology, building a topology diagram and configuring an optical switch to realize a physical topology. Monitor traffic changes, adjust task placement and topology to minimize total delay.

Benefits of technology

The dual minimization of communication overhead and topological reconstruction cost in distributed large-scale model training is achieved, communication efficiency is optimized, and network congestion and link bottlenecks are reduced.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120075122A_ABST
    Figure CN120075122A_ABST
Patent Text Reader

Abstract

The invention discloses a distributed large model training-oriented communication scheduling method, electronic equipment and a medium, and the method comprises the steps: obtaining the scale and degree of a server cluster and the transmission demand proportion of AllReduce and MP traffic, splitting the scale into AllReduce traffic and the degree of MP traffic sub-topology according to the two types of traffic demand proportions, constructing AllReduce sub-topology and MP sub-topology based on the two types of traffic demand proportions, combining the AllReduce sub-topology and MP sub-topology to obtain a topological graph, and carrying out the calculation of the topological graph. Physical topology is realized by using an optical switch; when the traffic changes, obtaining all distributed large model training task arrays, link arrays and candidate placement position arrays; according to an affinity graph corresponding to each candidate placement position constructed in the topological graph, calculating compatibility scores of all links in the affinity graph to obtain an optimal placement position; calculating the time delays of all connected sub-graphs in the affinity graph corresponding to the optimal placement position, taking the sum of the time delays as the total time delay, and taking the minimum total time delay as an optimization target; and performing distributed large model training on the physical topology according to the optimal placement position and the total time delay.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of network traffic scheduling in computer distributed large model training, and particularly relates to a communication scheduling method, an electronic device, and a medium for distributed large model training. Background Art

[0002] In the training environment of distributed deep learning, as the scale and complexity of the model continue to increase, the communication requirements and traffic patterns during the training process become more dynamically complex. Especially when training complex models such as graph neural networks (GNNs) and mixture of experts models (MoEs), different computing tasks and communication tasks will exhibit significant traffic variation characteristics in different iterations of the training. Traditional static network topologies often lack flexibility in coping with these dynamic changes, unable to adapt to frequent adjustments of communication requirements, resulting in low network resource utilization, high communication overhead, and limited model training speed.

[0003] To solve this problem, existing research has introduced dynamic network topology optimization to improve communication efficiency by reconstructing the network at runtime. However, frequent topology reconstruction may introduce additional overhead, involving significant reconstruction costs for switches and links. During the topology change process, for model parallel and data parallel tasks such as AllReduce synchronization operations, this latency will cause performance degradation. Summary of the Invention

[0004] The purpose of the present invention is to provide a communication scheduling method, an electronic device, and a medium for distributed large model training in view of the deficiencies of the prior art.

[0005] In a first aspect, an embodiment of the present invention provides a communication scheduling method for distributed large model training, the method comprising:

[0006] Obtain the number of servers for distributed large model training, the degree of each server, the transmission traffic requirements of AllReduce traffic, and the transmission traffic requirements of MP traffic;

[0007] Divide the degree of each server into the degrees of the AllReduce traffic sub-topology and the MP traffic sub-topology according to the ratio of the transmission requirements of AllReduce traffic and MP traffic;

[0008] Construct an AllReduce sub-topology and an MP sub-topology respectively according to the transmission requirements and degrees of AllReduce traffic and MP traffic; combine the AllReduce sub-topology and the MP sub-topology to obtain a topology graph; configure an optical switch according to the topology graph to implement a physical topology;

[0009] Monitor the traffic changes of all servers during the distributed large model training process. When the traffic changes, obtain all distributed large model training task arrays, link arrays, and candidate placement location arrays. Among them, the candidate placement location indicates placing a certain task on a certain server.

[0010] Traverse each candidate placement location. For each candidate placement location, construct a corresponding affinity graph according to the topology graph, calculate the compatibility scores of all links in the affinity graph under the condition that the bandwidth requirement satisfies the link capacity constraint, and take the candidate placement location with the highest compatibility score as the best placement location. Calculate the latency of all connected subgraphs in the affinity graph corresponding to the best placement location, and take the sum of the latencies of all connected subgraphs as the total latency, with the goal of minimizing the total latency. Conduct distributed large model training on the physical topology according to the best placement location and the total latency.

[0011] In a second aspect, an embodiment of the present invention provides a communication scheduling system for distributed large model training. The system includes:

[0012] A topology reconstruction component, which is used to obtain the number of servers for distributed large model training, the degree of each server, the transmission traffic demand of AllReduce traffic, and the transmission traffic demand of MP traffic; divide the degree of each server into the degrees of the AllReduce traffic sub-topology and the MP traffic sub-topology according to the ratio of the transmission demands of AllReduce traffic and MP traffic; construct an AllReduce sub-topology and an MP sub-topology respectively according to the transmission demands and degrees of AllReduce traffic and MP traffic; combine the AllReduce sub-topology and the MP sub-topology to obtain a topology graph; configure an optical switch according to the topology graph to implement a physical topology.

[0013] A traffic awareness component, which is used to monitor the traffic changes of all servers during the distributed large model training process. When the traffic changes, obtain all distributed large model training task arrays, link arrays, and candidate placement location arrays. Among them, the candidate placement location indicates placing a certain task on a certain server. Traverse each candidate placement location. For each candidate placement location, construct a corresponding affinity graph according to the topology graph, calculate the compatibility scores of all links in the affinity graph under the condition that the bandwidth requirement satisfies the link capacity constraint, and take the candidate placement location with the highest compatibility score as the best placement location. Calculate the latency of all connected subgraphs in the affinity graph corresponding to the best placement location, and take the sum of the latencies of all connected subgraphs as the total latency, with the goal of minimizing the total latency. Conduct distributed large model training on the physical topology according to the best placement location and the total latency.

[0014] In a second aspect, an embodiment of the present invention provides an electronic device, including a memory and a processor, where the memory is coupled to the processor; wherein, the memory is used to store program data, and the processor is used to execute the program data to implement the above-mentioned communication scheduling method for distributed large model training.

[0015] In a third aspect, an embodiment of the present invention provides a computer-readable storage medium, on which a computer program is stored, and when the program is executed by a processor, it implements the above-mentioned communication scheduling method for distributed large model training.

[0016] In a fourth aspect, an embodiment of the present invention provides a computer program product, including a computer program / instructions, and when the computer program / instructions are executed by a processor, they implement the above-mentioned communication scheduling method for distributed large model training.

[0017] Compared with the prior art, the beneficial effects of the present invention are as follows:

[0018] The present invention provides a communication scheduling method for distributed large model training, which is used for distributed large model training tasks and constructing a topology graph based on transmission traffic requirements, and configuring an optical switch according to the topology graph to implement a physical topology; at the same time, monitoring the traffic changes of all servers during the distributed large model training process, adjusting the latency of jobs, reducing network congestion and link bottlenecks, thereby reducing the need for physical topology reconstruction. The present invention combines intelligent scheduling and dynamic topology to achieve the dual minimization of communication overhead and topology reconstruction cost in distributed large model training, and optimize communication efficiency. BRIEF DESCRIPTION OF THE DRAWINGS

[0019] In order to more clearly illustrate the technical solutions in the embodiments of the present invention, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the following drawings are only some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.

[0020] Figure 1 It is a schematic diagram of the communication scheduling method for distributed large model training provided by an embodiment of the present invention;

[0021] Figure 2 It is a flowchart of constructing a topology graph provided by an embodiment of the present invention;

[0022] Figure 3 It is a flowchart of traffic awareness and task scheduling provided by an embodiment of the present invention;

[0023] Figure 4 It is a schematic diagram of a sharded cluster system provided by an embodiment of the present invention;

[0024] Figure 5 The logical schematic diagram of the bipartite affinity diagram provided by the embodiment of the present invention;

[0025] Figure 6 The schematic diagram of an electronic device provided by the embodiment of the present invention. Detailed implementation manners

[0026] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.

[0027] It should be noted that, without conflict, the features in the following embodiments and implementation manners can be combined with each other.

[0028] As Figure 1 shown, the embodiment of the present invention provides a communication scheduling method for distributed large model training, and the method includes the following steps:

[0029] Step S1, obtaining the number of servers for distributed large model training, the degree of each server, the transmission traffic demand of AllReduce traffic, and the transmission traffic demand of MP traffic.

[0030] Specifically, in this example, obtain the task information, load mode, and parallelization strategy of the distributed large model training task; among them, the task information of the distributed large model training task includes task type, computing requirements, and data scale; the load mode of the distributed large model training task includes: compute-intensive, data-intensive; the parallelization strategy of the distributed large model training task includes: data parallelism, model parallelism;

[0031] Obtain the number of servers for distributed large model training, the degree of each server, the transmission traffic demand of AllReduce traffic, and the transmission traffic demand of MP traffic according to the task information, load mode, and parallelization strategy of the distributed large model training task.

[0032] It should be noted that the MP traffic is related to model parallelism and includes the activation values and gradient data calculated during the forward calculation and backpropagation processes. The traffic characteristics are scattered and the traffic is small. The AllReduce traffic is generated during both data parallelism and traffic parallelism. It is mainly used to synchronize model parameters between different aggregators during the global reduction (AllReduce) operation after a training sample is trained.

[0033] Step S2: Divide the degree of each server into the degrees of the AllReduce traffic sub-topology and the MP traffic sub-topology according to the ratio of the transmission requirements of the AllReduce traffic and the MP traffic; construct the AllReduce sub-topology and the MP sub-topology according to the transmission requirements and degrees of the AllReduce traffic and the MP traffic respectively; combine the AllReduce sub-topology and the MP sub-topology to obtain a topology graph; configure an optical switch according to the topology graph to implement the physical topology.

[0034] Further, as Figure 2 shown, the step S2 specifically includes the following steps:

[0035] Step S201: Divide the degree of each server into the degrees of the AllReduce traffic sub-topology and the MP traffic sub-topology according to the ratio of the transmission requirements of the AllReduce traffic and the MP traffic.

[0036] Specifically, the expression is as follows:

[0037] T = T MP + T AllReduce

[0038] In the formula, T represents the total traffic, T MP represents the transmission requirement of the MP traffic, and T AllReduce represents the transmission requirement of the AllReduce traffic.

[0039] d A = (T AllReduce / T) × d

[0040] d MP = (T MP / T) × d

[0041] In the formula, d represents the degree of the server, d A represents the degree allocated to the AllReduce traffic sub-topology, and d MP represents the degree allocated to the MP traffic sub-topology.

[0042] Step S202: Construct the AllReduce sub-topology according to the transmission requirement and degree of the AllReduce traffic.

[0043] For each AllReduce group, allocate the corresponding degree for each AllReduce group according to the traffic corresponding to each AllReduce group;

[0044] Take the server as a node and use an empty set as the initially selected first candidate permutation set; calculate the Euler's totient function for the degree corresponding to each server, and the result output by the Euler's totient function represents different connection methods between the servers; fill the result output by the Euler's totient function into the second candidate permutation set;

[0045] In response to the number of servers, the degree corresponding to each AllReduce group, and the second candidate permutation set corresponding to each AllReduce group, construct the first candidate permutation corresponding to each AllReduce group, and iteratively optimize the first candidate permutation through the method of selected permutations and combinations to obtain the topology corresponding to each AllReduce group;

[0046] Merge the topologies corresponding to each AllReduce group to obtain the AllReduce sub-topology.

[0047] Exemplarily, in this example, take the k-th AllReduce group as an example to elaborate on the specific process of constructing the AllReduce sub-topology; including:

[0048] Obtain the server set, take the server as a node, mark each node, and use an empty set as the initially selected first candidate permutation set G AllReduce ={}.

[0049] For the degree of each server node, calculate its Euler's totient function. The result of the Euler's totient function is the number of positive integers relatively prime to the degree, that is, the possible connection combinations between the servers, that is, the different permutation methods between the servers, denoted as the second candidate permutation set P k .

[0050] In response to the number of servers n, the degree d corresponding to each AllReduce group k 、and the second candidate permutation set P corresponding to each AllReduce group k , initialize an empty set as the first candidate permutation G corresponding to the k-th AllReduce group k .

[0051] Take the first element P k in the second candidate permutation set P k [0] and assign it to q. q represents the smallest permutation in the second candidate permutation set P k as the initial selection.

[0052] Obtain the connection description of the permutation q through the GetConn(q) function and add these connections to the first candidate permutation G k as the starting point for constructing the topology corresponding to the k-th AllReduce group.

[0053] Calculate the ratio x of several sequences, , and this ratio x is used to select the next best permutation.

[0054] Determine the new candidate permutation q' by multiplying the previously selected permutation q by the ratio x, so as to generate candidate permutations based on geometric sequences.

[0055] Select the permutation closest to q' from the second set of candidate permutations P k such that the distance between the added permutation and the existing topology is minimized.

[0056] Obtain the connection description of the permutation q' through the GetConn(q) function and add it to the first set of candidate permutations G k .

[0057] Update the current permutation q to the permutation q', complete one iteration, and perform the iteration d k -1 times, each time selecting a new permutation and adding it to the first set of candidate permutations G k .

[0058] Return the finally constructed first set of candidate permutations G k as the sub-topology of the k-th AllReduce group, and merge the topologies corresponding to each AllReduce group to obtain the AllReduce sub-topology, denoted as G AllReduce ={G 1 , G 2 ,,,G k ,,, G K}, where K is the number of AllReduce groups.

[0059] Step S203, construct the MP sub-topology according to the transmission requirements of the MP traffic and the degree.

[0060] Initialize an empty set as the third set of candidate permutations G MP ={};

[0061] According to the transmission requirements of the MP traffic, perform maximum weight matching through the BlossomMaximumWeightMatching algorithm (Edmonds maximum weight matching algorithm), and add the matched edges to the third set of candidate permutations G MP so that the number of hops passed by the data transmission path in the parallel of the distributed large model is minimized;

[0062] Iteratively optimize the third set of candidate permutations d MP -1 times to obtain the MP sub-topology.

[0063] Step S204, combine the AllReduce sub-topology and the MP sub-topology to obtain a topology graph; and configure optical switches according to the topology graph to implement the physical topology (such asFigure 4 as shown

[0064] Furthermore, in this example, the Cion-change idea is also used to calculate the routing rules for the AllReduce topology, and the shortest path algorithm is used to calculate the routing rules for MP. The routes of these two parts are merged to form the final route R

[0065] Step S3: Monitor the traffic changes of all servers during the distributed large model training. When the traffic changes, obtain all distributed large model training task arrays, link arrays, and candidate placement location arrays; where the candidate placement location means placing a certain task on a certain server; traverse each candidate placement location, and for each candidate placement location, construct a corresponding affinity graph according to the topology graph, calculate the compatibility scores of all links in the affinity graph under the condition that the bandwidth requirement satisfies the link capacity constraint, and use the candidate placement location with the highest compatibility score as the best placement location; calculate the latency of all connected subgraphs in the affinity graph corresponding to the best placement location, and use the sum of the latencies of all connected subgraphs as the total latency, with the minimum total latency as the optimization goal; perform distributed large model training on the physical topology according to the best placement location and the total latency

[0066] Specifically, as Figure 3 shown, the step S3 specifically includes the following sub-steps

[0067] Step S301: When the traffic changes, obtain all distributed large model training task arrays Jobs, link arrays Links, and candidate placement location arrays Candidates

[0068] Step S302: Traverse the candidate placement locations. For each candidate placement location c, construct a corresponding affinity graph G c ={U c ,V c ,E c}, as Figure 5 shown, U c represents the task nodes, V c represents the link nodes, and E c represents the edge set. For each task j and link l, check whether task j shares the link with other tasks. If link l carries multiple tasks, add these tasks to the node set V c , if task j is passing through link l, create an edge. If there is a loop in the affinity graph G c , then remove this candidate location

[0069] Step S303: For all links, calculate the compatibility scores of tasks passing through the links through the compatibility optimization formula. The compatibility optimization formula is as follows

[0070]

[0071] Among them, represents the set of jobs competing for a certain link ; represents the bandwidth requirement of the total work task at the rotation angle of (determined according to the route R); represents the total capacity of the link bandwidth, represents a set of discrete angles , whose range is [0, 2π], represents the bandwidth requirement at a certain angle α on the unified circle, represents the number of iterations of the job on its unified circle, represents the job on the link rotation angle (in radians), represents the compatibility score of the shared job link .

[0072] Step S304, sort the candidate positions. For all links, sort the candidate placement positions according to the compatibility score, and select the candidate position with the highest score as the best placement position for the task. Obtain the affinity graph G = {U, V, E} corresponding to the best placement position. Through breadth-first traversal of each connected subgraph H = {U H , V H , E H}, for each subgraph, initialize an empty time_shift H to store the time shift in the subgraph.

[0073] Step S305, calculate the unique time shift. Find the edges of each task and link and , and calculate the final time shift and according to the weights of the edges, and use modulo operation to ensure that the time shift does not exceed the range of the iteration time . Add the calculated time shift to , and then add the time shift of each subgraph to . Based on the best task and link placement combination, and the time offset of the task, with the minimum total delay as the optimization goal.

[0074] Step S4, when the bandwidth requirement does not meet the link capacity constraint, reconstruct the AllReduce sub-topology and the MP sub-topology according to the current traffic change, update the topology graph; and reconfigure the optical switch according to the updated topology graph to implement the physical topology;

[0075] Update According to the optimal placement location and the total delay, continue the distributed large model training on the physical topology according to the updated optimal placement location and the total delay.

[0076] Further, an embodiment of the present invention provides a communication scheduling system for distributed large model training, and the system includes:

[0077] A topology reconstruction component, configured to obtain the number of servers for distributed large model training, the degree of each server, the transmission traffic requirements of AllReduce traffic, and the transmission traffic requirements of MP traffic; divide the degree of each server into the degrees of the AllReduce traffic sub-topology and the MP traffic sub-topology according to the ratio of the transmission requirements of AllReduce traffic and MP traffic; construct the AllReduce sub-topology and the MP sub-topology respectively according to the transmission requirements and degrees of AllReduce traffic and MP traffic; combine the AllReduce sub-topology and the MP sub-topology to obtain a topology graph; configure an optical switch according to the topology graph to implement the physical topology;

[0078] A traffic awareness component, configured to monitor the traffic changes of all servers during the distributed large model training process. When the traffic changes, obtain all distributed large model training task arrays, link arrays, and candidate placement location arrays; where the candidate placement location indicates placing a certain task on a certain server; traverse each candidate placement location, and for each candidate placement location, construct a corresponding affinity graph according to the topology graph, calculate the compatibility scores of all links in the affinity graph under the condition that the bandwidth requirement meets the link capacity constraint, and use the candidate placement location with the highest compatibility score as the optimal placement location; calculate the delays of all connected subgraphs in the affinity graph corresponding to the optimal placement location, and use the sum of the delays of all connected subgraphs as the total delay, with the minimum total delay as the optimization goal; perform distributed large model training on the physical topology according to the optimal placement location and the total delay.

[0079] Regarding the system in the above embodiments, the specific manners in which each module performs operations have been described in detail in the embodiments related to the method, and will not be elaborated here.

[0080] For the system embodiments, since they basically correspond to the method embodiments, the relevant parts can be referred to the partial descriptions of the method embodiments. The system embodiments described above are only illustrative. The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed to multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this application. A person of ordinary skill in the art can understand and implement it without creative efforts.

[0081] Correspondingly, this application also provides an electronic device, including: one or more processors; a memory for storing one or more programs; when the one or more programs are executed by the one or more processors, the one or more processors implement the communication scheduling method for distributed large model training as described above. As Figure 6 shown, it is a hardware structure diagram of any device with data processing capabilities where the communication scheduling method for distributed large model training provided by the embodiment of the present invention is located. In addition to Figure 6 the processors, memory, and network interfaces shown, any device with data processing capabilities where the device in the embodiment is located usually includes other hardware according to the actual functions of the device with data processing capabilities, which will not be elaborated here.

[0082] Correspondingly, this application also provides a computer-readable storage medium, on which computer instructions are stored. When the instructions are executed by a processor, the communication scheduling method for distributed large model training as described above is implemented. The computer-readable storage medium can be an internal storage unit of any device with data processing capabilities described in any of the foregoing embodiments, such as a hard disk or memory. The computer-readable storage medium can also be an external storage device, such as a plug-in hard disk, a smart media card (SMC), an SD card, a flash card, etc. equipped on the device. Further, the computer-readable storage medium can also include both the internal storage unit of any device with data processing capabilities and the external storage device. The computer-readable storage medium is used to store the computer program and other programs and data required by any device with data processing capabilities, and can also be used to temporarily store the data that has been output or will be output.

[0083] The above are only the preferred embodiments of the present invention and are not intended to limit the present invention. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principles of the present invention shall be included within the scope of protection of the present invention.

[0084] The above embodiments are only used to illustrate the design concept and features of the present invention, and the purpose is to enable those skilled in the art to understand the content of the present invention and implement it accordingly. The protection scope of the present invention is not limited to the above embodiments. Therefore, all equivalent changes or modifications made according to the principles and design concepts disclosed by the present invention are within the protection scope of the present invention.

Claims

1. A communication scheduling method for distributed large model training, characterized in that: The method comprises: Obtain the number of servers used for distributed large model training, the degree of each server, the transmission flow requirements of AllReduce traffic, and the transmission flow requirements of MP traffic; The degree of each server is divided into the degree of AllReduce traffic sub-topology and the degree of MP traffic sub-topology according to the ratio of the transmission requirements of AllReduce traffic and MP traffic; the AllReduce sub-topology and the MP sub-topology are constructed respectively according to the transmission requirements and degrees of AllReduce traffic and MP traffic; the AllReduce sub-topology and the MP sub-topology are combined to obtain a topology map; the optical switch is configured according to the topology map to realize the physical topology; Monitor the traffic changes of all servers during the distributed large model training process. When the traffic changes, obtain all distributed large model training task arrays, link arrays, and candidate placement arrays; where the candidate placement indicates placing a task on a certain server; Traverse each candidate placement location, and for each candidate placement location, build a corresponding affinity graph according to the topology graph, calculate the compatibility scores of all links in the affinity graph under the condition that the bandwidth demand meets the link capacity constraint, and take the candidate placement location with the highest compatibility score as the best placement location; calculate the delay of all connected subgraphs in the affinity graph corresponding to the best placement location, and take the sum of the delays of all connected subgraphs as the total delay, with the minimum total delay as the optimization goal; perform distributed large model training on the physical topology based on the best placement location and the total delay.

2. A communication scheduling method for distributed large model training according to claim 1, characterized in that: The process of obtaining the number of servers used for distributed large model training, the degree of each server, the transmission flow requirements of AllReduce traffic, and the transmission flow requirements of MP traffic includes: Obtain the task information, load mode, and parallelization strategy of the distributed large model training task; the task information of the distributed large model training task includes the task type, computing requirements, and data scale; the load modes of the distributed large model training task include: computing intensive and data intensive; the parallelization strategies of the distributed large model training task include: data parallelism and model parallelism; According to the task information, load pattern, and parallelization strategy of the distributed large model training task, the number of servers used for distributed large model training, the degree of each server, the transmission traffic requirements of AllReduce traffic, and the transmission traffic requirements of MP traffic are obtained.

3. A communication scheduling method for distributed large model training according to claim 1, characterized in that: The process of building the AllReduce sub-topology includes: For each AllReduce group, the degree corresponding to each AllReduce group is allocated according to the traffic corresponding to each AllReduce group; The server is used as a node, and an empty set is used as the first candidate permutation set for initialization; the Euler function of the degree corresponding to each server is calculated, and the result output by the Euler function represents different connection modes between the servers; the result output by the Euler function is filled into the second candidate permutation set; In response to the number of servers, the degree corresponding to each AllReduce group, and the second candidate permutation set corresponding to each AllReduce group, a first candidate permutation corresponding to each AllReduce group is constructed, and the first candidate permutation is iteratively optimized by selecting a permutation and combination method to obtain a topology corresponding to each AllReduce group; The topologies corresponding to each AllReduce group are merged to obtain the AllReduce sub-topology.

4. A communication scheduling method for distributed large model training according to claim 1, characterized in that: The process of building an MP sub-topology includes: Initialize an empty set as the third candidate permutation set; According to the transmission requirements of MP traffic, the maximum weight matching is performed through the Edmonds maximum weight matching algorithm, and the matched edges are added to the third candidate permutation set to minimize the number of hops of the data transmission path in the distributed large model parallel; The third candidate permutation set is iteratively optimized to obtain the MP sub-topology.

5. A communication scheduling method for distributed large model training according to claim 1, characterized in that: The process of calculating the compatibility scores of all links in the affinity graph under the condition that the bandwidth demand meets the link capacity constraint includes: When the bandwidth requirement of all distributed large model training tasks at a rotation angle of α is greater than the total capacity of the current link bandwidth, the compatibility score of the current link is the difference between the bandwidth requirement of all distributed large model training tasks at a rotation angle of α and the total capacity of the current link bandwidth; otherwise, the compatibility score of the current link is 0; Among them, the link capacity constraint is: the bandwidth requirement of all distributed large model training tasks at each angle α is less than or equal to the bandwidth requirement of all distributed large model training tasks when the rotation angle is α, and the ratio of the rotation angle of the jth job on the current link is greater than or equal to 0 and less than or equal to 2π to the number of iterations of the jth job on its unified circle.

6. A communication scheduling method for distributed large model training according to claim 1, characterized in that: The method further comprises: When the bandwidth demand does not meet the link capacity constraint, the AllReduce sub-topology and MP sub-topology are reconstructed according to the current traffic changes, and the topology map is updated; and the optical switch is reconfigured according to the updated topology map to implement the physical topology; The best placement location and total latency are updated, and distributed large model training continues on the physical topology based on the updated best placement location and total latency.

7. A communication scheduling system for distributed large model training, characterized in that: The system comprises: The topology reconstruction component is used to obtain the number of servers used for distributed large model training, the degree of each server, the transmission flow requirements of AllReduce traffic, and the transmission flow requirements of MP traffic; the degree of each server is divided into the degree of AllReduce traffic sub-topology and the degree of MP traffic sub-topology according to the ratio of the transmission requirements of AllReduce traffic and MP traffic; the AllReduce sub-topology and the MP sub-topology are constructed according to the transmission requirements and degrees of AllReduce traffic and MP traffic respectively; the AllReduce sub-topology and the MP sub-topology are combined to obtain a topology map; the optical switch is configured according to the topology map to realize the physical topology; The traffic perception component is used to monitor the traffic changes of all servers during the distributed large model training process. When the traffic changes, all distributed large model training task arrays, link arrays, and candidate placement arrays are obtained; among which, the candidate placement position means placing a certain task on a certain server; each candidate placement position is traversed, and for each candidate placement position, the corresponding affinity graph is constructed according to the topology graph, and the compatibility scores of all links in the affinity graph are calculated under the condition that the bandwidth demand meets the link capacity constraint. The candidate placement position with the highest compatibility score is taken as the optimal placement position; the delay of all connected subgraphs in the affinity graph corresponding to the optimal placement position is calculated, and the sum of the delays of all connected subgraphs is taken as the total delay, with the minimum total delay as the optimization goal; the distributed large model training is performed on the physical topology according to the optimal placement position and the total delay.

8. An electronic device, comprising a memory and a processor, characterized in that: The memory is coupled to the processor; wherein the memory is used to store program data, and the processor is used to execute the program data to implement the communication scheduling method for distributed large model training as described in any one of claims 1-6 above.

9. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the program is executed by the processor, it implements the communication scheduling method for distributed large model training as described in any one of claims 1-6.

10. A computer program product comprising a computer program / instructions, characterized in that When the computer program / instruction is executed by a processor, it implements the communication scheduling method for distributed large model training as described in any one of claims 1-6.

Citation Information

Patent Citations

  • Distributed cloud traffic scheduling method and device based on topology awareness

    CN117135128A

  • Distributed machine learning training system based on adaptive topology and auxiliary routing

    CN117474081A

  • Edge calculation method and device for distributed large model pipeline parallel training

    CN119829254A