A Collective Communication Control Method, Device and Medium for Distributed Training

By reducing the data in the first cluster to a specified computing node whose number is aligned with the number of computing nodes in the second cluster in distributed training, the asymmetry of the number of computing nodes caused by different AI chip computing power is solved, which reduces the overhead of the set communication and improves the cross-cluster reduction efficiency.

CN119336451BActive Publication Date: 2025-06-13ZHEJIANG LAB
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411863321.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-12-17
Publication Date
2025-06-13
Estimated Expiration
2044-12-17

AI Technical Summary

Technical Problem

In distributed training, due to the different computing power of different AI chips, the number of computing nodes participating in hybrid training between computing clusters is different, and network competition occurs, resulting in high set communication overhead and low cross-cluster reduction efficiency, reducing the efficiency of distributed training.

Method used

When the difference in the number of computing nodes between any two clusters is within a preset range in the cluster participating in the data reduction, the data in the first cluster is reduced to the designated computing node, and the number of specified computing nodes is the same as the number of computing nodes in the second cluster, and the second cluster is the cluster with the smallest number of computing nodes. Then, the designated computing node reduces data with the computing nodes in the second cluster, and the designated computing nodes are controlled to synchronize data with other nodes in the first cluster except the designated computing nodes.

Benefits of technology

By reducing data to a designated computing node with aligned number, the computing nodes are avoided from simultaneously reducing data with multiple computing nodes, reducing network competition, reducing set communication overhead, and improving the efficiency of cross-cluster reduction.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119336451B_ABST
    Figure CN119336451B_ABST
Patent Text Reader

Abstract

The present application discloses a collective communication control method, device and medium for distributed training. The method includes: when the difference in the number of computing nodes between any two clusters in a cluster participating in data reduction is within a preset range, reducing the data on all computing nodes in the first cluster to a specified computing node, and the number of specified computing nodes is the same as the number of computing nodes in the second cluster with the smallest number of computing nodes. Controlling the specified computing node to perform data reduction with the computing nodes in the second cluster; controlling the specified computing node to synchronize data with other nodes in the first cluster except the specified computing node. Thus, except for the cluster with the smallest number of computing nodes, other clusters first perform a reduction within the cluster, reducing the data to the specified computing nodes with the same number as the smallest number of nodes in each cluster, ensuring that the number of nodes in each cluster is the same during cross-cluster reduction, avoiding some nodes from performing reduction with multiple nodes simultaneously, and reducing the collective communication overhead.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of computer technology, and in particular, to a collective communication control method, apparatus, and medium for distributed training. Background Art

[0002] With the rapid development of big data and artificial intelligence (AI) models, distributed training technology has been widely applied to AI models to jointly complete complex computing tasks of AI models through the collaboration of multiple computing nodes in a computing cluster, that is, to decompose the tasks originally completed on a single computing node to multiple computing nodes to collaboratively complete complex computing tasks on the AI scale.

[0003] Collective communication refers to a specific form of communication between a group of processes or computing nodes. In the distributed training of AI models, optimizing the performance of collective communication is the main technical means to improve the efficiency of distributed training, and improving the efficiency of distributed training is crucial for ensuring that each computing node in the cluster fully exerts its computing performance. Currently, most collective communications are constructed based on homogeneous networks. When optimizing the performance of collective communication in a homogeneous network, data is evenly divided among computing nodes, and the data is reduced to each computing node through collective communication semantics to achieve the optimal communication bandwidth, thereby improving the efficiency of distributed training.

[0004] However, in order to meet the explosive growth of intelligent computing power requirements of AI models, various AI chips are continuously iterated and updated, resulting in a situation where different AI chips are mixedly deployed in AI models. Since different AI chips have different computing powers, the number of computing nodes participating in hybrid training between different computing clusters is different. Therefore, when performing data reduction across clusters, due to the asymmetry of the number of computing nodes between the clusters participating in the reduction, some computing nodes need to perform data reduction with multiple computing nodes simultaneously, resulting in network competition, high collective communication overhead, low cross-cluster reduction efficiency, and reduced distributed training efficiency.

[0005] Therefore, it can be seen that how to solve the problems of high collective communication overhead, low cross-cluster reduction efficiency, and improve the efficiency of distributed training is an urgent problem to be solved by those skilled in the art. Summary of the Invention

[0006] In view of this, one aspect of this application provides a collective communication control method for distributed training, and the method includes:

[0007] When the difference in the number of computing nodes between any two clusters in the cluster participating in data reduction is within a preset range, reduce the data distributed on all computing nodes in the first cluster to the specified computing nodes in the first cluster; where the number of the specified computing nodes is the same as the number of computing nodes in the second cluster, and the second cluster is the cluster with the smallest number of computing nodes among the clusters participating in data reduction.

[0008] Control the specified computing nodes and the computing nodes in the second cluster to perform data reduction.

[0009] Control the specified computing nodes and other nodes in the first cluster except the specified computing nodes to perform data synchronization.

[0010] Optionally, the reducing the data distributed on all computing nodes in the first cluster to the specified computing nodes in the first cluster includes:

[0011] Divide the computing nodes in the first cluster into m node groups; where each node group has k computing nodes, and k = floor(n / m), n is the number of computing nodes in the first cluster, and m is the number of computing nodes in the second cluster.

[0012] Select one computing node from each node group as the specified computing node.

[0013] Reduce the data distributed on all computing nodes in the first cluster to the specified computing nodes.

[0014] Optionally, the selecting one computing node from each node group as the specified computing node includes:

[0015] Sort the computing nodes in the first cluster in ascending order of device numbers.

[0016] Based on the sorting result, divide the k computing nodes into each node group in the way of dividing one computing node for each node group in turn.

[0017] Take the computing node with the smallest device number in each node group as the specified computing node.

[0018] Optionally, the reducing the data distributed on all computing nodes in the first cluster to the specified computing nodes includes:

[0019] Determine whether there are target computing nodes in the first cluster that have not been grouped.

[0020] If not, reduce the data within each of the node groups to the designated computing nodes within the corresponding group;

[0021] If so, execute the step of reducing the data within each of the node groups to the designated computing nodes within the corresponding group; and evenly divide the data on the target computing node into m portions of data, and send the m portions of data to different designated computing nodes in a one-to-one correspondence.

[0022] Optionally, the collective communication control method for distributed training further includes:

[0023] When the gap is not within the preset range, sort in ascending order according to the number of computing nodes;

[0024] Based on the sorting result, merge the first-ranked cluster and the second-ranked cluster into a target cluster;

[0025] Determine whether the gap between any two clusters among the target cluster and the unmerged clusters is within the preset range;

[0026] If so, use the target cluster and the unmerged clusters as the clusters participating in data reduction; and enter the step of reducing the data distributed on all computing nodes within the first cluster to the designated computing node within the first cluster, and execute the subsequent steps.

[0027] Optionally, if the gap between any two clusters among the target cluster and the unmerged clusters is not within the preset range, the method further includes:

[0028] Determine whether the difference between the number of computing nodes in the target cluster and the number of computing nodes in the designated cluster exceeds a threshold; wherein, the designated cluster is the cluster with the smallest number of computing nodes among the unmerged clusters;

[0029] If it exceeds the threshold, enter the step of using the target cluster and the unmerged clusters as the clusters participating in data reduction, and execute the subsequent steps;

[0030] If it does not exceed the threshold, determine whether the number of unmerged clusters is greater than 1; if it is greater, enter the step of sorting in ascending order according to the number of computing nodes, and execute the subsequent steps; if it is not greater, enter the step of using the target cluster and the unmerged clusters as the clusters participating in data reduction, and execute the subsequent steps.

[0031] Optionally, the gap between the numbers of computing nodes among the clusters participating in data reduction being within the preset range includes:

[0032] If the result of the specified calculation on the maximum and minimum values of the number of computing nodes in the cluster participating in data reduction is not greater than a preset value, then the gap is within the preset range; wherein, the specified calculation includes ratio calculation or difference calculation.

[0033] Another aspect of the present application provides a collective communication control device for distributed training, the device includes:

[0034] An intra-cluster reduction module, configured to reduce the data distributed on all computing nodes in the first cluster to the specified computing nodes in the first cluster when the gap between the number of computing nodes between any two clusters in the cluster participating in data reduction is within a preset range; wherein, the number of the specified computing nodes is the same as the number of computing nodes in the second cluster, and the second cluster is the cluster with the smallest number of computing nodes among the clusters participating in data reduction;

[0035] An inter-cluster reduction module, configured to control the specified computing nodes and the computing nodes in the second cluster to perform data reduction;

[0036] A data synchronization module, configured to control the specified computing nodes and other nodes in the first cluster except the specified computing nodes to perform data synchronization.

[0037] Another aspect of the present application provides a collective communication control device for distributed training, including a memory and a processor, where a computer program that can run on the processor is stored on the memory, and when the processor executes the program, the steps of the collective communication control method for distributed training are implemented.

[0038] Another aspect of the present application provides a computer-readable storage medium, on which a computer program is stored, and when the program is executed by a processor, the steps of the collective communication control method for distributed training are implemented.

[0039] The beneficial effects of the collective communication control method, device and medium for distributed training provided by the present application are as follows: In distributed training, when performing cross-cluster data reduction, if the number of computing nodes between clusters is different, except for the cluster with the smallest number of computing nodes, other clusters first perform a data reduction within the cluster to reduce the data to the specified computing nodes with the same number as the smallest number of computing nodes in each cluster, so as to ensure that when performing cross-cluster reduction, the number of computing nodes between different clusters is aligned, that is, the number of computing nodes is the same, avoiding the phenomenon that some computing nodes need to perform data reduction with multiple computing nodes at the same time, reducing the network competition, reducing the overhead of collective communication, and improving the efficiency of cross-cluster reduction. Description of the Drawings

[0040] Figure 1Schematic flowchart of a collective communication control method for distributed training provided by an embodiment of the present application;

[0041] Figure 2 Schematic principle diagram of a collective communication control method for distributed training provided by an embodiment of the present application;

[0042] Figure 3 Schematic flowchart of a collective communication control method for distributed training provided by another embodiment of the present application;

[0043] Figure 4 Schematic principle diagram of a collective communication control method for distributed training provided by another embodiment of the present application;

[0044] Figure 5 Schematic principle diagram of a collective communication control method for distributed training provided by still another embodiment of the present application;

[0045] Figure 6 Schematic flowchart of a collective communication control method for distributed training provided by still another embodiment of the present application;

[0046] Figure 7 Schematic structural diagram of a collective communication control device for distributed training provided by an embodiment of the present application;

[0047] Figure 8 Schematic structural diagram of a collective communication control device for distributed training provided by another embodiment of the present application.

[0048] Reference numerals are as follows: 80 is a memory, 81 is a processor, 82 is a display screen, 83 is an input / output interface, 84 is a communication interface, 85 is a power supply, 86 is a communication bus, 801 is a computer program, 802 is an operating system, and 803 is data. Detailed implementation manners

[0049] The terms used in the present application are only for the purpose of describing specific embodiments and are not intended to limit the present application. The singular forms "a", "the", and "said" used in the present application and the appended claims are also intended to include the plural forms unless the context clearly dictates otherwise. It should also be understood that the term "and / or" as used herein refers to and encompasses any and all possible combinations of one or more of the associated listed items.

[0050] It should be understood that although the terms first, second, third, etc. may be used in this application to describe various information, such information should not be limited to these terms. These terms are only used to distinguish information of the same type from each other. For example, without departing from the scope of this application, the first information may also be referred to as the second information, and similarly, the second information may also be referred to as the first information. Depending on the context, the word "if" as used herein may be interpreted as "when" or "while" or "in response to a determination".

[0051] Figure 1 The flowchart of a collective communication control method for distributed training provided by an embodiment of this application is as Figure 1 shown, and the method includes:

[0052] S10: When the difference in the number of computing nodes between any two clusters in the clusters participating in data reduction is within a preset range, reduce the data distributed on all computing nodes in the first cluster to the specified computing nodes in the first cluster; where the number of specified computing nodes is the same as the number of computing nodes in the second cluster, and the second cluster is the cluster with the smallest number of computing nodes among the clusters participating in data reduction;

[0053] It can be understood that in the distributed training with mixed deployment of different AI chips, due to the different computing powers of different AI chips, the number of computing nodes in the clusters participating in data reduction is different. For example, in a deep learning model for a complex image recognition training task, two types of AI chips are required to complete it together, including chip A and chip B, where the intelligent computing power of chip A is higher than that of chip B. Therefore, 4 chip A are deployed in cluster A and 10 chip B are deployed in cluster B to jointly complete the image recognition task of the deep learning model. Obviously, the number of AI chips in cluster A and cluster B is different, that is, the number of computing nodes is different.

[0054] In distributed training, communication within the same computing node can use high-speed interconnection technologies such as NVLink for communication, and different computing nodes within the same cluster can use high-speed interconnection technologies such as GDR (GPU Direct RDMA) for communication, while only relatively low-speed RoCE (RDMA over Converged Ethernet) interconnection technology can be used for communication between different clusters.

[0055] It can be seen that in the hybrid distributed training scenario with heterogeneous communication links, when performing collective communication optimization to improve the efficiency of distributed training, due to the different number of computing nodes between the clusters participating in data reduction, some computing nodes may need to perform data reduction with multiple computing nodes simultaneously when participating in data reduction, resulting in network competition and affecting the collective communication optimization effect.

[0056] Because, to solve the above technical problems, in an alternative embodiment, the number of computing nodes between the clusters participating in data reduction is aligned, that is, the computing nodes between the clusters participating in data reduction are the same.

[0057] Specifically, the data distributed on all computing nodes within the first cluster is reduced to the specified computing nodes within the first cluster, where the number of specified computing nodes is the same as the number of computing nodes within the second cluster, and the second cluster is the cluster with the smallest number of computing nodes among the clusters participating in data reduction. That is to say, the number of computing nodes of all clusters participating in data reduction is aligned with the cluster with the smallest number of computing nodes. For ease of understanding, an example will be given below.

[0058] For example, the clusters participating in data reduction include cluster A, cluster B, and cluster C, and the number of computing nodes in cluster A is 15, the number of computing nodes in cluster B is 20, and the number of computing nodes in cluster C is 23. Obviously, in this example, the first clusters are cluster B and cluster C, and the second cluster with the smallest number of computing nodes is cluster A, and the number of specified computing nodes should be aligned with the number of computing nodes in cluster A, that is, the number of specified computing nodes is 15.

[0059] When performing cross-cluster data reduction among cluster A, cluster B, and cluster C, to ensure that the number of computing nodes between the three clusters is the same, first control cluster B and cluster C within their respective clusters to reduce the data distributed on all computing nodes to 15 specified computing nodes.

[0060] S11: Control the specified computing nodes to perform data reduction with the computing nodes within the second cluster;

[0061] Furthermore, after the number of computing nodes of each cluster is aligned, perform cross-cluster reduction. Specifically, control the specified computing nodes to perform data reduction with the computing nodes within the second cluster. For example, in the above example, control the 15 specified computing nodes on cluster B and cluster C to perform pairwise cluster data reduction with the 15 computing nodes in cluster A. That is, it includes the reduction between cluster B and cluster A, cluster B and cluster C, and cluster A and cluster C.

[0062] It should be noted that if the difference in the number of computing nodes between any two clusters among the clusters participating in data reduction is large, it will lead to the phenomenon of insufficient utilization of the communication link and waste of resources. For example, if there are 5 computing nodes in cluster A and 50 computing nodes in cluster B, then during cross-cluster data reduction, cluster B first reduces the data to 5 specified computing nodes, and then uses these 5 computing nodes to perform data reduction with the 5 computing nodes in cluster A.

[0063] Obviously, this will result in many communication links being idle during cross-cluster reduction. Moreover, when reducing 5 pieces of data, due to the large amount of data, the efficiency is relatively low. Therefore, in an alternative embodiment, before reducing the data distributed on all computing nodes within the first cluster to the specified computing node within the first cluster, it is first determined whether the difference in the number of computing nodes between any two clusters in the clusters participating in the data reduction is within a preset range.

[0064] That is to say, before the cluster with more computing nodes performs internal data reduction, it first determines whether the number of computing nodes among all the clusters participating in the data reduction is relatively uniform, that is, the difference in the number of computing nodes should not be too large.

[0065] It should be noted that when determining whether the difference in the number of computing nodes between any two clusters is within a preset range, it can be judged based on the difference in the number of computing nodes between any two clusters, or determined based on the ratio of the number of computing nodes between any two clusters. There are also other ways to determine whether the number of computing nodes between any two clusters is relatively uniform. This application does not make any limitations in this regard.

[0066] Figure 2 FIG. is a schematic diagram of the principle of a collective communication control method for distributed training provided by an embodiment of this application. In an alternative embodiment, when the number of computing nodes among the clusters participating in the data reduction is the same, such as Figure 2 As shown, the number of computing nodes in each of the N clusters from cluster 1 to cluster N is 4. Then, when performing cross-cluster data reduction, the data is evenly split among the computing nodes within each cluster, and each computing node is only responsible for reducing its own corresponding small part of the data. Further, then, among the computing nodes, data reduction is completed among the computing nodes within the cluster through the collective communication semantics method of Reduce Scatter, that is, data exchange is completed. Then, cross-cluster reduction is performed to complete the data reduction between clusters.

[0067] S12: Control the specified computing node to synchronize data with other nodes in the first cluster except the specified computing node.

[0068] Further, after completing the cross-cluster data reduction, step S12 is used to control the specified computing node to synchronize data with other nodes in the first cluster except the specified computing node. For example, in the above example, among the 20 computing nodes in cluster B, after 15 specified computing nodes complete data reduction with cluster A and cluster C, these 15 specified computing nodes need to synchronize the data to the other 5 computing nodes that have not performed cross-cluster data reduction. Similarly, cluster C also needs to perform the same data synchronization operation. Since all the computing nodes in cluster A have participated in the cross-cluster data reduction, there is no need to perform data synchronization.

[0069] Thus, in the collective communication control method for distributed training provided by the embodiments of the present application, during distributed training, when performing cross-cluster data reduction, if the number of computing nodes between clusters is different, for clusters other than the cluster with the smallest number of computing nodes, a data reduction is first performed within the cluster to reduce the data to the specified computing nodes with the same number as the smallest number of computing nodes in each cluster, so as to ensure that when performing cross-cluster reduction, the number of computing nodes between different clusters is aligned, that is, the number of computing nodes is the same, avoiding the phenomenon that some computing nodes need to perform data reduction with multiple computing nodes simultaneously, reducing the overhead of collective communication, and improving the efficiency of cross-cluster reduction.

[0070] Figure 3 FIG. is a schematic flowchart of a collective communication control method for distributed training provided by another embodiment of the present application. In an alternative embodiment, as Figure 3 shown, reducing the data distributed on all computing nodes in the first cluster to the specified computing nodes in the first cluster includes:

[0071] S30: Divide the computing nodes in the first cluster into m node groups; where each node group has k computing nodes, and k = floor(n / m), n is the number of computing nodes in the first cluster, and m is the number of computing nodes in the second cluster;

[0072] In an alternative embodiment, when selecting the specified computing nodes in the first cluster, they can be selected according to certain rules. Specifically, first divide the computing nodes in the first cluster according to the number of computing nodes in the second cluster. Therefore, first count the number of computing nodes in the first cluster and the second cluster, record the number of computing nodes in the first cluster as n, and record the number of computing nodes in the second cluster as m. Obviously, n > m.

[0073] Furthermore, divide the n computing nodes in the first cluster into m node groups, and each node group includes k computing nodes. To ensure that the number of computing nodes in each node group is the same and that each node group has an integer number of computing nodes. Therefore, k = floor(n / m), that is, the number of computing nodes in each node group is equal to the number of computing nodes in the first cluster divided by the number of computing nodes in the second cluster and then rounded down. For example, when n = 10 and m = 3, k = floor(n / m) = 3.

[0074] S31: Select one computing node from each node group as the specified computing node;

[0075] S32: Reduce the data distributed on all computing nodes in the first cluster to the specified computing nodes.

[0076] Further, a computing node is selected from each node group as a designated computing node. When selecting, it can be randomly selected or selected according to the device serial number of the computing node, and the present application does not make any limitation on this.

[0077] It can be understood that the computing nodes in the first cluster are divided into m node groups, and a designated computing node is selected from each node group. Eventually, m designated computing nodes can be obtained. Furthermore, the data distributed on all computing nodes within the first cluster can be reduced to the designated computing nodes selected in this embodiment.

[0078] Thus, in the collective communication control method for distributed training provided by the embodiments of the present application, the number of computing nodes between the clusters participating in data reduction is aligned, that is, aligned with the cluster having the smallest number of computing nodes. This avoids the situation where some computing nodes need to perform reduction with multiple computing nodes simultaneously when performing data reduction across clusters, optimizes collective communication, and improves the efficiency of distributed training.

[0079] As an optional embodiment, selecting a computing node from each node group as a designated computing node includes:

[0080] Sorting the computing nodes in the first cluster in ascending order according to their device serial numbers;

[0081] Based on the sorting result, in the way of dividing one computing node into each node group at a time, k computing nodes are sequentially divided into each node group;

[0082] Taking the computing node with the smallest device serial number in each node group as the designated computing node.

[0083] It can be understood that different computing nodes have different device serial numbers. Computing nodes with the same device serial number can communicate using one switch, while computing nodes with different device serial numbers require multiple switches for communication.

[0084] Therefore, in an optional embodiment, in order to save resources, when selecting a computing node from each node group as a designated computing node, the computing node with the smallest device serial number in each node group is taken as the designated computing node.

[0085] Specifically, first sort the device serial numbers of each computing node in the first cluster in ascending order. After sorting, in the way of dividing one computing node into each node group at a time, k computing nodes are sequentially divided into each node group. That is to say, after sorting in ascending order, one computing node is placed into m node groups one by one, and the node group where the currently placed computing node is located is different from the node group of the previously placed computing node.

[0086] It can be understood that since the number of computing nodes in the second cluster is the smallest, the maximum value in the device numbers of the computing nodes is less than the maximum value of the device numbers in the first cluster. Therefore, in order to align with the device numbers in the second cluster and save the device resources of the switch, the computing node with the smallest device number in each node group is used as the designated computing node. Figure 4 FIG. Figure 4 is a schematic diagram of the principle of a collective communication control method for distributed training provided by another embodiment of the present application. For ease of understanding, an example will be given below.

[0087] For example, as Figure 4 shown, the clusters participating in data reduction include cluster A and cluster B, and the number of computing nodes in cluster A is 7, and the number of computing nodes in cluster B is 3. Obviously, in this example, the first cluster is cluster A, and the second cluster with the smallest number of computing nodes is cluster B, and the number of designated computing nodes is 3. The device numbers of the computing nodes in cluster B are from 0 to 2, the device numbers of the computing nodes in cluster A are from 0 to 6, and the number of designated computing nodes is 3, that is, the number of node groups m = 3, and the number of computing nodes in each node group k = 2.

[0088] When determining the designated computing nodes in cluster A, as Figure 4 shown, the 7 computing nodes in cluster A are divided into 3 node groups including node group A1, node group A2, and node group A3. The 7 computing nodes are sorted in ascending order according to their device numbers.

[0089] Further, according to the sorting result, the 7 computing nodes are sequentially divided into 3 node groups, one computing node is divided each time, and the node groups divided each time are different, that is, the node groups where the computing nodes divided before and after are located are different. Thus, as Figure 4 shown, the computing nodes with device numbers 0 and 3 are in node group A1, the computing nodes with device numbers 1 and 4 are in node group A2, and the computing nodes with device numbers 2 and 5 are in node group A5.

[0090] In order to ensure that the device numbers of the designated computing nodes selected in the 3 node groups are aligned with the device numbers in cluster B, therefore, the computing node with the smallest device number in each node group is used as the designated computing node, that is, the computing node with device number 0 in node group A1, the computing node with device number 1 in node group A2, and the computing node with device number 2 in node group A3 are used as the designated computing nodes.

[0091] Further, in an alternative embodiment, when controlling the specified computing node to perform data reduction with the computing nodes in the second cluster, the computing nodes in the specified computing node with the same device serial number as the computing nodes in the second cluster are subjected to corresponding data reduction. For example, in the above example, the computing node with device serial number 0 in node group A1 is reduced with the computing node with device serial number 0 in cluster B, and the same applies to other computing nodes.

[0092] Thus, in the collective communication control method for distributed training provided by the embodiments of the present application, the computing node with the smallest device serial number in each node group is selected as the specified computing node to align with the device serial numbers of the computing nodes in the second cluster, optimizing the collective communication to improve the distributed efficiency while saving device resources.

[0093] In an alternative embodiment, reducing the data distributed on all computing nodes in the first cluster to the specified computing node includes:

[0094] Determining whether there are target computing nodes in the first cluster that have not been grouped;

[0095] If not, reducing the data in each node group to the specified computing node in the corresponding group;

[0096] If so, performing the step of reducing the data in each node group to the specified computing node in the corresponding group; and evenly dividing the data on the target computing node into m parts of data, and sending the m parts of data to different specified computing nodes in a one-to-one correspondence.

[0097] It can be understood that when grouping the computing nodes in the first cluster, it may not be possible to divide all nodes into node groups. Among them, the number t of target computing nodes that have not been grouped is t = n - k * m. For example, in the Figure 4 example shown, after dividing the 7 clusters in cluster A into 3 node groups, there is still 1 computing node that cannot be divided into a node group.

[0098] At this time, in an alternative embodiment, when reducing the data distributed on all computing nodes in the first cluster to the specified computing node, it is necessary to first determine whether there are target computing nodes in the current first cluster that have not been grouped.

[0099] If so, in order to ensure that all data in the first cluster participates in the reduction during cross-cluster reduction, therefore, the data on the target computing node that has not participated in the grouping is divided into m parts of data, that is, the number of divided parts is the same as the number of specified computing nodes. Further, the m parts of data are sent to different specified computing nodes in a one-to-one correspondence. For ease of understanding, the following will be combined with Figure 4 for illustration.

[0100] As shown Figure 4 in the figure, the target computing node participating in the grouping is the node with device number 6, and the shaded part in different computing nodes is used to represent the data in the current computing node. For example, the data on the target computing node with device number 6 is evenly divided into 3 parts, including data P61, data P62, and data P63. Further, the data on the target computing node is sent (Pend) to different specified computing nodes in a one-to-one correspondence. That is, data P61 is sent to the computing node with device number 0, data P62 is sent to the computing node with device number 1, and data P63 is sent to the computing node with device number 2.

[0101] Of course, it should be noted that the correspondence relationship between the m parts of data divided by the target computing node and the specified computing node is not limited in this application, as long as different data is sent to different specified computing nodes.

[0102] In a specific embodiment, while the target computing node sends data to the specified computing node, the data within each node group also needs to be reduced to the specified computing node within the corresponding group through the collective communication semantics All Reduce.

[0103] For example, as shown Figure 4 in the figure, the data P3 on the computing node with device number 3 is reduced to the specified computing node with device number 0 within the corresponding group through the collective communication semantics AllReduce, the data P3 on the computing node with device number 4 is reduced to the specified computing node with device number 1 within the corresponding group, and the data P3 on the computing node with device number 5 is reduced to the specified computing node with device number 2 within the corresponding group.

[0104] It should be noted that the step of reducing the data within each node group to the specified computing node within the corresponding group and the step of evenly dividing the data on the target computing node into m parts of data and sending the m parts of data to different specified computing nodes in a one-to-one correspondence can be carried out simultaneously or in a sequential order, which is not limited in this application.

[0105] In another alternative embodiment, if there is no target computing node that has not been grouped among the computing nodes in the first cluster, it means that the number of computing nodes in the current first cluster can be evenly divided into m node groups. At this time, the data within each node group can be directly reduced to the specified computing node within the corresponding group.

[0106] Figure 5It is a schematic diagram of the principle of a collective communication control method for distributed training provided by another embodiment of the present application. Further, in an alternative embodiment, after data reduction is completed within the cluster, that is, when the data in the node group is reduced to the designated computing node, and the target computing nodes that have not performed data reduction also evenly send the data to each designated computing node, as Figure 4 shown, data reduction is performed between the designated computing nodes of each cluster through the collective communication semantics of ReduceScatter.

[0107] Further, as Figure 5 shown, after cross-cluster data reduction, data reduction is performed between the designated computing nodes corresponding to different node groups in Cluster A through the collective communication semantics of All Gather. Further, as Figure 5 shown, within the same node group, the designated computing node broadcasts the data to other computing nodes in the node group through the collective communication semantics of Broadcast.

[0108] In another alternative embodiment, when there are target computing nodes in the computing nodes within the first cluster that have not been grouped, as Figure 5 shown, after the designated computing node completes cross-cluster reduction, in addition to performing data reduction with other designated computing nodes within the node and broadcasting the data to other computing nodes in the same node group, the designated computing node also needs to broadcast the data to the target computing nodes that have not been grouped through the collective communication semantics of Broadcast, thus completing cross-cluster data reduction.

[0109] Figure 6 It is a flowchart of a collective communication control method for distributed training provided by another embodiment of the present application. As an alternative embodiment, as Figure 6 shown, for the collective communication control method for distributed training provided by the present application, when the difference is not within the preset range, it further includes:

[0110] S60: Sort in ascending order according to the number of computing nodes;

[0111] S61: Based on the sorting result, merge the first-ranked cluster and the second-ranked cluster into a target cluster;

[0112] In a specific embodiment, if the number of computing nodes between any two clusters participating in data reduction in a cluster varies greatly, it will lead to insufficient utilization of communication links and waste of resources. Therefore, to solve this technical problem, in an alternative embodiment, when the gap is not within the preset range, the clusters are sorted in ascending order of the number of computing nodes, and the first-ranked cluster and the second-ranked cluster are merged into a target cluster, that is, the cluster with the smallest number and the cluster with the second-smallest number are merged to obtain the target cluster.

[0113] S62: Determine whether the gap between any two clusters among the target cluster and the unmerged clusters is within the preset range; if so, proceed to step S63.

[0114] S63: Use the target cluster and the unmerged clusters as the clusters participating in data reduction; and proceed to step S10 in the above embodiment, and execute subsequent steps S11 and S12.

[0115] Furthermore, after the target cluster is obtained by merging, determine whether the gap between any two clusters among the target cluster and the unmerged clusters is within the preset range, that is, determine whether the number of computing nodes among the clusters is relatively uniform after the merge. If the gap in the number of computing nodes among the clusters after the merge is within an acceptable range, that is, relatively uniform, then cross-cluster data reduction can be performed, and step S63 is executed at this time.

[0116] Specifically, use the target cluster and each of the unmerged clusters as the new clusters participating in data reduction, and implement cross-cluster reduction through steps S10 to S12 in the above embodiment. For ease of understanding, an example will be given below.

[0117] For example, the clusters currently participating in data reduction include cluster 1, cluster 2, and cluster 3, and the number of computing nodes in cluster 1 is 5, the number of computing nodes in cluster 2 is 50, and the number of computing nodes in cluster 3 is 60. Obviously, when data reduction is performed on cluster 1, cluster 2, and cluster 3, cluster 2 and cluster 3 need to first reduce the data to 5 specified computing nodes within the cluster, and when cross-cluster reduction is performed, some communication links will be idle, and the computing nodes participating in the reduction will have low reduction efficiency due to the large amount of data.

[0118] Therefore, in the technical solution provided in the embodiment of the present application, since the number of computing nodes in the three clusters varies greatly, at this time, after sorting the number of computing nodes in the three clusters, cluster 1 and cluster 2 are merged to obtain a target cluster, and the obtained target cluster includes 55 computing nodes.

[0119] Further, taking both the target cluster and Cluster 3 as new clusters participating in data reduction, obviously, the difference in the number of computing nodes between 55 computing nodes and 60 computing nodes is relatively small, which can avoid a large number of idle links during cross-cluster reduction.

[0120] Therefore, in the collective communication control method for distributed training provided by the embodiments of the present application, when there is a large difference in the number of computing nodes between the clusters participating in data reduction, the clusters with the smallest and the second smallest number are merged to avoid a large number of idle links during cross-cluster, further optimizing the collective communication performance, thereby improving the distributed efficiency.

[0121] Based on the above embodiments, as an optional embodiment, as Figure 6 shown, if the difference between any two clusters in the target cluster and the unmerged clusters is not within the preset range, the method further includes:

[0122] S64: Determine whether the difference between the number of computing nodes in the target cluster and the number of computing nodes in the specified cluster exceeds a threshold; where the specified cluster is the cluster with the smallest number of computing nodes among the unmerged clusters; if it exceeds the threshold, go to step S63; if it does not exceed the threshold, go to step S65.

[0123] Based on the above embodiments, after merging the clusters with the smallest and the second smallest number to obtain the target cluster, if the difference in the number of computing nodes between any two clusters in the target cluster and the unmerged clusters is still large. That is to say, after one merge, the difference in the number of computing nodes between the clusters is still large.

[0124] At this time, it is necessary to first judge through step S64 whether the reason for this is that the difference between the number of computing nodes in the target cluster and the number of computing nodes in the specified cluster exceeds the threshold. That is to say, judge whether the large difference between the current clusters is caused by the excessive number of computing nodes in the merged target cluster.

[0125] If so, that is, the difference between the number of computing nodes in the target cluster and the number of computing nodes in the specified cluster exceeds the threshold, at this time, stop merging the clusters and return to step S63, that is, use the currently merged target cluster and the unmerged clusters as new clusters participating in reduction for data reduction.

[0126] S65: Determine whether the number of unmerged clusters is greater than 1; if it is greater, go to step S60 and execute the subsequent steps; if it is not greater, go to the step of using the target cluster and the unmerged clusters as the clusters participating in data reduction and execute the subsequent steps.

[0127] Of course, in another alternative embodiment, if the difference between the number of computing nodes in the target cluster and the number of computing nodes in the specified cluster does not exceed the threshold, it indicates that the uneven number of computing nodes among the current clusters is not caused by the excessive number of computing nodes in the target cluster obtained by merging.

[0128] At this time, step S65 is further used to determine whether the number of unmerged clusters is greater than 1. In fact, after determining that the uneven nodes among the current clusters are not caused by the target cluster, the clusters can be continuously merged to achieve the goal of even nodes among the clusters.

[0129] Specifically, step S65 is used to determine whether the current target cluster and the unmerged clusters are still greater than 2, that is, to determine whether the number of current clusters can still be merged. If so, return to step S60 to continue merging the clusters. However, it should be noted that each time returning to step S60, the merged target cluster and the unmerged clusters are sorted and merged.

[0130] Of course, if the number of current clusters cannot be merged continuously, that is, at this time, the number of clusters without data reduction is 1, that is, only 2 clusters are left. At this time, only step S63 can be returned, taking the target cluster and the unmerged cluster as the clusters participating in data reduction; and entering step S10 in the above embodiment, and executing subsequent steps S11 and S12.

[0131] As an alternative embodiment, the difference in the number of computing nodes among the clusters participating in data reduction is within a preset range, including:

[0132] If the result of the specified calculation of the maximum and minimum values of the number of computing nodes in the clusters participating in data reduction is not greater than the preset value, the difference is within the preset range; where the specified calculation includes ratio calculation or difference calculation.

[0133] In a specific embodiment, when determining whether the difference in the number of computing nodes among the clusters participating in data reduction is within the preset range, the number of computing nodes in each cluster can be counted, and the maximum value Umax of the number of computing nodes and the minimum value Umin of the number of computing nodes can be determined.

[0134] Further, perform a specified calculation on the maximum value Umax and the minimum value Umin. The specified calculation may include, but is not limited to, ratio calculation or difference calculation. When performing ratio calculation, determine whether the calculation result of Umax / Umin is greater than 1. If it is not greater than 1, it indicates that the gap in the number of computing nodes in each cluster is within an acceptable range. At this time, cross-cluster data reduction can be performed. If it is greater than 1, it indicates that the gap in the number of computing nodes in each cluster is too large. To avoid wasting too much communication link resources, the clusters need to be merged. For the specific merging method, refer to the description of the above embodiments.

[0135] In another alternative embodiment, it is also possible to calculate the difference between the maximum value Umax and the minimum value Umin. If the difference is less than a preset value, it indicates that the gap in the number of computing nodes in each cluster is within an acceptable range; otherwise, it is considered not within the acceptable range.

[0136] It should be noted that for different specified calculation methods, the setting of the corresponding preset value is different. In addition, it should also be noted that in addition to the ratio and difference calculation methods for the gap in the number of computing nodes in each cluster being within an acceptable range, other calculation methods may also be used, and the present application does not limit this. For example, calculate the differences in the number of computing nodes between all clusters, and further determine whether the difference between any two differences is less than a preset value. If it is less, it indicates that the gap in the number of computing nodes in each cluster is within an acceptable range; otherwise, it is considered not within the acceptable range.

[0137] In the above embodiments, the collective communication control method for distributed training has been described in detail. The present application also provides an embodiment corresponding to a collective communication control device for distributed training.

[0138] Figure 7 Shown in the following is a schematic structural diagram of a collective communication control device for distributed training provided by an embodiment of the present application. Figure 7 As shown, the device includes:

[0139] An intra-cluster reduction module 70, configured to, when the gap in the number of computing nodes between any two clusters among the clusters participating in data reduction is within a preset range, reduce the data distributed on all computing nodes in the first cluster to the specified computing nodes in the first cluster; wherein, the number of specified computing nodes is the same as the number of computing nodes in the second cluster, and the second cluster is the cluster with the smallest number of computing nodes among the clusters participating in data reduction;

[0140] A cross-cluster reduction module 71, configured to control the specified computing nodes and the computing nodes in the second cluster to perform data reduction;

[0141] A data synchronization module 72, configured to control the specified computing nodes and other nodes in the first cluster except the specified computing nodes to perform data synchronization.

[0142] In addition, the collective communication control device for distributed training provided by the embodiments of the present application further includes:

[0143] A node group division module, configured to divide the computing nodes in the first cluster into m node groups; where each node group has k computing nodes, and k = floor(n / m), n is the number of computing nodes in the first cluster, and m is the number of computing nodes in the second cluster;

[0144] A specified computing node selection module, configured to select one computing node from each node group as the specified computing node;

[0145] An in-cluster reduction sub-module, configured to reduce the data distributed on all computing nodes in the first cluster to the specified computing nodes;

[0146] A first sorting module, configured to sort in ascending order according to the device numbers of the computing nodes in the first cluster;

[0147] A computing node division module, configured to, based on the sorting result, divide k computing nodes into each node group in turn in the way of dividing one computing node for each node group;

[0148] A specified computing node determination module, configured to use the computing node with the smallest device number in each node group as the specified computing node.

[0149] A first processing module, configured to determine whether there are target computing nodes in the computing nodes in the first cluster that have not been grouped; if not, reduce the data in each node group to the specified computing node in the corresponding group; if so, execute the step of reducing the data in each node group to the specified computing node in the corresponding group; and evenly divide the data on the target computing node into m pieces of data, and send the m pieces of data to different specified computing nodes in a one-to-one correspondence manner.

[0150] A second sorting module, configured to sort in ascending order according to the number of computing nodes when the gap is not within the preset range;

[0151] A merging module, configured to merge the first cluster and the second cluster into a target cluster based on the sorting result;

[0152] A second processing module, configured to determine whether the gap between any two clusters in the target cluster and the unmerged clusters is within the preset range; if so, use the target cluster and the unmerged clusters as the clusters participating in data reduction; and enter the step of reducing the data distributed on all computing nodes in the first cluster to the specified computing node in the first cluster, and execute the subsequent steps.

[0153] A third processing module is configured to determine whether the difference between the number of computing nodes in the target cluster and the number of computing nodes in the specified cluster exceeds a threshold; wherein, the specified cluster is the cluster with the smallest number of computing nodes among the unmerged clusters; if it exceeds the threshold, proceed to the step of using the target cluster and the unmerged clusters as the clusters participating in data reduction, and execute subsequent steps; if it does not exceed the threshold, determine whether the number of unmerged clusters is greater than 1; if it is greater, proceed to the step of sorting in ascending order according to the number of computing nodes, and execute subsequent steps; if it is not greater, proceed to the step of using the target cluster and the unmerged clusters as the clusters participating in data reduction, and execute subsequent steps.

[0154] Figure 8 FIG. is a schematic structural diagram of a collective communication control device for distributed training provided by another embodiment of the present application, as Figure 8 shown, the collective communication control device for distributed training includes: a memory 80 for storing computer programs;

[0155] a processor 81 for implementing the steps of the collective communication control method for distributed training mentioned in the above embodiment when executing the computer program.

[0156] The collective communication control device for distributed training provided in this embodiment may include, but is not limited to, a laptop computer or a desktop computer, etc.

[0157] Wherein, the processor 81 may include one or more processing cores, such as a 4-core processor, an 8-core processor, etc. The processor 81 may be implemented in at least one hardware form of a digital signal processor (DSP for short), a field programmable gate array (FPGA for short), or a programmable logic array (PLA for short). The processor 81 may also include a main processor and a coprocessor. The main processor is a processor for processing data in the wake state, also known as a central processing unit (CPU for short); the coprocessor is a low-power processor for processing data in the standby state. In some embodiments, the processor 81 may be integrated with a graphics processing unit (GPU for short), and the GPU is responsible for rendering and drawing the content to be displayed on the display screen. In some embodiments, the processor 81 may also include an artificial intelligence (AI for short) processor, and the AI processor is used to process computational operations related to machine learning.

[0158] The memory 80 may include one or more computer-readable storage media, which may be non-transitory. The memory 80 may also include high-speed random access memory and non-volatile memory, such as one or more disk storage devices and flash storage devices. In this embodiment, the memory 80 is at least used to store the following computer program 801. After the computer program is loaded and executed by the processor 81, the relevant steps of the collective communication control method for distributed training disclosed in any of the foregoing embodiments can be implemented. In addition, the resources stored in the memory 80 may also include an operating system 802, data 803, etc., and the storage method may be transient storage or permanent storage. Among them, the operating system 802 may include Windows, Unix, Linux, etc. The data 803 may include, but is not limited to, relevant data involved in the collective communication control method for distributed training, etc.

[0159] In some embodiments, the collective communication control device for distributed training may further include a display screen 82, an input / output interface 83, a communication interface 84, a power supply 85, and a communication bus 86.

[0160] Those skilled in the art can understand that Figure 8 the structure shown in does not constitute a limitation on the collective communication control device for distributed training, and may include more or fewer components than those shown in the figure.

[0161] The collective communication control device for distributed training provided by the embodiments of the present application includes a memory and a processor. When the processor executes the program stored in the memory, the collective communication control method for distributed training in the above embodiments can be implemented.

[0162] It should be noted that although the operations are depicted in a specific order in the drawings, this should not be construed as requiring these operations to be performed in the specific order shown or sequentially, or requiring all of the illustrated operations to be performed to achieve the desired result. In some cases, multitasking and parallel processing may be advantageous. In addition, the separation of the various system modules and components in the above embodiments should not be understood as required in all embodiments, and it should be understood that the described program components and systems can generally be integrated together in a single software product or packaged into multiple software products.

Claims

1. A collective communication control method for distributed training, characterized in that: The method comprises: When the difference in the number of computing nodes between any two clusters in the clusters participating in data reduction is within a preset range, the data distributed on all computing nodes in the first cluster are reduced to the designated computing nodes in the first cluster; wherein the number of the designated computing nodes is the same as the number of computing nodes in the second cluster, and the second cluster is the cluster with the smallest number of computing nodes in the clusters participating in data reduction; Controlling the designated computing node and computing nodes in the second cluster to perform data reduction through the collective communication semantics of Reduce Scatter; Control the designated computing node to synchronize data with other nodes in the first cluster except the designated computing node.

2. The collective communication control method for distributed training according to claim 1, characterized in that: The reducing the data distributed on all computing nodes in the first cluster to a designated computing node in the first cluster includes: Divide the computing nodes in the first cluster into m node groups; wherein each of the node groups has k computing nodes, and k=floor(n / m), n is the number of computing nodes in the first cluster, and m is the number of computing nodes in the second cluster; Selecting a computing node from each of the node groups as the designated computing node; The data distributed on all computing nodes in the first cluster are reduced to the designated computing node.

3. The collective communication control method for distributed training according to claim 2, characterized in that: The selecting a computing node from each of the node groups as the designated computing node comprises: Sort the computing nodes in the first cluster according to their device serial numbers from small to large; Based on the sorting result, the k computing nodes are divided into the node groups in sequence in a manner that each node group divides one computing node at a time; The computing node with the smallest device serial number in each of the node groups is used as the designated computing node.

4. The collective communication control method for distributed training according to claim 2, characterized in that: The reducing the data distributed on all computing nodes in the first cluster to the designated computing node includes: Determine whether there is a target computing node that has not been grouped in the computing nodes in the first cluster; If not, the data in each of the node groups is reduced to the designated computing nodes in the corresponding group; If so, execute the step of reducing the data in each of the node groups to the designated computing nodes in the corresponding group; and divide the data on the target computing node into m portions of data on average, and send the m portions of data to different designated computing nodes in a one-to-one correspondence manner.

5. The collective communication control method for distributed training according to claim 1, characterized in that: The method further comprises: When the gap is not within the preset range, the computing nodes are sorted in ascending order of number; Based on the ranking results, the first and second clusters are merged into the target cluster; Determine whether the gap between any two clusters in the target cluster and the unmerged clusters is within a preset range; If yes, the target cluster and the unmerged cluster are used as the clusters participating in data reduction; and the step of reducing the data distributed on all computing nodes in the first cluster to the designated computing nodes in the first cluster is entered, and subsequent steps are executed.

6. The collective communication control method for distributed training according to claim 5, characterized in that: If the gap between any two clusters in the target cluster and the unmerged cluster is not within the preset range, the method further includes: Determine whether the difference between the number of computing nodes in the target cluster and the number of computing nodes in a specified cluster exceeds a threshold; wherein the specified cluster is the cluster with the smallest number of computing nodes in the unmerged clusters; If the threshold is exceeded, the step of taking the target cluster and the unmerged cluster as the clusters participating in data reduction is entered, and subsequent steps are performed; If the threshold is not exceeded, determine whether the number of the unmerged clusters is greater than 1; if it is greater, enter the step of sorting the computing nodes in ascending order, and execute subsequent steps; if it is not greater, enter the step of taking the target cluster and the unmerged cluster as the clusters participating in data reduction, and execute subsequent steps.

7. The collective communication control method for distributed training according to claim 1, characterized in that: The difference in the number of computing nodes between clusters participating in data reduction is within the preset range, including: If the result of a specified calculation of the maximum and minimum numbers of computing nodes in the cluster participating in data reduction is not greater than a preset value, the gap is within the preset range; wherein the specified calculation includes a ratio calculation or a difference calculation.

8. A collective communication control device for distributed training, characterized in that: The device comprises: The cluster internal reduction module is used to reduce the data distributed on all computing nodes in the first cluster to the designated computing nodes in the first cluster when the difference in the number of computing nodes between any two clusters in the clusters participating in the data reduction is within a preset range; wherein the number of the designated computing nodes is the same as the number of computing nodes in the second cluster, and the second cluster is the cluster with the smallest number of computing nodes in the clusters participating in the data reduction; A cross-cluster reduction module, used to control the designated computing node and the computing nodes in the second cluster to perform data reduction through the collective communication semantics of Reduce Scatter; The data synchronization module is used to control the designated computing node to synchronize data with other nodes in the first cluster except the designated computing node.

9. A distributed training collective communication control device, comprising a memory and a processor, wherein the memory stores a computer program that can be run on the processor, characterized in that: When the processor executes the program, the steps of the collective communication control method for distributed training described in any one of claims 1 to 7 are implemented.

10. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the program is executed by a processor, the steps of the collective communication control method for distributed training described in any one of claims 1 to 7 are implemented.

Citation Information

Patent Citations

  • Distributed set communication method and device, equipment and storage medium

    CN115776523A

  • Pod data backup and path optimization method based on Kubernetes multi-cluster

    CN117149519A