Method for Flexible and Fast All-Reduction for Arbitrary Tree Topologies
By applying a full reduction algorithm and a reduction-dispersive algorithm on the processor set of deep learning systems, the problem of low data communication across multiple GPUs is solved, and more efficient data transmission and system performance is achieved.
Patent Information
- Application Number
- CN202080018167.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2019-04-01
- Filing Date
- 2020-03-17
- Publication Date
- 2025-06-10
- Estimated Expiration
- 2040-03-17
AI Technical Summary
In deep learning systems, data communication across multiple GPUs is easily affected by network restrictions, resulting in severe slowdown in data transmission, which in turn affects system efficiency.
The full reduction algorithm is used to allocate data in the tree topology of the processor set, and the combined data is sent between subprocessors through an iterative reduction-dispersion algorithm, passing it step by step from the root node to the leaf node, and finally combining the data on all processors.
Through this method, the bottlenecks in data transmission can be effectively reduced and the overall efficiency of deep learning systems can be improved, especially in large-scale deep learning systems.
Smart Images

Figure CN113518973B_ABST
Abstract
Description
BACKGROUND OF THE INVENTION
[0001] One of the main uses of the present invention relates to the field of deep learning, and more particularly, to performing an all-reduce algorithm on an arbitrary tree topology in deep learning and / or high-performance computing implementations.
[0002] Deep learning refers to a class of machine learning algorithms based on learning multiple levels of features or representations of a dataset. Current deep learning methods include using multiple layers of cascaded non-linear processing units for feature extraction and transformation. Feature extraction refers to the process of receiving an initial set of measurement data and establishing derived values (or features) to facilitate subsequent learning and generalization steps. In many cases, higher-level features are derived from lower-level features to generate a hierarchical representation of the measurement data and the derived features.
[0003] Deep learning algorithms are based on distributed representations. Distributed representations operate under the assumption that the observed (or measured) data is the result of the interaction of one or more factors organized into one or more layers. Conceptually, deep learning introduces an additional assumption that the interaction is at the level of layer representations that abstract or compose the factors providing the measurement data. Under this assumption, multiple layers and the size of the layers correspond to different amounts of abstraction.
[0004] Any or all of the data utilized and created in a deep learning system can be transmitted across one or more networks and can subsequently be subject to any limitations of the one or more networks. In particular, with respect to large-scale deep learning systems in both local and cloud environments, any network communication can be bottlenecked due to the large number of learners, the frequency of data exchange across the network, and the amount of data being exchanged. Additionally, communication across multi-tiered or heterogeneous (e.g., in a cloud environment) networks can be largely inefficient because the weakest link or node in the network will largely determine how the network as a whole will perform.
[0005] One method of improving efficiency in large-scale deep learning systems is to employ a reduction (REDUCE) operation. The reduction operation is a classical concept from functional programming that typically and efficiently reduces an initial set of numbers into a smaller set of numbers in parallel via a function that is both associative and commutative. Some existing programming functions implement the reduction operation across multiple processes, where the result is in some cases returned to the root process or in other cases returned to all of the processes involved.
[0006] When training a deep learning system in parallel across multiple graphics processing units (GPUs), choices must be made regarding how to distribute batches or groups of data to be read and executed, and operations, across the available GPUs. Each GPU then runs the forward pass of the network on its own data, as well as the error backpropagation, to determine the gradient of the loss with respect to any existing network parameters. The GPUs then communicate with each other to compute the average gradient, and this communication can occur across various networks. Each time the communication occurs, it is vulnerable to any network limitations, which can result in a significant slowdown of data transfer within the deep learning system. Summary of the Invention
[0007] According to an embodiment, a method for sending data across processors to combine data on the processors is provided, including receiving a data set D (D1, D2, …, Dn) at a set of processors P (P1, P2, …, Pk), where the data Di is received at the processor Pi. The method includes allocating a target portion of the data set to a processor in the set of processors, where the set of processors is configured in a tree topology including a root and one or more leaves, and where the target portion is allocated based on the number of child processors connected to a parent node. The method also includes starting from one or more leaves, sending iteratively allocated combined data between child processors sharing the same parent node in each branch of the tree topology, and incrementing a level in the tree topology until reaching the root. The method further includes sending combined data between child processors from one branch to child processors in at least one other branch.
[0008] In another form, a system for implementing a method for sending data across processors to combine data on the processors is provided.
[0009] In another form, a computer program product for sending data across processors to combine data on the processors is provided.
[0010] After studying the following drawings and detailed description, other systems, methods, features, and advantages of the present invention will be or will become apparent to those of ordinary skill in the art. All such additional systems, methods, features, and advantages are included within this specification and the summary of the invention, are within the scope of the present invention, and are protected by the appended claims. Brief Description of the Drawings
[0011] The present invention can be better understood with reference to the following drawings and description. The components in the drawings are not necessarily to scale, with emphasis being placed on illustrating the principles of the present invention. Further, in the drawings, the same reference numerals denote corresponding parts in different views.
[0012] Figure 1 is a block diagram of an example embodiment of a computing system configured to perform a reduction algorithm.
[0013] Figure 2 is a representative view of an example embodiment of a computing system having a tree topology.
[0014] Figure 3 is a representative view of a process for allocating target portions to nodes in a computing system.
[0015] Figure 4 is a representative view of a process for performing a reduce-scatter algorithm among nodes in a computing system.
[0016] Figure 5 is a representative view of nodes in a computing system after a reduce-scatter operation.
[0017] Figure 6 is a representative view of a process for allocating target portions to nodes in a computing system.
[0018] Figure 7 is a representative view of a process for allocating target items to nodes in a computing system.
[0019] Figure 8 is a representative view of a process for performing another reduce-scatter algorithm among nodes in a computing system.
[0020] Figure 9 is a representative view of nodes in a computing system after another reduce-scatter operation.
[0021] Figure 10 is a representative view of nodes in a computing system after a gather operation.
[0022] Figure 11 is a flowchart of an example embodiment of a method for performing a reduction algorithm on nodes of a computing system.
[0023] Figure 12 is a block diagram of an example embodiment of a node of a computing system DETAILED DESCRIPTION
[0024] Aspects of the example embodiments described herein provide a method for performing a reduction operation on any network topology including both symmetric and asymmetric tree topologies. Now referring to Figure 1 , an example embodiment of a deep learning computing system 100 configured to execute a reduction algorithm for sending data across processors to combine data on the processors is shown. In this embodiment, the deep learning computing system 100 includes a set of processors or nodes 102. In at least one embodiment, at least one processor in the set of processors 102 is a graphics processing unit (GPU).
[0025] In some embodiments, the processors or set of nodes 102 of the deep learning computing system 100 can be arranged in any network topology. For example, the set of processors 102 can be configured in an asymmetric tree topology including a root and one or more leaves. In an embodiment having an asymmetric tree topology, the root is the top processor or node in the tree. Each branch of the tree from the root can include a parent node that is connected to one or more child processors or nodes at a lower level or branch of the tree. The branch can also include one or more leaves, which are processors or nodes that are not connected to any child processors or nodes. In other embodiments, the set of processors 102 can be configured in a symmetric tree topology.
[0026] Figure 2 An example embodiment showing the set of processors 102 of the deep learning computing system 100 arranged in a tree topology is shown. As Figure 2 shown, the tree topology of the set of processors 102 includes a root processor or node 200 at the top level. Next, moving down a level to the next branch, the set of processors 102 includes a first parent node 210 and a second parent node 212. In the example embodiment, the first parent node 210 and the second parent node 212 can be processors. However, in general, each parent node typically does not operate as a processor but rather as a connection between child processors or nodes or as a connection between subtrees (other connected parent nodes). During a reduction operation, no active computation occurs at these intermediate branch nodes and all computation is performed by the child processors.
[0027] In this embodiment, the first parent node 210 is connected to two child processors, including a first child processor 220 and a second child processor 222. Since the first child processor 220 and the second child processor 222 are not connected to any other child processors or nodes, the first child processor 220 and the second child processor 222 are leaves. In this embodiment, the second parent node 212 is connected to three child processors, including a third child processor 224, a fourth child processor 226, and a fifth child processor 228. Since the third child processor 224, the fourth child processor 226, and the fifth child processor 228 are not connected to any other child processors or nodes, the third child processor 224, the fourth child processor 226, and the fifth child processor 228 are also leaves.
[0028] As Figure 2As shown, the tree topology of the processor set 102 of the deep learning computing system 100 is asymmetric because the first parent node 210 has two child processors, the first child processor 220 and the second child processor 222, while the second parent node 212 has three child processors, the third child processor 224, the fourth child processor 226, and the fifth child processor 228. However, the techniques described herein apply to both asymmetric and symmetric tree topologies.
[0029] In some embodiments, a data set D (D1, D2, …, Dn) is received at a processor set P (P1, P2, …, Pk), where data Di is received at processor Pi. In this embodiment, the data set includes first data 230, second data 232, third data 234, fourth data 236, and fifth data 238. As Figure 2 shown, the first data 230 is received at the first child processor 220, the second data 232 is received at the second child processor 222, the third data 234 is received at the third child processor 224, the fourth data 236 is received at the fourth child processor 226, and the fifth data 238 is received at the fifth child processor 228. Further, as shown in this embodiment, each data in the data set is associated with a data item having twelve indices labeled from zero (0) to eleven (11).
[0030] In an example embodiment, Figure 2 represents the initial state of the processor set 102 of the deep learning computing system 100 when receiving the data sets 230, 232, 234, 236, 238. The techniques of the example embodiments presented herein provide a method for performing a reduction algorithm that is used to send data (i.e., data sets 230, 232, 234, 236, 238) across processors (i.e., processor set 102) to combine the data on the processors.
[0031] Now referring to Figure 3 , a process of allocating a target portion of a data set to the processor set 102 in the deep learning computing system 100 is shown. As shown in this embodiment, the target portions of the data sets 230, 232, 234, 236, 238 are allocated to the processor set 102 at the first level of the tree. In an example embodiment, the target portion is allocated based on the number of child processors connected to the parent node.
[0032] For example, as Figure 3As shown, the first parent node 210 is connected to two child processors, a first child processor 220 and a second child processor 222. At branch 300, based on the degree of the first parent node 210 (i.e., 2 in this case), each child processor is assigned a target portion of the data set. Each child processor connected to the first parent node 210 is assigned a target portion corresponding to 1 / 2 of the data. As shown in this embodiment, a first target portion 302 (e.g., 1 / 2) is assigned to the first child processor 220, and a second target portion 304 (e.g., 1 / 2) is assigned to the second child processor 222. The sum of the target portions assigned to all child processors connected to the first parent node 210 at branch 300 adds up to one (e.g., 1 / 2 + 1 / 2 = 1).
[0033] A similar process is used to assign target portions at branch 310. However, at branch 310, the second parent node 212 is connected to three child processors, a third child processor 224, a fourth child processor 226, and a fifth child processor 228. Thus, based on the degree of the second parent node 212 (i.e., 3 in this case), each child processor is assigned a target portion of the data set. Each child processor connected to the second parent node 212 is assigned a target portion corresponding to 1 / 3 of the data. As shown in this embodiment, a third target portion 312 (e.g., 1 / 3) is assigned to the third child processor 224, a fourth target portion 314 (e.g., 1 / 3) is assigned to the fourth child processor 226, and a fifth target portion 316 (e.g., 1 / 3) is assigned to the fifth child processor 228. The sum of the target portions assigned to all child processors connected to the second parent node 212 at branch 310 adds up to one (e.g., 1 / 3 + 1 / 3 + 1 / 3 = 1).
[0034] Figure 4 Illustrated is a process of performing a reduce-scatter algorithm among a set of processors 102 in a deep learning computing system 100 to send the allocated data among child processors. In an example embodiment, the reduce-scatter operation is performed among the leaves at each branch of the tree. The reduce-scatter algorithm causes the processors to exchange data such that each processor ends up with a fragment of the final result. During the reduce-scatter operation, each processor combines the data item it receives from another processor with the data items present in its own data item range, where the data items in its own data item range correspond to the assigned chunk or partition (i.e., target portion). This operation is iteratively performed until each processor contains at least one data element representing the aggregation of the corresponding data items from each partition (i.e., target portion). After the reduce-scatter operation is complete, each processor has an array including the values that are the contributions from each processor on the branch.
[0035] In this embodiment, the ring algorithm is executed among sub-processors sharing the same parent node in each branch of the tree topology. For example, a first sub-processor 220 and a second sub-processor 222 both connected to a first parent node 210 exchange target portions (e.g., 1 / 2) of their first data 230 and second data 232 between them. As Figure 4 shown, in a first operation 400, the first sub-processor 220 receives from the second sub-processor 222 the target portion of its second data 232 corresponding to data items 0 to 5 (i.e., six data items, which are 1 / 2 of the twelve data items included in the first data 230 and the second data 232). Similarly, in a second operation 402, the second sub-processor 222 receives from the first sub-processor 220 the target portion of its first data 230 corresponding to data items 6 to 11 (i.e., six data items, which are 1 / 2 of the twelve data items included in the first data 230 and the second data 232).
[0036] A similar reduction-scatter operation is performed on another branch of the tree descending from a second parent node 212. For example, a third sub-processor 224, a fourth sub-processor 226, and a fifth sub-processor 228 all connected to the second parent node 212 exchange target portions (e.g., 1 / 3) of their third data 234, fourth data 236, and fifth data 238 between them. As Figure 4 shown, in a third operation 404, the third sub-processor 224 receives from the fourth sub-processor 226 the target portion of its fourth data 236 corresponding to data items 0 and 1, and receives from the fifth sub-processor 228 the target portion of its fifth data 238 corresponding to data items 2 and 3 (i.e., four data items, which are 1 / 3 of the twelve data items included in the third data 234, the fourth data 236, and the fifth data 238). Similarly, in a fourth operation 406, the fourth sub-processor 226 receives from the third sub-processor 224 the target portion of its third data 234 corresponding to data items 4 and 5, and receives from the fifth sub-processor 228 the target portion of its fifth data 238 corresponding to data items 6 and 7 (i.e., four data items, which are 1 / 3 of the twelve data items included in the third data 234, the fourth data 236, and the fifth data 238).
[0037] Furthermore, in a fifth operation 408, the fifth sub-processor 228 receives from the third sub-processor 224 the target portion of its third data 234 corresponding to data items 8 and 9, and receives from the fourth sub-processor 226 the target portion of its fourth data 236 corresponding to data items 10 and 11 (i.e., four data items, which are 1 / 3 of the twelve data items included in the third data 234, the fourth data 236, and the fifth data 238).
[0038] In this manner, a reduce-scatter algorithm is used to start from one or more leaves and send iteratively allocated combined data between sub-processors sharing the same parent node in each branch of a tree topology, and increase the levels in the tree topology until reaching the root 200.
[0039] Now refer to Figure 5 , Figure 5 which shows the set of processors 102 in the deep learning computing system 100 after the reduce-scatter operation as shown in Figure 4 is performed. After the reduce-scatter operation is performed between sub-processors sharing the same parent node in each branch of the tree, each processor includes combined data associated with a portion of its data set. For example, as shown in Figure 5 , the first sub-processor 220 includes first combined data 500 corresponding to combined data items in the range from 0 to 5 (i.e., [0, 6)), the second sub-processor 222 includes second combined data 502 corresponding to combined data items in the range from 6 to 11 (i.e., [6, 12)). Similarly, the third sub-processor 224 includes third combined data 504 corresponding to combined data items in the range from 0 to 3 (i.e., [0, 4)), the fourth sub-processor 226 includes fourth combined data 506 corresponding to combined data items in the range from 4 to 7 (i.e., [4, 8)), and the fifth sub-processor 228 includes fifth combined data 508 corresponding to combined data items in the range from 8 to 11 (i.e., [8, 12)).
[0040] Figure 6 shows the process of allocating a target portion of the combined data to the set of processors 102 in the deep learning computing system 100 for the next level on a branch of the tree. In this embodiment, the next level is the root 200. The root 200 is connected to two parent nodes, a first parent node 210 and a second parent node 212. At branch 600, based on the degree of the root 200 (i.e., 2 in this case), each parent node connected to the root 200 is allocated a target portion. Moving down to the next level in the tree, each sub-processor connected to the first parent node 210 and the second parent node 212 is allocated a target portion corresponding to 1 / 2 of the data of its parent node.
[0041] As shown in this embodiment, a first target portion 602 (e.g., 1 / 4) is assigned to the first sub-processor 220, and a second target portion 604 (e.g., 1 / 4) is assigned to the second sub-processor 222. The sum of the target portions assigned to all sub-processors connected to the first parent node 210 adds up to the 1 / 2 portion (e.g., 1 / 4 + 1 / 4 = 1 / 2) assigned to the first parent node 210 at branch 600. A similar process is used to assign target portions to the sub-processors of the second parent node 212. As shown in this embodiment, a third target portion 606 (e.g., 1 / 6) is assigned to the third sub-processor 224, a fourth target portion 608 (e.g., 1 / 6) is assigned to the fourth sub-processor 226, and a fifth target portion 610 (e.g., 1 / 6) is assigned to the fifth sub-processor 228. The sum of the target portions assigned to all sub-processors connected to the second parent node 212 adds up to the 1 / 2 portion (e.g., 1 / 6 + 1 / 6 + 1 / 6 = 1 / 2) assigned to the second parent node 212 at branch 600.
[0042] Now refer to Figure 7 , Figure 7 which shows the process of assigning target data items to each processor in the processor set 102 of the deep learning computing system 100. In this embodiment, based on the determined target portions as described above with reference to Figure 6 , each sub-processor is assigned the target data items of its combined data. In one example embodiment, before assigning the target portions, the sub-processors in the processor set 102 are sorted from last to first based on the last data item index value of the combined data for each processor, where ties are decided by the first data item index. In this embodiment, since the fifth combined data 508 has a last data item index of 12 and a first data item index of 8, the fifth sub-processor 228 is assigned the last position (e.g., one-fifth of five). Since the second combined data 502 has a last data item index of 12 and a first data item index of 6, the second sub-processor 222 is assigned the fourth position. In this case, the fifth sub-processor 228 and the second sub-processor 222 have the same last data item index (i.e., 12), but since the fifth sub-processor has a larger first data item index (i.e., 8 compared to 6), the fifth sub-processor 228 is assigned the fifth position before the second sub-processor 222.
[0043] The fourth sub-processor 226 is assigned to the third position because the third combined data 506 has the next largest last data item index (i.e., 8), and then the first sub-processor 220 is assigned to the second position based on the last data item index 6 of the first combined data 500, and the third sub-processor 224 is assigned to the first position based on the last data item index 4 of the third combined data 504. At this point, the order of assigning the target parts to the processor set 102 from the first to the last is as follows: the third sub-processor 224, the first sub-processor 220, the fourth sub-processor 226, the second sub-processor 222, and the fifth sub-processor 228.
[0044] Then, each target part of the combined data is assigned to each sub-processor in the processor set 102 in this determined order. As Figure 7 shown, the third sub-processor 224 is first assigned the target part (i.e., 1 / 6) of its third combined data 504 corresponding to combined data items 0 and 1 (i.e., [0, 2)). Next, the first sub-processor 220 is assigned the target part (i.e., 1 / 4) of its first combined data 500 corresponding to combined data items 2 to 4 (i.e., [2, 5)). The fourth sub-processor 226 is the next processor to be assigned the target part (i.e., 1 / 6) of its fourth combined data 506 corresponding to combined data items 5 and 6 (i.e., [5, 7)). The second sub-processor 222 is assigned the target part (i.e., 1 / 4) of its second combined data 502 corresponding to combined data items 7 to 9 (i.e., [7, 10)), and finally, the fifth sub-processor 226 is assigned the target part (i.e., 1 / 6) of its fifth combined data 508 corresponding to combined data items 10 and 11 (i.e., [10, 12)).
[0045] Figure 8 illustrates the process of performing another reduction-scatter algorithm among the processor sets 102 in the deep learning computing system 100. In an example embodiment, a hierarchical algorithm is executed, where the assigned parts are reduced based on the degree of the current branch of the tree topology. In this embodiment, the reduction-scatter algorithm is performed between the sub-processors from one branch (i.e., the children of the first parent node 210) to the sub-processors in another branch (i.e., the children of the second parent node 212). During this reduction-scatter operation, data items within the same range are exchanged between the sub-processors in each branch.
[0046] In this embodiment, a reduce-scatter algorithm is performed among sub-processors having different parent nodes in a tree. For example, a first sub-processor 220 and a second sub-processor 222, both connected to a first parent node 210, exchange their target portions with a target portion assigned to a third sub-processor 224, a fourth sub-processor 226, and a fifth sub-processor 228, all of which are connected to a second parent node 212. Thus, in this embodiment, there is no in-branch exchange among sub-processors having the same parent node.
[0047] As Figure 8 shown, in a first operation 800, the third sub-processor 224 receives its target portion corresponding to data items 0 and 1 (i.e., two of the twelve data items 1 / 6) from the first sub-processor 220. In a second operation 802, the first sub-processor 220 receives its target portion corresponding to data items 2 to 4 (i.e., three of the twelve data items 1 / 4) from the third sub-processor 224 (sending data items 2 and 3) and the fourth sub-processor 226 (sending data item 4). In a third operation 804, the fourth sub-processor 226 receives its target portion corresponding to data items 5 and 6 (i.e., two of the twelve data items 1 / 6) from the first sub-processor 220 (sending data item 5) and the second sub-processor 222 (sending data item 6). Similarly, in a fourth operation 806, the second sub-processor 222 receives its target portion corresponding to data items 7 to 9 (i.e., three of the twelve data items 1 / 4) from the fourth sub-processor 226 (sending data item 7) and the fifth sub-processor 228 (sending data items 8 and 9).
[0048] Finally, in a fifth operation 808, the fifth sub-processor 228 receives its target portion corresponding to data items 10 and 11 (i.e., two of the twelve data items 1 / 6) from the second sub-processor 222. Once the fifth operation 808 ends, the reduce-scatter algorithm has completed combining data across branches. That is, six data items of the combined data are sent by sub-processors of each branch to sub-processors of the other branch, a total of twelve data items. In this way, the reduce-scatter algorithm is used to send combined data among sub-processors in each branch of a tree topology.
[0049] Figure 9 is shown as described above with reference to Figure 8The processor set 102 in the deep learning computing system 100 after performing the described reduce-scatter operation. In this embodiment, the data sets 230, 232, 234, 236, 238 have been combined across all processors in the processor set 102. Each processor now includes an allocated portion of the combined data. For example, the first sub-processor 220 includes the first combined data 900, and the first combined data 900 includes data items corresponding to the range from two to four (i.e., [2, 5)) from each processor in the processor set 102. The second sub-processor 222 includes the second combined data 902, and the second combined data 902 includes data items corresponding to the range from seven to nine (i.e., [7, 10)) from each processor in the processor set 102. The third sub-processor 224 includes the third combined data 904, and the third combined data 904 includes data items corresponding to the range from zero to one (i.e., [0, 2)) from each processor in the processor set 102. The fourth sub-processor 226 includes the fourth combined data 906, and the fourth combined data 906 includes data items corresponding to the range from five to six (i.e., [5, 7)) from each processor in the processor set 102. The fifth sub-processor 228 includes the fifth combined data 908, and the fifth combined data 908 includes data items corresponding to the range from ten to eleven (i.e., [10, 12)).
[0050] To combine the complete data set from each processor in the processor set 102 on each processor, a all-gather algorithm is executed. Using the all-gather algorithm, the combined data is collected from each processor in the processor set 102 in the reverse order of the order in which the combined data was sent, as described above with reference to Figure 8 This technique allows each part of the combined data to be combined together such that each processor in the processor set 102 in the deep learning computing system 100 includes the complete set of combined data.
[0051] Now referring to Figure 10 , Figure 10 illustrates the processor set 102 in the deep learning computing system 100 after performing the all-gather algorithm. As shown in this embodiment, the complete set of combined data 1000, which includes the data sets 230, 232, 234, 236, 238 previously received separately by each processor in the processor set 102, is now included at each processor. That is, by applying the techniques described herein, the complete set of combined data 1000 is included at the first sub-processor 220, the second sub-processor 222, the third sub-processor 224, the fourth sub-processor 226, and the fifth sub-processor 228.
[0052] Figure 11FIG. 0 is a flowchart of an example embodiment of a method 1100 for performing an all - reduce algorithm on nodes of a deep - learning computing system 100. In the example embodiment, the method 1100 may be executed by a set of processors 102 of the deep - learning computing system 100 to combine data sets across all processors. For example, at least one processor in the set of processors 102 may include a leader node or a coordinator node that implements or coordinates instructions among the processors in the set of processors 102.
[0053] In the example embodiment, the method 1100 may begin at operation 1102. At operation 1102, a data set D (D1, D2, …, Dn) is received at a set of processors P (P1, P2, …, Pk), where data Di is received at processor Pi. For example, as Figure 2 shown, data sets 230, 232, 234, 236, 238 are received at the set of processors 102. Next, the method 1100 proceeds to operation 1104. At operation 1104, a target portion of the data set is assigned to processors of the set of processors, where the set of processors is configured in a tree topology including a root and one or more leaves, and where the target portion is assigned based on the number of child processors connected to a parent node. For example, as described above with reference to Figure 3 FIG., target portions 302, 304, 306, 308 are assigned to child processors 220, 222, 224, 226, 228.
[0054] Once the target portions are assigned at operation 1104, the method 1100 includes operation 1106, where starting from one or more leaves, the combined data of the assignment is sent between child processors sharing the same parent node in each branch of the tree topology, and the level in the tree topology is incremented until the root is reached. For example, as Figure 4 shown, a reduce - scatter algorithm is executed such that the first child processor 220 and the second child processor 222 exchange the target portions of their data, and the third child processor 224, the fourth child processor 226, and the fifth child processor 228 exchange the target portions of their data.
[0055] In some embodiments, once the iteration of sending the combined data of the assignment executed at operation 1106 is complete, the method 1100 may further include assigning the target portions of the combined data in an order based on the last data index of the combined data assigned to each processor in the set of processors 102, for example, as Figure 6 and 7 shown.
[0056] Next, the method 1100 further includes operation 1108, where the combined data is sent between child processors from one branch to child processors in at least one other branch. For example, as described above with reference toFigure 8 as described by the operations of the reduction-scatter operation shown.
[0057] Additionally, in some embodiments, method 1100 may further include performing one or more gather operations that are part of a all-gather algorithm in an order opposite to the order in which the combined data is sent. This technique allows each part of the combined data to be combined together such that each processor in the processor set includes a complete set of the combined data, e.g., as shown above Figure 10 as shown.
[0058] Although the example embodiments described herein have been discussed with reference to a deep learning computing system (e.g., deep learning computing system 100), the techniques described herein can be used by processors of other systems where large data sets are to be combined across a set of processors.
[0059] Figure 12 FIG. shows a block diagram of components of a representative node 1200 that includes a processor 1202, according to an example embodiment. In one embodiment, processor 1202 may include a graphics processing unit (GPU). It should be understood that Figure 12 only an illustration of one implementation is provided and does not imply any limitation to the environments in which different embodiments can be implemented. Many modifications can be made to the described environments.
[0060] As Figure 12 shown, node 1200 includes a communication fabric 1204 that provides communication between the (one or more) processors 1202, memory 1206, persistent storage 1208, communication unit 1212, and the (one or more) input / output (I / O) interfaces 1214. The communication fabric 1204 can be implemented with any architecture designed to transfer data and / or control information between processors (such as microprocessors, communication and network processors, etc.), system memory, peripherals, and any other hardware components within the system. For example, the communication fabric 1204 can be implemented with one or more buses.
[0061] Memory 1206 and persistent storage 1208 are computer-readable storage media. In this embodiment, memory 1206 includes random access memory (RAM) 1216 and cache memory 1218. In general, memory 1206 can include any suitable volatile or non-volatile computer-readable storage media.
[0062] One or more programs may be stored in the persistent storage device 1208 for access and / or execution by one or more of the corresponding processors 1202 via one or more memories of the memory 1206. In this embodiment, the persistent storage device 1208 includes a magnetic hard disk drive. As an alternative or addition to the magnetic hard disk drive, the persistent storage device 1208 may include a solid state drive, a semiconductor storage device, a read-only memory (ROM), an erasable programmable read-only memory (EPROM), a flash memory, or any other computer-readable storage medium capable of storing program instructions or digital information.
[0063] The medium used by the persistent storage device 1208 may also be removable. For example, a removable hard disk drive may be used for the persistent storage device 1208. Other examples include optical discs and magnetic disks, thumb drives, and smart cards that are inserted into a drive for transfer to another computer-readable storage medium that is also part of the persistent storage device 1208.
[0064] In these examples, the communication unit 1212 provides communication with other processors, data processing systems, or devices (e.g., between processors such as the processor set 102 described above). In an example embodiment, the communication unit 1212 may include one or more network interface cards. The communication unit 1212 may provide communication by use of any one or both of physical and wireless communication links.
[0065] (Multiple) I / O interfaces 1214 allow for the input and output of data with other devices that may be connected to the node 1200. For example, the I / O interface 1214 may provide a connection to an external device 1220 such as a keyboard, keypad, touch screen, and / or some other suitable input device. The external device 1220 may also include a portable computer-readable storage medium such as a thumb drive, a portable optical disc or magnetic disk, and a memory card. The software and data used to practice embodiments of the present invention may be stored on such portable computer-readable storage media and may be loaded onto the persistent storage device 1208 via the (multiple) I / O interfaces 1214. The (multiple) I / O interfaces 1214 may also be connected to a display 1222. The display 1222 provides a mechanism for displaying data to a user (e.g., a user of the deep learning computing system 100) and may be, for example, a computer monitor.
[0066] The programs described herein are identified based on the applications in which they are implemented in particular embodiments of the present invention. However, it should be understood that any particular program terms herein are used for convenience only, and thus the present invention should not be limited to use in any particular application identified and / or implied by such terms.
[0067] The present invention can be a system, method, and / or computer program product at any possible level of integration of technical details. The computer program product can include a computer-readable storage medium (or media) having computer-readable program instructions thereon for causing a processor to execute aspects of the present invention.
[0068] A computer-readable storage medium can be a tangible device that is capable of storing and retaining instructions for use by an instruction execution device. A computer-readable storage medium can be, by way of example and not limitation, an electronic storage device, a magnetic storage device, an optical storage device, an electromagnetic storage device, a semiconductor storage device, or any suitable combination of the foregoing. A non-exhaustive list of more specific examples of the computer-readable storage medium includes the following: a portable computer diskette, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or Flash memory), a static random access memory (SRAM), a portable compact disc read-only memory (CD-ROM), a digital versatile disc (DVD), a memory stick, a floppy disk, a mechanical encoding device such as a punched card or raised structures in a groove having instructions recorded thereon, and any suitable combination of the foregoing. As used herein, a computer-readable storage medium should not be construed to be a transient signal per se, such as a radio wave or other freely propagating electromagnetic wave, an electromagnetic wave propagating through a waveguide or other transmission medium (e.g., a light pulse through an optical fiber cable), or an electrical signal transmitted through a wire.
[0069] The computer-readable program instructions described herein can be downloaded from a computer-readable storage medium to a corresponding computing / processing device, or to an external computer or external storage device via a network, such as the Internet, a local area network, a wide area network, and / or a wireless network. The network can include copper transmission cables, optical transmission fibers, wireless transmission, routers, firewalls, switches, gateway computers, and / or edge servers. A network adapter card or network interface in each computing / processing device receives the computer-readable program instructions from the network and forwards the computer-readable program instructions for storage in a computer-readable storage medium within the respective computing / processing device.
[0070] The computer-readable program instructions for performing the operations of the present invention may be assembly instructions, instruction set architecture (ISA) instructions, machine instructions, machine-related instructions, microcode, firmware instructions, state-setting data, configuration data for integrated circuits, or source code or object code written in any combination of one or more programming languages, including object-oriented programming languages such as Smalltalk, C++, etc., and procedural programming languages such as the "C" programming language or similar programming languages. The computer-readable program instructions may be executed entirely on the user's computer, partially on the user's computer, executed as a stand-alone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In the latter case, the remote computer may be connected to the user's computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or may be connected to an external computer (e.g., through the Internet using an Internet service provider). In some embodiments, in order to perform aspects of the present invention, an electronic circuit, including, for example, a programmable logic circuit, a field-programmable gate array (FPGA), or a programmable logic array (PLA), may execute the computer-readable program instructions by utilizing the state information of the computer-readable program instructions to personalize the electronic circuit.
[0071] Aspects of the present invention are described herein with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of the invention. It should be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer-readable program instructions.
[0072] These computer-readable program instructions may be provided to a processor of a general-purpose computer, a special-purpose computer, or other programmable data processing apparatus to produce a machine, such that the instructions executed via the processor of the computer or other programmable data processing apparatus create means for implementing the functions / acts specified in one or more blocks of the flowchart and / or block diagram. These computer-readable program instructions may also be stored in a computer-readable storage medium that can direct a computer, a programmable data processing apparatus, and / or other devices to function in a particular manner, such that the computer-readable storage medium in which the instructions are stored comprises an article of manufacture including instructions for implementing aspects of the functions / acts specified in one or more blocks of the flowchart and / or block diagram.
[0073] The computer-readable program instructions may also be loaded onto a computer, other programmable data processing apparatus, or other device to cause a series of operational steps to be performed on the computer, other programmable apparatus, or other device to produce a computer-implemented process, such that the instructions executed on the computer, other programmable apparatus, or other device implement the functions / acts specified in one or more blocks of the flowchart and / or block diagram.
[0074] Figures 1 - 12 The flowcharts and block diagrams in illustrate the possible architectures, functions, and operations of systems, methods, and computer program products according to various embodiments of the present invention. In this regard, each block in the flowchart or block diagram may represent a module, segment, or portion of instructions, which includes one or more executable instructions for implementing the specified (multiple) logical functions. In some alternative embodiments, the functions noted in the block may not occur in the order noted in the figures. For example, two blocks shown in succession may in fact be executed substantially simultaneously, or these blocks may sometimes be executed in the reverse order, depending on the functions involved. It will also be noted that each block of the block diagrams and / or flowchart illustrations, and combinations of blocks in the block diagrams and / or flowchart illustrations, can be implemented by a dedicated hardware-based system that performs the specified functions or acts or a combination of dedicated hardware and computer instructions.
[0075] The description of the various embodiments of the present invention has been presented for purposes of illustration, but is not intended to be exhaustive or limited to the disclosed embodiments. Many modifications and variations will be apparent to those of ordinary skill in the art without departing from the scope and spirit of the described embodiments. The terms used herein have been chosen to best explain the principles of the embodiments, the practical application, or improvements made to the technology found in the marketplace, or to enable other ordinary skill in the art to understand the embodiments disclosed herein.
Claims
1. A method for sending data across processors to combine the data on the processors, the method comprises: receiving a data set at a set of processors in a deep learning system, wherein the set of processors is configured in a tree topology including a root, a plurality of parent nodes with respective branches, and a plurality of sub-processors connected to respective parent nodes among the plurality of parent nodes; allocating a target portion of the data set to a plurality of respective sub-processors among a plurality of sub-processor sets; performing a first reduce-scatter operation to cause a plurality of sub-processors sharing the same parent node to exchange data items of the target portion, such that one sub-processor sharing the same parent node combines its own data items with data items received from one or more other sub-processors sharing the same parent node; performing a second reduce-scatter operation to cause data items to be exchanged between a plurality of sub-processors connected to different parent nodes, such that the data set is combined across all of the plurality of sub-processors, and each sub-processor among the plurality of sub-processors in all of the plurality of respective branches includes the combined data; and using a gather algorithm for each of the plurality of sub-processors to gather the combined data of other sub-processors in all of the plurality of respective branches, such that all of the plurality of sub-processors include a complete set of the combined data for deep learning.
2. The method according to claim 1, wherein when performing the first reduce-scatter operation, a ring algorithm is used to exchange the data items of the target portion among the plurality of sub-processors.
3. The method according to claim 1, wherein when performing the second reduce-scatter operation, a hierarchical algorithm is used to exchange the data items of the target portion among the plurality of sub-processors.
4. The method according to claim 1, wherein, at the level of the plurality of parent nodes, distributing the target portion of the data set to the plurality of sub-processors is based on the degrees of the plurality of parent nodes; wherein, at the level of the root, distributing the target portion of the data set to the plurality of sub-processors is based on the degrees of the root.
5. The method according to claim 1, further comprises: allocating a plurality of ranges for data item indices for the plurality of respective sub-processors among the plurality of sub-processors; and wherein, when performing the second reduce-scatter operation, a first sub-processor connected to a first parent node receives a data item from a second sub-processor connected to a second parent node, and the data item is within the range of the data item index of the first sub-processor.
6. The method according to claim 1, wherein at least one processor in the set of processors is a graphics processing unit GPU.
7. A computer program product for sending data across processors to combine the data on the processors, the computer program product comprising a computer-readable storage medium having program instructions embodied therewith, the program instructions being executable by a processor in a processor set to cause the processor to perform a method, the method comprising: Receiving a data set at a set of processors in a deep learning system, wherein the set of processors is configured in a tree topology including a root, a plurality of corresponding branches of a plurality of parent nodes, and a plurality of sub-processors connected to respective parent nodes of the plurality of parent nodes; Allocating a target portion of the data set to a plurality of respective sub-processors in a plurality of sub-processor sets; Performing a first reduce-scatter operation to cause a plurality of sub-processors sharing the same parent node to exchange data items of the target portion, such that one sub-processor sharing the same parent node combines its own data item with data items received from one or more other sub-processors sharing the same parent node; Performing a second reduce-scatter operation to cause data items to be exchanged between a plurality of sub-processors connected to different parent nodes, such that the data set is combined across all of the plurality of sub-processors, and each sub-processor in the plurality of sub-processors in all of the plurality of corresponding branches includes the combined data; and For each sub-processor in the plurality of sub-processors, using an all-gather algorithm to gather the combined data of other sub-processors in all of the plurality of corresponding branches, such that all of the plurality of sub-processors include a complete set of the combined data for deep learning.
8. The computer program product according to claim 7, wherein when performing the first reduce-scatter operation, the data items of the target portion are exchanged between the plurality of sub-processors using a ring algorithm.
9. The computer program product according to claim 7, wherein, At the level of the plurality of parent nodes, distributing the target portion of the data set to the plurality of sub-processors is based on the plurality of degrees of the plurality of parent nodes; wherein, at the level of the root, distributing the target portion of the data set to the plurality of sub-processors is based on the plurality of degrees of the root.
10. The computer program product according to claim 7, the method further comprising: Allocating a plurality of ranges for data item indices for the plurality of respective sub-processors in the plurality of sub-processors; and wherein, when performing the second reduce-scatter operation, a first sub-processor connected to a first parent node receives a data item from a second sub-processor connected to a second parent node, and the data item is within the range of the data item index of the first sub-processor.
11. The computer program product according to claim 7, wherein at least one processor in the processor set is a graphics processing unit GPU.
12. The computer program product according to claim 7, wherein when performing the second reduce-scatter operation, a hierarchical algorithm is used to exchange the data items of the target portion among the plurality of sub-processors.
13. A deep learning system including a set of processors in a tree topology configured with a root, a plurality of parent nodes including a plurality of respective branches, and a plurality of sub-processors connected to respective ones of the plurality of parent nodes, wherein the set of processors is configured to implement a method for sending data across the set of processors to combine the data on the processors, the method comprises: receiving a data set at the set of processors in the deep learning system; allocating a target portion of the data set to a plurality of respective sub-processors in a plurality of sub-processor sets; performing a first reduce-scatter operation to cause a plurality of sub-processors sharing the same parent node to exchange the data items of the target portion, such that one sub-processor sharing the same parent node combines its own data item with the data items received from one or more other sub-processors sharing the same parent node; performing a second reduce-scatter operation to cause data items to be exchanged among a plurality of sub-processors connected to different parent nodes, such that the data set is combined across all of the plurality of sub-processors and each of the plurality of sub-processors in all of the plurality of respective branches includes the combined data; and collecting the combined data of other sub-processors in all of the plurality of respective branches for each of the plurality of sub-processors using an all-gather algorithm, such that all of the plurality of sub-processors include a complete set of the combined data for deep learning.
14. The deep learning system according to claim 13, wherein when performing the first reduce-scatter operation, a ring algorithm is used to exchange the data items of the target portion among the plurality of sub-processors.
15. The deep learning system according to claim 13, the method further comprises: allocating a plurality of ranges for data item indices for the plurality of respective sub-processors among the plurality of sub-processors; and wherein when performing the second reduce-scatter operation, a first sub-processor connected to a first parent node receives a data item from a second sub-processor connected to a second parent node, and the data item is within the range of the data item index of the first sub-processor.
16. The deep learning system according to claim 13, wherein at least one processor in the set of processors is a graphics processing unit GPU.
17. The deep learning system according to claim 13, wherein when performing the second reduce-scatter operation, a hierarchical algorithm is used to exchange the data items of the target portion among the plurality of sub-processors.
18. The deep learning system according to claim 13, wherein, At the level of the plurality of parent nodes, distributing the target portion of the data set to the plurality of sub-processors is based on the plurality of degrees of the plurality of parent nodes; wherein, at the level of the root, distributing the target portion of the data set to the plurality of sub-processors is based on the plurality of degrees of the root.
Citation Information
Patent Citations
Distribution of operations to remote computers
US20040083475A1