Method and System for Improving the Efficiency of Distributed Data Parallel Training of Deep Learning Models
By improving the gradient merging strategy of the Horovod platform, the problems of high communication frequency and over-merging are solved, and more efficient deep learning training is achieved, suitable for multi-node distributed systems.
Patent Information
- Application Number
- CN202211656348.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-12-22
- Publication Date
- 2025-07-22
- Estimated Expiration
- 2042-12-22
AI Technical Summary
In the Horovod platform, the communication computing overlapping technology has a high communication frequency and a long startup duration. The gradient merging technology has excessive merging problems, resulting in low training efficiency.
Using an improved gradient merging strategy, iteratively determines whether the gradients of each layer of the deep learning model can communicate together in the Horovod system, and uses the merged group list storage and broadcast gradient merging strategy to perform gradient communication in group, reducing the number of parameter synchronization between nodes and excessive merging.
Reduces communication frequency and startup duration, improves the efficiency and convergence of deep learning training, and is suitable for distributed systems and deep learning models of any scale.
Smart Images

Figure CN115859117B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of deep learning, and more specifically, relates to a method and system for improving the distributed data parallel training efficiency of a deep learning model. Background Art
[0002] Horovod is an emerging deep learning parallel training platform. The parameter synchronization process between nodes implemented by its Ring-AllReduce algorithm is as follows: First, all N nodes form a logical ring according to their initial numbers in the Horovod system; then each node transmits gradients to its adjacent downstream node and adds the gradients transmitted from the adjacent upstream node to the local gradients. After transmitting N - 1 times, each node will store the sum of N gradients; finally, all nodes calculate the average value of the N gradients locally to update the parameters of the deep learning model.
[0003] Horovod uses two technologies to minimize the communication time in deep learning training: communication-computation overlap technology and gradient merging technology.
[0004] The communication-computation overlap technology is based on the hierarchical structure of the deep learning model. In each layer of the model, gradient communication can start immediately after the local layer computation and the gradient communication of its upper layer are completed, so as to hide the communication time of this layer under the computation time of other layers.
[0005] The gradient merging technology is based on the communication-computation overlap technology. The gradient merging technology of the Horovod platform selects appropriate multiple small gradients to merge according to the gradient magnitude. The so-called gradient merging means communicating the gradients of multiple layers of the deep learning model together, thereby saving the startup duration associated with each inter-node communication. These gradients form a merging group. For example, communicating a merging group with n layers can reduce the startup duration from n times to 1 time; Shi Shaohuai et al. from Hong Kong Baptist University proposed a merged-gradient wait-free backpropagation algorithm (MG-WFBP) based on gradient merging, which can adaptively merge small messages of multiple consecutive layers for communication, further reducing the communication time.
[0006] However, the above two methods used on the Horovod platform to reduce the communication time in deep learning training both have some non-negligible defects: First, the existing gradient merging technology is not intelligent enough and there is an over-merging problem, which results in wasting the overlap time that could be used to overlap communication with the backpropagation computation time although the startup duration is saved; Second, the existing communication-computation overlap technology makes full use of the backpropagation computation time to overlap the gradient communication time, but the communication frequency is high and the startup duration is large. Summary of the Invention
[0007] In view of the above deficiencies or improvement requirements of the prior art, the present invention provides a method and system for improving the distributed data parallel training efficiency of a deep learning model, aiming to solve the technical problems of high communication frequency and long startup duration in the existing communication-computation overlapping technology of the Horovod platform, and the technical problem that the existing gradient merging technology of the Horovod platform wastes the overlap time due to excessive merging problems.
[0008] To achieve the above object, according to one aspect of the present invention, a method for improving the distributed data parallel training efficiency of a deep learning model is provided, which is applied in a Horovod system including multiple nodes. The method includes the following steps:
[0009] (1) Each node obtains a data set;
[0010] (2) Each node initializes the deep learning model to obtain an initialized deep learning model;
[0011] (3) The first node performs n (n ranges from 5 to 10, preferably 5) iterative pre-trainings on the initialized deep learning model, and obtains the average value t of the forward propagation calculation time obtained from the n iterative pre-trainings f , the average value t of the backward propagation calculation time obtained from the n iterative pre-trainings b , the start timestamp of the backward propagation calculation of the l-th layer of the deep learning model and the gradient p of the l-th layer of the deep learning model (l) , and there is represents the forward propagation calculation time of the l-th layer in the deep learning model, represents the backward propagation calculation time of the l-th layer in the deep learning model, and l ∈ [1, L], where L is the number of layers of the deep learning model;
[0012] (4) The first node sets the counter k = 1.
[0013] (5) The first node calculates the gradient communication time of the k-th layer of the deep learning model α represents the startup duration of communication among all nodes in the Horovod system, β represents the time occupied by transmitting each byte-sized gradient during the communication between two nodes in the Horovod system, and sets k = k + 1;
[0014] (6) The first node determines whether k is greater than L. If so, it proceeds to step (7); otherwise, it proceeds to step (5);
[0015] (7) The first node sets the counter i = L - 1 and sets the start timestamp of the gradient communication for the L-th layer of the deep learning model.
[0016] (8) The first node obtains the start timestamp of the gradient communication for the i-th layer of the deep learning model. And sets the counter i = i - 1;
[0017] (9) The first node determines whether i is less than 1. If so, it proceeds to step (10); otherwise, it proceeds to step (8).
[0018] (10) The first node sets the counter j = 1, the counter g = 1, initializes the secondary list group[:] of the merging group and the primary list m[:].
[0019] (11) The first node determines Whether it holds. If so, it proceeds to step (12); otherwise, it proceeds to step (13).
[0020] (12) The first node adds the layer index number j of the j-th layer of the deep learning model to the primary list m[:], sets Sets And proceeds to step (14);
[0021] (13) The first node sets g = g + 1, adds the primary list m[:] to the secondary list group[:] of the merging group, clears the primary list m[:], and proceeds to step (14).
[0022] (14) The first node sets j = j + 1 and uses what is obtained in step (12) And what is obtained in step (3) And To update the start timestamps of the gradient communications for the j-th layer to the 1st layer in the deep learning model.
[0023] (15) The first node determines whether j < L. If so, it returns to step (11); otherwise, it proceeds to step (16).
[0024] (16) The first node sends the secondary list group[:] of the merging group to all other nodes in the Horovod cluster in a broadcast manner.
[0025] (17) Each node sequentially performs forward propagation calculations on the deep learning model from the 1st layer to the L-th layer, sequentially performs backward propagation calculations on the deep learning model from the L-th layer to the 1st layer to obtain the gradients of each layer, and uses the obtained gradients of each layer in the deep learning model to call the Ring-AllReduce communication function in the Horovod system to perform gradient communication.
[0026] Preferably, the initial value of the weight parameter in step (2) is a random value output using a truncated normal distribution with a standard deviation of 0.1, the initial value of the bias parameter is set to 0, the initial learning rate lr = 0.0001, a learning strategy of stochastic gradient descent is adopted, the step size stepsize = 200, and the weight gamma = 0.1, that is, the learning rate is multiplied by 0.1 every 200 iterations;
[0027] Preferably, step (14) includes the following sub-steps:
[0028] (14-1) Set the counter q = j,
[0029] (14-2) Determine whether q is equal to L. If so, set the start timestamp of the gradient communication of the L-th layer of the deep learning model to Then go to step (14-4), otherwise go to step (14-3);
[0030] (14-3) The first node obtains the start timestamp of the gradient communication of the q-th layer of the deep learning model
[0031] (14-4) Set the counter q = q - 1;
[0032] (14-5) Determine whether q is less than 1. If so, the process ends, otherwise return to step (14-3).
[0033] Preferably, step (17) includes the following sub-steps:
[0034] (17-1) Set the counter w = L;
[0035] (17-2) All nodes perform backpropagation calculation on the w-th layer of the deep learning model and store the obtained gradient of the w-th layer of the deep learning model in the local buffer;
[0036] (17-3) Determine whether the number of inner list indexes of the first-level list where the w-th layer of the deep learning model is located in the merged group secondary list group[:] is only 1. If so, go to step (17-4), otherwise go to step (17-5);
[0037] (17-4) All nodes perform gradient communication, that is, each node calls the Ring-AllReduce communication function in the Horovod system and uses the gradient in the local buffer as a parameter to perform gradient communication, and then go to step (17-6);
[0038] (17-5) Among all the nodes, it is judged whether the layer index number of the w-th layer of the deep learning model in the merged second-level list group[:] of the merge group is the minimum value of the index numbers in the corresponding first-level list. If so, go to step (17-4); otherwise, execute step (17-6).
[0039] (17-6) Set w = w - 1 and wait for the gradient communication of all nodes to end.
[0040] (17-7) Judge whether w < 1. If so, the process ends; otherwise, return to step (17-2).
[0041] According to another aspect of the present invention, a system for improving the distributed data parallel training efficiency of a deep learning model is provided, which is applied in a Horovod system including multiple nodes. The system includes:
[0042] The first module, which is set in each node and is used to obtain a data set.
[0043] The second module, which is set in each node and is used to initialize the deep learning model to obtain an initialized deep learning model.
[0044] The third module, which is set in the first node and is used to perform n (n ranges from 5 to 10, preferably 5) iterative pre-trainings on the initialized deep learning model, and obtain the average value t of the forward propagation calculation times obtained from the n iterative pre-trainings f and the average value t of the backward propagation calculation times obtained from the n iterative pre-trainings b as well as the start timestamp of the backward propagation calculation of the l-th layer of the deep learning model and the gradient p of the l-th layer of the deep learning model (l) , and there is represents the forward propagation calculation time of the l-th layer in the deep learning model, represents the backward propagation calculation time of the l-th layer in the deep learning model, and l ∈ [1, L], where L is the number of layers of the deep learning model;
[0045] The fourth module, which is set in the first node and is used to set the counter k = 1.
[0046] The fifth module, which is set in the first node and is used to calculate the gradient communication time of the k-th layer of the deep learning model α represents the start-up duration of communication among all nodes in the Horovod system, β represents the time occupied by transmitting each byte-sized gradient during the communication between two nodes in the Horovod system, and set k = k + 1;
[0047] The sixth module, which is set in the 1st node, is used to determine whether k is greater than L. If so, it transfers to the seventh module; otherwise, it transfers to the fifth module.
[0048] The seventh module, which is set in the 1st node, is used to set the counter i = L - 1 and set the start timestamp of the gradient communication of the Lth layer of the deep learning model.
[0049] The eighth module, which is set in the 1st node, is used to obtain the start timestamp of the gradient communication of the ith layer of the deep learning model. And set the counter i = i - 1;
[0050] The ninth module, which is set in the 1st node, is used to determine whether i is less than 1. If so, it transfers to the tenth module; otherwise, it transfers to the eighth module.
[0051] The tenth module, which is set in the 1st node, is used to set the counter j = 1, the counter g = 1, initialize the secondary list group[:] of the merging group and the primary list m[:].
[0052] The eleventh module, which is set in the 1st node, is used to determine Whether it holds. If so, it transfers to the twelfth module; otherwise, it transfers to the thirteenth module.
[0053] The twelfth module, which is set in the 1st node, is used to add the layer index number j of the jth layer of the deep learning model to the primary list m[:], set Set And transfer to the fourteenth module;
[0054] The thirteenth module, which is set in the 1st node, is used to set g = g + 1, add the primary list m[:] to the secondary list group[:] of the merging group, clear the primary list m[:], and transfer to the fourteenth module.
[0055] The fourteenth module, which is set in the 1st node, is used to set j = j + 1 and use what is obtained from the twelfth module And what is obtained from the third module And To update the start timestamps of the gradient communication from the jth layer to the 1st layer in the deep learning model.
[0056] The fifteenth module, which is set in the 1st node, is used to determine whether there is j < L. If so, it returns to the eleventh module; otherwise, it transfers to the sixteenth module.
[0057] The sixteenth module, which is set in the 1st node, is used to send the secondary list group[:] of the merging group to all other nodes in the Horovod cluster by broadcasting.
[0058] The seventeenth module is provided in each node and is used to perform forward propagation calculations on the deep learning model layer by layer from the first layer to the L-th layer, perform backward propagation calculations on the deep learning model layer by layer from the L-th layer to the first layer to obtain the gradients of each layer, and use the obtained gradients of each layer in the deep learning model to call the Ring-AllReduce communication function in the Horovod system to perform gradient communication.
[0059] Generally speaking, compared with the prior art, the above technical solutions conceived by the present invention can achieve the following beneficial effects:
[0060] (1) Since the present invention adopts step (17), which enables all nodes to perform gradient communication according to the deep learning model gradient merging strategy obtained in step (16), reducing the number of parameter synchronizations between nodes, it can solve the problems of high communication frequency and long startup duration in the existing communication-computation overlap technology of the Horovod platform.
[0061] (2) Since the present invention adopts steps (10) to (14), first, it iteratively determines whether the gradient of the k-th (1 ≤ k < L) layer of the deep learning model with a total of L layers should be communicated together with the gradient of the (k + 1)-th layer, rather than determining whether the gradient of the k-th layer of the deep learning model should be communicated together with the gradient of the (k - 1)-th layer, making the judgment condition in step (11) include the overlapping time wasted by gradient merging, that is, if the overlapping time wasted by gradient merging is higher than the startup duration saved by gradient merging, the gradient of the k-th layer of the deep learning model should not be communicated together with the gradient of the (k + 1)-th layer. Therefore, it can solve the problem of over-merging in the existing gradient merging technology of the Horovod platform; second, it timely updates the start timestamp of the gradient communication of the relevant layers of the deep learning model, so it can ensure the accuracy of the gradient merging strategy, thereby ensuring the convergence of deep learning training.
[0062] (3) Since the present invention adopts step (5) to calculate the gradient communication time of the k-th layer of the deep learning model, and the parameters it uses include the startup duration α of all node communications, the time β occupied by transmitting each byte-sized gradient during the communication between two nodes, and the scale of the gradient to be transmitted. In the case of a known specific deep learning model, any distributed system with N nodes can use step (5) to calculate the gradient communication time of the k-th layer of the deep learning model. Therefore, the present invention is applicable to distributed systems of any scale and any deep learning model. Description of the Drawings
[0063] Figure 1 is a flowchart of the method for improving the distributed data parallel training efficiency of the deep learning model of the present invention;
[0064] Figure 2 It is the classification accuracy curve of the deep neural network ResNet152 trained by the present invention. Specific Embodiments
[0065] In order to make the objectives, technical solutions and advantages of the present invention clearer, the present invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention. In addition, the technical features involved in the various embodiments of the present invention described below can be combined with each other as long as they do not conflict with each other.
[0066] The basic idea of the present invention is to provide a method and system for improving the distributed data parallel training efficiency of a deep learning model, which proposes a new gradient merging strategy for the deep learning model. The distributed data parallel training process of the new deep learning model is as follows: First, all nodes obtain the data set and the deep learning model parameters; then, pre-training is performed on the first node, and it is iteratively determined whether the gradient of the kth (1≤k<L) layer of the L-layer deep learning model can communicate with the gradient of the (k + 1)th layer according to the data generated by the pre-training, and the gradient merging strategy is stored and represented using a two-level list and a layer index number. The layer index numbers included in the first-level list in the two-level list indicate that the corresponding layers should communicate together. After the iteration is completed, the gradient merging strategy is broadcast to all other nodes; finally, all nodes start the formal training. In the parameter synchronization stage, all nodes perform grouped gradient communication according to the layer index numbers included in the first-level list in the two-level list. The present invention will be further described below with reference to the accompanying drawings and embodiments.
[0067] As Figure 1 shown, the present invention provides a method for improving the distributed data parallel training efficiency of a deep learning model, which is applied to a Horovod system including multiple nodes. The method includes the following steps:
[0068] (1) Each node obtains the data set;
[0069] (2) Each node initializes the deep learning model to obtain the initialized deep learning model;
[0070] Specifically, the initial value of the weight parameter is a random value output using a truncated normal distribution with a standard deviation of 0.1, the initial value of the bias parameter is set to 0, the initial learning rate lr = 0.0001, and the learning strategy of stochastic gradient descent is adopted, with a step size of stepsize = 200 and a weight gamma = 0.1, that is, the learning rate is multiplied by 0.1 every 200 iterations;
[0071] (3) The first node performs n (where n ranges from 5 to 10, preferably 5) iterative pre-trainings on the initialized deep learning model and obtains the average value t of the forward propagation calculation times obtained from the n iterative pre-trainings f and the average value t of the backward propagation calculation times obtained from the n iterative pre-trainings b , the start timestamp of the backward propagation calculation of the l-th layer of the deep learning model and the gradient p of the l-th layer of the deep learning model (l) , and there is represents the forward propagation calculation time of the l-th layer in the deep learning model, represents the backward propagation calculation time of the l-th layer in the deep learning model, and l ∈ [1, L], where L is the number of layers of the deep learning model;
[0072] The advantage of this step is that the constants t b , p (l) necessary for the subsequent steps can be calculated, and the error from the formal training is reduced by taking the average value.
[0073] (4) The first node sets the counter k = 1;
[0074] (5) The first node calculates the gradient communication time of the k-th layer of the deep learning model α represents the startup duration of communication among all nodes in the Horovod system, β represents the time occupied by transmitting gradients of each byte size during communication between two nodes in the Horovod system, and sets k = k + 1;
[0075] (6) The first node determines whether k is greater than L. If so, it proceeds to step (7); otherwise, it proceeds to step (5);
[0076] The advantage of the above steps (4) to (6) is that the gradient communication times of each layer of the deep learning model necessary for the subsequent steps can be calculated
[0077] (7) The first node sets the counter i = L - 1 and sets the start timestamp of the gradient communication of the L-th layer of the deep learning model
[0078] (8) The first node obtains the start timestamp of the gradient communication of the i-th layer of the deep learning model and sets the counter i = i - 1;
[0079] (9) The first node determines whether i is less than 1. If so, it proceeds to step (10); otherwise, it proceeds to step (8);
[0080] The advantages of the above steps (7) to (9) are that the start timestamps of the gradient communications of each layer of the L-layer deep learning model required for the subsequent steps can be calculated.
[0081] (10) The first node sets the counter j = 1, the counter g = 1, initializes the secondary list group[:] of the merging group and the primary list m[:].
[0082] (11) The first node determines Whether it holds. If it is, go to step (12); otherwise, go to step (13).
[0083] The advantage of this step is to determine whether the gradient of the j-th layer of the deep learning model should be communicated together with the gradient of the (j + 1)-th layer. This determination condition includes the overlapping time wasted by gradient merging, that is, if the overlapping time wasted by gradient merging is higher than the startup duration saved by gradient merging, the gradient of the k-th layer of the deep learning model should not be communicated together with the gradient of the (k + 1)-th layer, thus avoiding excessive gradient merging.
[0084] (12) The first node adds the layer index number j of the j-th layer of the deep learning model to the primary list m[:], sets Sets And goes to step (14).
[0085] (13) The first node sets g = g + 1, adds the primary list m[:] to the secondary list group[:] of the merging group, clears the primary list m[:], and goes to step (14).
[0086] The advantages of the above steps (10) to (13) are that the gradient merging grouping strategy of all layers of the deep learning model can be stored in the secondary list in the form of layer index numbers, that is, the layers of the deep learning model corresponding to the layer index numbers included in the primary list in the secondary list should be communicated together, so that the subsequent steps perform gradient communication according to the secondary list.
[0087] (14) The first node sets j = j + 1, and uses the Obtained in step (12) and the And To update the start timestamps of the gradient communications of the j-th layer to the first layer in the deep learning model;
[0088] Specifically, in this step, the Obtained in step (12) and the And The process of updating the start timestamp of gradient communication from the j-th layer to the 1st layer in the deep learning model includes the following sub-steps:
[0089] (14-1) Set the counter q = j,
[0090] (14-2) Determine whether q is equal to L. If so, set the start timestamp of gradient communication of the L-th layer of the deep learning model to Then go to step (14-4); otherwise, go to step (14-3).
[0091] (14-3) The first node obtains the start timestamp of gradient communication of the q-th layer of the deep learning model
[0092] (14-4) Set the counter q = q - 1;
[0093] (14-5) Determine whether q is less than 1. If so, the process ends; otherwise, return to step (14-3).
[0094] The advantages of steps (14-1) to (14-5) are that, according to the change in step (12), update the start timestamp of gradient communication from the j-th layer to the 1st layer of the deep learning model in a timely manner. If it transfers to step (11) under the judgment condition of step (15), use the updated start timestamp for judgment, so as to ensure the accuracy of the gradient merging strategy and the convergence of deep learning training.
[0095] (15) The first node determines whether j < L. If so, return to step (11); otherwise, go to step (16);
[0096] (16) The first node sends the merged group secondary list group[:] to all other nodes in the Horovod cluster in a broadcast manner;
[0097] (17) Each node sequentially performs forward propagation calculations on the deep learning model from the 1st layer to the L-th layer, and sequentially performs backward propagation calculations on the deep learning model from the L-th layer to the 1st layer to obtain the gradients of each layer, and uses the obtained gradients of each layer in the deep learning model to call the Ring-AllReduce communication function in the Horovod system to perform gradient communication.
[0098] Specifically, in this step, the deep learning model sequentially performs backward propagation calculations from the L-th layer to the 1st layer to obtain the gradients of each layer. At the same time, according to the deep learning model layer index numbers in the merged group secondary list group[:], the process of performing grouped communication on the obtained gradients includes the following sub-steps:
[0099] (17-1) Set the counter w = L;
[0100] (17-2) All nodes perform backpropagation calculations on the w-th layer of the deep learning model and store the gradients of the w-th layer of the obtained deep learning model in the local buffer;
[0101] (17-3) Determine whether the number of inner index numbers of the first-level list where the w-th layer of the deep learning model is located in the merged group secondary list group[:] is only 1. If so, go to step (17-4); otherwise, go to step (17-5);
[0102] The advantage of this step is that if the judgment result is "yes", the gradient communication of the w-th layer of the deep learning model can be performed in a timely manner without communicating with other layers, thus avoiding wasting the overlap time.
[0103] (17-4) All nodes perform gradient communication, that is, each node calls the Ring-AllReduce communication function in the Horovod system and uses the gradients in the local buffer as parameters to perform gradient communication, and then go to step (17-6);
[0104] (17-5) All nodes determine whether the layer index number of the w-th layer of the deep learning model in the merged group secondary list group[:] is the minimum value of the index numbers in the first-level list where it is located. If so, go to step (17-4); otherwise, perform step (17-6);
[0105] The advantage of this step is that if the judgment result is "yes", the layers corresponding to all the index numbers included in the first-level list where the layer index number of the w-th layer of the deep learning model is located are communicated together; if the judgment result is "no", it is necessary to wait for all the layers corresponding to the index numbers in the first-level list where the layer index number of the w-th layer of the deep learning model is located to complete the backpropagation calculation before performing gradient communication, so as to perform gradient communication for multiple layers together to reduce the communication frequency and save the startup duration.
[0106] (17-6) Set w = w - 1 and wait for the gradient communication of all nodes to end;
[0107] (17-7) Determine whether w < 1. If so, the process ends; otherwise, return to step (17-2).
[0108] The advantages of steps (17-1) to (17-7) are that gradient communication in a grouped manner is performed for all layers of the deep learning model, thereby reducing the communication frequency and saving the startup duration.
[0109] In summary, through the above description of the present invention, the main advantages of the present invention are as follows: A method for improving the training efficiency of a deep learning model is proposed, which weighs the overlapping time wasted by gradient merging and the startup time saved, and can save the startup time without wasting the overlapping time to reduce the time of each iteration, thereby improving the deep learning training efficiency.
[0110] Test results
[0111] Apply the distributed communication method based on communication-computation overlap and gradient merging to ResNet50, ResNet152, DenseNet161, DenseNet201, and GoogLeNet deep neural network models. Before the formal training starts, obtain the grouping results of each model, and perform grouped communication at each iteration. As shown in Table 1, it can be seen that compared with the MG-WFBP method proposed in the "Background Art" of the present invention, the optimization rate of the single-iteration time of the distributed DL task based on the communication-computation overlap technology native to the Horovod platform by applying the present invention is higher.
[0112] Table 1
[0113]
[0114] Figure 2 Verify the accuracy of the gradient merging grouping strategy, that is, verify whether the gradients of the kth (1 ≤ k ≤ L) layer of the L-layer deep learning model update the parameters of the kth layer of the deep learning model correspondingly. During the training process of 17 rounds in the figure, the accuracy slightly decreases in the 3rd round because the learning rate is too large, and then the accuracy curve rises smoothly as the learning rate decreases, indicating that the model is gradually converging, thus verifying the accuracy of the gradient merging grouping strategy.
[0115] Those skilled in the art can easily understand that the above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, equivalent replacements, and improvements made within the spirit and principle of the present invention should be included in the protection scope of the present invention.
Claims
1. A method for improving the efficiency of distributed data parallel training of a deep learning model, which is applied in a Horovod system including multiple nodes, is characterized in that The method includes the following steps: (1) Each node obtains a data set; (2) Each node initializes a deep learning model to obtain an initialized deep learning model; (3) The first node performs n - iteration pre - training on the initialized deep - learning model and obtains the average value t of the forward - propagation calculation times obtained from the n - iteration pre - training f , the average value t of the backward - propagation calculation times obtained from the n - iteration pre - training b , the start timestamp of the backward - propagation calculation of the l - th layer of the deep - learning model and the gradient p of the l - th layer of the deep - learning model (l) , and there are represents the forward - propagation calculation time of the l - th layer in the deep - learning model, represents the backward - propagation calculation time of the l - th layer in the deep - learning model, and l ∈ [1, L], where L is the number of layers of the deep - learning model; (4) The first node sets the counter k = 1; (5) The first node calculates the gradient communication time of the k-th layer of the deep learning model α represents the startup duration of communication among all nodes in the Horovod system, β represents the time occupied by transmitting gradients of each byte size during communication between two nodes in the Horovod system, and set k = k + 1; (6) The first node determines whether k is greater than L. If so, it proceeds to step (7); otherwise, it proceeds to step (5); (7) The first node sets the counter i = L - 1 and sets the start timestamp of the gradient communication of the L-th layer of the deep learning model (8) The first node obtains the start timestamp of the gradient communication of the i-th layer of the deep learning model And set the counter i = i - 1; (9) The first node determines whether i is less than 1. If so, it proceeds to step (10); otherwise, it proceeds to step (8); (10) The first node sets the counter j = 1, the counter g = 1, initializes the merged group secondary list group[:] and the primary list m[:]; (11) The first node determines whether it holds. If so, go to step (12); otherwise, go to step (13). (12) The first node adds the layer index number j of the j-th layer of the deep learning model to the first-level list m[:], sets Set and proceeds to step (14); (13) The first node sets g = g + 1, adds the primary list m[:] to the merged group secondary list group[:], clears the primary list m[:], and proceeds to step (14); (14) The first node sets j = j + 1 and uses what is obtained in step (12) and what is obtained in step (3) as well as to update the start timestamps of gradient communication for the j-th layer to the first layer in the deep learning model; (15) The first node determines whether there is j < L. If so, it returns to step (11); otherwise, it proceeds to step (16); (16) The first node broadcasts the merged group secondary list group[:] to all other nodes in the Horovod cluster; (17) Each node sequentially performs forward propagation calculations on the deep learning model from the first layer to the Lth layer, performs backward propagation calculations on the deep learning model from the Lth layer to the first layer in sequence to obtain the gradients of each layer, and uses the gradients of each layer in the obtained deep learning model to call the Ring-AllReduce communication function in the Horovod system to perform gradient communication.
2. The method for improving the distributed data parallel training efficiency of the deep learning model according to claim 1, wherein In step (2), the initial values of the weight parameters are random values output using a truncated normal distribution with a standard deviation of 0.1, the initial value of the bias parameter is set to 0, the initial learning rate lr = 0.0001, the stochastic gradient descent learning strategy is adopted, the step size stepsize = 200, and the weight gamma = 0.1, that is, the learning rate is multiplied by 0.1 every 200 iterations.
3. The method for improving the distributed data parallel training efficiency of a deep learning model according to claim 1 or 2, characterized in that Step (14) includes the following sub-steps: (14-1) Set the counter q = j, (14-2) Determine whether q is equal to L. If so, set the start timestamp of gradient communication for the L-th layer of the deep learning model to Then go to step (14-4); otherwise, go to step (14-3). (14-3) The first node obtains the start timestamp of the gradient communication of the q-th layer of the deep learning model (14-4) Set the counter q = q - 1; (14-5) Determine whether q is less than 1. If so, the process ends; otherwise, return to step (14-3).
4. The method for improving the distributed data parallel training efficiency of the deep learning model according to claim 3, wherein, Step (17) includes the following sub-steps: (17-1) Set the counter w = L; (17-2) All nodes perform backward propagation calculations on the wth layer of the deep learning model and store the obtained gradients of the wth layer of the deep learning model in the local buffer; (17-3) Determine whether the number of inner index numbers of the primary list where the wth layer of the deep learning model is located in the merged group secondary list group[:] is only 1. If so, proceed to step (17-4); otherwise, proceed to step (17-5); (17-4) All nodes perform gradient communication, that is, each node calls the Ring-AllReduce communication function in the Horovod system and uses the gradients in the local buffer as parameters to perform gradient communication, and then proceeds to step (17-6); In all nodes, judge whether the layer index number of the w-th layer of the deep learning model in the secondary list group[:] of the merge group is the minimum value of the index numbers in the corresponding primary list. If so, go to step (17-4); otherwise, execute step (17-6). Set w = w - 1 and wait for the gradient communication of all nodes to end. Judge whether w < 1. If so, end the process; otherwise, return to step (17-2).
5. A system for improving the distributed data parallel training efficiency of a deep learning model, which is applied in a Horovod system including multiple nodes, is characterized in that The system includes: The first module, which is set in each node and is used to obtain the data set. The second module, which is set in each node and is used to initialize the deep learning model to obtain the initialized deep learning model. The third module, which is set in the first node, is used to perform n - iteration pre - training on the initialized deep - learning model and obtain the average value t of the forward - propagation calculation times obtained from the n - iteration pre - training f , the average value t of the back - propagation calculation times obtained from the n - iteration pre - training b , the start timestamp of the back - propagation calculation of the l - th layer of the deep - learning model and the gradient p of the l - th layer of the deep - learning model (l) , and there is indicating the forward - propagation calculation time of the l - th layer in the deep - learning model, indicating the back - propagation calculation time of the l - th layer in the deep - learning model, and l ∈ [1, L], where L is the number of layers of the deep - learning model; The fourth module, which is set in the first node and is used to set the counter k = 1. The fifth module is arranged in the first node and is used to calculate the gradient communication time of the k-th layer of the deep learning model α represents the startup duration of communication among all nodes in the Horovod system, β represents the time taken to transmit gradients of each byte size during communication between two nodes in the Horovod system, and k = k + 1 is set; The sixth module, which is set in the first node and is used to judge whether k is greater than L. If so, go to the seventh module; otherwise, go to the fifth module. The seventh module, which is set in the first node, is used to set the counter i = L - 1 and set the start timestamp of the gradient communication of the L-th layer of the deep learning model The eighth module, which is set in the first node, is used to obtain the start timestamp of the gradient communication of the i-th layer of the deep learning model And set the counter i = i - 1; The ninth module, which is set in the first node and is used to judge whether i is less than 1. If so, go to the tenth module; otherwise, go to the eighth module. The tenth module, which is set in the first node and is used to set the counter j = 1, the counter g = 1, initialize the secondary list group[:] of the merge group and the primary list m[:]. The eleventh module, which is set in the 1st node, is used to judge whether it holds. If so, it transfers to the twelfth module; otherwise, it transfers to the thirteenth module. The twelfth module, which is set in the first node, is used to add the layer index number j of the j-th layer of the deep learning model to the first-level list m[:], and then set transfer to the fourteenth module; The thirteenth module, which is set in the first node and is used to set g = g + 1, add the primary list m[:] to the secondary list group[:] of the merge group, clear the primary list m[:], and then go to the fourteenth module. The fourteenth module, which is set in the first node, is used to set j = j + 1 and use what is obtained by the twelfth module and what is obtained by the third module as well as to update the start timestamps of gradient communication for the j-th layer to the first layer in the deep learning model; The fifteenth module, which is set in the first node and is used to judge whether j < L. If so, return to the eleventh module; otherwise, go to the sixteenth module. The sixteenth module, which is set in the first node and is used to broadcast the secondary list group[:] of the merge group to all other nodes in the Horovod cluster. The seventeenth module, which is set in each node and is used to perform forward propagation calculations on the deep learning model layer by layer from the first layer to the L-th layer, perform backward propagation calculations on the deep learning model layer by layer from the L-th layer to the first layer to obtain the gradients of each layer, and use the obtained gradients of each layer in the deep learning model to call the Ring-AllReduce communication function in the Horovod system to perform gradient communication.
Citation Information
Patent Citations
Model parameter training method, server, system and storage medium
CN113283596A
Image quality in spin echo based imaging with parallel imaging
US20200191894A1