Parameter Synchronization Method and System for Deep Learning Model Distributed Training System
In the deep learning model distributed training system, the management node querys the parameter server for the completion status of the training task and obtains model parameters, and solves the problem of poor model effectiveness caused by incomplete parameters, and achieves complete acquisition and improvement of model parameters.
Patent Information
- Application Number
- CN202210311440.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-03-28
- Publication Date
- 2025-07-18
- Estimated Expiration
- 2042-03-28
AI Technical Summary
During the distributed training process of deep learning models, the distributed model parameters obtained by the management node at the end of training are incomplete, resulting in poor results in the trained model.
When the training task is completed, the management node sends query instructions to the target parameter server to determine whether the number of nodes completing the training task is equal to the preset value. If so, the trained model parameters are obtained from the parameter server and saved in the distributed file system; if not, it enters a dormant state and try regularly.
Ensure the integrity of the distributed model parameters obtained by the management node, and improve the effectiveness of the trained model.
Smart Images

Figure CN114625514B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of deep learning models, and in particular to a method and system for parameter synchronization of a distributed training system of a deep learning model. Background Art
[0002] When solving problems such as image detection, text classification, and speech recognition, deep learning distributed training is often used to classify massive amounts of image data and text data. For example, Google's AlphaGo program needs to be trained with chess records in the order of millions or even tens of millions.
[0003] In the scenario of deep learning distributed training, there is a typical distributed training of a parameter server. As Figure 1 shown, the distributed training of the parameter server includes multiple training roles, mainly including three types: a management node Master, an execution node Worker, and a parameter server node Parameter Server (hereinafter simply referred to as PS). Among them, the management node Master can also be replaced by one of the execution nodes Worker, such as replacing it with the execution node Worker with index 0.
[0004] Among them, both the management node Master and the execution node Worker are used to train the deep learning model, and the PS is used to save the parameters of the deep learning model. There is only one management node Master, there is one or more execution nodes Worker, and there is one or more PS. The management node Master (or the execution node Worker with index 0) not only participates in the training, but also controls the entire training process. When the training ends, it will save the entire trained model parameters. During the training process, the management node Master / execution node Worker will read the deep learning model parameters from the PS each time, perform training, and then update the parameters to the PS.
[0005] As Figure 2 shown, in the distributed training of the deep learning model, the Master / Worker will read the training data set and perform training. In order to prevent the Master / Worker from repeatedly training the data, the data set will be independently segmented. For example, if there are 100 data set training files, 1 Master and 9 Workers, then these 100 data set training files can be segmented into 10 parts, each part having 10 files. Then the Master and 9 Workers can independently read and train these 10 parts of data set training files, and each data set file has 10. Moreover, these data set training files will be saved in the distributed file system.
[0006] As Figure 3 shown, in addition to participating in the training of the dataset, Master also participates in the control of the training process and the saving of model parameters. After completing all the dataset training, Master will read out the finally trained model parameters from all the PSs and save them as a model in the distributed file system.
[0007] During the distributed training process, Master and Worker independently train the data assigned to them respectively, and the entire training process will not be synchronized with each other. Moreover, during the actual physical cluster training, the machines where Master and Worker are located may, due to various reasons, such as different CPU loads, different network states, etc., result in different training speeds for Master and Worker. It is possible that Master is faster, or Worker is faster, or Master is faster than some Workers but slower than some other Workers.
[0008] If Master is much faster than Worker, resulting in Master exiting only after completing the training of the dataset assigned to itself and only obtaining and saving the currently trained model parameters from the PS, at this time, in fact, Worker still has some datasets not trained, thus resulting in poor model performance.
[0009] In addition, due to distributed training, the startup order of Master and Worker will be different. If the total amount of the training dataset is relatively small, it is even possible that the Worker task has not started yet, and Master has already completed the training of its own dataset, which will also result in poor model performance.
[0010] In the prior art, for example, the invention patent "A Method for Updating Global Parameters in a Parameter Server" with the publication (announcement) number: CN109032630A (CN109032630B) and the applicant being the University of Electronic Science and Technology of China proposes to dynamically update global parameters according to the latency of weight parameters, thereby reducing the impact of high latency on the algorithm. However, this method is complex, will additionally occupy the computing resources of the execution nodes, and requires high-performance execution nodes for hardware support.
[0011] Another example is the invention patent "Distributed Job Coordination Control Method, Device, Computer Equipment and Storage Medium" with the publication (announcement) number: CN109101342A (CN109101342B) and the applicant being Beijing Baidu Netcom Science and Technology Co., Ltd., which proposes to achieve the coordination control of distributed jobs based on containers. However, this method requires improvements in both the parameter server node and the execution node, and the algorithm is relatively complex.
[0012] For another example, the invention patent "Method, Apparatus and System for Training Deep Neural Network" with the publication (announcement) number: CN107688493A (CN107688493B) and the applicant being Alibaba Group Holding Limited proposes to train the training data of each mapping layer through a parameter server and at least one computing node, but this method cannot solve the problem of synchronous training of each computing node. Summary of the Invention
[0013] To at least overcome the problems existing in the related art to a certain extent, the present invention provides a parameter synchronization method and system for a distributed training system of a deep learning model to solve the problem that in the prior art, when a deep learning model is distributedly trained, the distributed model parameters obtained by the management node at the end of the training are incomplete, resulting in a poor model effect.
[0014] According to the first aspect of the embodiments of the present invention, there is provided a parameter synchronization method for a distributed training system of a deep learning model. The system includes a management node, at least one execution node, and at least one parameter server. The method is applicable to the management node and includes:
[0015] In response to a training task completion instruction, send a query instruction to the target parameter server, where the query instruction carries a training task completion identifier;
[0016] Receive the query result sent by the target parameter server, where the query result carries data representing the number of management nodes and execution nodes that have completed the training task;
[0017] Judge whether the number of nodes that have completed the training task is equal to a preset value; if so, obtain the trained model parameters from the target parameter server and save the model parameters as a model in a distributed file system; if not, enter a completion sleep state;
[0018] The target parameter server is used to receive the training task completion identifier sent by the management node and the node identifier of the management node itself, and receive the training task completion identifier sent by the execution node and the node identifier of the execution node itself, and store the node identifiers according to the training task completion identifier, and query the number of node identifiers of the completed training task stored, and generate the query result.
[0019] Further, after entering the completion sleep state, it further includes:
[0020] Every first preset time period, send the query instruction to the target parameter server;
[0021] Receive the query result sent by the target parameter server;
[0022] When the number of nodes that have completed the training task shown in the query result is equal to the preset value, stop sending the query instruction to the target parameter server.
[0023] Further, before sending the query instruction to the target parameter server, it further includes:
[0024] In response to the training task completion instruction, send the training task completion flag and its own node serial number to the target parameter server, so that the target parameter server updates the stored node serial numbers of the nodes that have completed the training task according to the training task completion flag.
[0025] Further, the method further includes:
[0026] In response to the training task having been accurately started instruction, send a query instruction to the target parameter server, and the query instruction carries a training task ready-to-start flag;
[0027] Receive the query result sent by the target parameter server, and the query result carries data indicating the management nodes and the number of execution nodes that are ready to start the training task;
[0028] Judge whether the number of nodes that are ready to start the training task is equal to the preset value; if so, obtain the training data set and model parameters from the parameter server for training; if not, enter the sleep state.
[0029] Further, after entering the sleep state, it further includes:
[0030] Every second preset time period, send the query instruction to the target parameter server;
[0031] Receive the query result sent by the target parameter server;
[0032] When the query result shows that the number of nodes that are ready to start the training task is equal to the preset value, stop sending the query instruction to the target parameter server.
[0033] Further, before sending the query instruction to the target parameter server, it further includes:
[0034] In response to the training task having been accurately started instruction, send the training task ready-to-start flag and its own node serial number to the target parameter server, so that the target parameter server updates the stored node serial numbers of the nodes that start the training task according to the training task ready-to-start flag.
[0035] According to the second aspect of the embodiments of the present invention, a management node is provided, which is applicable to a deep learning model distributed training system and includes:
[0036] A memory for storing computer programs that can be called;
[0037] A processor that executes the parameter synchronization method of the above-mentioned deep learning model distributed training system by calling the computer program stored in the memory.
[0038] According to the third aspect of the embodiments of the present invention, a parameter synchronization method for a deep learning model distributed training system is provided. The system includes a management node, at least one execution node, and at least one parameter server. The method is applicable to the parameter server and includes:
[0039] Receiving a query instruction sent by the management node, where the query instruction carries a training task completion flag;
[0040] Receiving the training task completion flag sent by the management node and the node identifier of the management node itself, and receiving the training task completion flag sent by the execution node and the node identifier of the execution node itself, and storing the node identifiers according to the training task completion flag;
[0041] Querying the data of the management nodes and the number of execution nodes that have completed the training task stored according to the training task completion flag, and generating a query result;
[0042] Sending the query result to the management node, so that the management node determines that when the number of nodes that have completed the training task is equal to a preset value, it sends a model parameter acquisition instruction;
[0043] Receiving the model parameter acquisition instruction sent by the management node, and sending the trained model parameters to the management node, so that the management node saves the model parameters as a model in the distributed file system.
[0044] Further, the method further includes:
[0045] Receiving the training task completion flag sent by the management node and the node serial number of the management node itself, and at the same time, receiving the training task completion flag sent by the execution node and the node serial number of the execution node itself;
[0046] Storing the node serial numbers in a training task completion queue according to the training task completion flag;
[0047] The training task completion queue is pre-deployed in the parameter server.
[0048] Further, the method further includes:
[0049] Receiving a query instruction sent by the management node or the execution node, where the query instruction carries a training task ready to start flag;
[0050] Based on the training task's ready-to-start flag, query the stored data of the management nodes and the number of execution nodes that are ready to start the training task, and generate a query result;
[0051] Send the query result to the management node or the execution node, so that when the management node or the execution node determines that the number of nodes ready to start the training task is equal to the preset value based on the query result, send an instruction to obtain the training dataset and model parameters;
[0052] Receive the instruction to obtain the training dataset and model parameters sent by the management node or the execution node, and send the training dataset and model parameters to the corresponding management node or execution node, so that the management node or the execution node starts training.
[0053] Further, the method further includes:
[0054] Receive the training task's ready-to-start flag and the node serial number of the management node itself sent by the management node, and at the same time, receive the training task's ready-to-start flag and the node serial number of the execution node itself sent by the execution node;
[0055] Based on the training task's ready-to-start flag, store the node serial number in the training task start queue;
[0056] The training task start queue is pre-deployed in the parameter server.
[0057] According to the fourth aspect of the embodiments of the present invention, a parameter server is provided, which is applicable to a deep learning model distributed training system, and includes:
[0058] A memory for storing a computer program that can be called;
[0059] A processor, by calling the computer program stored in the memory, executes the parameter synchronization method of the above-mentioned deep learning model distributed training system.
[0060] According to the fifth aspect of the embodiments of the present invention, a deep learning model distributed training system is provided, including:
[0061] The above-mentioned management node, at least one execution node and at least one of the above-mentioned parameter servers, where,
[0062] The management node and the execution node are respectively communicatively connected to the parameter server.
[0063] According to the sixth aspect of the embodiments of the present invention, a parameter synchronization method for a deep learning model distributed training system is provided. The system includes a management node, at least one execution node and at least one parameter server. The method includes:
[0064] Upon receiving the training task completion instruction, the management node sends a query instruction to the target parameter server, and the query instruction carries a training task completion flag.
[0065] The parameter server receives the training task completion flag sent by the management node and the node identifier of the management node itself, and receives the training task completion flag sent by the execution node and the node identifier of the execution node itself, and stores the node identifiers according to the training task completion flag; and receives the query instruction sent by the management node, and queries the data of the number of management nodes and execution nodes that have completed the training task stored according to the training task completion flag, and generates a query result.
[0066] The management node receives the query result sent by the target parameter server, and determines whether the number of nodes that have completed the training task is equal to a preset value; if so, obtains the trained model parameters from the target parameter server, and saves the model parameters as a model in the distributed file system; if not, enters the completion sleep state.
[0067] The technical solution provided by the embodiments of the present invention may include the following beneficial effects:
[0068] When the training task of the management node ends, the management node sends a query instruction to the target parameter server. According to the query result, it determines whether the number of nodes that have completed the training task is equal to a preset value; if so, obtains the trained model parameters from the target parameter server, and saves the model parameters as a model in the distributed file system, thereby ensuring the integrity of the distributed model parameters obtained by the management node, and solving the problem that in the prior art, when the deep learning model is distributedly trained, due to the incomplete distributed model parameters obtained by the management node at the end of the training, the trained model has poor performance.
[0069] It should be understood that the above general description and the following detailed description are only exemplary and explanatory, and cannot limit the present invention. BRIEF DESCRIPTION OF THE DRAWINGS
[0070] The accompanying drawings herein are incorporated into the specification and constitute a part of the specification, showing embodiments consistent with the present invention, and are used together with the specification to explain the principles of the present invention.
[0071] Figures 1 to 3 is a schematic diagram of the distributed training process of the deep learning model shown according to the background art;
[0072] Figure 4 is a flowchart of a parameter synchronization method for a distributed training system of a deep learning model shown according to an exemplary embodiment;
[0073] Figure 5It is a flowchart of a parameter synchronization method for a distributed training system of a deep learning model shown according to another exemplary embodiment;
[0074] Figure 6 It is a flowchart of a parameter synchronization method for a distributed training system of a deep learning model shown according to another exemplary embodiment;
[0075] Figures 7 to 12 It is a schematic diagram of the process of a parameter synchronization method for a distributed training system of a deep learning model shown according to another exemplary embodiment. Detailed implementation manners
[0076] Here, the exemplary embodiments will be described in detail, and the examples are shown in the drawings. When the following description refers to the drawings, unless otherwise indicated, the same numbers in different drawings represent the same or similar elements. The implementation manners described in the following exemplary embodiments do not represent all implementation manners consistent with the present invention. On the contrary, they are merely examples of devices and methods consistent with some aspects of the present invention as detailed in the appended claims.
[0077] As described in the background art above, in the prior art, when training a deep learning model in a distributed manner, there is a problem that the distributed model parameters obtained by the management node at the end of the training are incomplete, resulting in a poor model effect. To solve this technical problem, the present invention provides the following several exemplary embodiments:
[0078] Embodiment 1
[0079] Figure 4 It is a flowchart of a parameter synchronization method for a distributed training system of a deep learning model shown according to an exemplary embodiment. The system includes a management node, at least one execution node, and at least one parameter server. As Figure 4 shown, this method is applicable to the management node and includes:
[0080] Step S11: In response to a training task completion instruction, send a query instruction to the target parameter server, where the query instruction carries a training task completion flag;
[0081] Step S12: Receive the query result sent by the target parameter server, where the query result carries data representing the number of management nodes and execution nodes that have completed the training task;
[0082] Step S13: Determine whether the number of nodes that have completed the training task is equal to a preset value; if so, obtain the trained model parameters from the target parameter server and save the model parameters as a model in the distributed file system; if not, enter a sleep state of completion.
[0083] It should be noted that the technical solution provided in this embodiment is applicable to the application scenario of distributed training of deep learning models. In specific practice, the management node of the deep learning model distributed training system can be selected from multiple execution nodes, and generally there is one management node.
[0084] The management node and the execution nodes operate independently. When the management node completes the training task assigned to it by the system, it does not know whether other execution nodes have also completed the training task. Therefore, in step S11, when the management node completes the training task assigned to it by the system, it will send a query instruction to the target parameter server. The query instruction carries a training task completion flag, which is to inform the target parameter server that the management node currently wants to query the number of nodes that have completed the training task.
[0085] Since the number of parameter servers for different training tasks is different, in specific practice, the deep learning model distributed training system includes one or more parameter servers. When there are multiple parameter servers, the target parameter server can be selected from multiple parameter servers.
[0086] Preferably, selecting the parameter server with the index number 001 as the target parameter server can ensure that the target parameter server always exists regardless of the number of parameter servers. In this way, the provided technical solution can be applicable to distributed training of various data scales, with a wide application scenario and strong compatibility.
[0087] It should be noted that in step S13, the preset value is set according to the user's needs, and generally set as: preset value = the number of management nodes + the number of execution nodes. The advantage of this setting is that the target parameter server stores the number of management nodes and execution nodes that have completed the training task. When the number of management nodes and execution nodes that have completed the training task is equal to the preset value, it means that all nodes have completed the training. At this time, obtaining the trained model parameters from the target parameter server and saving the model parameters as a model in the distributed file system can fully ensure the integrity of the obtained distributed model parameters.
[0088] It can be understood that in the technical solution provided in this embodiment, when the training task of the management node ends, it sends a query instruction to the target parameter server. According to the query result, it judges whether the number of nodes that have completed the training task is equal to the preset value; if so, it obtains the trained model parameters from the target parameter server and saves the model parameters as a model in the distributed file system, thereby ensuring the integrity of the distributed model parameters obtained by the management node and solving the problem in the prior art that when training a deep learning model distributively, due to the incomplete distributed model parameters obtained by the management node at the end of training, the trained model has poor performance.
[0089] Example Two
[0090] Figure 5 is a flowchart of a parameter synchronization method for a distributed training system of a deep learning model shown according to another exemplary embodiment. The system includes a management node, at least one execution node, and at least one parameter server. As Figure 5 shown, this method is applicable to the management node and includes:
[0091] Step S20: In response to a training task completion instruction, send a training task completion flag and its own node number to the target parameter server, so that the target parameter server updates the stored node numbers of the completed training tasks according to the training task completion flag;
[0092] Step S21: In response to a training task completion instruction, send a query instruction to the target parameter server, and the query instruction carries a training task completion flag;
[0093] Step S22: Receive the query result sent by the target parameter server, and the query result carries data representing the number of management nodes and execution nodes that have completed the training task;
[0094] Step S23: Determine whether the number of nodes that have completed the training task is equal to a preset value; if so, obtain the trained model parameters from the target parameter server and save the model parameters as a model in the distributed file system; if not, enter a sleep state;
[0095] Step S24: Every first preset duration, send the query instruction to the target parameter server;
[0096] Step S25: Receive the query result sent by the target parameter server;
[0097] Step S26: When the query result shows that the number of nodes that have completed the training task is equal to the preset value, stop sending the query instruction to the target parameter server.
[0098] It should be noted that the technical solution provided in this embodiment is applicable to the application scenario of distributed training of a deep learning model. In specific practice, the management node of the deep learning model distributed training system can be selected from multiple execution nodes, and generally there is one management node.
[0099] Compared with Embodiment One, this embodiment has additional steps S20, S24, S25, and S26. In Embodiment One, there is no introduction on how the target parameter server obtains the number of management nodes and execution nodes that have completed the training task. Step S20 in this embodiment discloses one of the implementation methods.
[0100] In this embodiment, after the management node completes the training task assigned to it by the system, it will send a training task completion flag and its own node serial number to the target parameter server, so that the target parameter server can update the stored node serial numbers of the completed training tasks according to the training task completion flag.
[0101] Since the execution node and the management node have the same nature, after the execution node completes the training task assigned to it by the system, it will also send a training task completion flag and its own node serial number to the target parameter server, so that the target parameter server can update the stored node serial numbers of the completed training tasks according to the training task completion flag.
[0102] Generally speaking, in specific practice, the target parameter server will perform the following steps:
[0103] Receive the training task completion flag sent by the management node and the management node's own node serial number. At the same time, receive the training task completion flag sent by the execution node and the execution node's own node serial number;
[0104] According to the training task completion flag, store the node serial number in the training task completion queue;
[0105] The training task completion queue is pre-deployed in the target parameter server.
[0106] It can be understood that since the management node and the execution node carry the training task completion flag when sending their own node serial numbers, the target parameter server will automatically store the node serial numbers with the task completion flag in the training task completion queue.
[0107] When the target parameter server receives a query instruction with a training task completion flag sent by the management node, it directly queries the number of stored node serial numbers in the training task completion queue and packages the queried number of nodes as a query result and feedbacks it to the management node.
[0108] In specific practice, step S20 and step S21 can overlap or be independent of each other. For example, the management node first sends its own node serial number to the target parameter server, and then sends a query instruction at a preset time interval (at this time, it no longer sends its own node serial number); or, when the management node sends its own node serial number to the target parameter server, it sends a query instruction once, and then sends a query instruction at a preset time interval (at this time, it no longer sends its own node serial number). The above steps S20 and S21 are only explanations of the step content and do not represent a limitation on the execution order.
[0109] In the newly added step S24 in the above compared with Embodiment 1, the first preset time length is set according to user needs. For example, it is set to 2 seconds.
[0110] After the management node enters the sleep state, then performing steps S24 to S26 is to prevent the management node from frequently accessing the parameter server, causing broadband occupation and additional query expenses. It can also enable the management node to obtain the trained model parameters from the target parameter server in a timely manner after other execution nodes complete the training tasks, and save the model parameters as a model stored in the distributed file system, improving the system response speed, reducing the user waiting time, and enhancing the user experience.
[0111] Embodiment III
[0112] A parameter synchronization method for a distributed training system of a deep learning model according to another exemplary embodiment. The system includes a management node, at least one execution node, and at least one parameter server. This method is applicable to the management node. Compared with Embodiment II, the additional steps include:
[0113] Step S30: In response to the accurate start instruction of the training task, send a training task ready-to-start flag and its own node serial number to the target parameter server, so that the target parameter server updates the stored node serial number for starting the training task according to the training task ready-to-start flag;
[0114] Step S31: In response to the accurate start instruction of the training task, send a query instruction to the target parameter server, and the query instruction carries the training task ready-to-start flag;
[0115] Step S32: Receive the query result sent by the target parameter server, and the query result carries data representing the number of management nodes and execution nodes that are ready to start the training task;
[0116] Step S33: Determine whether the number of nodes ready to start the training task is equal to a preset value; if so, obtain the training dataset and model parameters from the parameter server for training; if not, enter the sleep state;
[0117] Step S34: Send the query instruction to the target parameter server every second preset duration;
[0118] Step S35: Receive the query result sent by the target parameter server;
[0119] Step S36: When the query result shows that the number of nodes ready to start the training task is equal to the preset value, stop sending the query instruction to the target parameter server.
[0120] It should be noted that the technical solution provided in this embodiment is applicable to the application scenario of distributed training of deep learning models. In specific practice, the management node of the deep learning model distributed training system can be selected from multiple execution nodes, and generally there is one management node.
[0121] Compared with Embodiment 2, the newly added steps S30 to S36 in this embodiment are to ensure that the management node and the execution nodes start their respective training tasks synchronously, reduce the time for the management node to wait for other execution nodes to complete their respective training tasks when the training task ends, and ensure the integrity of the distributed model parameters obtained by the management node while ensuring the acquisition timeliness of the distributed model parameters.
[0122] In this embodiment, step S30 introduces an implementation manner of how the target parameter server obtains the management node and the number of execution nodes that are ready to start the training task.
[0123] In this embodiment, when the management node is ready to start the training task assigned to it by the system, it will send a training task ready-to-start flag and its own node number to the target parameter server, so that the target parameter server updates the stored node numbers of the nodes that are ready to start the training task according to the training task ready-to-start flag.
[0124] Since the execution nodes and the management node have the same nature, when the execution nodes are ready to start the training task assigned to them by the system, they will also send a training task ready-to-start flag and their own node numbers to the target parameter server, so that the target parameter server updates the stored node numbers of the nodes that are ready to start the training task according to the training task ready-to-start flag.
[0125] Generally speaking, in specific practice, the target parameter server will perform the following steps:
[0126] Receive the training task ready-to-start flag sent by the management node and the management node's own node number, and at the same time, receive the training task ready-to-start flag sent by the execution nodes and the execution nodes' own node numbers;
[0127] According to the training task ready-to-start flag, store the node numbers in the training task start queue;
[0128] The training task start queue is pre-deployed in the parameter server.
[0129] It can be understood that since the management node and the execution nodes carry the training task ready-to-start flag when sending their own node numbers, the target parameter server will automatically store the node numbers carrying the training task ready-to-start flag in the training task completion queue.
[0130] When the target parameter server receives a query instruction carrying the training task ready-to-start flag sent by the management node, it directly queries the number of stored node serial numbers in the training task start queue, and packages the queried number of nodes as a query result and feeds it back to the management node.
[0131] In specific practice, step S30 and step S31 can overlap or be independent of each other. For example, the management node first sends its own node serial number to the target parameter server, and then sends a query instruction every preset time interval (at this time, it no longer sends its own node serial number); or, when the management node sends its own node serial number to the target parameter server, it sends a query instruction once, and then sends a query instruction every preset time interval (at this time, it no longer sends its own node serial number). The above steps S30 and S31 are only explanations of the step content and do not represent a limitation on the execution order.
[0132] In this embodiment, the second preset time interval in step S34 is set according to user needs, and can be the same as or different from the first preset time interval in Embodiment 2. For example, it is set to 1 second.
[0133] Embodiment 4
[0134] A management node shown according to an exemplary embodiment is applicable to a deep learning model distributed training system. The management node includes:
[0135] A memory for storing computer programs that can be called;
[0136] A processor that, by calling the computer programs stored in the memory, executes the parameter synchronization method of the above-mentioned deep learning model distributed training system.
[0137] It can be understood that for the technical solution provided in this embodiment, when the training task of the management node ends, it sends a query instruction to the target parameter server. According to the query result, it judges whether the number of nodes that have completed the training task is equal to a preset value; if so, it obtains the trained model parameters from the target parameter server and saves the model parameters as a model in the distributed file system, thereby ensuring the integrity of the distributed model parameters obtained by the management node and solving the problem in the prior art that when a deep learning model is distributedly trained, the distributed model parameters obtained by the management node at the end of training are incomplete, resulting in a poor model training effect.
[0138] Embodiment 5
[0139] Figure 6 is a flowchart of a parameter synchronization method for a deep learning model distributed training system according to another exemplary embodiment. The system includes a management node, at least one execution node, and at least one parameter server, asFigure 6 As shown in Figure 6 , the method is applicable to a parameter server and includes:
[0140] Step S41: Receive a query instruction sent by a management node, where the query instruction carries a training task completion flag.
[0141] Step S42: According to the training task completion flag, query the stored data of the management node and the number of execution nodes that have completed the training task, and generate a query result.
[0142] Step S43: Send the query result to the management node, so that when the management node determines that the number of nodes that have completed the training task is equal to a preset value according to the query result, it sends a model parameter acquisition instruction.
[0143] Step S44: Receive the model parameter acquisition instruction sent by the management node, and send the trained model parameters to the management node, so that the management node saves the model parameters as a model in a distributed file system.
[0144] It should be noted that the technical solution provided in this embodiment is applicable to the application scenario of distributed training of a deep learning model, and is particularly applicable to the target parameter server mentioned in Embodiments 1 to 3. In specific practice, the management node of the deep learning model distributed training system can be selected from multiple execution nodes, and generally there is one management node.
[0145] In specific practice, the parameter server acquires and stores the management node and the number of execution nodes that have completed the training task. One implementation method can be:
[0146] Receive the training task completion flag sent by the management node and the node serial number of the management node itself. At the same time, receive the training task completion flag sent by the execution node and the node serial number of the execution node itself.
[0147] According to the training task completion flag, store the node serial numbers in a training task completion queue.
[0148] The training task completion queue is pre-deployed in the parameter server.
[0149] It can be understood that for the technical solution provided in this embodiment, when the training task of the management node ends, the management node sends a query instruction to the target parameter server. According to the query result, it is determined whether the number of nodes that have completed the training task is equal to the preset value. If so, the trained model parameters are obtained from the target parameter server and the model parameters are saved as a model in the distributed file system, thus ensuring the integrity of the distributed model parameters obtained by the management node and solving the problem in the prior art that when the deep learning model is distributedly trained, due to the incomplete distributed model parameters obtained by the management node at the end of the training, the trained model has poor performance.
[0150] Embodiment Six
[0151] A method for parameter synchronization of a deep learning model distributed training system according to another exemplary embodiment. The system includes a management node, at least one execution node, and at least one parameter server. This method is applicable to the parameter server. Compared with Embodiment Five, the additional steps include:
[0152] Step S51: Receive a query instruction sent by the management node or the execution node. The query instruction carries a training task ready-to-start flag.
[0153] Step S52: According to the training task ready-to-start flag, query the stored data of the management nodes and the number of execution nodes that are ready to start the training task, and generate a query result.
[0154] Step S53: Send the query result to the management node or the execution node, so that the management node or the execution node determines whether the number of nodes that are ready to start the training task is equal to the preset value according to the query result, and then sends a training dataset and a model parameter acquisition instruction.
[0155] Step S54: Receive the training dataset and the model parameter acquisition instruction sent by the management node or the execution node, and send the training dataset and the model parameters to the corresponding management node or execution node, so that the management node or the execution node starts the training.
[0156] In specific practice, the parameter server obtains and stores the number of management nodes and execution nodes that are ready to start the training task. One implementation method can be:
[0157] Receive the training task ready-to-start flag sent by the management node and the node serial number of the management node itself. At the same time, receive the training task ready-to-start flag sent by the execution node and the node serial number of the execution node itself.
[0158] According to the training task ready-to-start flag, store the node serial numbers in the training task start queue.
[0159] The training task startup queue is pre-deployed in the parameter server.
[0160] Compared with Embodiment 5, the newly added steps S51 to S54 in this embodiment are to ensure that the management node and the execution nodes start their respective training tasks synchronously, reduce the time for the management node to wait for other execution nodes to complete their respective training tasks when the training tasks end, and ensure the integrity of the distributed model parameters obtained by the management node while ensuring the acquisition timeliness of the distributed model parameters.
[0161] Embodiment 7
[0162] A parameter server shown according to an exemplary embodiment is applicable to a distributed training system of a deep learning model and includes:
[0163] A memory for storing computer programs that can be called;
[0164] A processor, by calling the computer programs stored in the memory, executes the parameter synchronization method of the above-mentioned distributed training system of the deep learning model.
[0165] It should be noted that the technical solution provided in this embodiment is applicable to the application scenario of distributed training of a deep learning model, especially applicable to the target parameter servers mentioned in Embodiments 1 to 3. In specific practice, the management node of the distributed training system of the deep learning model can be selected from multiple execution nodes, and generally there is one management node.
[0166] It can be understood that in the technical solution provided in this embodiment, when the management node finishes its own training task, it sends a query instruction to the target parameter server. According to the query result, it judges whether the number of nodes that have completed the training task is equal to the preset value; if so, it obtains the trained model parameters from the target parameter server and saves the model parameters as a model in the distributed file system, thereby ensuring the integrity of the distributed model parameters obtained by the management node and solving the problem in the prior art that when a deep learning model is distributedly trained, the distributed model parameters obtained by the management node at the end of training are incomplete, resulting in poor model training effects.
[0167] Embodiment 8
[0168] A distributed training system of a deep learning model shown according to an exemplary embodiment includes:
[0169] The above-mentioned management node, at least one execution node and at least one of the above-mentioned parameter servers, where
[0170] The management node and the execution nodes are respectively communicatively connected to the parameter server.
[0171] In specific practice, the parameter synchronization method of the deep learning model distributed training system includes:
[0172] Step S61: The management node responds to the training task completion instruction and sends a query instruction to the target parameter server. The query instruction carries a training task completion flag.
[0173] Step S62: The parameter server receives the query instruction sent by the management node, and according to the training task completion flag, queries the data of the management nodes and the number of execution nodes that have completed the training task stored, and generates a query result.
[0174] Step S63: The management node receives the query result sent by the target parameter server, and judges whether the number of nodes that have completed the training task is equal to the preset value; if so, obtains the trained model parameters from the target parameter server, and saves the model parameters as a model in the distributed file system; if not, enters the completion sleep state.
[0175] To facilitate the understanding of the parameter synchronization method of the deep learning model distributed training system provided in this embodiment, the management node is now represented as Master, the execution node is represented as Worker, the parameter server is represented as ParameterServer (hereinafter referred to as PS for short), the node serial number is represented as ID, and it is assumed that the Worker with index 0 is selected as Master, and there are still n Workers, n≥1; the PS with index 001 is selected as the target parameter server PS-1, combined with Figures 7 to 12 , the following is an example:
[0176] Preparation work: Deploy two ID queues on PS-1: the training task start queue (corresponding to the Start queue in Figures 7 to 12 ) and the training task completion queue (corresponding to the End queue in Figures 7 to 12 ).
[0177] Define two queue operations:
[0178] The push one operation means saving an ID to the queue.
[0179] The check equal operation means judging whether the length of the ID queue is equal to the expected length.
[0180] Start execution:
[0181] Step S71. In response to the accurate start instruction of the training task, the Master and the Workers that are ready to start the training task will send the training task ready-to-start flag and their own IDs to the target parameter server, perform the push one operation, and push their respective ID numbers into the Start queue (see Figure 7 ).
[0182] It should be noted that during the process of the Master and the Workers that are ready to start the training task pushing their respective IDs into the Start queue, since the start orders of different training tasks may be different, the IDs in the Start queue will not have a fixed order. Due to the technical solution provided in this embodiment, ultimately, the number of IDs in the Start queue is judged, rather than the order of the IDs in the Start queue. Therefore, this does not affect the implementation effect of this embodiment.
[0183] Step S72. The Master and the Workers that are ready to start the training task immediately perform the checkequal operation on the Start queue to determine whether the length of the ID queue meets n + 1.
[0184] Step S73. If the length of the Start queue is not equal to n + 1, then the Master and the Workers that are ready to start the training task enter the sleep state and re - execute the check equal operation every second preset duration until the length of the Start queue is n + 1 (see Figure 8 ). This step ensures that the Master and all Workers have been started and are ready to start training.
[0185] Step S74. After starting the training, the Master and all Workers read their respective independent training datasets and model parameters from their corresponding PSs for training, and update the model gradients to the PSs (see Figure 9 ).
[0186] Step S75. After their respective training tasks are completed, the Master and all Workers will perform the push one operation and push their respective ID numbers into the End queue (see Figure 10 ).
[0187] It should be noted that during the process of the Master and all Workers saving their respective IDs to the End queue, since the end orders of different training tasks may be different, the IDs in the End queue will not have a fixed order. Due to the technical solution provided in this embodiment, ultimately, the number of IDs in the End queue is judged, rather than the order of the IDs in the End queue. Therefore, this does not affect the implementation effect of this embodiment.
[0188] Step S76: When the training ends, the Worker will immediately exit its own training task (because the Worker has completed its own training task and no other operations need to be performed).
[0189] Step S77: The Master starts to perform a check equal operation on the End queue to determine whether the length of the End queue meets n + 1; if the length of the End queue is not equal to n + 1, then the Master enters the sleep state and re - executes the check equal operation every first preset duration until the length of the End queue is n + 1 (see Figure 11 ), this step ensures that the Master and all Workers have completed training.
[0190] Step S78: After completing the check equal operation on the End queue, the Master reads out the trained model parameters from each PS, and then saves the model to the distributed file system, so as to ensure that the finally saved model is the result of training with all samples (see Figure 12 ).
[0191] It can be understood that for the technical solution provided in this embodiment, when the management node finishes its own training task, it sends a query instruction to the target parameter server, and according to the query result, determines whether the number of nodes that have completed the training task is equal to the preset value; if so, it obtains the trained model parameters from the target parameter server and saves the model parameters as a model in the distributed file system, thus ensuring the integrity of the distributed model parameters obtained by the management node and solving the problem in the prior art that when a deep - learning model is distributedly trained, due to the incomplete distributed model parameters obtained by the management node at the end of the training, the trained model has a poor effect.
[0192] It can be understood that the same or similar parts in the above embodiments can be referred to each other, and the content not detailed in some embodiments can be seen in the same or similar content of other embodiments.
[0193] It should be noted that in the description of the present invention, terms such as "first", "second", etc. are only used for descriptive purposes and cannot be understood as indicating or implying relative importance. In addition, in the description of the present invention, unless otherwise stated, the meaning of "a plurality of" is at least two.
[0194] Any process or method description depicted in a flowchart or otherwise described herein can be understood to represent a module, segment, or portion of code including one or more executable instructions for implementing a specific logical function or process. The scope of the preferred embodiments of the present invention includes additional implementations where functions may be executed not in the order shown or discussed, including in a substantially simultaneous manner according to the involved functions or in a reverse order, which should be understood by those skilled in the art to which the embodiments of the present invention pertain.
[0195] It should be understood that various parts of the present invention can be implemented by hardware, software, firmware, or a combination thereof. In the above embodiments, multiple steps or methods can be implemented by software or firmware stored in a memory and executed by a suitable instruction execution system. For example, if implemented in hardware, as in another embodiment, any one or a combination of the following techniques well known in the art can be used: discrete logic circuits with logic gate circuits for implementing logical functions on data signals, application specific integrated circuits with suitable combinational logic gate circuits, programmable gate arrays (PGAs), field programmable gate arrays (FPGAs), etc.
[0196] Those of ordinary skill in the art of the present technology can understand that all or part of the steps carried by the methods of the above embodiments can be completed by instructing relevant hardware through a program, and the said program can be stored in a computer-readable storage medium. When the program is executed, it includes one or a combination of the steps of the method embodiments.
[0197] In addition, in each embodiment of the present invention, the functional units can be integrated into one processing module, or each unit can exist physically alone, or two or more units can be integrated into one module. The above integrated module can be implemented in the form of hardware or in the form of a software functional module. When the integrated module is implemented in the form of a software functional module and sold or used as an independent product, it can also be stored in a computer-readable storage medium.
[0198] The above-mentioned storage medium can be a read-only memory, a magnetic disk, an optical disk, etc.
[0199] In the description of this specification, the description with reference to terms such as "one embodiment", "some embodiments", "example", "specific example", or "some examples" means that the specific features, structures, materials, or characteristics described in connection with the embodiment or example are included in at least one embodiment or example of the present invention. In this specification, the schematic representations of the above terms do not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials, or characteristics described can be combined in a suitable manner in any one or more embodiments or examples.
[0200] Although the embodiments of the present invention have been shown and described above, it can be understood that the above embodiments are exemplary and should not be construed as limiting the present invention. Those of ordinary skill in the art can make changes, modifications, substitutions, and variations to the above embodiments within the scope of the present invention.
Claims
1. A parameter synchronization method for a distributed training system of a deep learning model, the system comprising a management node, at least one execution node and at least one parameter server, characterized in that, The method is applicable to a management node and includes: Responding to a training task completion instruction, sending a query instruction to a target parameter server, where the query instruction carries a training task completion flag; Receiving a query result sent by the target parameter server, where the query result carries data representing the number of management nodes and execution nodes that have completed the training task; Judging whether the number of nodes that have completed the training task is equal to a preset value; if so, obtaining the trained model parameters from the target parameter server and saving the model parameters as a model in a distributed file system; if not, entering a completion sleep state; The target parameter server is used to receive the training task completion flag and the node identifier of the management node itself sent by the management node, and receive the training task completion flag and the node identifier of the execution node itself sent by the execution node, and store the node identifiers according to the training task completion flag, and query the number of node identifiers of the completed training tasks stored, and generate the query result.
2. The method according to claim 1, wherein After entering the completion sleep state, it further includes: Sending the query instruction to the target parameter server every first preset time period; Receiving a query result sent by the target parameter server; When the query result shows that the number of nodes that have completed the training task is equal to the preset value, stopping sending the query instruction to the target parameter server.
3. The method according to claim 1, wherein Before sending the query instruction to the target parameter server, it further includes: Responding to a training task completion instruction, sending a training task completion flag and its own node number to the target parameter server, so that the target parameter server updates the stored node numbers of the completed training tasks according to the training task completion flag.
4. The method according to claim 1, wherein It further includes: Responding to a training task accurately started instruction, sending a query instruction to the target parameter server, where the query instruction carries a training task ready to start flag; Receiving a query result sent by the target parameter server, where the query result carries data representing the number of management nodes and execution nodes that are ready to start the training task; Judging whether the number of nodes that are ready to start the training task is equal to the preset value; if so, obtaining a training data set and model parameters from the parameter server for training; if not, entering an opening sleep state.
5. The method according to claim 4, characterized in that, After entering the opening sleep state, it further includes: Sending the query instruction to the target parameter server every second preset time period; Receiving a query result sent by the target parameter server; When the query result shows that the number of nodes that are ready to start the training task is equal to the preset value, stopping sending the query instruction to the target parameter server.
6. The method according to claim 4, characterized in that Before sending the query instruction to the target parameter server, it further includes: Responding to a training task accurately started instruction, sending a training task ready to start flag and its own node number to the target parameter server, so that the target parameter server updates the stored node numbers of the started training tasks according to the training task ready to start flag.
7. A management node, characterized in that, It is applicable to a deep learning model distributed training system and includes: A memory for storing computer programs that can be called; A processor, by invoking a computer program stored in the memory, executes the parameter synchronization method of the deep learning model distributed training system according to any one of claims 1 to 6.
8. A parameter synchronization method for a distributed training system of a deep learning model, the system comprising a management node, at least one execution node, and at least one parameter server, characterized in that, The method is applicable to a parameter server and includes: Receiving a query instruction sent by a management node, where the query instruction carries a training task completion flag; Receiving the training task completion flag sent by the management node and the node identifier of the management node itself, and receiving the training task completion flag sent by an execution node and the node identifier of the execution node itself, and storing the node identifiers according to the training task completion flag; Querying the stored data on the number of management nodes and execution nodes that have completed the training task according to the training task completion flag, and generating a query result; Sending the query result to the management node, so that the management node, according to the query result, determines that when the number of nodes that have completed the training task is equal to a preset value, sends a model parameter acquisition instruction; Receiving the model parameter acquisition instruction sent by the management node, and sending the trained model parameters to the management node, so that the management node saves the model parameters as a model in a distributed file system.
9. The method according to claim 8, wherein It further includes: Receiving the training task completion flag sent by the management node and the node serial number of the management node itself, and at the same time, receiving the training task completion flag sent by the execution node and the node serial number of the execution node itself; Storing the node serial numbers in a training task completion queue according to the training task completion flag; The training task completion queue is pre-deployed in the parameter server.
10. The method according to claim 8, wherein It further includes: Receiving a query instruction sent by a management node or an execution node, where the query instruction carries a training task ready-to-start flag; Querying the stored data on the number of management nodes and execution nodes that are ready to start the training task according to the training task ready-to-start flag, and generating a query result; Sending the query result to the management node or the execution node, so that the management node or the execution node, according to the query result, determines that when the number of nodes that are ready to start the training task is equal to a preset value, sends a training data set and model parameter acquisition instruction; Receiving the training data set and model parameter acquisition instruction sent by the management node or the execution node, and sending the training data set and model parameters to the corresponding management node or execution node, so that the management node or the execution node starts training.
11. The method according to claim 10, characterized in that, It further includes: Receiving the training task ready-to-start flag sent by the management node and the node serial number of the management node itself, and at the same time, receiving the training task ready-to-start flag sent by the execution node and the node serial number of the execution node itself; Storing the node serial numbers in a training task start queue according to the training task ready-to-start flag; The training task start queue is pre-deployed in the parameter server.
12. A parameter server, characterized in that, Applicable in a deep learning model distributed training system, it includes: A memory for storing a computer program that can be invoked; A processor, by invoking the computer program stored in the memory, executes the parameter synchronization method of the deep learning model distributed training system according to any one of claims 8 to 11.
13. A distributed training system for a deep learning model, characterized in that, It includes: The management node, at least one execution node, and at least one parameter server according to claim 12, wherein, The management node and the execution node are respectively communicatively connected to the parameter server.
14. A parameter synchronization method for a distributed training system of a deep learning model, the system comprising a management node, at least one execution node and at least one parameter server, characterized in that, The method includes: The management node, in response to a training task completion instruction, sends a query instruction to the target parameter server, and the query instruction carries a training task completion flag; The parameter server receives the training task completion flag sent by the management node and the node identifier of the management node itself, and receives the training task completion flag sent by the execution node and the node identifier of the execution node itself, and stores the node identifiers according to the training task completion flag; and receives the query instruction sent by the management node, and queries the data of the number of management nodes and execution nodes that have completed the training task stored according to the training task completion flag, and generates a query result; The management node receives the query result sent by the target parameter server, and determines whether the number of nodes that have completed the training task is equal to a preset value; if so, obtains the trained model parameters from the target parameter server, and saves the model parameters as a model in the distributed file system; if not, enters a completion sleep state.
Citation Information
Patent Citations
Method, device and system for training deep neural networks
CN107688493A
Methods, apparatus and systems for training deep neural networks
CN107688493B
A method for updating global parameters in a parameter server
CN109032630A
A method for updating global parameters in a parameter server
CN109032630B
Distributed job coordination control method, device, computer device and storage medium
CN109101342A