State data storage method, state data recovery method and state data recovery equipment for model training tasks

By dividing the model training task into subtasks in a distributed training system and executing it in parallel between computing nodes, and backing up state data in node memory, the problem of low recovery efficiency of model training tasks is solved, and more efficient task recovery and training efficiency is achieved.

CN120066861AActive Publication Date: 2025-05-30INSPUR SUZHOU INTELLIGENT TECH CO LTD
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
CN202510538834.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-27
Publication Date
2025-05-30
Estimated Expiration
2045-04-27

AI Technical Summary

Technical Problem

In distributed training systems, the recovery efficiency of model training tasks is low, resulting in frequent training interruption problems, affecting model training efficiency.

Method used

By dividing the model training task into multiple subtraining tasks and executing it in parallel between multiple computing nodes, the method of state data being backup in local memory and the memory of associated computing nodes is used to ensure that it can be quickly restored when task abnormalities are performed.

Benefits of technology

It improves the recovery efficiency of model training tasks, reduces the time for training interruptions, and enhances the timeliness and reliability of task recovery.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120066861A_ABST
    Figure CN120066861A_ABST
Patent Text Reader

Abstract

The invention discloses a state data storage method, a state data recovery method and state data recovery equipment of a model training task, and relates to the technical field of artificial intelligence computing, comprising the following steps: on one hand, after state data of a first sub-training task is stored in a memory of a computing node, when the first sub-training task or a second sub-training task is abnormal, the first sub-training task or the second sub-training task is abnormal; state data for task recovery can be obtained from a memory. As the data transmission efficiency of the memory is high, the recovery efficiency of the model training task can be greatly improved. And on the other hand, after the state data of the first sub-training task are mutually backed up in the memory (namely the local memory) of the first computing node and the memory of the second computing node, if one computing node is in an abnormal state, the state data can still be acquired from the memory of the other computing node, so that the timeliness and the reliability of task recovery are ensured. In conclusion, the technical scheme of the invention can solve the problem of low recovery efficiency of the model training task in related technologies.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of artificial intelligence computing technologies, and particularly to a method and device for saving and restoring status data of a model training task. Background Art

[0002] With the development of artificial intelligence technologies, both the scale of models and training data is gradually increasing, and a single device or Graphics Processing Unit (GPU) may not be able to meet the training requirements. In view of this, distributed training systems have emerged. In a distributed training system, a model or training data can be divided into multiple hardware nodes for parallel training, greatly improving the model training efficiency. However, with the expansion of the hardware scale, the problem of training interruption caused by hardware failures (such as node downtime, silent data errors, etc.) is also gradually increasing, and it is necessary to frequently restore the model training task during the model training process.

[0003] Currently, in some technologies, the restoration efficiency of model training tasks is low, which in turn affects the model training efficiency. Summary of the Invention

[0004] This application provides a method and device for saving and restoring status data of a model training task, a status data saving device, a restoration device, an electronic device, a computer-readable storage medium, and a computer program product, so as to at least solve the problem of low restoration efficiency of model training tasks in related technologies.

[0005] This application provides a method for saving status data of a model training task. The model training task is divided into multiple sub-training tasks, and multiple computing nodes execute the sub-training tasks in parallel. The method is applied to a first computing node among the multiple computing nodes. The method includes: Receiving task configuration information allocated by a management node, where the task configuration information includes a first sub-training task to be executed and a second computing node associated with the first computing node; Executing the first sub-training task, and saving the status data of the first sub-training task to the local memory and the memory of the second computing node. The status data in the local memory and the status data in the memory of the second computing node are backed up to each other. The status data is used to restore the first sub-training task when the first sub-training task is abnormal, and / or to restore the second sub-training task when the second sub-training task running on the second computing node is abnormal.

[0006] The present application provides a method for resuming a model training task, where the model training task is divided into multiple sub-training tasks, and the sub-training tasks are executed in parallel by multiple computing nodes. The method is applied to a management node of the multiple computing nodes; the method includes: Monitoring an abnormal target sub-training task and a first target computing node that runs the target sub-training task. Among them, in the memory of the first target computing node and in the memory of a second target computing node associated with the first target computing node, status data of the target sub-training task is saved; If the first target computing node is in a normal state, the target sub-training task is resumed based on the status data in the memory of the first target computing node; If the first target computing node is in a faulty state, the target sub-training task is resumed based on the status data in the memory of the second target computing node.

[0007] The present application further provides a status data saving device for a model training task. The model training task is divided into multiple sub-training tasks, and multiple computing nodes execute the sub-training tasks in parallel. The method is applied to a first computing node among the multiple computing nodes; the device includes: An information receiving module, configured to receive task configuration information allocated by a management node. The task configuration information includes a first sub-training task to be executed and a second computing node associated with the first computing node; A task execution module, configured to execute the first sub-training task and save the status data of the first sub-training task to local memory and the memory of the second computing node. The status data in the local memory and the status data in the memory of the second computing node are backed up to each other. The status data is used to resume the first sub-training task when the first sub-training task is abnormal, and / or to resume a second sub-training task running on the second computing node when the second sub-training task is abnormal.

[0008] The present application provides a model training task resuming device. The model training task is divided into multiple sub-training tasks, and the sub-training tasks are executed in parallel by multiple computing nodes. The method is applied to a management node of the multiple computing nodes; the device includes: A monitoring module, configured to monitor an abnormal target sub-training task and a first target computing node that runs the target sub-training task. Among them, in the memory of the first target computing node and in the memory of a second target computing node associated with the first target computing node, status data of the target sub-training task is saved; The first recovery module is configured to, if the first target computing node is in a normal state, recover the target sub-training task based on the state data in the memory of the first target computing node; The second recovery module is configured to, if the first target computing node is in a faulty state, recover the target sub-training task based on the state data in the memory of the second target computing node.

[0009] This application also provides an electronic device, including: a memory for storing a computer program; a processor for implementing the steps of any of the above state data saving methods when executing the computer program.

[0010] This application also provides a computer-readable storage medium storing a computer program, wherein the computer program, when executed by a processor, implements the steps of any of the above annotation methods.

[0011] This application also provides a computer program product including a computer program, which implements the steps of any of the above annotation methods when executed by a processor.

[0012] In the technical solutions of some embodiments of this application, on the one hand, after saving the state data of the first sub-training task to the memory of the computing node, when the first sub-training task or the second sub-training task is abnormal, the state data for task recovery can be obtained from the memory. Since the data transmission efficiency of the memory is high, the recovery efficiency of the model training task can be greatly improved. On the other hand, after mutually backing up the state data of the first sub-training task in the memory of the first computing node (i.e., local memory) and the memory of the second computing node, if one of the computing nodes is in an abnormal state, the state data can still be obtained from the memory of the other computing node, ensuring the timeliness and reliability of task recovery. Based on the above two aspects, the technical solutions of this application can solve the problem of low recovery efficiency of model training tasks in related technologies. BRIEF DESCRIPTION OF THE DRAWINGS

[0013] To more clearly illustrate the embodiments of this application, the following will briefly introduce the drawings required in the embodiments. Obviously, the drawings in the following description are only some embodiments of this application. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained based on these drawings.

[0014] Figure 1 It is a schematic diagram of modules of some distributed training systems; Figure 2 It is a flowchart of the state data saving method provided by some embodiments of this application; Figure 3Schematic diagram of the execution process of the model training task provided by some embodiments of the present application; Figure 4 Schematic diagram of the recovery method provided by some embodiments of the present application; Figure 5 Schematic diagram of the modules of the status data storage device provided by some embodiments of the present application; Figure 6 Schematic diagram of the modules of the recovery device provided by some embodiments of the present application; Figure 7 Schematic diagram of the modules of the electronic device provided by some embodiments of the present application. Detailed implementation manners

[0015] Next, the technical solutions in the embodiments of the present application will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. Based on the embodiments of the present application, all other embodiments obtained by those of ordinary skill in the art without creative efforts belong to the protection scope of the present application.

[0016] It should be noted that in the description of the present application, the terms "include", "comprise" or any other variation thereof are intended to cover a non-exclusive inclusion, so that a process, method, article or device including a series of elements not only includes those elements but also includes other elements not expressly listed, or further includes elements inherent to such process, method, article or device. The terms "first", "second", etc. in the present application are used to distinguish similar objects and not to describe a specific order or sequence.

[0017] In order to enable those skilled in the art of the present technology to better understand the solution of the present application, the present application will be further described in detail below in conjunction with the accompanying drawings and specific implementation manners.

[0018] Refer to Figure 1 , which is a schematic diagram of the modules of some distributed training systems. Figure 1 In, the distributed training system includes a management node 11 and multiple computing nodes 12. Each computing node 12 respectively includes a central processing unit 122, a video memory 123, a memory 124, and multiple graphics processing units 121. Among them, the memory 124 is the data storage space for the central processing unit 122, and the video memory 123 is the data storage space for the graphics processing unit 121.

[0019] In the management node 11, the model training task can be divided into multiple sub-training tasks, and these sub-training tasks are assigned to each computing node 12. Multiple computing nodes 12 can execute the sub-training tasks in parallel, thereby improving the model training efficiency. For example, the model training task is divided into sub-training tasks A, B, and C, and sub-training task A is assigned to computing node a, sub-training task B is assigned to computing node b, and sub-training task C is assigned to computing node c. In this way, computing nodes a, b, and c can run sub-training tasks A, B, and C in parallel, thereby improving the model training efficiency.

[0020] In each computing node 12, the graphics processor 121 can be used to execute the model training task, and the central processing unit 122 can be used to cooperate in executing the model training task. For example, the central processing unit 122 can read the model training data from the disk or memory, convert the model training data into the format required for model training, and then input the data into the graphics processor 121. The graphics processor 121 can perform specific model training calculations, such as the forward propagation calculation of the neural network model, the gradient calculation in the backpropagation process, etc.

[0021] In each computing node 12, according to the number of graphics processors 121 it includes, the sub-training task can be further divided, that is, the graphics processors 121 in the same computing node 12 can execute the sub-training task assigned to the computing node 12 in parallel. For example, assuming that the above-mentioned computing node a includes 8 graphics processors G1, G2,..., G8, then the sub-training task A assigned to computing node a can be further divided into training tasks A1, A2,..., A8. Graphics processor G1 can execute training task A1, graphics processor G2 can execute training task A2, and so on.

[0022] The following explains the division of the model training task.

[0023] The management node 11 can divide the model training task or the sub-training tasks of each computing node 12 according to task division strategies such as data parallelism, tensor parallelism, and pipeline parallelism.

[0024] Data parallelism means that the training data set of the model to be trained is divided into multiple sub-training sets, and each sub-training set can be assigned to a graphics processor 121 for model training. When dividing the model training task according to the data parallelism strategy, the training results of each graphics processor 121 can be aggregated, and the model parameters can be updated based on the aggregated results.

[0025] Tensor parallelism means dividing the parameters or tensors of the model to be trained into multiple parameter sets or tensor sets, and each parameter set or tensor set can be assigned to a graphics processing unit 121 for model training. When dividing the model training tasks according to the tensor parallelism strategy, the graphics processing units 121 can communicate with each other for data synchronization and parameter update.

[0026] Pipeline parallelism means dividing the model training tasks by stages, and each computing node can execute the training tasks of one of the stages. For example, in a neural network model, one or more neural network layers can be regarded as a stage, and each computing node can execute the training tasks of one of the stages, that is, execute the training tasks of one or more neural network layers. Multiple computing nodes can execute the assigned sub-training tasks in sequence according to the data propagation direction, and there are dependencies between the sub-training tasks of the computing nodes. For example, assume that computing node a executes the training tasks of neural network layers 0 to 3, and computing node b executes the training tasks of neural network layers 4 to 7. After computing node a executes the training tasks of neural network layers 0 to 3, the obtained results can be input into computing node b, and computing node b executes the training tasks of neural network layers 4 to 7 based on the output data of computing node a.

[0027] The task division strategy between computing nodes 12 can be different from the task division strategy between graphics processing units 121. For example, according to the pipeline parallelism strategy, after assigning the training tasks of neural network layers 0 to 3 to computing node a, in computing node a, according to the tensor parallelism task division strategy, the training tasks of neural network layers 0 to 3 can be further divided into multiple sub-training tasks, so that the graphics processing units 121 in computing node a can execute the training tasks of neural network layers 0 to 3 in parallel.

[0028] Although the distributed training system can improve the model training efficiency, with the expansion of the hardware scale, the problem of training interruption caused by hardware failures (such as node downtime, silent data errors, etc.) is gradually increasing, and it is necessary to frequently resume the model training tasks during the model training process.

[0029] Currently, in some technologies, model snapshots of the model to be trained are recorded at one or more specified time points during the model training process. The model snapshots can be saved in a remote non-volatile storage. When the model training is interrupted, the model snapshot at the latest time point can be obtained from the remote non-volatile storage, and based on the information in the model snapshot, the model training tasks can be resumed starting from the time when the model snapshot was recorded. In these technologies, the data volume of the model snapshots is relatively large. If data compression is performed, it may cause damage to the model training progress. If data compression is not performed, limited by the network transmission bandwidth, it will be time-consuming to obtain the model snapshots from the remote storage, and the recovery efficiency of the model training tasks is low, which affects the model training efficiency.

[0030] In view of this, the present application first provides a method for saving state data of a model training task, which can solve the problem of low recovery efficiency of model training tasks in some technologies. The state data saving method can be applied to Figure 1 the first computing node among the multiple computing nodes 12 shown in the figure. Referring to Figure 2 together, it is a schematic flowchart of the state data saving method provided by some embodiments of the present application. Figure 2 In [the figure], the state data saving method includes the following steps: Step S201: Receive the task configuration information allocated by the management node 11. The task configuration information includes the first sub-training task to be executed and the second computing node associated with the first computing node.

[0031] Regarding the sub-training task, reference can be made to Figure 1 the relevant description, which will not be elaborated here.

[0032] In the task configuration information, it may include the node information of the second computing node associated with the first computing node, such as node identification, IP address, etc.

[0033] In this embodiment, the management node 11 may use the computing nodes that meet the following conditions 11) to 13) as the second computing node.

[0034] 11) There is data interaction with the first computing node during the model training process. For example, assume that the training process of the model to be trained includes a forward propagation process and a backward propagation process. Computing node a executes the training tasks of neural network layers 0 to 3, and computing node b executes the training tasks of neural network layers 4 to 7. During the forward propagation process, computing node b needs to receive the output data of computing node a. During the backward propagation process, computing node a needs to receive the gradient data output by computing node b. Therefore, when computing node a is the first computing node, computing node b can be regarded as the computing node with data interaction with the first computing node.

[0035] 12) The remaining memory space is greater than the space threshold.

[0036] 13) Communicatively connected to the first computing node.

[0037] In practical applications, the computing nodes that only meet the above conditions 12) to 13) can also be used as the second computing node. Or, on the basis of the above conditions 12) to 13), other conditions can also be defined according to actual needs, except for the above condition 11). The present application does not limit this.

[0038] Step S202: Execute the first sub-training task and save the status data of the first sub-training task to the local memory and the memory of the second computing node. The status data in the local memory and the status data in the memory of the second computing node are backups of each other. The status data is used to recover the first sub-training task when the first sub-training task is abnormal and / or to recover the second sub-training task when the second sub-training task running on the second computing node is abnormal.

[0039] Specifically, the status data of the first sub-training task is used to track the execution process of the first sub-training task.

[0040] In this embodiment, the status data of the first sub-training task can be one of the following data: 21) The model snapshot of the model being trained during the execution of the first sub-training task.

[0041] 22) The execution log of the first sub-training task.

[0042] Among them, the model snapshot can represent the complete training status of the model being trained at a specified time point, including but not limited to the model parameters, optimizer status, training progress information (such as the current number of epochs, batch number) of the model being trained at a specific time point, etc. For example, during the execution of the first sub-training task, at time point T1 when training reaches the 50th epoch, the model snapshot P1 can be saved. The model snapshot P1 can include the weights and biases of the model being trained at time point T1, the optimizer learning rate, momentum settings, training progress information (such as the number of completed epochs and total batch number). At time point T2 when training reaches the 100th epoch, the model snapshot P2 can be saved. The model snapshot P2 can include the weights and biases of the model being trained at time point T2, the optimizer learning rate, momentum settings, training progress information.

[0043] The execution log can represent the key events, information flow, and timestamps during the execution of the first sub-training task. Among them, the key events and information flow include but not limited to the intermediate parameters of the model being trained during the execution of the first sub-training task (such as the output data of a specific neural network layer, gradient information, etc.), system information (such as memory occupancy rate, etc.), the identification of the graphics processor, etc. For example, when starting to train the first neural network layer based on the 1st epoch, the start training timestamp and the initial memory occupancy rate can be recorded; after the training of the first neural network layer based on the 1st epoch is completed, the end training timestamp and the output data of the first neural network layer can be recorded.

[0044] When the first sub-training task or the second sub-training task is abnormal, the status data of the first sub-training task can be used to recover the first sub-training task or the second sub-training task. For example, when the output data of the first sub-training task is used as the input data of the second sub-training task, the second sub-training task can be recovered based on the output data of the first sub-training task.

[0045] Specifically, when the first sub-training task is abnormal, if the first computing node is in a normal state, the status data saved in the memory of the first computing node (i.e., local memory) can be used to recover the first sub-training task. If the first computing node is in an abnormal state, the status data saved in the memory of the second computing node can be used to recover the first sub-training task. Similarly, when the second sub-training task is abnormal, if the second computing node is in a normal state, the status data saved in the memory of the second computing node can be used to recover the second sub-training task. If the second computing node is in an abnormal state, the status data saved in the memory of the first computing node (i.e., local memory) can be used to recover the second sub-training task.

[0046] Since the status data in the memory of the first computing node (i.e., local memory) and the status data in the memory of the second computing node are backed up with each other, when one of the computing nodes fails, the status data for recovering the sub-training task point can also be obtained from the memory of the neighboring computing node. Compared with some technologies that save the model snapshot in the remote non-volatile storage, the present application saves the status data in the memories of multiple computing nodes, which can greatly improve the transmission efficiency of the status data, and then can improve the recovery efficiency of the first sub-training task or the second sub-training task, and further improve the model training efficiency.

[0047] In summary, in the technical solutions of some embodiments of the present application, on the one hand, after the status data of the first sub-training task is saved to the memory of the computing node, when the first sub-training task or the second sub-training task is abnormal, the status data for task recovery can be obtained from the memory. Since the data transmission efficiency of the memory is high, the recovery efficiency of the model training task can be greatly improved. On the other hand, after the status data of the first sub-training task is backed up with each other in the memory of the first computing node (i.e., local memory) and the memory of the second computing node, if one of the computing nodes is in an abnormal state, the status data can still be obtained from the memory of the other computing node, ensuring the timeliness and reliability of task recovery. Based on the above two aspects, the technical solution of the present application can solve the problem of low recovery efficiency of the model training task in the related art.

[0048] In some embodiments, considering that there is a risk of data loss in the memory during the abnormal or restart process of the computing node, the number of second computing nodes can be multiple, that is, the status data of the first sub-training task can be saved in the local memory and the memories of multiple second computing nodes. Performing multiple backups of the status data can reduce the degree of data loss risk, and thus improve the reliability of task recovery.

[0049] Furthermore, in some embodiments, the status data can also be saved to non-volatile storage. Compared with the memory, the status data in non-volatile storage will not be lost due to reasons such as device restart, but its data transfer efficiency is relatively slow. In view of this, when at least one of the first computing node and the second computing node is in a normal state, task recovery can be based on the status data in the memory of the first computing node (i.e., local memory) or the memory of the second computing node. In this way, the recovery efficiency of the model training task can be improved. When both the first computing node and the second computing node are abnormal, task recovery can be based on the status data in non-volatile storage. In this way, the reliability of model training task recovery can be ensured.

[0050] In some embodiments, in step S202, since the status data of the first sub-training task can be one of the model snapshot and the execution log, the task configuration information can also include status data parameters. The status data parameters are used to specify the target data as the status data. Based on the status data parameters, the execution log of the first sub-training task can be used as the status data, or the model snapshot of the model being trained during the execution of the first sub-training task can be used as the status data. For example, when the status data parameters specify that the execution log is used as the status data, the execution log of the first sub-training task can be saved to the local memory and the memories of the second computing nodes; when the status data parameters specify that the model snapshot is used as the status data, the execution log of the first sub-training task can be saved to the local memory and the memories of the second computing nodes.

[0051] On the management node 11 side, before the model training task starts to execute, the total data volume of the execution log of the model training task and the data volume of the model snapshot can be evaluated, and the status data parameters can be set according to the total data volume of the execution log and the data volume of the model snapshot to specify the target data as the status data. Specifically, when the total data volume of the execution log of the model training task is greater than the data volume of the model snapshot, the management node 11 can specify the model snapshot as the status data through the status data parameters. When the total data volume of the execution log is less than the data volume of the model snapshot, the management node 11 can specify the execution log of the first sub-training task as the status data through the status data parameters. In this way, the amount of status data saved to the memory can be relatively small, avoiding the occupancy rate of the memory space by the status data being too high and affecting other services.

[0052] Continue to refer to Figure 1 . In some embodiments, the first computing node includes a first central processing unit, a first graphics processing unit, and local video memory. Based on some specific task partitioning strategies (such as pipeline parallelism), the first graphics processing unit in the first computing node can execute the model training task only in some stages of model training. That is, during the entire process of model training, the first graphics processing unit can have idle time, which is also called a pipeline bubble. For example, assume that the first computing node is used to execute the training task of neural network layers 0 to 3 of the model to be trained (i.e., the first sub-training task). Then, when training neural network layers 4 to 7 of the model to be trained, the first graphics processing unit is in an idle state.

[0053] Based on a specific task partitioning strategy, the first computing node can save the state data of the first sub-training task to the local memory and the memory of the second computing node in the following manner: when the first graphics processing unit is in a task execution state, save the state data to the local video memory; when the first graphics processing unit is in an idle state, save the state data in the local video memory to the local memory through the first graphics processing unit, and save the state data in the local memory to the memory of the second computing node through the first central processing unit.

[0054] Specifically, during the execution of the first sub-training task, the first graphics processing unit can obtain the state data of the first sub-training task. These state data can be first saved in the local video memory. When the first graphics processing unit finishes executing the first sub-training task and is in an idle state, then save the state data in the local video memory to the local memory through the first graphics processing unit. The first central processing unit can save the state data in the local memory to the memory of the second computing node through a background task.

[0055] This asynchronous state data saving scheme can prevent the operation of saving state data from preempting resources with the first sub-training task, and thus can ensure the training efficiency of the first sub-training task. For example, assume that during the execution of the first sub-training task, the first graphics processing unit needs to communicate with the first central processing unit to obtain training data required for training or send the training result to the first central processing unit. If the state data is saved to the first memory during the task execution, then the operation of saving state data will preempt communication resources with the first sub-training task, thereby affecting the training efficiency of the first sub-training task.

[0056] The following further elaborates on the solution of this application in combination with some specific application scenarios. Combine and refer to Figure 3 , which is a schematic diagram of the execution process of the model training task provided by some embodiments of this application. Figure 3Among them, the trained model includes multiple neural network layers. According to the task allocation strategy of pipeline parallelism, different training tasks of different neural network layers are assigned to different computing nodes. For example, computing node A executes the training tasks of neural network layers 0 to 2, computing node B executes the training task of neural network layer 3, and so on. The training process of the trained model may include the forward propagation process and the backward propagation process of data. In the forward propagation process, data is transmitted in the direction indicated by the solid arrow, and the direction indicated by the solid arrow is also called the forward propagation direction. In the backward propagation process, data is transmitted in the direction indicated by the dashed arrow, and the direction indicated by the dashed arrow is also called the backward propagation direction. The data transmitted in the forward propagation process may include training data, output data of each computing node, etc.; the data transmitted in the backward propagation process may include gradient data of each computing node.

[0057] Considering that when a computing node is abnormal, the probability of a processor or a graphics processor in the computing node malfunctioning is relatively small. Therefore, when taking the execution log of the first sub-training task as status data, only the communication information between computing nodes can be recorded. In this way, the data volume of the execution log can be reduced, and thus the occupation of memory space can be reduced. Specifically, when taking the execution log of the first sub-training task as status data, the method of the present application may further include: In the forward propagation process, calculate the output data output to the third computing node in the forward propagation direction of the data, and take the output data as the execution log; In the backward propagation process, calculate the gradient data output to the fourth computing node in the backward propagation direction of the data, and take the gradient data as the execution log.

[0058] Specifically, the third computing node is the next computing node of the first computing node in the forward propagation direction, and the fourth computing node is the next computing node of the first computing node in the backward propagation direction. For example, with reference to Figure 3 . Assume that computing node C is taken as the first computing node, then computing node D is the third computing node, and computing node B is the fourth computing node.

[0059] Since there is data interaction between the third computing node, the fourth computing node and the first computing node, the third computing node and the fourth computing node can be associated with the first computing node. That is, the second computing node associated with the first computing node can include the third computing node and the fourth computing node. During the forward propagation process, the third computing node needs to execute the sub-training task assigned to it based on the output data of the first computing node. After the sub-training task executed by the third computing node is abnormal, the sub-training task on the third computing node can be restored based on the output data of the first computing node. During the backward propagation process, the fourth computing node needs to execute the sub-training task assigned to it based on the output data of the first computing node. After the sub-training task executed by the fourth computing node is abnormal, the sub-training task on the fourth computing node can be restored based on the output data of the first computing node.

[0060] Based on the above description, in some embodiments, the third computing node includes a second graphics processor, and the fourth computing node includes a third graphics processor.

[0061] During the forward propagation process, the first computing node can save the output data to the local memory and the second graphics processor through the first graphics processor, and save the output data in the local memory to the memory of the third computing node through the first central processing unit. Specifically, during the execution of the first sub-training task by the first graphics processor, the output data can be sent to the second graphics processor. Based on the received data, the second graphics processor can execute the sub-training task assigned to the third computing node. During the idle time after the first graphics processor executes the first sub-training task, the output data can be saved to the local memory of the first computing node. The first central processing unit saves the output data to the memory of the third computing node through a background task. In this way, after the sub-training task on the third computing node is abnormal, if the third computing node is in a normal state, the sub-training task can be restored based on the output data in the memory of the third computing node; if the third computing node is in an abnormal state, the sub-training task can be restored based on the output data in the memory of the first computing node.

[0062] Similarly, during the backward propagation process, the first computing node can save the gradient data to the local memory and the third graphics processor through the first graphics processor, and save the gradient data in the local memory to the memory of the fourth computing node through the first central processing unit. The related principle is similar to the above forward propagation process and will not be elaborated here.

[0063] Based on the above description, with reference to Figure 3It can be understood that during the forward propagation process, the first computing node can receive and save the output data of the fourth computing node (i.e., computing node B), and during the backward propagation process, the first computing node can receive and save the gradient data of the third computing node (i.e., computing node D). After the first subtraining task in the first computing node is abnormal, if the first computing node is in a normal state, the first subtraining task can be restored based on the output data or gradient data in the memory of the first computing node. If the first computing node is in an abnormal state, the first subtraining task can be restored based on the output data in the memory of the fourth computing node or the gradient data in the memory of the third computing node.

[0064] In some embodiments, when saving the output data and gradient data, auxiliary information can also be saved simultaneously. The auxiliary information can include, but is not limited to, the timestamp of the output data, the graphics processor identifier of the first graphics processor, etc. In this way, based on the auxiliary information, the generation time of the output data and the graphics processor that generated the output data can be determined.

[0065] Furthermore, with reference to Figure 3 In some embodiments, after all of the first computing node (i.e., computing node C), the third computing node (i.e., computing node D), and the fourth computing node (i.e., computing node B) are abnormal, the next computing node (i.e., computing node E) of the third computing node in the forward propagation direction can be found as the first auxiliary computing node, and the next computing node (i.e., computing node A) of the fourth computing node in the backward propagation direction can be found as the second auxiliary computing node. Based on the gradient data in the memory of the first auxiliary computing node and the output data in the memory of the second auxiliary computing node, the subtraining tasks on the third computing node and the fourth computing node can be restored first, and then the subtraining task on the first computing node can be restored. In this way, when multiple computing nodes are in an abnormal state, data acquisition from a remote non-volatile storage can be avoided, thereby improving the recovery efficiency of the model training task.

[0066] When performing task recovery based on the execution log, the model training task can be partially recovered, that is, only the abnormal subtraining task can be restored. In this way, the task recovery time can be reduced and the task recovery efficiency can be improved.

[0067] Continue to refer to Figure 3. In some embodiments, when the model snapshot of the trained model is used as the state data, if the first computing node trains the first neural network layer and the second neural network layer, the first computing node can also perform the following operations: for any target neural network layer in the first neural network layer and the second neural network layer, when the forward propagation process and the backward propagation process of the target neural network layer are both completed, the parameters of the target neural network layer are saved as the model snapshot. Simply put, the first computing node can save the state data layer by layer, that is, after each neural network layer is trained, the saving of the state data is triggered. In this way, the amount of state data saved at one time is avoided from being too large.

[0068] Furthermore, considering that when restoring the task, the parameters of the neural network layer that was last trained are usually referred to, so only the parameters of the neural network layer that was last trained can be saved in the local memory and the memory of the second computing node. Here, the second computing node includes at least one of the above-mentioned third computing node and fourth computing node. In view of this, in some embodiments, when saving the parameters of the first neural network layer to the local memory and the memory of the second computing node, if the parameters of the third neural network layer have been saved in the local memory and the memory of the second computing node, the parameters of the third neural network layer are deleted. In this way, the consumption of memory space is reduced.

[0069] Corresponding to the state data saving method, the present application also provides a method for restoring a model training task, which can solve the problem of low restoration efficiency of the model training task in some technologies. The restoration method can be applied to Figure 1 the management node 11 shown. Referring to Figure 4 , it is a schematic flowchart of the restoration method provided by some embodiments of the present application. Figure 4 In Step S401, monitor the abnormal target sub-training task and the first target computing node running the target sub-training task, where the state data of the target sub-training task is saved in the memory of the first target computing node and the memory of the second target computing node associated with the first target computing node.

[0070] Step S402, if the first target computing node is in a normal state, restore the target sub-training task based on the state data in the memory of the first target computing node.

[0071] Step S403, if the first target computing node is in a faulty state, restore the target sub-training task based on the state data in the memory of the second target computing node.

[0072] Regarding the above steps S401 to S403, reference can be made to the relevant descriptions of the above state data saving method, which will not be elaborated here.

[0073] In some embodiments, before monitoring the target sub-training task, the management node 11 may further perform the following operations: Generate task configuration information, which includes status data parameters, sub-training tasks to be executed by each computing node, and the association relationship between computing nodes. Among them, the status data parameters are used to specify the target data as status data; Send the task configuration information to each computing node.

[0074] In some embodiments, the status data parameters are determined based on the following method: Calculate the total data volume of the execution logs of the model training task and the data volume of the model snapshot of the model to be trained; If the total data volume of the execution logs is greater than the data volume of the model snapshot, set the status data parameter to the first value to specify that the model snapshot is used as the status data; If the total data volume of the execution logs is less than the data volume of the model snapshot, set the status data parameter to the second value to specify that the execution logs of each sub-training task are used as the status data.

[0075] In some embodiments, when the status data parameter specifies that the model snapshot is used as the status data, the sub-training tasks can be divided according to the division logic that the execution duration difference of each sub-training task does not exceed the duration threshold. Specifically, assume that the training duration of the i-th neural network layer is , and it is necessary to divide the model training task into M sub-training tasks, then the execution duration of a single sub-training task can be shown as in expression (1): (1) where T is the execution duration of a single sub-training task, N is the maximum number of layers of the neural network layer, the values of i and N are positive integers, and M is the number of computing nodes.

[0076] In some embodiments, when dividing the sub-training tasks according to the execution duration T, if the task data volume corresponding to the sub-training task (such as the sum of the model parameter quantity and the training data volume) is greater than x times the video memory capacity of a single computing node, a prompt of task allocation failure can be returned. Among them, the value of x is a numerical value between 0 and 1.

[0077] In some embodiments, when the status data parameter specifies that the execution logs of each sub-training task are used as the status data, the sub-training tasks can be divided according to the division logic that the data volume difference of the status data generated by each computing node does not exceed the data volume threshold. Specifically, assume that the total data volume of the model to be trained is shown as in expression (2): (2) where, Represents the total number of parameters of the trained model, Represents the data volume of the i-th neural network layer, where N is the maximum number of neural network layers.

[0078] The data volume of a single sub-training task can be as shown in expression (3): (3) Where, Represents the data volume of a single sub-training task, and M is the number of computing nodes.

[0079] In some embodiments, when dividing sub-training tasks according to the data volume If the data volume is greater than x1 times the video memory capacity of a single computing node or greater than x2 times the memory capacity of a single computing node, a task allocation failure prompt can be returned. Where the values of x1 and x2 are numerical values between 0 and 1.

[0080] In some embodiments, according to the principles of the above expressions (1), (2), and (3), the sub-training tasks executed by a single computing node can be further divided, which will not be elaborated here.

[0081] The recovery method has the same technical features as the above state data saving method, that is, the state data of the target sub-training task is saved in the memory of the first target computing node and the memory of the second target computing node associated with the first target computing node. Therefore, it has the same beneficial effects as the above state data saving method, which will not be elaborated here.

[0082] Through the description of the above embodiments, those skilled in the art can clearly understand that the method according to the above embodiments can be implemented by means of software plus a necessary general hardware platform. Of course, it can also be implemented by hardware, but in many cases, the former is a better implementation method.

[0083] The embodiments of the present application also provide a state data saving device for model training tasks. Referring to Figure 5 , it is a schematic diagram of the modules of the state data saving device provided in some embodiments of the present application. Figure 5 In The information receiving module 501 is used to receive the task configuration information allocated by the management node. The task configuration information includes the first sub-training task to be executed and the second computing node associated with the first computing node; The task execution module 502 is configured to execute the first sub-training task and save the status data of the first sub-training task to the local memory and the memory of the second computing node. The status data in the local memory and the status data in the memory of the second computing node are backed up to each other. The status data is used to recover the first sub-training task when the first sub-training task is abnormal, and / or to recover the second sub-training task when the second sub-training task running on the second computing node is abnormal.

[0084] In some embodiments, the task configuration information further includes status data parameters, and the status data parameters are used to specify the target data as the status data; the task execution module 502 is further configured to: Based on the status data parameters, use the execution log of the first sub-training task as the status data, or use the model snapshot of the model being trained during the execution of the first sub-training task as the status data; Wherein, when the total data volume of the execution log of the model training task is greater than the data volume of the model snapshot, the management node specifies the model snapshot as the status data through the status data parameters, and when the total data volume of the execution log is less than the data volume of the model snapshot, the management node specifies the execution log of the first sub-training task as the status data through the status data parameters.

[0085] In some embodiments, the model being trained includes multiple neural network layers, and the training process of the model being trained includes a forward propagation process and a backward propagation process of data. The first computing node executes the first sub-training task to train at least one neural network layer; When using the execution log of the first sub-training task as the status data, the task execution module 502 is further configured to: During the forward propagation process, calculate the output data output to the third computing node in the forward propagation direction of the data, and use the output data as the execution log; During the backward propagation process, calculate the gradient data output to the fourth computing node in the backward propagation direction of the data, and use the gradient data as the execution log.

[0086] In some embodiments, the first computing node is associated with the third computing node and the fourth computing node. The first computing node includes a first central processing unit and a first graphics processing unit. The third computing node includes a second graphics processing unit, and the fourth computing node includes a third graphics processing unit; the task execution module 502 is further configured to: During the forward propagation process, through the first graphics processing unit, save the output data to the local memory and the second graphics processing unit, and through the first central processing unit, save the output data in the local memory to the memory of the third computing node; During the backpropagation process, the gradient data is saved to the local memory and the third graphics processing unit through the first graphics processing unit, and the gradient data in the local memory is saved to the memory of the fourth computing node through the first central processing unit.

[0087] In some embodiments, the trained model includes multiple neural network layers. The training process of the trained model includes a forward propagation process and a backpropagation process. The first computing node executes a first subtraining task to train at least one neural network layer. When using the model snapshot of the trained model as the state data, if the first computing node trains the first neural network layer and the second neural network layer, the task execution module 502 is further configured to: For any target neural network layer among the first neural network layer and the second neural network layer, when the training of both the forward propagation process and the backpropagation process of the target neural network layer is completed, the parameters of the target neural network layer are saved as the model snapshot.

[0088] In some embodiments, the task execution module 502 is further configured to: When saving the parameters of the first neural network layer to the local memory and the memory of the second computing node, if the parameters of the third neural network layer have been saved in the local memory and the memory of the second computing node, the parameters of the third neural network layer are deleted.

[0089] In some embodiments, the first computing node includes a first central processing unit, a first graphics processing unit, and a local video memory. The task execution module 502 is specifically configured to: When the first graphics processing unit is in the task execution state, save the state data to the local video memory. When the first graphics processing unit is in the idle state, save the state data in the local video memory to the local memory through the first graphics processing unit, and save the state data in the local memory to the memory of the second computing node through the first central processing unit.

[0090] In some embodiments, the task execution module 502 is further configured to: Save the state data to the non-volatile storage. The state data in the non-volatile storage is used to restore the first subtraining task or the second subtraining task when both the first computing node and the second computing node are abnormal.

[0091] The embodiments of the present application further provide a recovery device for a model training task. Referring to Figure 6 which is the schematic diagram of the modules of the recovery device provided in some embodiments of the present application. Figure 6 In A monitoring module 601, configured to monitor abnormal target sub-training tasks and a first target computing node that runs the target sub-training tasks, wherein status data of the target sub-training tasks is stored in the memory of the first target computing node and in the memory of a second target computing node associated with the first target computing node; A first recovery module 602, configured to, if the first target computing node is in a normal state, recover the target sub-training task based on the status data in the memory of the first target computing node; A second recovery module 603, configured to, if the first target computing node is in a faulty state, recover the target sub-training task based on the status data in the memory of the second target computing node.

[0092] In some embodiments, before monitoring the target sub-training task, the monitoring module 601 is further configured to: Generate task configuration information, where the task configuration information includes status data parameters, sub-training tasks to be executed by each computing node, and association relationships between the computing nodes, and the status data parameters are used to specify target data as status data; Send the task configuration information to each computing node.

[0093] In some embodiments, the monitoring module 601 determines the status data parameters based on the following method: Calculate the total data volume of the execution logs of the model training task and the data volume of the model snapshot of the model to be trained; If the total data volume of the execution logs is greater than the data volume of the model snapshot, set the status data parameters to a first value to specify that the model snapshot is used as status data; If the total data volume of the execution logs is less than the data volume of the model snapshot, set the status data parameters to a second value to specify that the execution logs of each sub-training task are used as status data.

[0094] In some embodiments, the monitoring module 601 determines the sub-training tasks to be executed by each computing node based on the following method: When the status data parameters specify that the model snapshot is used as status data, divide the sub-training tasks according to the division logic that the execution duration difference of each sub-training task does not exceed a duration threshold; When the status data parameters specify that the execution logs of each sub-training task are used as status data, divide the sub-training tasks according to the division logic that the data volume difference of the status data generated by each computing node does not exceed a data volume threshold.

[0095] For the descriptions of the features in the embodiments corresponding to the status data storage device, reference may be made to the relevant descriptions of the embodiments corresponding to the status data storage method. For the descriptions of the features in the embodiments corresponding to the recovery device, reference may be made to the relevant descriptions of the embodiments corresponding to the recovery method, which will not be elaborated here one by one.

[0096] With reference to Figure 7 , an embodiment of the present application further provides an electronic device, including a memory 10 and a processor 20. A computer program is stored in the memory 10, and the processor 20 is configured to run the computer program to execute the steps in any of the above-described state data saving method embodiments.

[0097] An embodiment of the present application further provides a computer-readable storage medium, in which a computer program is stored. The computer program is configured to execute the steps in any of the above-described state data saving method embodiments when running.

[0098] In an exemplary embodiment, the above computer-readable storage medium may include, but is not limited to: various media such as USB flash drives, read-only memories (ROM for short), random access memories (RAM for short), external hard drives, magnetic disks, or optical discs that can store computer programs.

[0099] An embodiment of the present application further provides a computer program product. The above computer program product includes a computer program, and when the computer program is executed by a processor, it implements the steps in any of the above-described state data saving method embodiments.

[0100] An embodiment of the present application further provides another computer program product, including a non-volatile computer-readable storage medium. The non-volatile computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, it implements the steps in any of the above-described state data saving method embodiments.

[0101] Those skilled in the art can further realize that the units and algorithm steps of each example described in combination with the embodiments disclosed herein can be implemented by electronic hardware, computer software, or a combination of the two. To clearly illustrate the interchangeability of hardware and software, the components and steps of each example have been generally described according to their functions in the above description. Whether these functions are executed in a hardware or software manner depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of the present application.

[0102] The above has introduced in detail a method, apparatus, device, and storage medium for saving status data of a model training task provided by this application. Specific examples are used in this article to elaborate on the principle and implementation manner of this application. The description of the above embodiments is only used to help understand the method and its core idea of this application. It should be noted that for those of ordinary skill in the art of this technology, without departing from the principle of this application, several improvements and modifications can still be made to this application, and these improvements and modifications also fall within the protection scope of the claims of this application.

Claims

1. A method for preserving state data of a model training task, characterized in that: The model training task is divided into a plurality of sub-training tasks, and a plurality of computing nodes execute the sub-training tasks in parallel. The method is applied to a first computing node among the plurality of computing nodes; the method comprises: Receive task configuration information assigned by a management node, where the task configuration information includes a first sub-training task to be performed and a second computing node associated with the first computing node; Execute the first sub-training task, and save the state data of the first sub-training task to the local memory and the memory of the second computing node, the state data in the local memory and the state data in the memory of the second computing node back up each other, and the state data is used to recover the first sub-training task when the first sub-training task is abnormal, and / or to recover the second sub-training task running on the second computing node when the second sub-training task is abnormal.

2. The method according to claim 1, characterized in that The task configuration information further includes a state data parameter, and the state data parameter is used to specify target data as state data; the method further includes: Based on the state data parameter, using the execution log of the first sub-training task as the state data, or using a model snapshot of the trained model during the execution of the first sub-training task as the state data; Among them, when the total data volume of the execution log of the model training task is greater than the data volume of the model snapshot, the management node specifies the model snapshot as the status data through the status data parameter; when the total data volume of the execution log is less than the data volume of the model snapshot, the management node specifies the execution log of the first sub-training task as the status data through the status data parameter.

3. The method according to claim 2, characterized in that The trained model includes multiple neural network layers, the training process of the trained model includes a forward propagation process and a backward propagation process of data, and the first computing node executes the first sub-training task to train at least one neural network layer; When the execution log of the first sub-training task is used as the state data, the method further includes: During the forward propagation process, output data output to a third computing node in a forward propagation direction of the data is calculated, and the output data is used as the execution log; During the back propagation process, gradient data output to a fourth computing node in a back propagation direction of the data is calculated, and the gradient data is used as the execution log.

4. The method according to claim 3, characterized in that: The first computing node is associated with the third computing node and the fourth computing node, the first computing node includes a first central processing unit and a first graphics processing unit, the third computing node includes a second graphics processing unit, and the fourth computing node includes a third graphics processing unit; the method further includes: During the forward propagation process, the output data is saved to the local memory and the second graphics processor through the first graphics processor, and the output data in the local memory is saved to the memory of the third computing node through the first central processor; During the back propagation process, the gradient data is saved to the local memory and the third graphics processor through the first graphics processor, and the gradient data in the local memory is saved to the memory of the fourth computing node through the first central processing unit.

5. The method according to claim 2, characterized in that: The trained model includes multiple neural network layers, the training process of the trained model includes a forward propagation process and a backward propagation process of data, and the first computing node executes the first sub-training task to train at least one neural network layer; When the model snapshot of the trained model is used as the state data, if the first computing node trains the first neural network layer and the second neural network layer, the method further includes: For any target neural network layer in the first neural network layer and the second neural network layer, when both the forward propagation process and the backward propagation process of the target neural network layer have completed training, the parameters of the target neural network layer are saved as the model snapshot.

6. The method according to claim 5, characterized in that The method further comprises: When saving the parameters of the first neural network layer to the local memory and the memory of the second computing node, if the parameters of the third neural network layer are already saved in the local memory and the memory of the second computing node, the parameters of the third neural network layer are deleted.

7. The method according to claim 1 or 2, characterized in that: The first computing node includes a first central processing unit, a first graphics processing unit and a local video memory; Saving the state data of the first sub-training task to the local memory and the memory of the second computing node includes: When the first graphics processor is in a task execution state, the state data is saved in the local video memory; when the first graphics processor is in an idle state, the state data in the local video memory is saved in the local memory through the first graphics processor, and the state data in the local memory is saved in the memory of the second computing node through the first central processing unit.

8. The method according to claim 1, characterized in that The method further comprises: The state data is saved in a non-volatile storage, and the state data in the non-volatile storage is used to restore the first sub-training task or the second sub-training task when both the first computing node and the second computing node are abnormal.

9. A method for recovering a model training task, characterized in that: The model training task is divided into a plurality of sub-training tasks, and the sub-training tasks are executed in parallel by a plurality of computing nodes. The method is applied to a management node of the plurality of computing nodes; the method comprises: Monitoring an abnormal target sub-training task and a first target computing node running the target sub-training task, wherein state data of the target sub-training task is stored in a memory of the first target computing node and in a memory of a second target computing node associated with the first target computing node; If the first target computing node is in a normal state, restoring the target sub-training task based on the state data in the memory of the first target computing node; If the first target computing node is in a faulty state, the target sub-training task is restored based on the state data in the memory of the second target computing node.

10. The method according to claim 9, characterized in that Before monitoring the target sub-training task, the method further includes: Generate task configuration information, the task configuration information including state data parameters, sub-training tasks to be performed by each computing node, and associations between computing nodes, wherein the state data parameters are used to specify target data as state data; The task configuration information is sent to each of the computing nodes.

11. The method according to claim 10, characterized in that The state data parameters are determined based on the following method: Calculate the total data volume of the execution log of the model training task and the data volume of the model snapshot of the trained model; If the total data volume of the execution log is greater than the data volume of the model snapshot, setting the state data parameter to a first value to specify that the model snapshot is used as the state data; If the total data volume of the execution log is less than the data volume of the model snapshot, the state data parameter is set to a second value to specify that the execution log of each sub-training task is used as the state data.

12. The method according to claim 11, characterized in that The sub-training tasks to be performed by each computing node are divided based on the following method: When the state data parameter specifies that the model snapshot is used as the state data, the sub-training tasks are divided according to a division logic in which a difference in execution time of each sub-training task does not exceed a time threshold; The state data parameter specifies that the execution log of each sub-training task is used as the state data, and the sub-training tasks are divided according to the division logic that the data amount difference of the state data generated by each computing node does not exceed the data amount threshold.

13. A computer-readable storage medium, characterized in that: The computer-readable storage medium is used to store a computer program, and when the computer program is executed by a processor, the method according to any one of claims 1 to 12 is implemented.

14. An electronic device, characterized in that: The electronic device comprises a processor and a memory, wherein the memory is used to store a computer program, and when the computer program is executed by the processor, the method according to any one of claims 1 to 12 is implemented.

15. A computer program product comprising a computer program, characterized in that When the computer program is executed by a processor, the method according to any one of claims 1 to 12 is implemented.

Citation Information

Patent Citations

  • Model training method and device

    CN115033292A

  • Breakpoint storage and recovery method and device for large model training scene

    CN117851453A

  • Large model training fault recovery method and device based on distributed memory management

    CN119473732A

  • Model training methods and apparatus

    US20250111226A1