Method, Recovery Method and Device for Saving Status Data of Model Training Task

By backing up state data in the memory of the computing node in the distributed training system, the problem of low recovery efficiency of model training tasks is solved, fast and reliable task recovery is achieved, and model training efficiency is improved.

CN120066861BActive Publication Date: 2025-07-08INSPUR SUZHOU INTELLIGENT TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510538834.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-04-27
Publication Date
2025-07-08
Estimated Expiration
2045-04-27

AI Technical Summary

Technical Problem

In distributed training systems, the recovery efficiency of model training tasks is low, which affects the model training efficiency.

Method used

The model training task is divided into multiple sub-training tasks and is executed in parallel in the memory of multiple compute nodes. The state data is saved by backups in local memory and associated compute node memory, so as to quickly recover in abnormal situations.

Benefits of technology

It improves the recovery efficiency and reliability of model training tasks, reduces the impact of training interruptions caused by hardware failures, and improves the overall training efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120066861B_ABST
    Figure CN120066861B_ABST
Patent Text Reader

Abstract

The present application discloses a method and device for saving and restoring status data of a model training task, relating to the field of artificial intelligence computing technology. On the one hand, after saving the status data of a first sub-training task to the memory of a computing node, when the first sub-training task or the second sub-training task is abnormal, the status data for task restoration can be obtained from the memory. Since the data transmission efficiency of the memory is high, the restoration efficiency of the model training task can be greatly improved. On the other hand, after mutually backing up the status data of the first sub-training task in the memory of a first computing node (i.e., local memory) and the memory of a second computing node, if one of the computing nodes is in an abnormal state, the status data can still be obtained from the memory of the other computing node, ensuring the timeliness and reliability of task restoration. In summary, the technical solution of the present application can solve the problem of low restoration efficiency of the model training task in the related art.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of artificial intelligence computing technology, and particularly to a method, a recovery method, and a device for saving state data of a model training task. Background Art

[0002] With the development of artificial intelligence technology, both the scale of models and training data is gradually increasing, and a single device or a Graphics Processing Unit (GPU) may not be able to meet the training requirements. In view of this, a distributed training system has emerged. In a distributed training system, a model or training data can be divided into multiple hardware nodes for parallel training, which greatly improves the model training efficiency. However, with the expansion of the hardware scale, the problem of training interruption caused by hardware failures (such as node downtime, silent data errors, etc.) is also gradually rising, and it is necessary to frequently recover the model training task during the model training process.

[0003] Currently, in some technologies, the recovery efficiency of model training tasks is low, which in turn affects the model training efficiency. Summary of the Invention

[0004] This application provides a method for saving state data of a model training task, a recovery method, a state data saving device, a recovery device, an electronic device, a computer-readable storage medium, and a computer program product, so as to at least solve the problem of low recovery efficiency of model training tasks in related technologies.

[0005] This application provides a method for saving state data of a model training task. The model training task is divided into multiple sub-training tasks, and multiple computing nodes execute the sub-training tasks in parallel. The method is applied to a first computing node among the multiple computing nodes. The method includes:

[0006] Receiving task configuration information allocated by a management node, where the task configuration information includes a first sub-training task to be executed and a second computing node associated with the first computing node;

[0007] Executing the first sub-training task, and saving the state data of the first sub-training task to the local memory and the memory of the second computing node. The state data in the local memory and the state data in the memory of the second computing node are backed up to each other. The state data is used to recover the first sub-training task when the first sub-training task is abnormal, and / or to recover the second sub-training task when the second sub-training task running on the second computing node is abnormal.

[0008] The present application provides a method for restoring a model training task, where the model training task is divided into multiple sub-training tasks, and the sub-training tasks are executed in parallel by multiple computing nodes. The method is applied to a management node of the multiple computing nodes; the method includes:

[0009] Monitoring an abnormal target sub-training task and a first target computing node that runs the target sub-training task. Among them, in the memory of the first target computing node and in the memory of a second target computing node associated with the first target computing node, status data of the target sub-training task is saved;

[0010] If the first target computing node is in a normal state, restoring the target sub-training task based on the status data in the memory of the first target computing node;

[0011] If the first target computing node is in a fault state, restoring the target sub-training task based on the status data in the memory of the second target computing node.

[0012] The present application further provides a status data saving device for a model training task. The model training task is divided into multiple sub-training tasks, and multiple computing nodes execute the sub-training tasks in parallel. The method is applied to a first computing node among the multiple computing nodes; the device includes:

[0013] An information receiving module, configured to receive task configuration information allocated by a management node, where the task configuration information includes a first sub-training task to be executed and a second computing node associated with the first computing node;

[0014] A task execution module, configured to execute the first sub-training task, and save the status data of the first sub-training task to the local memory and the memory of the second computing node. The status data in the local memory and the status data in the memory of the second computing node are backed up to each other. The status data is used to restore the first sub-training task when the first sub-training task is abnormal, and / or to restore a second sub-training task running on the second computing node when the second sub-training task is abnormal.

[0015] The present application provides a restoration device for a model training task. The model training task is divided into multiple sub-training tasks, and the sub-training tasks are executed in parallel by multiple computing nodes. The method is applied to a management node of the multiple computing nodes; the device includes:

[0016] A monitoring module, configured to monitor an abnormal target sub-training task and a first target computing node that runs the target sub-training task, wherein status data of the target sub-training task is stored in the memory of the first target computing node and in the memory of a second target computing node associated with the first target computing node;

[0017] A first recovery module, configured to, if the first target computing node is in a normal state, recover the target sub-training task based on the status data in the memory of the first target computing node;

[0018] A second recovery module, configured to, if the first target computing node is in a faulty state, recover the target sub-training task based on the status data in the memory of the second target computing node.

[0019] The present application further provides an electronic device, including: a memory, configured to store a computer program; a processor, configured to implement the steps of any of the above state data saving methods when executing the computer program.

[0020] The present application further provides a computer-readable storage medium, in which a computer program is stored, and wherein the computer program implements the steps of any of the above annotation methods when executed by a processor.

[0021] The present application further provides a computer program product, including a computer program, and the computer program implements the steps of any of the above annotation methods when executed by a processor.

[0022] In the technical solutions of some embodiments of the present application, on the one hand, after saving the status data of the first sub-training task to the memory of the computing node, when the first sub-training task or the second sub-training task is abnormal, the status data for task recovery can be obtained from the memory. Since the data transmission efficiency of the memory is high, the recovery efficiency of the model training task can be greatly improved. On the other hand, after mutually backing up the status data of the first sub-training task in the memory of the first computing node (i.e., the local memory) and the memory of the second computing node, if one of the computing nodes is in an abnormal state, the status data can still be obtained from the memory of the other computing node, ensuring the timeliness and reliability of task recovery. Based on the above two aspects, the technical solutions of the present application can solve the problem of low recovery efficiency of the model training task in the related art. BRIEF DESCRIPTION OF THE DRAWINGS

[0023] To more clearly illustrate the embodiments of the present application, the following will briefly introduce the drawings required in the embodiments. Obviously, the drawings in the following description are only some embodiments of the present application. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.

[0024] Figure 1 Schematic diagrams of modules of some distributed training systems;

[0025] Figure 2 Schematic flowchart of a status data saving method provided by some embodiments of the present application;

[0026] Figure 3 Schematic flowchart of the execution process of a model training task provided by some embodiments of the present application;

[0027] Figure 4 Schematic flowchart of a recovery method provided by some embodiments of the present application;

[0028] Figure 5 Schematic diagrams of modules of a status data saving device provided by some embodiments of the present application;

[0029] Figure 6 Schematic diagrams of modules of a recovery device provided by some embodiments of the present application;

[0030] Figure 7 Schematic diagrams of modules of an electronic device provided by some embodiments of the present application. Detailed implementation manners

[0031] Next, the technical solutions in the embodiments of the present application will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative efforts belong to the protection scope of the present application.

[0032] It should be noted that in the description of the present application, the terms "include", "comprise" or any other variant thereof are intended to cover a non-exclusive inclusion, so that a process, method, article or device including a series of elements not only includes those elements, but also includes other elements not expressly listed, or further includes elements inherent to such process, method, article or device. The terms "first", "second", etc. in the present application are used to distinguish similar objects, rather than to describe a specific order or sequence.

[0033] To enable those skilled in the art of the present technology to better understand the solution of the present application, the present application will be further described in detail below in conjunction with the accompanying drawings and specific implementation manners.

[0034] With reference to Figure 1 , which are schematic diagrams of modules of some distributed training systems. Figure 1In it, the distributed training system includes a management node 11 and multiple computing nodes 12. Each computing node 12 respectively includes a central processing unit 122, a video memory 123, a memory 124, and multiple graphics processing units 121. Among them, the memory 124 is the data storage space for the central processing unit 122, and the video memory 123 is the data storage space for the graphics processing unit 121.

[0035] In the management node 11, the model training task can be divided into multiple sub-training tasks and these sub-training tasks are assigned to each of the computing nodes 12. The multiple computing nodes 12 can execute the sub-training tasks in parallel, thereby improving the model training efficiency. For example, the model training task is divided into sub-training tasks A, B, and C, the sub-training task A is assigned to the computing node a, the sub-training task B is assigned to the computing node b, and the sub-training task C is assigned to the computing node c. In this way, the computing nodes a, b, and c can run the sub-training tasks A, B, and C in parallel, thereby improving the model training efficiency.

[0036] In each computing node 12, the graphics processing unit 121 can be used to execute the model training task, and the central processing unit 122 can be used to cooperate in executing the model training task. For example, the central processing unit 122 can read the model training data from the disk or the memory, convert the model training data into the format required for model training, and then input the data into the graphics processing unit 121. The graphics processing unit 121 can perform specific model training calculations, such as the forward propagation calculation of a neural network model, the gradient calculation in the backpropagation process, etc.

[0037] In each computing node 12, according to the number of graphics processing units 121 it includes, the sub-training task can be further divided, that is, the graphics processing units 121 in the same computing node 12 can execute the sub-training tasks assigned to the computing node 12 in parallel. For example, assume that the above-mentioned computing node a includes 8 graphics processing units G1, G2,..., G8, then the sub-training task A assigned to the computing node a can be further divided into training tasks A1, A2,..., A8. The graphics processing unit G1 can execute the training task A1, the graphics processing unit G2 can execute the training task A2, and so on.

[0038] The following explains the division of the model training task.

[0039] The management node 11 can divide the model training task or the sub-training tasks of each computing node 12 according to task division strategies such as data parallelism, tensor parallelism, and pipeline parallelism.

[0040] Data parallelism means dividing the training data set of the model to be trained into multiple sub-training data sets, and each sub-training data set can be assigned to a graphics processor 121 for model training. When dividing the model training tasks according to the data parallelism strategy, the training results of each graphics processor 121 can be aggregated, and the model parameters can be updated based on the aggregated results.

[0041] Tensor parallelism means dividing the parameters or tensors of the model to be trained into multiple parameter sets or tensor sets, and each parameter set or tensor set can be assigned to a graphics processor 121 for model training. When dividing the model training tasks according to the tensor parallelism strategy, the graphics processors 121 can communicate with each other to perform data synchronization and parameter update.

[0042] Pipeline parallelism means dividing the model training tasks by stage, and each computing node can execute the training tasks of one of the stages. For example, in a neural network model, one or more neural network layers can be regarded as one stage, and each computing node can execute the training tasks of one of the stages, that is, execute the training tasks of one or more neural network layers. Multiple computing nodes can execute the assigned sub-training tasks in sequence according to the data propagation direction, and there is a dependency relationship between the sub-training tasks of the computing nodes. For example, assume that computing node a executes the training tasks of neural network layers 0-3, and computing node b executes the training tasks of neural network layers 4-7. After computing node a executes the training tasks of neural network layers 0-3, the obtained results can be input into computing node b, and computing node b executes the training tasks of neural network layers 4-7 based on the output data of computing node a.

[0043] The task division strategy between computing nodes 12 and the task division strategy between graphics processors 121 can be different. For example, after assigning the training tasks of neural network layers 0-3 to computing node a according to the pipeline parallelism strategy, in computing node a, according to the tensor parallelism task division strategy, the training tasks of neural network layers 0-3 can be further divided into multiple sub-training tasks, so that the graphics processors 121 in computing node a can execute the training tasks of neural network layers 0-3 in parallel.

[0044] Although the distributed training system can improve the model training efficiency, with the expansion of the hardware scale, the problem of training interruption caused by hardware failures (such as node downtime, silent data errors, etc.) is also gradually increasing, and it is necessary to frequently resume the model training tasks during the model training process.

[0045] Currently, in some technologies, model snapshots of the model being trained are recorded at one or more specified time points during the model training process. The model snapshots can be saved in a remote non-volatile storage. When the model training is interrupted, the model snapshot at the most recent time point can be obtained from the remote non-volatile storage, and based on the information in the model snapshot, the model training task can be resumed starting from the time when the model snapshot was recorded. In these technologies, the data volume of the model snapshots is relatively large. If data compression is performed, it may cause damage to the model training progress. If data compression is not performed, limited by the network transmission bandwidth, it will be time-consuming to obtain the model snapshot from the remote storage, and the recovery efficiency of the model training task is relatively low, affecting the model training efficiency.

[0046] In view of this, the present application first provides a method for saving the status data of a model training task, which can solve the problem of relatively low recovery efficiency of the model training task in some technologies. The status data saving method can be applied to Figure 1 the first computing node among the multiple computing nodes 12 shown in Figure 2 . Referring to Figure 2 , the flow diagram of the status data saving method provided for some embodiments of the present application is shown.

[0047] Step S201: Receive the task configuration information assigned by the management node 11. The task configuration information includes the first subtraining task to be executed and the second computing node associated with the first computing node.

[0048] Regarding the subtraining task, reference can be made to Figure 1 the relevant description, which will not be elaborated here.

[0049] In the task configuration information, the node information of the second computing node associated with the first computing node can be included, such as the node identifier, IP address, etc.

[0050] In this embodiment, the management node 11 can use the computing nodes that meet the following conditions 11) to 13) as the second computing node.

[0051] 11) There is data interaction with the first computing node during the model training process. For example, assume that the training process of the model being trained includes a forward propagation process and a backward propagation process. Computing node a executes the training tasks of neural network layers 0 to 3, and computing node b executes the training tasks of neural network layers 4 to 7. During the forward propagation process, computing node b needs to receive the output data of computing node a. During the backward propagation process, computing node a needs to receive the gradient data output by computing node b. Therefore, when computing node a is the first computing node, computing node b can be regarded as a computing node that has data interaction with the first computing node.

[0052] 12) The amount of remaining memory space is greater than the space amount threshold.

[0053] 13) Communicate and connect with the first computing node.

[0054] In practical applications, a computing node that only meets the above conditions 12) to 13) can also be used as the second computing node. Or, on the basis of the above conditions 12) to 13), other conditions except the above condition 11) can also be defined according to actual requirements. This application does not limit this.

[0055] Step S202, execute the first sub-training task, and save the status data of the first sub-training task to the local memory and the memory of the second computing node. The status data in the local memory and the status data in the memory of the second computing node are backed up to each other. The status data is used to recover the first sub-training task when the first sub-training task is abnormal, and / or to recover the second sub-training task when the second sub-training task running on the second computing node is abnormal.

[0056] Specifically, the status data of the first sub-training task is used to track the execution process of the first sub-training task.

[0057] In this embodiment, the status data of the first sub-training task can be one of the following data:

[0058] 21) A model snapshot of the model being trained during the execution of the first sub-training task.

[0059] 22) The execution log of the first sub-training task.

[0060] Among them, the model snapshot can represent the complete training state of the model being trained at a specified time point, including but not limited to the model parameters, optimizer state, training progress information (such as the current number of epochs, batch number) of the model being trained at a specific time point, etc. For example, during the execution of the first sub-training task, at time point T1 when training reaches the 50th epoch, a model snapshot P1 can be saved. The model snapshot P1 can include the weights and biases of the model being trained at time point T1, the optimizer learning rate, momentum settings, training progress information (such as the number of completed epochs and total batch number). At time point T2 when training reaches the 100th epoch, a model snapshot P2 can be saved. The model snapshot P2 can include the weights and biases of the model being trained at time point T2, the optimizer learning rate, momentum settings, training progress information.

[0061] The execution log can characterize the key events, information flow, and timestamps during the execution of the first sub-training task. Among them, the key events and information flow include, but are not limited to, the intermediate parameters of the model being trained during the execution of the first sub-training task (such as the output data of a specific neural network layer, gradient information, etc.), system information (such as memory occupancy rate, etc.), the identifier of the graphics processing unit, etc. For example, when starting to train the first neural network layer based on the 1st epoch, the timestamp of the start of training and the initial memory occupancy rate can be recorded; after the training of the first neural network layer based on the 1st epoch is completed, the timestamp of the end of training and the output data of the first neural network layer can be recorded.

[0062] When the first sub-training task or the second sub-training task is abnormal, the status data of the first sub-training task can be used to recover the first sub-training task or the second sub-training task. For example, when the output data of the first sub-training task is used as the input data of the second sub-training task, the second sub-training task can be recovered based on the output data of the first sub-training task.

[0063] Specifically, when the first sub-training task is abnormal, if the first computing node is in a normal state, the status data saved in the memory of the first computing node (i.e., local memory) can be used to recover the first sub-training task; if the first computing node is in an abnormal state, the status data saved in the memory of the second computing node can be used to recover the first sub-training task. Similarly, when the second sub-training task is abnormal, if the second computing node is in a normal state, the status data saved in the memory of the second computing node can be used to recover the second sub-training task; if the second computing node is in an abnormal state, the status data saved in the memory of the first computing node (i.e., local memory) can be used to recover the second sub-training task.

[0064] Since the status data in the memory of the first computing node (i.e., local memory) and the status data in the memory of the second computing node are backed up to each other, when one of the computing nodes fails, the status data for recovering the sub-training task can also be obtained from the memory of the neighboring computing node. Compared with some technologies that save the model snapshot in a remote non-volatile storage, the present application saves the status data in the memories of multiple computing nodes, which can greatly improve the transmission efficiency of the status data, and thus can improve the recovery efficiency of the first sub-training task or the second sub-training task, and further improve the model training efficiency.

[0065] In summary, in the technical solutions of some embodiments of the present application, on the one hand, after saving the status data of the first sub-training task to the memory of the computing node, when the first sub-training task or the second sub-training task is abnormal, the status data for task recovery can be obtained from the memory. Since the data transmission efficiency of the memory is high, the recovery efficiency of the model training task can be greatly improved. On the other hand, after mutually backing up the status data of the first sub-training task in the memory of the first computing node (i.e., local memory) and the memory of the second computing node, if one of the computing nodes is in an abnormal state, the status data can still be obtained from the memory of the other computing node, ensuring the timeliness and reliability of task recovery. Based on the above two aspects, the technical solution of the present application can solve the problem of low recovery efficiency of the model training task in the related art.

[0066] In some embodiments, considering that there is a risk of data loss in the memory during the abnormal or restart process of the computing node, therefore, the number of second computing nodes can be multiple, that is, the status data of the first sub-training task can be saved in the local memory and the memories of multiple second computing nodes. Performing multiple backups of the status data can reduce the degree of data loss risk, and thus improve the reliability of task recovery.

[0067] Furthermore, in some embodiments, the status data can also be saved to non-volatile storage. Compared with the memory, the status data in non-volatile storage will not be lost due to reasons such as device restart, but its data transmission efficiency is relatively slow. In view of this, when at least one of the first computing node and the second computing node is in a normal state, task recovery can be based on the status data in the memory of the first computing node (i.e., local memory) or the memory of the second computing node. In this way, the recovery efficiency of the model training task can be improved. When both the first computing node and the second computing node are abnormal, task recovery can be based on the status data in non-volatile storage. In this way, the reliability of model training task recovery can be ensured.

[0068] In some embodiments, in step S202, since the status data of the first sub-training task can be one of the model snapshot and the execution log, the task configuration information can also include status data parameters. The status data parameters are used to specify the target data as the status data. Based on the status data parameters, the execution log of the first sub-training task can be used as the status data, or the model snapshot of the model being trained during the execution of the first sub-training task can be used as the status data. For example, when the status data parameters specify that the execution log is used as the status data, the execution log of the first sub-training task can be saved to the local memory and the memory of the second computing node; when the status data parameters specify that the model snapshot is used as the status data, the execution log of the first sub-training task can be saved to the local memory and the memory of the second computing node.

[0069] On the management node 11 side, before the model training task starts to execute, the total data volume of the execution log of the model training task and the data volume of the model snapshot can be evaluated, and the status data parameter can be set according to the total data volume of the execution log and the data volume of the model snapshot to specify the target data as the status data. Specifically, when the total data volume of the execution log of the model training task is greater than the data volume of the model snapshot, the management node 11 can specify the model snapshot as the status data through the status data parameter. When the total data volume of the execution log is less than the data volume of the model snapshot, the management node 11 can specify the execution log of the first sub-training task as the status data through the status data parameter. In this way, the amount of status data saved in the memory can be relatively small, avoiding the excessive occupancy rate of the memory space by the status data and affecting other services.

[0070] Continue to refer to Figure 1 In some embodiments, the first computing node includes a first central processing unit, a first graphics processing unit, and a local video memory. Based on some specific task partitioning strategies (such as pipeline parallelism), the first graphics processing unit in the first computing node can execute the model training task only in some stages of the model training, that is, during the entire process of the model training, the first graphics processing unit can have idle time, which is also called a pipeline bubble. For example, assume that the first computing node is used to execute the training task of the 0-3 neural network layers of the model to be trained (i.e., the first sub-training task), then when the 4-7 neural network layers of the model to be trained are being trained, the first graphics processing unit is in an idle state.

[0071] Based on a specific task partitioning strategy, the first computing node can save the status data of the first sub-training task to the local memory and the memory of the second computing node in the following way: when the first graphics processing unit is in the task execution state, save the status data to the local video memory, and when the first graphics processing unit is in the idle state, save the status data in the local video memory to the local memory through the first graphics processing unit, and save the status data in the local memory to the memory of the second computing node through the first central processing unit.

[0072] Specifically, during the execution of the first sub-training task by the first graphics processing unit, the status data of the first sub-training task can be obtained. These status data can be first saved in the local video memory. When the first graphics processing unit finishes executing the first sub-training task and is in the idle state, then save the status data in the local video memory to the local memory through the first graphics processing unit. The first central processing unit can save the status data in the local memory to the memory of the second computing node through a background task.

[0073] This asynchronous state data saving solution can prevent the operation of saving state data from preempting resources with the first subtraining task, thereby ensuring the training efficiency of the first subtraining task. For example, assume that during the execution of the first subtraining task, the first graphics processing unit needs to communicate with the first central processing unit to obtain training data required for training or send the training results to the first central processing unit. If the state data is saved to the first memory during the task execution, then the operation of saving state data will preempt communication resources with the first subtraining task, thus affecting the training efficiency of the first subtraining task.

[0074] The following further elaborates on the solution of this application in combination with some specific application scenarios. Referring to Figure 3 is a schematic flowchart of the execution process of the model training task provided by some embodiments of this application. Figure 3 In, the model to be trained includes multiple neural network layers. According to the task allocation strategy of pipeline parallelism, training tasks of different neural network layers are assigned to different computing nodes. For example, computing node A executes the training tasks of neural network layers 0 to 2, computing node B executes the training tasks of neural network layer 3, and so on. The training process of the model to be trained may include the forward propagation process and the backward propagation process of data. In the forward propagation process, data is transmitted in the direction indicated by the solid arrow, and the direction indicated by the solid arrow is also called the forward propagation direction. In the backward propagation process, data is transmitted in the direction indicated by the dashed arrow, and the direction indicated by the dashed arrow is also called the backward propagation direction. The data transmitted in the forward propagation process may include training data, output data of each computing node, etc.; the data transmitted in the backward propagation process may include gradient data of each computing node.

[0075] Considering that when a computing node is abnormal, the probability of a processor or a graphics processing unit in the computing node failing is relatively low. Therefore, when using the execution log of the first subtraining task as state data, only the communication information between computing nodes can be recorded. In this way, the data volume of the execution log can be reduced, and thus the memory space occupancy can be reduced. Specifically, when using the execution log of the first subtraining task as state data, the method of this application may further include:

[0076] In the forward propagation process, calculate the output data output to the third computing node in the forward propagation direction of the data, and use the output data as the execution log;

[0077] In the backward propagation process, calculate the gradient data output to the fourth computing node in the backward propagation direction of the data, and use the gradient data as the execution log.

[0078] Specifically, the third computing node is the next computing node of the first computing node in the forward propagation direction, and the fourth computing node is the next computing node of the first computing node in the backward propagation direction. For example, with reference to Figure 3 . Assuming that computing node C is the first computing node, then computing node D is the third computing node, and computing node B is the fourth computing node.

[0079] Since there is data interaction between the third computing node, the fourth computing node and the first computing node, the third computing node and the fourth computing node can be associated with the first computing node. That is, the second computing nodes associated with the first computing node can include the third computing node and the fourth computing node. During the forward propagation process, the third computing node needs to execute the sub-training task assigned to it based on the output data of the first computing node. After the sub-training task executed by the third computing node is abnormal, the sub-training task on the third computing node can be restored based on the output data of the first computing node. During the backward propagation process, the fourth computing node needs to execute the sub-training task assigned to it based on the output data of the first computing node. After the sub-training task executed by the fourth computing node is abnormal, the sub-training task on the fourth computing node can be restored based on the output data of the first computing node.

[0080] Based on the above description, in some embodiments, the third computing node includes a second graphics processor, and the fourth computing node includes a third graphics processor.

[0081] During the forward propagation process, the first computing node can save the output data to the local memory and the second graphics processor through the first graphics processor, and save the output data in the local memory to the memory of the third computing node through the first central processing unit. Specifically, during the process of the first graphics processor executing the first sub-training task, the output data can be sent to the second graphics processor. Based on the received data, the second graphics processor can execute the sub-training task assigned to the third computing node. During the idle time after the first graphics processor executes the first sub-training task, the output data can be saved to the local memory of the first computing node. The first central processing unit saves the output data to the memory of the third computing node through a background task. In this way, after the sub-training task on the third computing node is abnormal, if the third computing node is in a normal state, the sub-training task can be restored based on the output data in the memory of the third computing node. If the third computing node is in an abnormal state, the sub-training task can be restored based on the output data in the memory of the first computing node.

[0082] Similarly, during the backpropagation process, the first computing node can save the gradient data to the local memory and the third graphics processing unit through the first graphics processing unit, and save the gradient data in the local memory to the memory of the fourth computing node through the first central processing unit. The relevant principle is similar to the above forward propagation process and will not be elaborated here.

[0083] Based on the above description, with reference to Figure 3 . It can be understood that during the forward propagation process, the first computing node can receive and save the output data of the fourth computing node (i.e., computing node B), and during the backpropagation process, the first computing node can receive and save the gradient data of the third computing node (i.e., computing node D). After the first subtask in the first computing node is abnormal, if the first computing node is in a normal state, the first subtask can be restored based on the output data or gradient data in the memory of the first computing node. If the first computing node is in an abnormal state, the first subtask can be restored based on the output data in the memory of the fourth computing node or the gradient data in the memory of the third computing node.

[0084] In some embodiments, when saving the output data and gradient data, auxiliary information can also be saved simultaneously. The auxiliary information can include, but is not limited to, the timestamp of the output data, the graphics processing unit identifier of the first graphics processing unit, etc. In this way, based on the auxiliary information, the generation time of the output data and the graphics processing unit that generated the output data can be determined.

[0085] Furthermore, with reference to Figure 3 . In some embodiments, after all of the first computing node (i.e., computing node C), the third computing node (i.e., computing node D), and the fourth computing node (i.e., computing node B) are abnormal, the next computing node (i.e., computing node E) in the forward propagation direction of the third computing node can be found as the first auxiliary computing node, and the next computing node (i.e., computing node A) in the backpropagation direction of the fourth computing node can be found as the second auxiliary computing node. Based on the gradient data in the memory of the first auxiliary computing node and the output data in the memory of the second auxiliary computing node, the subtasks on the third computing node and the fourth computing node can be restored first, and then the subtask on the first computing node can be restored. In this way, when multiple computing nodes are in an abnormal state, data acquisition from a remote non-volatile storage can be avoided, thereby improving the recovery efficiency of the model training task.

[0086] When recovering tasks based on the execution log, local recovery of the model training task can be performed, that is, only the abnormal subtasks can be restored. In this way, the task recovery time can be reduced and the task recovery efficiency can be improved.

[0087] Continue to refer to Figure 3. In some embodiments, when using the model snapshot of the trained model as the state data, if the first computing node trains the first neural network layer and the second neural network layer, the first computing node may further perform the following operations: for any target neural network layer among the first neural network layer and the second neural network layer, when the forward propagation process and the backward propagation process of the target neural network layer are both completed, the parameters of the target neural network layer are saved as the model snapshot. Simply put, the first computing node can save the state data layer by layer, that is, after each neural network layer is trained, the saving of the state data is triggered. In this way, the amount of state data saved at one time is avoided from being too large.

[0088] Furthermore, considering that when restoring the task, the parameters of the neural network layer that was last trained are usually referred to, so only the parameters of the neural network layer that was last trained can be saved in the local memory and the memory of the second computing node. Here, the second computing node includes at least one of the above-mentioned third computing node and the fourth computing node. In view of this, in some embodiments, when saving the parameters of the first neural network layer to the local memory and the memory of the second computing node, if the parameters of the third neural network layer are already saved in the local memory and the memory of the second computing node, the parameters of the third neural network layer are deleted. In this way, the consumption of memory space is reduced.

[0089] Corresponding to the state data saving method, the present application also provides a method for restoring a model training task, which can solve the problem of low restoration efficiency of the model training task in some technologies. The restoration method can be applied to Figure 1 the management node 11 shown. Referring to Figure 4 , it is a schematic flowchart of the restoration method provided by some embodiments of the present application. Figure 4 In it, the restoration method includes the following steps:

[0090] Step S401, monitor the abnormal target sub-training task and the first target computing node running the target sub-training task, where the state data of the target sub-training task is saved in the memory of the first target computing node and the memory of the second target computing node associated with the first target computing node.

[0091] Specifically, the first target computing node includes a first central processing unit, a first graphics processing unit, and local video memory. During the execution of the target sub-training task by the first graphics processing unit, the state data of the target sub-training task is obtained; the first target computing node saves the state data of the target sub-training task to the memory of the first target computing node and the second target computing node based on the following method:

[0092] During the execution of the target sub-training task by the first graphics processor, the status data is saved to the local video memory. When the first graphics processor finishes executing the target sub-training task and is in an idle state, the status data in the local video memory is saved to the local memory through the first graphics processor, and the status data in the local memory is saved to the memory of the second target computing node through the first central processing unit.

[0093] Step S402, if the first target computing node is in a normal state, the target sub-training task is restored based on the status data in the memory of the first target computing node.

[0094] Step S403, if the first target computing node is in a faulty state, the target sub-training task is restored based on the status data in the memory of the second target computing node.

[0095] Regarding the above steps S401 to S403, reference can be made to the relevant descriptions of the above status data saving method, which will not be elaborated here.

[0096] In some embodiments, before monitoring the target sub-training task, the management node 11 can also perform the following operations:

[0097] Generate task configuration information, which includes status data parameters, the sub-training tasks to be executed by each computing node, and the association relationship between computing nodes. Among them, the status data parameters are used to specify the target data as the status data;

[0098] Send the task configuration information to each computing node.

[0099] In some embodiments, the status data parameters are determined based on the following method:

[0100] Calculate the total data volume of the execution logs of the model training task and the data volume of the model snapshots of the model to be trained;

[0101] If the total data volume of the execution logs is greater than the data volume of the model snapshots, set the status data parameters to the first value to specify that the model snapshots are used as the status data;

[0102] If the total data volume of the execution logs is less than the data volume of the model snapshots, set the status data parameters to the second value to specify that the execution logs of each sub-training task are used as the status data.

[0103] In some embodiments, when the status data parameters specify that the model snapshots are used as the status data, the sub-training tasks can be divided according to the division logic that the execution duration difference of each sub-training task does not exceed the duration threshold. Specifically, assume that the training duration of the i-th neural network layer is , and it is necessary to divide the model training task into M sub-training tasks. Then, the execution duration of a single sub-training task can be shown as in Expression (1):

[0104]

[0105] Where T is the execution duration of a single sub-training task, N is the maximum number of neural network layers, the values of i and N are positive integers, and M is the number of computing nodes.

[0106] In some embodiments, when dividing the sub-training tasks according to the execution duration T, if the task data volume corresponding to the sub-training task (such as the sum of the model parameter quantity and the training data volume) is greater than x times the video memory capacity of a single computing node, a prompt of task allocation failure can be returned. Where the value of x is a numerical value between 0 and 1.

[0107] In some embodiments, when the status data parameter specifies that the execution logs of each sub-training task are used as status data, the sub-training tasks can be divided according to the division logic that the data volume difference of the status data generated by each computing node does not exceed the data volume threshold. Specifically, assume that the total data volume of the model to be trained is as shown in Expression (2):

[0108]

[0109] Where represents the total number of parameters of the model to be trained, represents the data volume of the i-th neural network layer, and N is the maximum number of neural network layers.

[0110] The data volume of a single sub-training task can be shown as in Expression (3):

[0111]

[0112] Where represents the data volume of a single sub-training task, and M is the number of computing nodes.

[0113] In some embodiments, when dividing the sub-training tasks according to the data volume if the data volume is greater than x1 times the video memory capacity of a single computing node or greater than x2 times the memory capacity of a single computing node, a prompt of task allocation failure can be returned. Where the values of x1 and x2 are numerical values between 0 and 1.

[0114] In some embodiments, according to the principles of the above Expressions (1), (2), and (3), the sub-training tasks executed by a single computing node can be further divided, which will not be elaborated here.

[0115] The recovery method has the same technical features as the above-mentioned state data saving method, that is, the state data of the target sub-training task is saved in the memory of the first target computing node and the memory of the second target computing node associated with the first target computing node. Therefore, it has the same beneficial effects as the above-mentioned state data saving method and will not be elaborated here.

[0116] Through the description of the above embodiments, those skilled in the art can clearly understand that the method according to the above embodiments can be implemented by means of software plus a necessary general hardware platform. Of course, it can also be implemented by hardware, but in many cases the former is a better implementation method.

[0117] The embodiments of the present application also provide a device for saving the state data of a model training task. Referring to Figure 5 , it is a schematic diagram of the modules of the state data saving device provided in some embodiments of the present application. Figure 5 In

[0118] The information receiving module 501 is used to receive the task configuration information allocated by the management node. The task configuration information includes the first sub-training task to be executed and the second computing node associated with the first computing node;

[0119] The task execution module 502 is used to execute the first sub-training task and save the state data of the first sub-training task to the local memory and the memory of the second computing node. The state data in the local memory and the state data in the memory of the second computing node are backed up to each other. The state data is used to recover the first sub-training task when the first sub-training task is abnormal, and / or to recover the second sub-training task when the second sub-training task running on the second computing node is abnormal.

[0120] In some embodiments, the task configuration information further includes state data parameters, and the state data parameters are used to specify the target data as the state data; the task execution module 502 is further used for:

[0121] Based on the state data parameters, taking the execution log of the first sub-training task as the state data, or taking the model snapshot of the model being trained during the execution of the first sub-training task as the state data;

[0122] Among them, when the total data volume of the execution log of the model training task is greater than the data volume of the model snapshot, the management node specifies the model snapshot as the state data through the state data parameters. When the total data volume of the execution log is less than the data volume of the model snapshot, the management node specifies the execution log of the first sub-training task as the state data through the state data parameters.

[0123] In some embodiments, the trained model includes multiple neural network layers. The training process of the trained model includes a forward propagation process and a backward propagation process of data. The first computing node executes a first sub-training task to train at least one neural network layer;

[0124] When using the execution log of the first sub-training task as status data, the task execution module 502 is further configured to:

[0125] In the forward propagation process, calculate the output data output to the third computing node in the forward propagation direction of the data, and use the output data as the execution log;

[0126] In the backward propagation process, calculate the gradient data output to the fourth computing node in the backward propagation direction of the data, and use the gradient data as the execution log.

[0127] In some embodiments, the first computing node is associated with the third computing node and the fourth computing node. The first computing node includes a first central processing unit and a first graphics processing unit. The third computing node includes a second graphics processing unit, and the fourth computing node includes a third graphics processing unit. The task execution module 502 is further configured to:

[0128] In the forward propagation process, through the first graphics processing unit, save the output data to the local memory and the second graphics processing unit, and through the first central processing unit, save the output data in the local memory to the memory of the third computing node;

[0129] In the backward propagation process, through the first graphics processing unit, save the gradient data to the local memory and the third graphics processing unit, and through the first central processing unit, save the gradient data in the local memory to the memory of the fourth computing node.

[0130] In some embodiments, the trained model includes multiple neural network layers. The training process of the trained model includes a forward propagation process and a backward propagation process. The first computing node executes a first sub-training task to train at least one neural network layer;

[0131] When using the model snapshot of the trained model as status data, if the first computing node trains the first neural network layer and the second neural network layer, the task execution module 502 is further configured to:

[0132] For any target neural network layer in the first neural network layer and the second neural network layer, when the training of both the forward propagation process and the backward propagation process of the target neural network layer is completed, save the parameters of the target neural network layer as the model snapshot.

[0133] In some embodiments, the task execution module 502 is further configured to:

[0134] When saving the parameters of the first neural network layer to the local memory and the memory of the second computing node, if the parameters of the third neural network layer are already saved in the local memory and the memory of the second computing node, the parameters of the third neural network layer are deleted.

[0135] In some embodiments, the first computing node includes a first central processing unit, a first graphics processing unit, and local video memory;

[0136] The task execution module 502 is specifically configured to:

[0137] When the first graphics processing unit is in a task execution state, save the status data to the local video memory. When the first graphics processing unit is in an idle state, save the status data in the local video memory to the local memory through the first graphics processing unit, and save the status data in the local memory to the memory of the second computing node through the first central processing unit.

[0138] In some embodiments, the task execution module 502 is further configured to:

[0139] Save the status data to non-volatile storage. The status data in the non-volatile storage is used to restore the first sub-training task or the second sub-training task when both the first computing node and the second computing node are abnormal.

[0140] Embodiments of the present application further provide a recovery device for a model training task. Referring to Figure 6 is a schematic diagram of the modules of the recovery device provided in some embodiments of the present application. Figure 6 In, the recovery device includes:

[0141] A monitoring module 601, configured to monitor an abnormal target sub-training task and a first target computing node running the target sub-training task. Among them, in the memory of the first target computing node and the memory of a second target computing node associated with the first target computing node, status data of the target sub-training task is saved;

[0142] A first recovery module 602, configured to, if the first target computing node is in a normal state, restore the target sub-training task based on the status data in the memory of the first target computing node;

[0143] A second recovery module 603, configured to, if the first target computing node is in a faulty state, restore the target sub-training task based on the status data in the memory of the second target computing node.

[0144] In some embodiments, before monitoring the target sub-training task, the monitoring module 601 is further configured to:

[0145] Generate task configuration information, where the task configuration information includes status data parameters, sub-training tasks to be executed by each computing node, and the association relationships between the computing nodes. Among them, the status data parameters are used to specify the target data as status data;

[0146] Send the task configuration information to each computing node.

[0147] In some embodiments, the monitoring module 601 determines the status data parameters based on the following method:

[0148] Calculate the total amount of execution log data of the model training task and the amount of data of the model snapshot of the model to be trained;

[0149] If the total amount of execution log data is greater than the amount of data of the model snapshot, set the status data parameter to the first value to specify that the model snapshot is used as the status data;

[0150] If the total amount of execution log data is less than the amount of data of the model snapshot, set the status data parameter to the second value to specify that the execution logs of each sub-training task are used as the status data.

[0151] In some embodiments, the monitoring module 601 determines the sub-training tasks to be executed by each computing node based on the following method:

[0152] When the status data parameter specifies that the model snapshot is used as the status data, divide the sub-training tasks according to the division logic that the execution duration difference of each sub-training task does not exceed the duration threshold;

[0153] When the status data parameter specifies that the execution logs of each sub-training task are used as the status data, divide the sub-training tasks according to the division logic that the data volume difference of the status data generated by each computing node does not exceed the data volume threshold.

[0154] For the description of the features in the embodiments corresponding to the status data saving device, reference can be made to the relevant descriptions in the embodiments corresponding to the status data saving method. For the description of the features in the embodiments corresponding to the recovery device, reference can be made to the relevant descriptions in the embodiments corresponding to the recovery method, which will not be elaborated here one by one.

[0155] With reference to Figure 7 , an embodiment of the present application further provides an electronic device, including a memory 10 and a processor 20. A computer program is stored in the memory 10, and the processor 20 is configured to run the computer program to execute the steps in any of the above embodiments of the status data saving method.

[0156] An embodiment of the present application further provides a computer-readable storage medium, in which a computer program is stored. The computer program is configured to execute the steps in any of the above embodiments of the status data saving method when running.

[0157] In an exemplary embodiment, the above computer-readable storage medium may include, but is not limited to: various media capable of storing computer programs such as USB flash drives, read-only memory (ROM for short), random access memory (RAM for short), mobile hard disks, magnetic disks, or optical discs.

[0158] The embodiments of the present application further provide a computer program product. The above computer program product includes a computer program, and when the computer program is executed by a processor, the steps in any of the above embodiments of the state data saving method are implemented.

[0159] The embodiments of the present application further provide another computer program product, including a non-volatile computer-readable storage medium. The non-volatile computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the steps in any of the above embodiments of the state data saving method are implemented.

[0160] Those skilled in the art can further realize that the units and algorithm steps of each example described in combination with the embodiments disclosed herein can be implemented by electronic hardware, computer software, or a combination of the two. To clearly illustrate the interchangeability of hardware and software, the composition and steps of each example have been generally described according to functions in the above description. Whether these functions are executed in a hardware or software manner depends on the specific application and design constraints of the technical solution. Skilled professionals can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of the present application.

[0161] The above has introduced in detail a method, apparatus, device, and storage medium for saving state data of a model training task provided by the present application. Specific examples are used herein to elaborate on the principle and implementation manner of the present application. The description of the above embodiments is only used to help understand the method and its core idea of the present application. It should be noted that for those of ordinary skill in the art in the technical field, without departing from the principle of the present application, several improvements and modifications can be made to the present application, and these improvements and modifications also fall within the protection scope of the claims of the present application.

Claims

1. A method for saving state data of a model training task, characterized in that, The model training task is divided into multiple sub-training tasks, and multiple computing nodes execute the sub-training tasks in parallel. The method is applied to a first computing node among the multiple computing nodes; the method includes: Receiving task configuration information assigned by a management node, where the task configuration information includes a first sub-training task to be executed and a second computing node associated with the first computing node; Executing the first sub-training task, and saving the status data of the first sub-training task to local memory and the memory of the second computing node. The status data in the local memory and the status data in the memory of the second computing node are backed up to each other. The status data is used to recover the first sub-training task when the first sub-training task is abnormal, and / or to recover the second sub-training task when the second sub-training task running on the second computing node is abnormal; Wherein, the first computing node includes a first central processing unit, a first graphics processing unit, and local video memory. During the process of executing the first sub-training task by the first graphics processing unit, the status data of the first sub-training task is obtained; saving the status data of the first sub-training task to local memory and the memory of the second computing node includes: During the process of executing the first sub-training task by the first graphics processing unit, saving the status data to local video memory. When the first graphics processing unit finishes executing the first sub-training task and is in an idle state, saving the status data in the local video memory to local memory through the first graphics processing unit, and saving the status data in the local memory to the memory of the second computing node through the first central processing unit.

2. The method according to claim 1, wherein The task configuration information further includes status data parameters, and the status data parameters are used to specify target data as status data; the method further includes: Based on the status data parameters, using the execution log of the first sub-training task as the status data, or using a model snapshot of the model being trained during the execution process of the first sub-training task as the status data; Wherein, when the total data volume of the execution log of the model training task is greater than the data volume of the model snapshot, the management node specifies the model snapshot as the status data through the status data parameters. When the total data volume of the execution log is less than the data volume of the model snapshot, the management node specifies the execution log of the first sub-training task as the status data through the status data parameters.

3. The method according to claim 2, wherein The model being trained includes multiple neural network layers. The training process of the model being trained includes a forward propagation process and a backward propagation process of data. The first computing node executes the first sub-training task to train at least one neural network layer; When using the execution log of the first sub-training task as the status data, the method further includes: During the forward propagation process, calculating output data output to a third computing node in the forward propagation direction of the data, and using the output data as the execution log; During the backpropagation process, calculate the gradient data output to the fourth computing node in the reverse propagation direction of the data, and use the gradient data as the execution log.

4. The method according to claim 3, characterized in that The first computing node is associated with the third computing node and the fourth computing node. The first computing node includes a first central processing unit and a first graphics processing unit. The third computing node includes a second graphics processing unit, and the fourth computing node includes a third graphics processing unit. The method further includes: During the forward propagation process, through the first graphics processing unit, save the output data to the local memory and the second graphics processing unit, and through the first central processing unit, save the output data in the local memory to the memory of the third computing node. During the backpropagation process, through the first graphics processing unit, save the gradient data to the local memory and the third graphics processing unit, and through the first central processing unit, save the gradient data in the local memory to the memory of the fourth computing node.

5. The method according to claim 2, wherein The trained model includes multiple neural network layers. The training process of the trained model includes a forward propagation process and a backpropagation process of data. The first computing node executes the first sub-training task to train at least one neural network layer. When using the model snapshot of the trained model as the state data, if the first computing node trains the first neural network layer and the second neural network layer, the method further includes: For any target neural network layer among the first neural network layer and the second neural network layer, when the forward propagation process and the backpropagation process of the target neural network layer are both completed, save the parameters of the target neural network layer as the model snapshot.

6. The method according to claim 5, wherein The method further includes: When saving the parameters of the first neural network layer to the local memory and the memory of the second computing node, if the parameters of the third neural network layer have been saved in the local memory and the memory of the second computing node, delete the parameters of the third neural network layer.

7. The method according to claim 1, characterized in that The method further includes: Save the state data to non-volatile storage. The state data in the non-volatile storage is used to restore the first sub-training task or the second sub-training task when both the first computing node and the second computing node are abnormal.

8. A method for resuming a model training task, characterized in that, The model training task is divided into multiple sub-training tasks, and the sub-training tasks are executed in parallel by multiple computing nodes. The method is applied to the management node of the multiple computing nodes. The method includes: Monitor the abnormal target sub-training task and the first target computing node running the target sub-training task. Among them, in the memory of the first target computing node and the memory of the second target computing node associated with the first target computing node, the state data of the target sub-training task is saved. If the first target computing node is in a normal state, restore the target sub-training task based on the state data in the memory of the first target computing node. If the first target computing node is in a fault state, the target sub-training task is restored based on the state data in the memory of the second target computing node; Wherein, the first target computing node includes a first central processing unit, a first graphics processing unit and a local video memory. During the execution of the target sub-training task by the first graphics processing unit, the state data of the target sub-training task is obtained; the first target computing node saves the state data of the target sub-training task to the memories of the first target computing node and the second target computing node based on the following method: During the execution of the target sub-training task by the first graphics processing unit, the state data is saved to the local video memory. When the first graphics processing unit finishes executing the target sub-training task and is in an idle state, the state data in the local video memory is saved to the local memory through the first graphics processing unit, and the state data in the local memory is saved to the memory of the second target computing node through the first central processing unit.

9. The method according to claim 8, characterized in that Before monitoring the target sub-training task, the method further includes: Generating task configuration information, where the task configuration information includes state data parameters, sub-training tasks to be executed by each computing node, and association relationships between computing nodes. The state data parameters are used to specify target data as state data; Sending the task configuration information to each of the computing nodes.

10. The method according to claim 9, wherein The state data parameters are determined based on the following method: Calculating the total data volume of the execution logs of the model training task and the data volume of the model snapshot of the model to be trained; If the total data volume of the execution logs is greater than the data volume of the model snapshot, the state data parameters are set to a first value to specify the model snapshot as the state data; If the total data volume of the execution logs is less than the data volume of the model snapshot, the state data parameters are set to a second value to specify the execution logs of each sub-training task as the state data.

11. The method according to claim 10, wherein The sub-training tasks to be executed by each of the computing nodes are divided based on the following method: When the state data parameters specify the model snapshot as the state data, the sub-training tasks are divided according to the division logic that the execution duration differences of each sub-training task do not exceed a duration threshold; When the state data parameters specify the execution logs of each sub-training task as the state data, the sub-training tasks are divided according to the division logic that the data volume differences of the state data generated by each computing node do not exceed a data volume threshold.

12. A computer-readable storage medium, characterized in that, The computer-readable storage medium is used to store a computer program, and when the computer program is executed by a processor, the method described in any one of claims 1 to 11 is implemented.

13. An electronic device, characterized in that, The electronic device includes a processor and a memory. The memory is used to store a computer program, and when the computer program is executed by the processor, the method described in any one of claims 1 to 11 is implemented.

14. A computer program product, comprising a computer program, characterized in that, When the computer program is executed by a processor, the method described in any one of claims 1 to 11 is implemented.

Citation Information

Patent Citations

  • Model training method and device

    CN115033292A

  • Breakpoint storage and recovery method and device for large model training scene

    CN117851453A