Model training method and electronic equipment

By storing checkpoint data in the training node memory and utilizing multi-model group redundant backup, the problem of model training interruption caused by GPU card failure is solved, and efficient fault recovery and low-cost GPU computing power utilization are achieved.

CN120705568APending Publication Date: 2025-09-26HUAWEI CLOUD COMPUTING TECHNOLOGIES CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202410853974.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2024-03-26
Filing Date
2024-06-28
Publication Date
2025-09-26

AI Technical Summary

Technical Problem

In large-scale distributed training clusters, GPU card failures lead to model data loss, resulting in model training interruptions and waste of GPU computing power. Existing technologies cannot effectively avoid interruptions and reduce costs.

Method used

Checkpoint data is directly stored in the memory of the training node, and the redundant backup feature of multiple model groups is used for data recovery, which reduces communication and remote storage dependence and achieves fast fault recovery through cluster memory collaborative management.

Benefits of technology

It significantly reduces the performance and bandwidth requirements for remote storage, reduces the waste of GPU card computing power, and improves the efficiency and reliability of model training.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120705568A_ABST
    Figure CN120705568A_ABST
Patent Text Reader

Abstract

The embodiment of the invention provides a model training method and device and a computing device cluster. The method is applied to a training cluster, the training cluster comprises a first model training group and a second model training group of a plurality of model training groups, and the first model training group and the second model training group are used for training a neural network model in parallel in a data parallel mode. The method comprises the steps of generating check point data of a first node and a second node after a round of training is finished; the first training node stores the check point data into a memory of the first training node, and the second training node stores the check point data into a memory of the second training node; and if the first training node breaks down in the process of training the model, the first training node obtains the check point data from the memory of the second training node. According to the model training method provided by the embodiment of the invention, the consumption of communication resources can be reduced, the requirements on the performance and bandwidth of remote storage are reduced, and the cost of the whole system is remarkably reduced.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of computer technology, and in particular to a model training method and electronic equipment. Background Art

[0002] The vast potential of large artificial intelligence (AI) models has led manufacturers to turn to the research and use of large models. Training large AI models with hundreds of billions of parameters often requires the use of thousands of graphics processing unit (GPU) cards for distributed cluster training. In such scenarios, frequent GPU card failures will cause the loss of model data being trained. The loss of model data will lead to the interruption of model training, and the model training needs to be re-performed, resulting in a high cost of wasted GPU card computing power. Therefore, a method is needed to restore the model training status when model data is lost to avoid the waste of GPU card computing power caused by GPU card failure, thereby improving model training efficiency. Summary of the Invention

[0003] To address the problem of how to avoid GPU card failure leading to interruption of model training and waste of GPU computing power, the present application provides a model training method, apparatus, and computing device cluster. The present application also provides a computer program product and a computer-readable storage medium.

[0004] The embodiments of this application adopt the following technical solutions:

[0005] In a first aspect, the present application provides a model training method, which is applied to a training cluster, wherein the training cluster includes multiple model training groups, the multiple model training groups include a first model training group and a second model training group, the first model training group and the second model training group are used to train a neural network model in parallel in a data-parallel manner, the first model training group and the second model training group include the same node structure, the first model training group includes a first training node, and the second model training group includes a second training node, wherein, based on the node structure, the first training node corresponds to the second training node, and the method includes:

[0006] Generate checkpoint data after a round of training, wherein the checkpoint data is the checkpoint data of the first training node and the second training node;

[0007] The first training node saves the checkpoint data into a memory of the first training node, and the second training node saves the checkpoint data into a memory of the second training node;

[0008] If the first training node fails during model training, the first training node obtains the checkpoint data from the memory of the second training node to perform data recovery on the first training node.

[0009] According to the model training method of the first aspect, the checkpoint data of the training node is backed up using the training node's own memory. During the checkpoint data backup process of the training node, the training nodes do not need to communicate with each other, which greatly reduces the consumption of communication resources.

[0010] According to the model training method of the first aspect, the training node's own memory is used to back up the checkpoint data of the training node, and the checkpoint data of the failed training node is retransmitted and restored by utilizing the redundant backup characteristics of the multi-model group. The checkpoint data backup and recovery process does not require the participation of the remote storage service. The remote storage service is moved out of the critical path, reducing the performance and bandwidth requirements of the remote storage, and significantly reducing the cost of the overall system.

[0011] In an implementation of the first aspect, the method further includes:

[0012] For the second training node that has not failed, data recovery is performed on the second training node using the checkpoint data stored in the memory of the second training node itself.

[0013] In an implementation of the first aspect, generating checkpoint data after a round of training includes:

[0014] At the end of a round of training, the model parameters of all training nodes in the training cluster are averaged and the calculated result is used as the checkpoint data for this round of training.

[0015] In an implementation of the first aspect, the method further includes:

[0016] Download initial checkpoint data using the first training node;

[0017] The first training node sends the initial checkpoint data to the second training node.

[0018] According to the above implementation method, the initial checkpoint data is downloaded to a model training group, and the initial checkpoint data is sent to other model training groups through one model training group, which reduces the bandwidth requirement for the initial checkpoint data and improves the download efficiency of the initial checkpoint data.

[0019] In an implementation of the first aspect, the method further includes:

[0020] Use the first model training group to upload the checkpoint data stored in the memory of each training node in the first model training group to a remote storage service.

[0021] According to the above implementation method, a solution is adopted to move the remote storage out of the critical path, and the cluster memory data is asynchronously persisted to the remote end, reducing dependence on remote performance; and, according to the method of one embodiment of the present application, the temporary storage of CKPT by memory and the persistent storage of CKPT by the remote storage service are distinguished, without increasing the space consumption of the remote storage.

[0022] In an implementation of the first aspect, during the process of uploading checkpoint data by the first model training group, it is prohibited to delete the checkpoint data saved in the memory of each training node in the first model training group, and it is prohibited to use new checkpoint data to overwrite the saved checkpoint data.

[0023] According to the above implementation method, it is possible to prevent the generation of new checkpoint data from interfering with the uploading of checkpoint data, thereby ensuring that the checkpoint data can be successfully uploaded to the remote storage service.

[0024] In an implementation of the first aspect, while the first model training group is uploading checkpoint data, saving new checkpoint data into the memory of the training node of the first model training group is stopped.

[0025] In an implementation of the first aspect, after the first model training group completes uploading the checkpoint data, the new checkpoint data stored in the memory of the training node of the second model training group is synchronized to the memory of the training node of the first model training group.

[0026] According to the above implementation method, the generation of new checkpoint data can be prevented from interfering with the upload of checkpoint data, ensuring that the checkpoint data can be successfully uploaded to the remote storage service; and it can be ensured that the checkpoint data version of the training node that uploads the checkpoint data is consistent with that of the training node that does not upload the checkpoint data.

[0027] In a second aspect, the present application provides a model training device, which is applied to a training cluster, wherein the training cluster includes multiple model training groups, the multiple model training groups include a first model training group and a second model training group, the first model training group and the second model training group are used to train a neural network model in parallel in a data-parallel manner, the first model training group and the second model training group include the same node structure, the first model training group includes a first training node, and the second model training group includes a second training node, wherein, based on the node structure, the first training node corresponds to the second training node; the training cluster generates checkpoint data after a round of training, wherein the checkpoint data is the checkpoint data of the first training node and the second training node;

[0028] The device comprises:

[0029] a storage control module, configured to instruct the first training node to save the checkpoint data into a memory of the first training node, and to instruct the second training node to save the checkpoint data into a memory of the second training node;

[0030] A recovery control module is used to instruct the first training node to obtain the checkpoint data from the memory of the second training node to perform data recovery on the first training node if the first training node fails during the training of the model.

[0031] In an implementation of the second aspect, the recovery control module is further configured to:

[0032] Instruct the second training node to use the checkpoint data stored in the memory of the second training node itself to perform data recovery on the second training node.

[0033] In an implementation of the second aspect, at the end of a round of training, the training cluster calculates an average value of the model parameters of all training nodes in the training cluster, and uses the calculation result as checkpoint data for the round of training.

[0034] In an implementation of the second aspect, the apparatus further includes an initialization module, configured to:

[0035] instructing the first training node to download initial checkpoint data;

[0036] Instruct the first training node to broadcast the initial checkpoint data to the second training node.

[0037] In an implementation of the second aspect, the apparatus further includes a data uploading module, wherein the data uploading module is configured to:

[0038] Instruct the first model training group to upload the checkpoint data stored in the memory of each training node in the first model training group to a remote storage service.

[0039] In an implementation of the second aspect, the storage control module is further configured to:

[0040] During the process of uploading checkpoint data by the first model training group, it is prohibited to delete the checkpoint data saved in the memory of each training node in the first model training group, and it is prohibited to use new checkpoint data to overwrite the saved checkpoint data.

[0041] In an implementation of the second aspect, the storage control module is further configured to:

[0042] During the process of uploading the checkpoint data by the first model training group, stop saving the new checkpoint data to the memory of the training node of the first model training group.

[0043] In an implementation of the second aspect, the storage control module is further configured to:

[0044] After the first model training group completes uploading the checkpoint data, the new checkpoint data stored in the memory of the training node of the second model training group is synchronized to the memory of the training node of the first model training group.

[0045] In a third aspect, the present application provides a computing device cluster, comprising at least one computing device, each computing device comprising a processor and a memory;

[0046] The processor of the at least one computing device is configured to execute instructions stored in the memory of the at least one computing device, so that the computing device cluster executes the method according to the first aspect.

[0047] In a fourth aspect, the present application provides a computer program product comprising instructions, which, when executed by a computing device cluster, causes the computing device cluster to execute the method described in the first aspect.

[0048] In a fifth aspect, the present application provides a computer-readable storage medium comprising computer program instructions. When the computer program instructions are executed by a computing device cluster, the computing device cluster executes the method described in the first aspect. BRIEF DESCRIPTION OF THE DRAWINGS

[0049] Figure 1 Schematic diagram of a model training cluster reading and writing CKPT according to an embodiment of the present application is shown;

[0050] Figure 2Schematic diagram of direct remote storage reading and writing data according to an embodiment of the present application;

[0051] Figure 3 The figure shows the distributed cache reading and writing data process according to an embodiment of the present application;

[0052] Figure 4 The figure shows the process of reading and writing CKPT data using memory according to an embodiment of the present application;

[0053] Figure 5 The figure shows a flow chart of a model training method according to an embodiment of the present application;

[0054] Figure 6 FIG2 is a schematic diagram of a training cluster structure according to an embodiment of the present application;

[0055] Figure 7 The figure shows a flow chart of a model training method according to an embodiment of the present application;

[0056] Figure 8 Shown is a schematic diagram of a model training device according to an embodiment of the present application;

[0057] Figure 9 FIG2 is a schematic diagram of a training node deployment scenario according to an embodiment of the present application;

[0058] Figure 10 FIG2 is a schematic diagram of a training node deployment scenario according to an embodiment of the present application;

[0059] Figure 11 FIG2 is a schematic diagram of a computing device according to an embodiment of the present application;

[0060] Figure 12 Shown is a schematic diagram of a computing device cluster according to an embodiment of the present application;

[0061] Figure 13 FIG2 is a schematic diagram of a network connection of a computing device cluster according to an embodiment of the present application. DETAILED DESCRIPTION

[0062] To make the purpose, technical solutions, and advantages of this application more clear, the technical solutions of this application will be clearly and completely described below in conjunction with the specific embodiments of this application and the corresponding drawings. Obviously, the embodiments described are only part of the embodiments of this application, not all of them. Based on the embodiments in this application, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of this application.

[0063] The terms used in the implementation section of this application are only used to explain the specific embodiments of this application and are not intended to limit this application.

[0064] To solve the problem of how to avoid GPU card failure causing interruption of model training and waste of GPU computing power, a feasible technical solution is to use the method of regularly storing checkpoints (CKPT) to avoid loss of training data during training.

[0065] For example, a persistent storage service is provided on a cloud server, storing model data in the cloud server every three hours. If a GPU card fails in the fourth hour, the training cluster can load the model data stored on the cloud server last time, eliminating the need to start training from scratch.

[0066] However, storing and restoring CKPT in large-scale distributed training clusters suffers from low performance and high cost. Specifically, periodic checkpoint storage places high demands on the network bandwidth and performance of remote storage. In large-scale distributed training clusters, this can become a bottleneck for the entire system, leading to lengthy recovery times and significant waste of GPU resources.

[0067] A distributed training cluster will split a large model and place it on multiple nodes for parallel training. All the nodes that make up a complete model are called a model group.

[0068] Taking the training of a 70 billion parameter model with 512 nodes as an example, in order to improve the parallel training efficiency of the model, the 512 nodes can be divided into 32 model groups to train 32 models in parallel, and 16 nodes form a model group to jointly train a single model.

[0069] For example, Figure 1 The figure shows a schematic diagram of reading and writing CKPT in a model training cluster according to an embodiment of the present application.

[0070] Any two of the 32 model groups such as Figure 1 The model group 101 and the model group 102 are shown.

[0071] A 70 billion parameter model has a CKPT size of 1 TB. Because such a large model cannot fit on a single GPU, it is split into 128 parts and trained on 128 GPUs. Model group 101 contains nodes 0 to 15, and model group 102 contains nodes 16 to 31. Each node contains eight GPU cards (GPUs 0 to 7).

[0072] At the beginning of each round of training, the 32 model groups load the same CKPT to ensure the same initialization state. That is, at the beginning of each round of training, the parameters between the model groups are the same.

[0073] During the training process, 32 models are grouped on 8 GPU cards in pipeline parallelism, optimizer parallelism, and data parallelism.

[0074] At the end of a round of training, the training results of each model group are summarized to obtain the average value as the CKPT result of this round of training. The CKPT results are synchronized to all model groups before the next round of training.

[0075] After each round of training, the CKPT results of this round of training are uploaded to the remote Object Storage Service (OBS).

[0076] At the beginning of each training round, all 32 model groups need to pull data from OBS. Therefore, OBS needs to provide 32 TB of CKPT data download traffic to guarantee data service for all 32 model groups. If the download is required to be completed within 30 minutes, OBS needs to provide at least 32 TB / 30 minutes = 145.64 Gbps of bandwidth. In some application scenarios, OBS also needs to provide 16 TB of data set download traffic. If the download is required to be completed within 30 minutes, OBS needs to provide at least 48 TB / 30 minutes = 218.453 Gbps of bandwidth. This places a very high demand on the network bandwidth of remote storage. As download times shorten and the number of model parameters increases, the bandwidth demand for remote storage will increase, significantly increasing the cost of remote storage.

[0077] Furthermore, while all nodes are storing and downloading CKPT, GPU training cards are idle, resulting in a significant waste of computing power. Furthermore, GPU cards are billed by time, and tens of thousands of cards can cost tens of thousands of yuan per minute. It's worth noting that all computational results are lost from the time the CKPT is saved until the cluster failure occurs. This waste of GPU computing power during this period is also a source of cost. Therefore, the frequency of model saves significantly impacts the performance and cost of the training cluster.

[0078] To improve the performance of storing and reading CKPT and reduce its cost, a feasible solution is to use direct remote storage.

[0079] Figure 2 FIG2 is a schematic diagram of directly reading and writing data from remote storage according to an embodiment of the present application.

[0080] The most direct way to improve the storage and reading of CKPT is to improve the performance of the remote storage service. Therefore, the performance of the remote storage service can be improved by increasing the network bandwidth of the remote storage service, placing the remote storage service on a network closer to the training cluster, or developing an efficient network file system.

[0081] like Figure 2As shown in the figure, the remote storage service is placed on a network closer to the nodes of the training cluster. This approach allows training nodes to use higher bandwidth and shorter network paths to store and read CKPT. Representative technical solutions include Amazon S3, Huawei Cloud OBS, and Huawei Cloud Scalable File Service (SFS). However, this approach does not have significant advantages in terms of performance and cost, as it still requires a large amount of traffic to traverse the external network, which has a much longer network path than within the cluster. Furthermore, the higher the network bandwidth, the more expensive it is.

[0082] To improve the performance of storing and reading CKPT and reduce its cost, another feasible solution is to use distributed cache.

[0083] Figure 3 The figure shows the distributed cache reading and writing data process according to an embodiment of the present application.

[0084] Distributed caching places the path for storing and reading CKPT in the cluster where the model is trained.

[0085] like Figure 3 As shown, each node's local solid-state drive (SSD) disks are connected to form a cache pool to expand storage capacity. CKPTs are stored in the cache pool, and the complete CKPTs are evenly distributed within the cache pool to ensure network load balancing within the cluster. In this way, cluster nodes prioritize reading from the cluster's cache pool when downloading CKPT data, thereby leveraging the inter-node network to reduce network dependence on remote storage services. Representative solutions include JuiceFs, Alluxio, and MemArts. However, the performance of these solutions relies on advanced prefetching algorithms to improve the cache pool's hit rate and advanced local affinity scheduling algorithms to reduce network latency within the cluster. Their performance still has significant room for improvement.

[0086] To improve the performance of storing and reading CKPT and reduce its cost, another feasible solution is to use memory storage.

[0087] Figure 4 The figure shows the process of reading and writing CKPT data using memory according to an embodiment of the present application.

[0088] like Figure 4As shown, using memory storage and access to store CKPT directly in memory improves CKPT access performance through fast memory storage and access, reduces cluster nodes' reliance on the network, and thus reduces the impact of network latency and bandwidth, significantly improving the performance of training cluster storage and access to CKPT. However, the disadvantage of this solution is that memory data becomes unavailable in the event of a system crash. Only by ensuring that data stored in memory remains accessible in the event of a system crash can the cluster avoid accessing remote storage networks.

[0089] In one in-memory storage solution (e.g., Gemini), cluster nodes are grouped for backup. Nodes in the same group back up each other's data, increasing the probability that cluster nodes will remain available in the event of a system crash, thereby preventing in-memory data from becoming unavailable. However, this in-memory storage solution requires at least half of the memory to back up the data of other nodes, resulting in high memory consumption. Furthermore, to back up data from one node to other nodes, the parameter plane network is scheduled during training to transmit the backed-up CKPT data, potentially interfering with normal model training traffic and affecting training efficiency.

[0090] The above solution cannot achieve both high performance and low cost. High performance refers to minimizing the impact of GPU card failures during training, quickly recovering to the training state, and improving training efficiency. Low cost refers to requiring less storage space, lowering network bandwidth requirements, and minimizing waste of GPU card computing power. A detailed analysis is shown in Table 1.

[0091] Table 1

[0092]

[0093] Direct remote storage relies too much on remote storage services, resulting in high network latency and high bandwidth requirements in large model training clusters. In the event of a system failure, it takes a long time to read remote data, resulting in poor recovery performance. Increasing the storage frequency of CKPT consumes significant storage space. Furthermore, during a failure, a large number of cards remain idle, awaiting recovery, resulting in significant waste of computing power.

[0094] Distributed caches spread data across different nodes in the cluster. Therefore, their performance during fault recovery relies on sophisticated prefetching algorithms to ensure cache hits and local affinity scheduling algorithms to ensure data hits on the local node. Increasing the CKPT storage frequency consumes significant storage space while reducing bandwidth requirements on remote networks. However, since increasing the CKPT storage frequency is difficult, it cannot be restored to a more recent training state, resulting in significant waste of card computing power.

[0095] Gemini uses an in-memory storage solution, but it relies on the memory of each node to back up other nodes, resulting in high memory consumption and traffic interference. During fault recovery, it is necessary to ensure that the data of the memory node is still available, and Gemini relies on the data of the backup node to be available. Its performance is high but unstable. It stores temporary data in memory, which can significantly increase the storage frequency of CKPT. This solution accesses data from memory, reducing the demand for remote network bandwidth. When a fault occurs, this solution can recover to a more recent training state, with less waste of card computing power, but it relies on the recovered data being available in the node memory, so its performance is unstable.

[0096] To address the low performance and high cost issues of checkpoint storage in large model training clusters, an embodiment of the present application proposes a model training method, which includes a training cluster checkpoint storage process.

[0097] In a large model training cluster, the memory of each node is often more than three times the memory of all the graphics cards in the node, which is sufficient to store multiple copies of CKPT data. Therefore, in the model training method of one embodiment of the present application, the CKPT data for each round of training is directly stored in the node's memory.

[0098] In addition, distributed training clusters often use multiple model groups for parallel training, and the multiple model groups aggregate and synchronize data at the end of each round of training. During each round of data synchronization, the data of the multiple model groups is the same. Therefore, in the model training method of one embodiment of the present application, the data redundancy characteristics of multiple model groups are used to provide backup redundancy for memory data.

[0099] In order to achieve mutual backup of CKPT data among multiple model groups, in the model training method of one embodiment of the present application, the backup data among multiple model groups is ensured to be in the same state by tracking the CKPT version of each node to ensure data consistency.

[0100] Specifically, the network in which GPU cards between nodes communicate directly and are used for training parameter interaction is called a parameter plane network. In the model training method of one embodiment of the present application, distributed cluster memory collaborative management is adopted, and GPU cards between nodes are connected to the parameter plane network. The storage and recovery of the model CKPT are perceived and controlled through the cluster memory collaborative management control, and the cluster nodes are coordinated through the cluster memory collaborative management control to store the CKPT of each node directly in the local memory to realize the storage of the CKPT; by tracking and coordinating the CKPT versions of all nodes, the CKPT required for fault recovery is transmitted using the local memory and the parameter plane network to realize the recovery of the cluster training state. The storage status of the CKPT during the training process is perceived through the cluster memory collaborative management control, including information such as generation time, storage location, size, version number, etc., and the CKPT is directly stored in the local memory or storage device to realize the fast storage of the CKPT.

[0101] Figure 5 The figure shows a flow chart of a model training method according to an embodiment of the present application.

[0102] S500: Start the model training process.

[0103] Specifically, in S500 , a training cluster is used to perform model training.

[0104] The model training method of one embodiment of the present application is applied to a training cluster for performing model training.

[0105] A training cluster contains multiple model training groups, which are used to train neural network models in parallel using data parallelism. Each model training group targets a complete model structure and contains multiple training nodes. Different model training groups have the same node structure, meaning that different model training groups contain the same number of training nodes, and the training nodes in different model training groups correspond one to one.

[0106] For example, a training cluster includes at least two model training groups (a first model training group and a second model training group), and each of the first model training group and the second model training group corresponds to a complete model structure. The first model training group includes at least a first training node, and the second model training group includes at least a second training node. Based on the node structure, the first training node corresponds to the second training node.

[0107] It is understandable that the first model training group may further include a third training node, and the second model training group may further include a fourth training node. Based on the node structure, the third training node corresponds to the fourth training node.

[0108] S501: After a round of training is completed, a model parameter training result of the current round of training is generated, that is, a CKPT of the current round of training is generated.

[0109] The CKPT of the current training round corresponds to a complete model. That is, the CKPT of the current training round contains the CKPTs corresponding to all training nodes in a model training group.

[0110] Specifically, the CKPT of the current round of training includes at least CKPT A, which is the CKPT of the first training node and the second training node. It is understandable that the CKPT of the current round of training may also include CKPT B, which is the CKPT of the third training node and the fourth training node.

[0111] Specifically, in S501, trained parameters are obtained by exchanging data with different training nodes. At the end of a training round, the model parameters of all nodes are aggregated to the central node for average calculation. The calculation result is used as the model parameter training result for that round of training, that is, the CKPT of the current training round. The central node can be any training node in the training cluster or a combination of any number of training nodes. The central node can also be any computing node other than all training nodes in the training intensive.

[0112] S510: Synchronize the CKPT of the current training round to all model training groups (all training nodes).

[0113] Specifically, in S510, all model training groups load the same CKPT (the CKPT for the current training round), and the training nodes in the same model training group load their corresponding CKPTs. That is, based on the node structure, the portion of the CKPT corresponding to a training node in the CKPT for the current training round is loaded onto the corresponding training node in the model training group.

[0114] Specifically, CKPT A is synchronized to the first training node and the second training node, and CKPT B is synchronized to the third training node and the fourth training node.

[0115] S520: Save the CKPT of each training node to the memory of the training node.

[0116] Specifically, in S520, the first training node saves CKPT A to the memory of the first training node, and the second training node saves CKPT A to the memory of the second training node. It is understandable that the third training node saves CKPT B to the memory of the third training node, and the fourth training node saves CKPT B to the memory of the second training node.

[0117] Figure 6 FIG. 1 is a schematic diagram of a training cluster structure according to an embodiment of the present application.

[0118] like Figure 6 As shown, in one embodiment, the training cluster includes at least a model training group 61 and a model training group 62 .

[0119] The model training group 61 includes training nodes 600 - 615 , and the model training group 62 includes training nodes 616 - 631 .

[0120] In S500, the training nodes in model training group 61 and model training group 62 exchange data for model training. After a round of training is completed, the trained model parameters of model training group 61 (training nodes 600-615) and model training group 62 (training nodes 616-631) are aggregated and averaged to obtain the model parameter training results for the current training round, i.e., the CKPT for the current training round. The CKPT for the current training round includes CKPT00-15.

[0121] The CKPT for the current training round is synchronized to model training group 61 and model training group 62. Training nodes 600-615 of model training group 61 load their corresponding CKPTs from the CKPT for the current training round, while training nodes 616-631 of model training group 62 load their corresponding CKPTs from the CKPT for the current training round. That is, CKPTs 00-15 are synchronized to training nodes 600-615, and CKPTs 00-15 are synchronized to training nodes 616-631, respectively.

[0122] After CKPT synchronization, the CKPTs loaded by training nodes 600 to 615 are consistent with the CKPTs loaded by training nodes 616 to 631. For example, the CKPT loaded by training node 600 is consistent with the CKPT loaded by training node 616.

[0123] Further, CKPT00-15 are saved in the memory of training nodes 600-615 respectively, and CKPT00-15 are saved in the memory of training nodes 616-631 respectively.

[0124] According to the model training method of the embodiment of the present application, the CKPT of the training node is backed up using the training node's own memory. During the backup process, the training nodes do not need to communicate with each other, which greatly reduces the consumption of communication resources.

[0125] Optionally, in one embodiment, the CKPT of the latest round of training is saved in the memory of the training node.

[0126] Optionally, in another embodiment, a CKPT of multiple rounds of training including the latest round of training is stored in the memory of the training node. Specifically, a key value (KV) is maintained on the training node, where the Key is the timestamp-based version number of the CKPT and the Value is metadata related to the CKPT, including the size and storage location of the CKPT.

[0127] Optionally, in one embodiment, in order to avoid memory overflow, the size of the CKPT in the memory of the training node is monitored, and when the size of the CKPT in the memory of the training node exceeds a preset value, the oldest version of the CKPT in the memory is deleted.

[0128] Optionally, in one embodiment, to avoid memory overflow, the number of CKPT versions stored in the training node's memory is limited. When the number of CKPT versions in the training node's memory reaches a preset value, the oldest CKPT version in the memory is deleted.

[0129] Considering that the memory capacity of the training node is always limited, in order to expand the memory capacity, optionally, in one embodiment, part or all of the CKPT stored in the memory of the training node is saved to other storage media outside the memory (for example, SSD).

[0130] When there are a large number of model training groups, greater model parallelism can be achieved. However, because the training nodes of different model training groups store the same CKPT data in memory, this can lead to significant backup redundancy of the CKPT data. To reduce backup redundancy, in one embodiment, the model training groups can be grouped, with multiple model training groups within a group backing up a single copy of the CKPT data. For example, four model groups can back up the same copy of the CKPT data, allowing the cluster memory to store more model versions.

[0131] According to the model training method of the embodiment of the present application, the training node's memory is used to store CKPT data. The CKPT data in the memory will soon be replaced by a new version and is temporarily stored data. However, in actual use, it is necessary to persist a certain version of CKPT at regular intervals for further analysis and use in the future.

[0132] In response to the above situation, in one embodiment, the CKPT data in the memory of all training nodes of one model training group among multiple model training groups is uploaded to a remote storage service according to the requirement of CKPT persistence.

[0133] According to the method of one embodiment of the present application, a solution of moving remote storage out of the critical path is adopted, and cluster memory data is asynchronously persisted to the remote end, reducing dependence on remote performance; and, according to the method of one embodiment of the present application, a distinction is made between the temporary storage of CKPT by memory and the persistent storage of CKPT by the remote storage service, without increasing the space consumption of remote storage.

[0134] In one embodiment, to prevent memory overflow, new checkpoint data overwrites older checkpoint data when it is saved to memory. Alternatively, memory usage is monitored and older checkpoint data is deleted when insufficient space is available. However, during the data upload phase, it is possible that a certain version of checkpoint data may be deleted or overwritten by a newer version midway through uploading. To prevent this, in one embodiment, the model training group's memory data is prevented from being deleted or replaced by a newer version while the CKPT data is being uploaded to the remote storage service.

[0135] According to the method of an embodiment of the present application, it is possible to prevent the generation of new checkpoint data from interfering with the uploading of checkpoint data, thereby ensuring that the checkpoint data can be successfully uploaded to the remote storage service.

[0136] Specifically, in one embodiment, for a model training group that is uploading CKPT data to a remote storage service, deleting the stored CKPT data in memory is prohibited, and overwriting the stored CKPT data with new CKPT data is prohibited. That is, the stored CKPT data in memory is frozen during the entire data upload process. However, during the CKPT data upload process, new CKPT data can be saved to unused space in memory.

[0137] Furthermore, in order to avoid overflow of memory storage space, in another embodiment, S520 is not executed for the model training group that is uploading CKPT data to the remote storage service. That is, during the entire data uploading process, new CKPT data is not saved in the memory, so that the CKPT data saved in the memory will not be overwritten by the new CKPT data, and the saved CKPT data will not be deleted due to insufficient storage space. During the entire data uploading process, the CKPT data saved in the memory is guaranteed to remain unchanged. After the CKPT data is uploaded to the remote storage service, it is synchronized from other model training groups (for example, adjacent model training groups), and during the CKPT data uploading period, the latest version of the CKPT data that is not saved in the memory.

[0138] According to the method of one embodiment of the present application, the generation of new checkpoint data can be prevented from interfering with the uploading of checkpoint data, ensuring that the checkpoint data can be successfully uploaded to the remote storage service; and it can be ensured that the checkpoint data version of the training node that uploads the checkpoint data is consistent with that of the training node that does not upload the checkpoint data.

[0139] Based on the training cluster checkpoint storage process proposed in one embodiment of the present application, the model training method in one embodiment of the present application also includes a training cluster checkpoint recovery process.

[0140] In the model training method of one embodiment of the present application, the CKPT and parameter plane network stored in the memory of the training node are used for rapid data recovery.

[0141] Specifically, when a training node fails, the failed training node is repaired and the status of all training nodes is restored to the most recent CKPT to ensure that all nodes are in a consistent state.

[0142] For training nodes that have not experienced any failures, the most recent CKPT data is read from the local memory first. This method is highly efficient in restoring CKPT due to the high speed of reading memory.

[0143] For the training node whose fault has been repaired, there is no previously stored CKPT in its memory. Therefore, the CKPT state is restored by reading the CKPT data in the memory of the corresponding training node in other model training groups through the parameter surface network.

[0144] Figure 7 The figure shows a flow chart of a model training method according to an embodiment of the present application.

[0145] S700: Start the model training process. Refer to S500.

[0146] S710: Determine whether there is a training node failure.

[0147] If there is no training node failure, the process returns to S700.

[0148] If a training node failure occurs, execute S711.

[0149] S711, interrupt the training operation of all training nodes.

[0150] S720: Repair the faulty training node.

[0151] S730, for training nodes other than the failed training node, use the latest version of CKPT data stored in the training node's own memory to recover the training node data.

[0152] S731, for the failed training node, use the latest version of CKPT data stored in the training node memory corresponding to the failed training node based on the node structure in other model groups to recover the training node data.

[0153] For example, if the first training node in the first model training group fails, the second training node in the second model training group will use the latest version of CKPT data stored in the second training node's memory to recover the training node data. For the first training node in the first model training group, the first training node retrieves the checkpoint data from the memory of the second training node to recover the first training node's data.

[0154] like Figure 6 As shown, in one embodiment, if training node 615 fails, the training of training nodes 600-615 and training nodes 616-631 is interrupted. The failure of training node 615 is repaired by restoring training nodes 600-614 using the latest versions of CKPT data stored in the memory of training nodes 600-614, restoring training nodes 616-631 using the latest versions of CKPT data stored in the memory of training nodes 616-631, and restoring training node 615 using the latest version of CKPT data stored in the memory of training node 631.

[0155] According to the method of the embodiment of the present application, in a large model training cluster, the problems of low performance and high cost in storing and reading CKPT can be solved, the time for storing and reading CKPT can be reduced, the storage frequency of CKPT can be increased without increasing the consumption of storage space, the bandwidth demand for remote storage services can be reduced, and the waste of computing power of GPU cards can be reduced.

[0156] According to the method of one embodiment of the present application, in the event of a failure, the multi-model group redundancy feature is utilized to retransmit and recover data. In this way, the remote storage service is removed from the critical path, reducing the performance and bandwidth requirements of the remote storage and significantly reducing the overall system cost.

[0157] Furthermore, before model training, when the training nodes of the training cluster load CKPT in the initial state, the local memory of all training nodes does not save CKPT data. If all training nodes are required to pull CKPT data from the remote storage service, the instantaneous bandwidth demand of the remote storage will be very large, and the performance will be seriously degraded.

[0158] In response to the above situation, in one embodiment, when a training node of a training cluster loads CKPT in an initial state, one of the multiple model training groups is instructed to pull the CKPT from a remote storage service. The training node of the model training group that downloads the CKPT from the remote storage service sends (e.g., by broadcasting) the CKPT to the training nodes of other model training groups via the parameter plane network.

[0159] For example, the first training node in the first model training group is used to download the initial checkpoint data corresponding to the training node; the first training node in the first model training group broadcasts the initial checkpoint data downloaded by the first training node to the second training node in the second model training group.

[0160] Optionally, during the broadcast process, the model group that has received the broadcast data may also be allowed to broadcast it to more model groups.

[0161] For example, the training nodes in the first model training group are used to download the initial checkpoint data; the training nodes in the first model training group broadcast the downloaded initial checkpoint data to the training nodes in the second model training group; and the training nodes in the second model training group broadcast the downloaded initial checkpoint data to the training nodes in the third model training group.

[0162] According to the method of one embodiment of the present application, the initial checkpoint data is downloaded to a model training group, and the initial checkpoint data is sent to other model training groups through one model training group, thereby reducing the bandwidth requirement for the initial checkpoint data and improving the download efficiency of the initial checkpoint data.

[0163] Furthermore, according to the model training method proposed in the embodiment of the present application, an embodiment of the present application also proposes a model training device, which is applied to the training cluster. The method steps performed by the model training device can refer to the method steps in the previous method embodiment.

[0164] Figure 8 Shown is a schematic diagram of a model training device according to an embodiment of the present application.

[0165] like Figure 8 As shown, in one embodiment, the model training apparatus 800 is applied to a training cluster 810 .

[0166] The training cluster 810 includes at least a first model training group 811 and a second model training group 812 . The first model training group 811 and the second model training group 812 are respectively for complete model structures. The first model training group 811 and the second model training group 812 include the same node structure.

[0167] like Figure 8As shown, the first model training group 811 includes at least a first training node 813 , and the second model training group 812 includes at least a second training node 814 , wherein, based on the node structure, the first training node 813 corresponds to the second training node 814 .

[0168] The first model training group 811 further includes a third training node 815 , and the second model training group 812 further includes a fourth training node 816 , wherein, based on the node structure, the third training node 815 corresponds to the fourth training node 816 .

[0169] After a round of training, the training cluster 810 generates the model parameter training results for the current round of training, that is, generates the CKPT for the current round of training. Specifically, at the end of a round of training, the training cluster 810 calculates the average value of the model parameters of all training nodes, and uses the calculated result as the checkpoint data for that round of training.

[0170] Specifically, after a round of training, the training cluster 810 generates at least CKPT A, which is the CKPT of the first training node 813 and the second training node 814. It is understandable that after a round of training, the training cluster 810 also generates CKPTB, which is the CKPT of the third training node 815 and the fourth training node 816.

[0171] Training cluster 810 synchronizes the CKPT of the current training round to each model training group within training cluster 810. Specifically, based on the node structure, training cluster 810 synchronizes the CKPT of a single training node within the CKPT of the current training round to the corresponding training node in each model training group. That is, CKPT A is synchronized to the first training node 813 and the second training node 814, and CKPT B is synchronized to the third training node 815 and the fourth training node 816.

[0172] In one embodiment, the model training device 800 includes a storage control module 801, which is used to instruct each training node in the training cluster 810 to save the CKPT obtained by itself into its own memory.

[0173] That is, the storage control module 801 instructs the first training node 813 to save the CKPT A obtained by the first training node 813 into the memory of the first training node 813 , and the storage control module 801 instructs the second training node 814 to save the CKPT A obtained by the second training node 814 into the memory of the second training node 814 ;

[0174] The storage control module 801 also instructs the third training node 815 to save the CKPT B obtained by the third training node 815 into the memory of the third training node 815, and the storage control module 801 also instructs the fourth training node 816 to save the CKPT B obtained by the fourth training node 816 into the memory of the fourth training node 816.

[0175] During the model training process, if a training node in the training cluster 810 fails, the training cluster 810 interrupts the training operations of all training nodes and repairs the failed training node.

[0176] In one embodiment, the model training device 800 also includes a recovery control module 802, which is used to instruct the failed training node to use the latest version of CKPT data stored in the training node memory corresponding to the failed training node based on the node structure in other model groups to recover the training node data when a training node in the training cluster 810 fails; and to instruct training nodes other than the failed training node to use the latest version of CKPT data stored in the training node's own memory to recover the training node data.

[0177] For example, if the first training node 813 fails during model training, the recovery control module 802 instructs the first training node 813 to retrieve checkpoint data from the memory of the second training node 814 to recover data on the first training node 813. Furthermore, the recovery control module 802 instructs the second training node 814 to retrieve checkpoint data from the memory of the second training node 814 to recover data on the second training node 814; instructs the third training node 815 to retrieve checkpoint data from the memory of the third training node 815 to recover data on the third training node 815; and instructs the fourth training node 816 to retrieve checkpoint data from the memory of the fourth training node 816 to recover data on the fourth training node 816.

[0178] Optionally, in one embodiment, the model training device 800 further includes an initialization module 803. The initialization module 803 is used to, before the training cluster 810 starts training the model, instruct a model training group in the training cluster 810 to download the initial checkpoint data corresponding to the complete model structure, that is, each training node in the model training group downloads the initial checkpoint data corresponding to the training node; and instruct the model training group that has completed downloading the initial checkpoint data to send the downloaded initial checkpoint data to other model training groups, that is, the training node that has completed downloading the initial checkpoint data sends the downloaded initial checkpoint data to the corresponding training nodes in the other model training groups.

[0179] Specifically, initialization module 803 instructs first training node 813 to download initial checkpoint data corresponding to first training node 813; and instructs first training node 813 to broadcast the downloaded initial checkpoint data to second training node 814. It is understandable that initialization module 803 also instructs third training node 815 to download initial checkpoint data corresponding to third training node 815; and instructs third training node 815 to broadcast the downloaded initial checkpoint data to fourth training node 816.

[0180] In one embodiment, the model training device 800 also includes a data upload module 804, which is used to instruct a model training group in the training cluster 810 to upload the checkpoint data stored in the memory of each training node in the model training group to the remote storage service; in the process of a model training group uploading the checkpoint data stored in the memory of each training node in the model training group to the remote storage service, other model training groups do not upload the checkpoint data to the remote storage service.

[0181] For example, the data upload module 804 instructs the first model training group 811 to upload the checkpoint data stored in the memory of each training node in the first model training group 811 to the remote storage service. At the same time, the second model training group 812 does not upload the checkpoint data to the remote storage service.

[0182] Optionally, in one embodiment, the storage control module 801 is also used to prohibit the deletion of checkpoint data saved in the memory of each training node in a certain model training group during the process of uploading checkpoint data to a remote storage service, and to prohibit the use of new checkpoint data to overwrite the checkpoint data saved in the memory of each training node in the model training group.

[0183] For example, during the process of uploading checkpoint data in the first model training group 811, the storage control module 801 prohibits deleting the checkpoint data saved in the memory of each training node in the first model training group 811, and prohibits using new checkpoint data to overwrite the checkpoint data saved in the memory of each training node in the first model training group 811.

[0184] Optionally, in another embodiment, the storage control module 801 is also used to stop saving new checkpoint data into the memory of the training node of a model training group during the process of a model training group uploading checkpoint data to a remote storage service.

[0185] For example, during the process of the first model training group 811 uploading checkpoint data, the storage control module 801 stops saving new checkpoint data into the memory of the training node of the first model training group 811.

[0186] Furthermore, in one embodiment, the storage control module 801 is also used to, after a model training group completes uploading the checkpoint data, synchronize the new checkpoint data stored in the memory of the training node of another model training group (the new checkpoint data is the checkpoint data that failed to be saved in the memory of the training node of the model training group during the process of the model training group uploading the checkpoint data) to the memory of the training node of the model training group that uploaded the checkpoint data.

[0187] For example, after the first model training group 811 completes uploading the checkpoint data, the new checkpoint data stored in the memory of the training node of the second model training group 812 is synchronized to the memory of the training node of the first model training group 811.

[0188] In the description of the embodiments of the present application, for the convenience of description, the functions are divided into various modules and described separately. The division of each module is merely a division of logical functions. When implementing the embodiments of the present application, the functions of each module can be implemented in the same or multiple software and / or hardware.

[0189] Specifically, the device proposed in the embodiment of the present application can be fully or partially integrated into a physical entity (for example, a GPU or other type of processor) during actual implementation, or it can be physically separated. And these modules can all be implemented in the form of software calling through a processing element; they can also all be implemented in the form of hardware; some modules can also be implemented in the form of software calling through a processing element, and some modules can be implemented in the form of hardware. For example, the detection module can be a separately established processing element, or it can be integrated in a chip of an electronic device. The implementation of other modules is similar. In addition, these modules can be fully or partially integrated together, or they can be implemented independently. During the implementation process, each step of the above method or each of the above modules can be completed by an integrated logic circuit of hardware in the processor element or an instruction in the form of software.

[0190] For example, in the above embodiment, the storage control module 801, the recovery control module 802, the initialization module 803, and the data upload module 804 can all be implemented via software or hardware. By way of example, the implementation of the storage control module 801 will be described below using the storage control module 801 as an example. Similarly, the implementation of the recovery control module 802, the initialization module 803, and the data upload module 804 can refer to the implementation of the storage control module 801.

[0191] As an example of a software functional unit, the storage control module 801 may include code running on a computing instance. The computing instance may include at least one of a physical host (computing device), a virtual machine, and a container. Furthermore, the computing instance may be one or more. For example, the storage control module 801 may include code running on multiple hosts / virtual machines / containers. It should be noted that the multiple hosts / virtual machines / containers used to run the code may be distributed in the same region or in different regions. Furthermore, the multiple hosts / virtual machines / containers used to run the code may be distributed in the same availability zone (AZ) or in different AZs, each AZ including one data center or multiple geographically close data centers. Typically, a region may include multiple AZs.

[0192] Similarly, multiple hosts / virtual machines / containers running the code can be distributed within the same virtual private cloud (VPC) or across multiple VPCs. Typically, a VPC is set up within a region. Cross-region communication between two VPCs within the same region, or between VPCs in different regions, requires a communication gateway within each VPC to interconnect the VPCs.

[0193] As an example of a hardware functional unit, the storage control module 801 may include at least one computing device, such as a server. Alternatively, the storage control module 801 may be implemented using an application-specific integrated circuit (ASIC) or a programmable logic device (PLD). The PLD may be a complex programmable logical device (CPLD), a field-programmable gate array (FPGA), a generic array logic (GAL), or any combination thereof.

[0194] The multiple computing devices included in the storage control module 801 can be distributed in the same region or in different regions. The multiple computing devices included in the storage control module 801 can be distributed in the same AZ or in different AZs. Similarly, the multiple computing devices included in the storage control module 801 can be distributed in the same VPC or in multiple VPCs. The multiple computing devices can be any combination of servers, ASICs, PLDs, CPLDs, FPGAs, GALs, and other computing devices.

[0195] It should be noted that, in other embodiments, the storage control module 801 can be used to execute any step in the XXX method, the recovery control module 802 can be used to execute any step in the XXX method, and the initialization module 803 can be used to execute any step in the XXX method. The steps that the storage control module 801, the recovery control module 802, and the initialization module 803 are responsible for implementing can be specified as needed. The full functions of the YY device can be realized by respectively implementing different steps in the XXX method through the storage control module 801, the recovery control module 802, and the initialization module 803.

[0196] Specifically, in one embodiment, one or more of the storage control module 801, the recovery control module 802, the initialization module 803 and the data upload module 804 can be implemented based on the model training framework.

[0197] In the embodiments of this application, the code library used by developers to develop deep learning models is called a model training framework, which can complete the entire process of model definition, training, and inference. The model training framework provides memory management for training nodes to enable backing up CKPT to memory and restoring training nodes from CKPT in memory.

[0198] Figure 9 Shown is a schematic diagram of a training node deployment scenario according to an embodiment of the present application.

[0199] like Figure 9 As shown in the figure, each training node is equipped with multiple graphics cards for model training and large-capacity memory. In addition to the traditional Transmission Control Protocol (TCP) network between training nodes, the graphics cards of the training nodes also include a high-bandwidth parameter surface network for model parameter transmission.

[0200] A training framework for model training and a software package of a method according to an embodiment of the present application are installed in the operating system of each training node, wherein a software development kit (SDK) that can be called by the training framework is provided in the software package of the method according to an embodiment of the present application. The SDK is used as an interface for interaction between the training framework and the software package of the method according to an embodiment of the present application, and is used to call the function of storing and restoring CKPT according to the method according to an embodiment of the present application.

[0201] The remote storage service outside the training cluster provides data persistence services. The remote storage service provides the Portable Operating System Interface of UNIX (POSIX).

[0202] According to the method of one embodiment of the present application, cluster memory collaborative management is adopted, and the node memory stores the entire CKPT, thereby improving the frequency of CKPT storage; and the cluster memory control manages multi-node CKPT and supports group redundant backup; and the CKPT is directly stored in the memory, which greatly reduces the latency of storing CKPT.

[0203] According to the method of one embodiment of the present application, local memory and parameter plane network are used for collaborative recovery, and data in local memory is directly read, which greatly shortens the recovery delay of CKPT; and the faulty node reads data through the parameter plane network, which greatly shortens the recovery delay.

[0204] The methods and devices provided in the embodiments of the present application involve products including services that provide CKPT storage and download services for large-model distributed training clusters, which may be sold in the form of services that accelerate CKPT storage and access efficiency, and are bound to large-model training clusters for sale.

[0205] Specifically, in the products involved in the methods and devices provided in the embodiments of the present application:

[0206] The training node memory stores the complete CKPT data of the node: In computer system architecture, there are many designs and products that effectively use memory to store or cache data. According to the method of one embodiment of the present application, a major feature of using memory is to use the node's memory to store the complete CKPT data of the node, rather than caching partial data. The advantage of this is that when the system needs to restore CKPT data, as long as the corresponding version of the data is saved in the node, the corresponding data can be read completely from the memory of the node without having to read it elsewhere;

[0207] Collaborative fast recovery between the parameter plane network and local memory: Unlike solutions that read CKPT from remote storage, the method according to one embodiment of the present application prioritizes reading data from local memory during fault recovery. Because each node stores exactly the complete CKPT data for that node, this solution can achieve fault recovery in seconds. If the local node happens to be the node being replaced after a failure, data is read from the redundant backup node in the cluster via the parameter plane network. For an AI training cluster with 32 model groups, each node can have up to 32 redundant backups. This large number of redundant backups greatly improves the availability of node memory backup data.

[0208] Distinguish between temporary storage and persistent storage for CKPT data: According to the method of one embodiment of the present application, during the training process, the CKPT data of each round is saved in the memory as temporary storage, and the memory data is asynchronously persisted to the remote storage according to the persistence requirements, and the remote storage is moved out of the critical path, taking into account the cost of storage and network bandwidth and the performance of storage and recovery.

[0209] Figure 10 Shown is a schematic diagram of a training node deployment scenario according to an embodiment of the present application.

[0210] The training cluster includes at least a model training group 901 and a model training group 902. The model training group 901 includes training nodes 911-914, and the model training group 902 includes training nodes 915-918.

[0211] The training cluster is installed with a training framework and related components, which are run in the operating system environment of the training cluster to provide training framework services 920 .

[0212] Training nodes 911 to 914 and training nodes 915 to 918 are installed with SDKs with relevant functions to connect with the training framework.

[0213] During the model training process, the training framework service 920 will call the SDK's storage function interface for CKPT storage.

[0214] A remote storage service 930 is deployed outside the training cluster. During the model training process, the training framework service 920 regularly uploads the CKPT stored in the memory of the training nodes 911-914 or the training nodes 915-918 to the remote storage service 930.

[0215] Specifically, before starting model training, the training framework service 920 calls the loading interface provided by the SDK of the training node. When it is detected that the training node does not save the CKPT data, it selects one (model training group 901 or 902) from all model training groups and allows the model training group 901 or 902 to download the CKPT from the remote storage service 920.

[0216] The model training group 901 or 902 that downloads CKPT data from the remote storage service 920 broadcasts the data to other training nodes (training nodes of the model training group 902 or 901) through the parameter plane network.

[0217] The training framework service 920 starts the training process and obtains the trained parameters by interacting with different training nodes. At the end of a round of training, the model parameters of all nodes are converged to the central node for average calculation as the model parameter training result of this round of training, namely CKPT.

[0218] The central node may be any training node in the training cluster, or the central node may be an independent node other than the training nodes configured in the training cluster.

[0219] After a round of training is completed and the CKPT is obtained, the central node will synchronize the final CKPT of this round to all nodes to prepare for the next round of training. At the same time, the training framework service 920 calls the storage CKPT function of the SDK of each training node to save the CKPT to the memory of each training node.

[0220] When the training node SDK receives a call to store a temporary CKPT, it notifies the training node to check whether its local memory has sufficient space to store the current CKPT. If local memory is insufficient, it checks whether the overflow to disk option has been configured. If overflow to disk has been configured, subsequent CKPT storage processes will be directly stored on the local disk. Otherwise, the oldest CKPT version in local memory will be deleted first.

[0221] The training node stores its CKPT in pre-allocated local memory or disk and maintains a KV locally, where the Key is the timestamp-based version number of the CKPT and the Value is the metadata related to the CKPT, including the size and storage location of the CKPT.

[0222] During multiple rounds of model training, at each preset time interval (e.g., 3 hours) or training round (e.g., 30 rounds), the training framework service 920 places the current CKPT into the remote storage service 930 for further analysis and use in the future.

[0223] Specifically, when the training framework service 920 is started, it will configure the time for persistence to the remote storage service 930 (for example, 3 hours), and when the conditions for persistent storage are met, the training framework service 920 will call the persistent storage interface through the SDK of the training node for persistent storage.

[0224] The training framework service 920 selects one of the model training groups as the model training group for persistent storage and instructs the model training group to upload the CKPT stored in its local memory to the remote storage service. The latest version of the CKPT in the model training group is locked during the upload and will not be deleted. After the selected model training group persists the CKPT to the remote storage, the CKPT persistence operation is completed and the CKPT is unlocked.

[0225] During fault recovery, the training framework service 920 calls the SDK of the training node to load the model CKPT in preparation for resuming training. Specifically, the training framework service 920 determines the latest saved CKPT version, and the training node searches its own memory to see whether the CKPT corresponding to that version is saved. If the CKPT corresponding to that version is saved in the training node's own memory, the training node directly reads the latest CKPT data from the memory, loads it onto the graphics card for the next round of model training, and ends the recovery operation. If the CKPT corresponding to that version is not saved in the training node's own memory, the training node reads the CKPT data of that version from the corresponding training node of the other model training group through the parameter surface network, loads it onto the graphics card for the next round of model training, and ends the recovery operation.

[0226] The present application also provides a computing device.

[0227] Figure 11 FIG. 1 is a schematic diagram of a computing device according to an embodiment of the present application.

[0228] like Figure 11 As shown, computing device 1100 includes a bus 1102, a processor 1104, a memory 1106, and a communication interface 1108. Processor 1104, memory 1106, and communication interface 1108 communicate with each other via bus 1102. Computing device 1100 can be a server or a terminal device. It should be understood that this application does not limit the number of processors and memories in computing device 1100.

[0229] The bus 1102 may be a peripheral component interconnect (PCI) bus or an extended industry standard architecture (EISA) bus. The bus may be divided into an address bus, a data bus, a control bus, etc. For ease of representation, Figure 11 The bus 1102 may include a path for transmitting information between various components of the computing device 1100 (eg, memory 1106, processor 1104, communication interface 1108).

[0230] The processor 1104 may include any one or more processors such as a central processing unit (CPU), a graphics processing unit (GPU), a microprocessor (MP), or a digital signal processor (DSP).

[0231] The memory 1106 may include volatile memory, such as random access memory (RAM). The processor 1104 may also include non-volatile memory, such as read-only memory (ROM), flash memory, a hard disk drive (HDD), or a solid state drive (SSD).

[0232] Memory 1106 stores executable program code, and processor 1104 executes the executable program code to implement the functions of the aforementioned storage control module 801, recovery control module 802, initialization module 803, and data upload module 804, thereby implementing the model training method proposed in the embodiment of the present application. In other words, memory 1106 stores instructions for executing the model training method proposed in the embodiment of the present application.

[0233] The communication interface 1108 uses a transceiver module such as, but not limited to, a network interface card or a transceiver to implement communication between the computing device 1100 and other devices or a communication network.

[0234] Embodiments of the present application also provide a computing device cluster. The computing device cluster includes at least one computing device. The computing device can be a server, such as a central server, an edge server, or a local server in a local data center. In some embodiments, the computing device can also be a terminal device such as a desktop computer, a laptop computer, or a smartphone.

[0235] Figure 12 FIG2 is a schematic diagram of a computing device cluster according to an embodiment of the present application.

[0236] like Figure 12 As shown, the computing device cluster includes at least one computing device 1100. The memory 1106 in one or more computing devices 1100 in the computing device cluster may store the same instructions for executing the model training method proposed in the embodiment of the present application.

[0237] In some possible implementations, the memory 1106 of one or more computing devices 1100 in the computing device cluster may also respectively store partial instructions for executing the model training method proposed in the embodiments of the present application. In other words, the combination of one or more computing devices 1100 can jointly execute instructions for executing the model training method proposed in the embodiments of the present application.

[0238] It should be noted that the memory 1106 in different computing devices 1100 in the computing device cluster can store different instructions, each used to perform part of the functions of the model training device proposed in the embodiments of the present application. In other words, the instructions stored in the memory 1106 in different computing devices 1100 can implement the functions of one or more modules among the storage control module 801, the recovery control module 802, the initialization module 803, and the data upload module 804.

[0239] In some possible implementations, one or more computing devices in a computing device cluster may be connected via a network, which may be a wide area network or a local area network.

[0240] Figure 13 FIG2 is a schematic diagram of a network connection of a computing device cluster according to an embodiment of the present application.

[0241] like Figure 13 As shown, two computing devices 1100A and 1100B are connected via a network, specifically, via a communication interface in each computing device.

[0242] Figure 13 The connection method between the computing device clusters shown can be based on the fact that the model method provided in this application requires interaction between multiple training nodes, data aggregation of multiple training nodes, and data synchronization of multiple training nodes.

[0243] It should be understood that Figure 13 The functionality of the computing device 1100A shown in FIG. 1 may also be implemented by multiple computing devices 1100. Similarly, the functionality of the computing device 1100B may also be implemented by multiple computing devices 1100.

[0244] The present application embodiment also provides another computing device cluster. The connection relationship between the computing devices in the computing device cluster can be similarly referred to as Figure 12 and Figure 13 The connection mode of the computing device cluster is different in that the memory 1106 in one or more computing devices 1100 in the computing device cluster may store the same instructions for executing the model training method.

[0245] In some possible implementations, the memory 1106 of one or more computing devices 1100 in the computing device cluster may also store partial instructions for executing the model training method. In other words, the combination of one or more computing devices 1100 can jointly execute instructions for executing the model training method.

[0246] Those skilled in the art will appreciate that the embodiments of the present application may be provided as methods, apparatus, or computer program products. Thus, the present invention may take the form of an entirely hardware embodiment, an entirely software embodiment, or an embodiment combining software and hardware. Furthermore, the present invention may take the form of a computer program product implemented on one or more computer-usable storage media containing computer-usable program code.

[0247] In the several embodiments provided in this application, if any function is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this application, or the part that contributes to the prior art, or part of the technical solution, can be embodied in the form of a software product, which is stored in a storage medium and includes a number of instructions for enabling a computer device (which can be a personal computer, server, or network device, etc.) to perform all or part of the steps of the method described in each embodiment of this application.

[0248] Specifically, embodiments of the present application further provide a computer program product comprising instructions. The computer program product may be software or a program product comprising instructions that can be run on a computing device or stored in any available medium. When the computer program product is run on at least one computing device, the at least one computing device executes the model training method proposed in embodiments of the present application.

[0249] The embodiment of the present application also provides a computer-readable storage medium. The computer-readable storage medium can be any available medium that can be stored by a computing device or a data storage device such as a data center containing one or more available media. The available medium can be a magnetic medium (e.g., a floppy disk, a hard disk, a tape), an optical medium (e.g., a DVD), or a semiconductor medium (e.g., a solid-state drive). The computer-readable storage medium includes instructions that instruct the computing device to execute the model training method proposed in the embodiment of the present application.

[0250] The description of the embodiments in this application is described with reference to the flowcharts and / or block diagrams of the methods, devices (apparatus), and computer program products according to the embodiments of the application. It should be understood that each process and / or box in the flowchart and / or block diagram, as well as the combination of the processes and / or boxes in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing device to produce a machine, so that the instructions executed by the processor of the computer or other programmable data processing device generate instructions for implementing the steps in the process. Figure 1 a process or multiple processes and / or boxes Figure 1 A device that provides the functions specified in a block or multiple blocks.

[0251] These computer program instructions may also be stored in a computer readable memory that can direct a computer or other programmable data processing device to work in a specific manner, so that the instructions stored in the computer readable memory produce an article of manufacture comprising an instruction device, which implements the process Figure 1 a process or multiple processes and / or boxes Figure 1 The function specified in one or more boxes.

[0252] These computer program instructions can also be loaded onto a computer or other programmable data processing device so that a series of operational steps are executed on the computer or other programmable device to produce a computer-implemented process, thereby providing the instructions executed on the computer or other programmable device for implementing the process. Figure 1 a process or multiple processes and / or boxes Figure 1 A step that specifies a function in one or more boxes.

[0253] It should also be noted that, in the embodiments of the present application, "at least one" refers to one or more, and "more than one" refers to two or more. "And / or" describes the association relationship of associated objects, indicating that three relationships may exist. For example, A and / or B can represent the existence of A alone, the existence of A and B at the same time, and the existence of B alone. Among them, A and B can be singular or plural. The character " / " generally indicates that the previous and subsequent associated objects are in an "or" relationship. "At least one of the following" and similar expressions refer to any combination of these items, including any combination of single or plural items. For example, at least one of a, b and c can represent: a, b, c, a and b, a and c, b and c or a and b and c, where a, b, c can be single or multiple.

[0254] In the embodiments of the present application, the terms "comprises," "includes," or any other variations thereof are intended to encompass non-exclusive inclusion, such that a process, method, commodity, or apparatus comprising a series of elements includes not only those elements but also other elements not explicitly listed, or elements inherent to such process, method, commodity, or apparatus. In the absence of further limitations, an element defined by the phrase "comprising a ..." does not preclude the presence of additional identical elements in the process, method, commodity, or apparatus comprising the element.

[0255] The present application may be described in the general context of computer-executable instructions executed by a computer, such as program modules. Generally, program modules include routines, programs, objects, components, data structures, etc. that perform specific tasks or implement specific abstract data types. The present application may also be practiced in distributed computing environments where tasks are performed by remote processing devices connected through a communications network. In a distributed computing environment, program modules may be located in local and remote computer storage media, including storage devices.

[0256] The various embodiments in this application are described in a progressive manner. Similar parts between the various embodiments can be referred to in conjunction with each other. Each embodiment focuses on the differences from other embodiments. In particular, the device embodiments are generally similar to the method embodiments, so the description is relatively simple. For relevant parts, refer to the partial description of the method embodiments.

[0257] Those skilled in the art will appreciate that the various units and algorithm steps described in the embodiments of the present application can be implemented using a combination of electronic hardware, computer software, and electronic hardware. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Professionals and technicians can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.

[0258] Those skilled in the art will clearly understand that, for the convenience and brevity of description, the specific working processes of the above-described devices, apparatuses and units can refer to the corresponding processes in the aforementioned method embodiments and will not be repeated here.

[0259] The above description is merely a specific embodiment of the present application. Any person skilled in the art may easily conceive of variations or substitutions within the technical scope disclosed in this application, and such variations or substitutions shall be within the scope of protection of this application. The scope of protection of this application shall be subject to the scope of protection of the claims.

Claims

1. A model training method, characterized in that: The method is applied to a training cluster, the training cluster including multiple model training groups, the multiple model training groups including a first model training group and a second model training group, the first model training group and the second model training group are used to train a neural network model in parallel in a data-parallel manner, the first model training group and the second model training group have the same node structure, the first model training group includes a first training node, and the second model training group includes a second training node, wherein based on the node structure, the first training node corresponds to the second training node, and the method includes: Generate checkpoint data after a round of training, wherein the checkpoint data is the checkpoint data of the first training node and the second training node; The first training node saves the checkpoint data into a memory of the first training node, and the second training node saves the checkpoint data into a memory of the second training node; If the first training node fails during model training, the first training node obtains the checkpoint data from the memory of the second training node to perform data recovery on the first training node.

2. The method according to claim 1, characterized in that The method further comprises: For the second training node that has not failed, data recovery is performed on the second training node using the checkpoint data stored in the memory of the second training node itself.

3. The method according to claim 1, characterized in that The checkpoint data generated after a round of training includes: At the end of a round of training, the model parameters of all training nodes in the training cluster are averaged and the calculated result is used as the checkpoint data for this round of training.

4. The method according to claim 1, wherein The method further comprises: Download initial checkpoint data using the first training node; The initial checkpoint data is broadcasted by the first training node to the second training node.

5. The method according to any one of claims 1 to 4, characterized in that The method further comprises: Use the first model training group to upload the checkpoint data stored in the memory of each training node in the first model training group to a remote storage service.

6. The method according to claim 5, characterized in that During the process of uploading checkpoint data by the first model training group, it is prohibited to delete the checkpoint data saved in the memory of each training node in the first model training group, and it is prohibited to use new checkpoint data to overwrite the saved checkpoint data.

7. The method according to claim 5, characterized in that During the process of uploading the checkpoint data by the first model training group, stop saving the new checkpoint data to the memory of the training node of the first model training group.

8. The method according to claim 7, characterized in that After the first model training group completes uploading the checkpoint data, the new checkpoint data stored in the memory of the training node of the second model training group is synchronized to the memory of the training node of the first model training group.

9. A model training device, characterized in that: The device is applied to a training cluster, the training cluster includes multiple model training groups, the multiple model training groups include a first model training group and a second model training group, the first model training group and the second model training group are used to parallel train a neural network model in a data-parallel manner, the first model training group and the second model training group include the same node structure, the first model training group includes a first training node, and the second model training group includes a second training node, wherein, based on the node structure, the first training node corresponds to the second training node; the training cluster generates checkpoint data after a round of training, wherein the checkpoint data is the checkpoint data of the first training node and the second training node; The device comprises: a storage control module, configured to instruct the first training node to save the checkpoint data into a memory of the first training node, and to instruct the second training node to save the checkpoint data into a memory of the second training node; A recovery control module is used to instruct the first training node to obtain the checkpoint data from the memory of the second training node to perform data recovery on the first training node if the first training node fails during the training of the model.

10. The device according to claim 9, characterized in that The recovery control module is further configured to: Instruct the second training node to use the checkpoint data stored in the memory of the second training node itself to perform data recovery on the second training node.

11. The device according to claim 9, characterized in that At the end of a round of training, the training cluster calculates the average value of the model parameters of all training nodes in the training cluster, and uses the calculated result as the checkpoint data of this round of training.

12. The device according to claim 9, characterized in that The device further includes an initialization module, wherein the initialization module is configured to: instructing the first training node to download initial checkpoint data; Instruct the first training node to broadcast the initial checkpoint data to the second training node.

13. The device according to any one of claims 9 to 12, characterized in that The device further includes a data uploading module, which is configured to: Instruct the first model training group to upload the checkpoint data stored in the memory of each training node in the first model training group to a remote storage service.

14. The device according to claim 13, characterized in that The storage control module is further configured to: During the process of uploading checkpoint data by the first model training group, it is prohibited to delete the checkpoint data saved in the memory of each training node in the first model training group, and it is prohibited to use new checkpoint data to overwrite the saved checkpoint data.

15. The device according to claim 13, characterized in that The storage control module is further configured to: During the process of uploading the checkpoint data by the first model training group, stop saving the new checkpoint data to the memory of the training node of the first model training group.

16. The device according to claim 15, characterized in that The storage control module is further configured to: After the first model training group completes uploading the checkpoint data, the new checkpoint data stored in the memory of the training node of the second model training group is synchronized to the memory of the training node of the first model training group.

17. A computing device cluster, characterized in that: comprising at least one computing device, each computing device including a processor and a memory; The processor of the at least one computing device is configured to execute instructions stored in a memory of the at least one computing device, so that the computing device cluster executes the method according to any one of claims 1 to 8.

18. A computer program product comprising instructions, characterized in that When the instructions are executed by a computing device cluster, the computing device cluster is caused to perform the method according to any one of claims 1 to 8.

19. A computer-readable storage medium, characterized in that The method comprises computer program instructions, and when the computer program instructions are executed by a computing device cluster, the computing device cluster performs the method according to any one of claims 1 to 8.