Erasure code-based checkpoint management method, and distributed training system

WO2026200673A1PCT designated stage Publication Date: 2026-10-01HUAWEI TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
PCT/CN2026/084479
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2025-03-26
Filing Date
2026-03-19
Publication Date
2026-10-01

Smart Images

  • Figure CN2026084479_01102026_PF_FP_ABST
    Figure CN2026084479_01102026_PF_FP_ABST
Patent Text Reader

Abstract

An erasure code-based checkpoint management method and a distributed training system, relating to the technical field of neural network models. During a process in which the distributed training system trains a neural network model by means of a plurality of training nodes, checkpoint data is saved in CUP memories of the plurality of training nodes, thereby avoiding the impact of network bandwidth when saving the checkpoint data, helping to improve the speed of saving the checkpoint data, shortening the time required for saving the checkpoint data, and further helping to improve the frequency of saving the checkpoint data. In addition, by means of configuring the checkpoint data to comprise K pieces of raw data and M pieces of verification data, and separately storing the K pieces of raw data and the M pieces of verification data onto different training nodes, the K pieces of raw data can be obtained from any M pieces of data among the K pieces of raw data and the M pieces of verification data, such that the distributed training system can tolerate simultaneous failure of up to M training nodes, thereby helping to improve the fault tolerance of the distributed training system.
Need to check novelty before this filing date? Find Prior Art

Description

An erasure coding-based checkpoint management method and a distributed training system

[0001] This application claims priority to Chinese patent application filed on March 26, 2025, with application number 202510378322.8 and entitled "Checkpoint Management Method and Distributed Training System Based on Erasure Coding", the entire contents of which are incorporated herein by reference. Technical Field

[0002] This application relates to the field of neural network model technology, and in particular to a checkpoint management method and distributed training system based on erasure coding. Background Technology

[0003] Because failures frequently occur during the training of neural network models, related technologies have proposed saving checkpoints in a remote storage system. This allows the training state of the neural network model to be restored using these checkpoints in case of a failure. However, saving checkpoints to a remote storage system is time-consuming, resulting in a low frequency of checkpoint saving. In this case, when a failure occurs during neural network model training, restoring the training state using remote checkpoints can only restore the training state to an earlier point in time, leading to the need for significant computational resources and time for retraining. Summary of the Invention

[0004] This application provides a checkpoint management method and a distributed training system based on erasure coding, which not only helps to shorten the time spent saving checkpoints, but also helps to improve the fault tolerance of the distributed training system.

[0005] Firstly, a checkpoint management method based on erasure coding is provided, applied to a distributed training system. The distributed training system trains a neural network model using multiple training nodes. Each training node includes a computing unit, a central processing unit (CPU), and a first memory of the CPU. The computing unit includes a second memory. The method includes: writing first training state data from the second memory of each training node into the first memory of each training node, wherein the first training state data is obtained by the distributed training system training the neural network; performing erasure coding on the second training state data in the first memory of each training node to obtain N checkpoint data, wherein the second training state data is obtained by the distributed training system training the neural network, and includes the first training state data; the N checkpoint data includes K original data and M verification data, wherein the K original data includes the second training state data of each training node, and any K checkpoint data from the N checkpoint data are used to obtain the second training state data of each training node; storing the N checkpoint data in the first memory of the multiple training nodes, wherein different checkpoint data from the N checkpoint data are stored in the first memory of different training nodes.

[0006] In the above scheme, by storing checkpoints (i.e., N checkpoint data) locally on the training nodes, the impact of network bandwidth on checkpoint storage is avoided, thus improving the speed and time of checkpoint storage. Furthermore, since checkpoints are stored in the first memory of the training nodes, and memory has a high read / write speed, the storage speed is further improved, reducing storage time. This reduced storage time facilitates more frequent checkpoint storage, enabling recovery to a more recent training state in case of failure, thus avoiding the need for extensive retraining. Moreover, the reduced checkpoint storage time also shortens the time spent training the neural network model during checkpoint storage, minimizing its impact on training throughput. This allows for increased checkpoint storage frequency while maintaining training throughput (i.e., throughput during neural network model training). Additionally, because this scheme incorporates erasure coding operations to obtain N checkpoint data, the distributed training system can tolerate up to M simultaneous training node failures, thus improving the fault tolerance of the distributed training system. Based on the foregoing, this scheme achieves good fault tolerance while balancing training throughput and checkpoint retention frequency, thus reducing recovery costs in the event of a failure. Especially in large-scale distributed training scenarios, setting a larger M can significantly improve the training performance of the distributed training system. Furthermore, storing checkpoints in the local memory of the training nodes, compared to storing them in remote storage devices, helps the distributed training system retrieve checkpoint data faster in the event of a failure, thereby improving the speed at which the distributed training system recovers its training state.

[0007] In one possible implementation, the second training state data includes the values ​​of tensor key-value data.

[0008] In this implementation, erasure coding is performed only on the values ​​of the tensor key data, thereby achieving non-serialization encoding through a non-serialization encoding protocol. This helps reduce the overhead of storing checkpoints, increase the frequency of storing checkpoints, and improve the efficiency of restoring the training state. Furthermore, when subsequently restoring the training state of the neural network model through checkpoints, a non-serialization decoding protocol can be used to achieve non-serialization decoding, further reducing the overhead of restoring the training state and improving its efficiency.

[0009] In another possible implementation, the multiple training nodes include a first training node, and the method further includes: storing the third training state data in the first memory of the first training node into the first memory of the second training node, wherein the second training node includes at least some of the training nodes other than the first training node, and the third training state data includes the training metadata and tensor key-value data obtained by the distributed training system from training the neural network.

[0010] In this implementation, the third training state data stored in the first memory of the first training node is transferred to the first memory of the second training node, thereby achieving multiple backups of the third training state data in the first memory of the first training node. In this way, as long as some training nodes do not fail, the third training state data of each training node can be obtained, thus helping to improve the fault tolerance of the third training state data.

[0011] In another possible implementation, storing the N checkpoint data in the first memory of multiple training nodes includes: determining K third training nodes as K data nodes, and determining M verification nodes from the training nodes other than the K third training nodes. The number of communications required to store the K original data to the K data nodes and the M verification data to the M verification nodes when the third training nodes are data nodes is less than the number of communications required to store the K original data to the K data nodes and the M verification data to the M verification nodes when the third training nodes are verification nodes. The K original data are stored in the first memory of the K data nodes, and the M verification data are stored in the first memory of the M verification nodes.

[0012] In this implementation, by determining K third training nodes as K data nodes and storing K raw data on the K data nodes, the number of communication operations for data transmission is reduced, thereby reducing the time overhead of data transmission operations.

[0013] In another possible implementation, the distributed training system includes multiple training processes deployed on multiple training nodes for training neural network models. K third training nodes among the multiple training nodes are identified as K data nodes. This includes: obtaining physical interval groups for the multiple training processes, each physical interval group comprising multiple physical sub-interval groups, where training processes in a physical sub-interval group reside on the same training node; obtaining data interval groups for the multiple training processes, where each data interval group comprises K data sub-interval groups; identifying a first physical sub-interval group from the multiple physical sub-interval groups, wherein the number of identical training processes in the first physical sub-interval group and the target data sub-interval group among the K data sub-interval groups is greater than the number of identical training processes in the second physical sub-interval group and the target data sub-interval group, the second physical sub-interval group comprising at least some of the physical sub-interval groups other than the first physical sub-interval group; and identifying the training node where the training processes in the first physical sub-interval group reside as a data node, wherein the training node where the training processes in the first physical sub-interval group reside is a third training node.

[0014] In this implementation, K data nodes are determined by physical interval groups and data interval groups, which helps to improve the convenience and diversity of methods for determining data nodes.

[0015] In another possible implementation, erasure coding is performed on the second training state data in the first memory of each training node, including: encoding the second training state data of each training process according to the encoding matrix of each training process in multiple training processes to obtain M encoding blocks for each training process; dividing the multiple training processes into R XOR groups, where each XOR group includes K training processes; and performing an XOR operation on the encoding blocks of different training processes in each XOR group to obtain R*M parity blocks, where the R*M parity blocks are M parity data.

[0016] In this implementation, M verification data are obtained through encoding and XOR operations, which optimizes the calculation of erasure coding operations, thereby simplifying the complexity of erasure coding operations, improving the efficiency of erasure coding operations, and thus achieving efficient acquisition of M verification data.

[0017] In another possible implementation, the method further includes: determining an XOR target process for each XOR group in a plurality of XOR groups, wherein the XOR target process of each XOR group is used to perform an XOR operation on the encoding blocks of different training processes of each XOR group, and the training node where the XOR target process of each XOR group is located is used to store M parity blocks of each XOR group. When the M parity blocks of each XOR group are stored in the training node where the XOR target process of each XOR group is located, the number of communications from the M parity blocks stored in the M parity blocks to the M parity nodes satisfies the target condition.

[0018] In this implementation, by setting the verification block to be stored in the first memory of the training node where the target XOR process is located, it helps to avoid the communication resources occupied by the process of storing the verification block, thereby helping to reduce the total communication resources occupied when performing the XOR operation.

[0019] In another possible implementation, when the first XOR group in a plurality of XOR groups includes training processes on M verification nodes, the target XOR process of the first XOR group is the training process on M verification nodes.

[0020] In this implementation, by selecting the training processes on M verification nodes as the XOR target processes, the verification blocks can be directly stored on the verification nodes, which helps to reduce the number of subsequent data transmissions and thus helps to reduce the communication resources occupied by subsequent data transmissions.

[0021] In another possible implementation, based on the encoding matrix of each training process in multiple training processes, an encoding operation is performed on the second training state data of each training process to obtain M encoding blocks for each training process. This includes: writing the second training state data of each training process in the first XOR group of multiple XOR groups into the data buffer of each training process in the first XOR group to obtain each data block of each training process in the first XOR group; and performing an encoding operation on each data block of each training process in the first XOR group based on the encoding matrix of each training process in the first XOR group to obtain each encoding block of each training process in the first XOR group.

[0022] In this implementation, multiple data blocks for each training process are obtained by writing the second training state data of each training process into a data buffer. Encoding operations are then performed on each of these data blocks, decomposing the encoding task of each training process into multiple sub-encoding tasks. Each sub-encoding task corresponds to one data block in the data buffer, and by executing multiple sub-encoding tasks, M encoding blocks are obtained. After obtaining an encoding block, an XOR operation can be performed on it, allowing for pipelined execution of encoding and XOR operations. This improves the efficiency of erasure coding operations and reduces their time overhead.

[0023] In another possible implementation, the encoding operation is performed on each data block of each training process in the first XOR group, including: obtaining multiple threads of each training process in the first XOR group; and performing the encoding operation in parallel on each data block of each training process in the first XOR group through the multiple threads of each training process in the first XOR group.

[0024] In this implementation, by assigning each sub-encoding task to multiple threads for simultaneous execution, the multiple cores of the CPU can be fully utilized to accelerate the encoding operation, thereby helping to reduce the time overhead of the encoding operation.

[0025] In another possible implementation, an encoding operation is performed on each data block of each training process in the first XOR group, including: writing each encoded block of each training process in the first XOR group into the encoding buffer of each training process in the first XOR group; and performing an XOR operation on the encoded blocks in the encoding buffers of different training processes in the first XOR group.

[0026] In this implementation, by writing the encoded blocks into the encoding buffer and performing an XOR operation on the encoded blocks in the encoding buffer, the encoding operation of the data blocks and the XOR operation of the encoded blocks can be executed in a pipeline manner, which helps to improve the speed of erasure coding operation and reduce the time consumption of erasure coding operation.

[0027] In another possible implementation, the erasure coding operation performed on the second training state data in the first memory of each training node further includes: dividing multiple logical blocks of multiple training processes into K logical segments, wherein one logical block is used to indicate the second training state data obtained by a training process in training a neural network model; performing encoding operations on the K logical segments according to the erasure coding encoding matrix to obtain M logical check data, wherein the M logical check data are used to indicate M check data; and determining the encoding matrix of each training process in multiple training processes based on the M logical check data.

[0028] In this implementation, determining the encoding matrix for each training process using a global encoding result matrix helps improve the accuracy and reliability of the final M checksums. Furthermore, representing the second training state data of a training process using logical blocks, compared to actually acquiring the second training state data of other training processes, not only helps reduce the communication resources consumed during erasure coding operations but also improves the execution efficiency of erasure coding operations, thereby helping to reduce the time spent storing checkpoints.

[0029] In another possible implementation, the method further includes: when the training node storing the verification data fails, and each training node storing K original data is in a live state, the training state of the neural network model is restored using the K original data.

[0030] In this implementation, the training state of the neural network model is restored by using K original data from the second memory of the surviving node. This not only helps to improve the speed of restoring the training state and thus reduce the time overhead of restoring the training state, but also helps to reduce the communication overhead of restoring the training state.

[0031] In another possible implementation, the method further includes: when the training node storing the original data fails, and there are more than or equal to M training nodes in the surviving state among the N training nodes storing N checkpoint data, the training state of the neural network model is restored by using the K checkpoint data on the surviving training nodes.

[0032] In this implementation, when the original data is incomplete, the training state of the neural network model is restored through K checkpoint data, thereby improving the fault tolerance of the distributed training system and thus improving the training performance of the distributed training system.

[0033] In another possible implementation, erasure coding is performed on the second training state data in the first memory of each training node, including: during the training of a neural network model through multiple training nodes in a distributed training system, erasure coding is performed on the second training state data in the first memory of each training node using the CPU resources of each training node.

[0034] In this implementation, using CPU resources to perform erasure coding operations helps to execute the training and erasure coding operations of the neural network model in parallel, thereby helping to reduce the impact of storage checkpoints on the throughput of the trained neural network model.

[0035] In another possible implementation, storing the N checkpoint data in the first memory of multiple training nodes includes: during the training of a neural network model through multiple training nodes in a distributed training system, storing the N checkpoint data in the first memory of multiple training nodes using the CPU resources of each training node.

[0036] In this implementation, the erasure coding operation is performed using CPU resources, which helps to execute the training and data transfer operations of the neural network model in parallel, thereby helping to reduce the impact of storage checkpoints on the throughput of the trained neural network model.

[0037] In another possible implementation, an XOR operation is performed on the encoding blocks of different training processes in each of the multiple XOR groups, including: when the different training nodes of different training processes in the first XOR group are in a network idle state, an XOR operation is performed on the encoding blocks of different training processes in the first XOR group.

[0038] In this implementation, by setting the XOR operation to be performed when different training nodes are in a network idle state, it helps to avoid network resource competition between the XOR operation and the operation of training the neural network model, thereby helping to avoid affecting communication during the training process of the neural network model, and thus helping to avoid affecting the training throughput.

[0039] In another possible implementation, the N checkpoint data are stored in the first memory of multiple training nodes, including: when the fourth training node and the fifth training node are in a network idle state, the original data or verification data on the fourth training node is stored in the fifth training node.

[0040] In this implementation, by setting the data transmission operation to be performed when different training nodes are in a network idle state, it helps to avoid network resource competition between the XOR operation and the operation of training the neural network model, thereby helping to avoid affecting the communication during the training process of the neural network model, and thus helping to avoid affecting the training throughput.

[0041] Secondly, a distributed training system is provided, comprising: functional units for executing any of the methods provided in the first aspect, wherein the actions performed by each functional unit are implemented by hardware or by hardware executing corresponding software. For example, the distributed training system may include: a writing module, an erasure coding module, and a storage module; the writing module is used to write first training state data from the second memory of each training node into the first memory of each training node, wherein the first training state data is obtained by the distributed training system training a neural network; the erasure coding module is used to perform erasure coding operations on the second training state data in the first memory of each training node to obtain N checkpoint data, wherein the second training state data is data obtained by the distributed training system training a neural network, the second training state data includes the first training state data, the N checkpoint data includes K original data and M verification data, the K original data includes the second training state data of each training node, and any K checkpoint data from the N checkpoint data are used to obtain the second training state data of each training node; the storage module is used to store the N checkpoint data in the first memory of multiple training nodes, wherein different checkpoint data from the N checkpoint data are stored in the first memory of different training nodes among the multiple training nodes.

[0042] Thirdly, a processor is provided that can be used to execute any of the methods provided in the first aspect above.

[0043] Fourthly, a chip is provided, comprising: a processor and a power supply circuit; the power supply circuit can be used to supply power to the chip; the processor can be used to execute any of the methods provided in the first aspect above.

[0044] Fifthly, a computing device is provided, comprising: a processor, a memory, and computer programs / instructions stored in the memory; the processor executes the computer programs / instructions to cause the computing device to perform any of the methods provided in the first aspect above.

[0045] A sixth aspect provides a computing device cluster, comprising: at least one computing device, each computing device including a processor, a memory, and computer programs / instructions stored in the memory; the processor of each computing device executes the computer programs / instructions to enable the computing device cluster to implement any of the methods provided in the first aspect above.

[0046] In a seventh aspect, a computer program product is provided, comprising a computer program / instructions that, when executed by a computing device, implement any of the methods provided in the first aspect above.

[0047] Eighthly, a computer-readable storage medium is provided, on which a computer program / instructions are stored, which, when executed by a computing device, implement any of the methods provided in the first aspect above.

[0048] The technical effects of any of the implementation methods in aspects two through eight can be seen in the technical effects of different implementation methods in aspect one above, and will not be repeated here. Attached Figure Description

[0049] Figure 1 is a schematic diagram of a system architecture provided in this application;

[0050] Figure 2 is a schematic diagram of a training node provided in this application;

[0051] Figure 3 is a flowchart of a checkpoint management method based on erasure coding provided in this application;

[0052] Figure 4 is one of the schematic diagrams of a checkpoint management method provided in this application;

[0053] Figure 5 is a second schematic diagram of a checkpoint management method provided in this application;

[0054] Figure 6 is a schematic diagram of the third type of checkpoint management method provided in this application;

[0055] Figure 7 is a fourth schematic diagram of a checkpoint management method provided in this application;

[0056] Figure 8 is a fifth schematic diagram of a checkpoint management method provided in this application;

[0057] Figure 9 is a schematic diagram of a checkpoint management method provided in this application (the sixth one).

[0058] Figure 10 is a schematic diagram of a distributed training system provided in this application;

[0059] Figure 11 is a schematic diagram of a computing device provided in this application;

[0060] Figure 12 is a schematic diagram of a computing device cluster provided in this application;

[0061] Figure 13 is a schematic diagram of the connection of a computing device cluster provided in this application. Detailed Implementation

[0062] To facilitate understanding, a brief introduction to the relevant terms used in this application will be provided first.

[0063] Serialization is the process of converting data into a storable or transmissible sequence of bytes. Performing serialization on data incurs additional computational overhead. For example, data needs to be serialized before writing it to a file, sending it over a network, or persistently storing it.

[0064] The state dictionary (state_dict) is a collection of all data that needs to be stored during the training of a neural network model in a distributed training system. It includes model parameters, optimizer state, data loader state, and training metadata (such as checkpoint version, iteration count, etc.). The state dictionary can be used to reconstruct the training state of the neural network model, and state dictionaries at different points in time are used to reconstruct the training state at different points in time. During the training of a neural network model in a distributed training system, each training process maintains a sharded state dictionary. For ease of description, a shard of the state dictionary can be called a state dictionary shard.

[0065] A checkpoint is a state dictionary that is periodically saved during the training of a neural network model in a distributed training system. If the distributed training system fails or training is interrupted, checkpoints can be used to resume training, avoiding the need to retrain the neural network model from scratch.

[0066] Erasure coding (EC) is a data redundancy technique that generates several redundant blocks for the original data, thereby enabling the recovery of all the original data even if some of the original data is lost.

[0067] The technical solution provided in this application will be described in detail below with reference to the accompanying drawings.

[0068] In recent years, with the rapid development of deep learning technology, the emergence of large-scale deep learning models has brought new possibilities to fields such as natural language processing and computer vision. For example, the introduction of models such as GPT-4, LLaMA, and PaLM-E has significantly improved the performance of models in complex tasks. However, the training of these models often requires massive computing resources. Taking GPT-4 and LLaMA as examples, these models typically contain hundreds of billions or even trillions of parameters, and their training process requires thousands or even tens of thousands of GPUs to run continuously for weeks or even months. This extremely time-consuming training process inevitably encounters various failures, such as software problems and hardware failures. For instance, during the training of the LLaMA 3.1405B model, a total of 419 unexpected failures occurred within 54 days, involving various issues such as GPU failures, network infrastructure failures, server hardware failures, and software vulnerabilities. These failures occurred on average once every 3 hours, and 78% of them were attributed to hardware problems. These types of failures usually lead to training interruptions, data loss, etc., which necessitates restarting training and resulting in high recovery costs.

[0069] To address the issue of failures during the training of neural network models, checkpointing techniques have been proposed. For example, during the neural network modeling process, after each training process in a distributed training system performs synchronous blocking training, the distributed training system serializes the state dictionary, obtaining a serialized result. After writing the serialized result to a remote storage system, each training process in the distributed training system continues training. Checkpointing techniques periodically save the state dictionary to the remote storage system. When a failure occurs, the most recently saved checkpoint is loaded from the remote storage system, thereby restoring the training state to the corresponding checkpoint and achieving post-failure training state recovery.

[0070] However, because saving checkpoints to a remote storage system each time is time-consuming, the frequency of checkpoint saving is low. For example, checkpoints are only saved every few hours. In this case, if a failure occurs during the training of the neural network model, and the training state is restored using remote checkpoints, it can only restore the training state to an earlier point in time, resulting in a significant expenditure of computing resources and time for retraining.

[0071] In view of this, this application provides a checkpoint management method based on erasure coding, applied to a distributed training system. The distributed training system trains a neural network model using multiple training nodes. Each training node includes a computing unit, a central processing unit (CPU), and a first memory of the CPU. The computing unit includes a second memory. When a checkpoint needs to be saved, the distributed training system saves the checkpoint locally on the training node. This avoids the impact of network bandwidth on checkpoint saving, thereby improving the speed and time of checkpoint saving. Furthermore, by saving the checkpoint in the first memory of the training node, where memory has a high read / write speed, the speed of checkpoint saving can be further improved, reducing the time spent saving checkpoints. Reducing the time spent saving checkpoints helps increase the frequency of checkpoint saving, which facilitates recovery to a more recent training state in the event of a failure, thus helping to avoid consuming significant computing resources and time for retraining. Furthermore, since this scheme reduces the time spent saving checkpoints, it also reduces the time spent training the neural network model during checkpoint saving. This helps to reduce the impact of checkpoint saving on training throughput, thereby increasing the frequency of checkpoint saving while maintaining training throughput (i.e., the throughput during the training of the neural network model). Additionally, because this scheme incorporates erasure coding operations to obtain N checkpoint data during checkpoint saving, the distributed training system can tolerate a maximum of M training nodes failing simultaneously, thus improving the fault tolerance of the distributed training system. Moreover, storing checkpoints in the local memory of the training nodes, compared to storing them in remote storage devices, helps to increase the speed at which the distributed training system retrieves checkpoint data in the event of a failure, thereby improving the rate at which the distributed training system recovers its training state.

[0072] Next, the system architecture involved in the technical solution provided in this application will be further described with reference to the accompanying drawings.

[0073] This application provides a distributed training system applied to the aforementioned erasure coding-based checkpoint management method. The distributed training system can be used to train a neural network model using multiple training nodes. Any two training nodes can communicate with each other.

[0074] For example, as shown in Figure 1, multiple training nodes include training node 0, ..., training node Z. Here, Z is a positive integer greater than or equal to 1.

[0075] It should be noted that in this application, "multiple" includes two or more, which will not be elaborated further.

[0076] For example, the training node in this application can be a physical machine, a virtual machine, etc.

[0077] It should be noted that this application does not limit the type of training nodes; the above is merely an illustrative example.

[0078] It should be noted that this application does not impose any restrictions on the distributed training framework used in the distributed training system. For example, the distributed training framework used in the distributed training system can be Megatron-LM, etc.

[0079] For example, during the training of neural network models in the distributed training system of this application, tensor parallelism (TP), pipeline parallelism (PP), data parallelism (DP), and fully sharded data parallelism (FSDP) can be used.

[0080] For example, the sharding method of the state dictionary is related to the parallelism strategy adopted by the distributed training system. When the distributed training system employs a tensor parallelism strategy, the state dictionary is sharded based on the tensor dimension; each shard of the state dictionary includes a portion of the tensor key-value data, for example, it can be split along the hidden layer dimension or attention head dimension of the neural network model. When the distributed training system employs a pipelined parallelism strategy, the state dictionary is sharded based on the neural network model; each shard of the state dictionary includes parameters from one or more consecutive layers. When the distributed training system employs a fully sharded data parallelism strategy, the model parameters are fully sharded across the various training processes of the distributed training system.

[0081] For example, the neural network model trained by the distributed training system can be a deep neural network (DNN) model, a convolutional neural network (CNN) model, a recurrent neural network (RNN) model, etc. For example, the deep neural network model can be a model such as GPT, BERT, T5, etc.

[0082] It should be noted that this application does not limit the type of neural network model trained by the distributed training system; the above is merely an illustrative example.

[0083] For example, each training node can deploy at least one training process of a distributed training system, and each training process can be used to train a neural network model. For instance, different training processes within the at least one training process can train the neural network model in parallel, which helps improve training efficiency. Here, a training process can be considered the smallest computational unit of the distributed training system.

[0084] It should be noted that in this application, "at least one" includes one or more, which will not be elaborated further.

[0085] In this application, the distributed training system includes a checkpoint module, which is used to implement the erasure coding-based checkpoint management method provided in this application. For example, during the training of a neural network model, each training process can call the checkpoint module when it is necessary to save checkpoints, thereby implementing the erasure coding-based checkpoint management method provided in this application.

[0086] In one example, multiple training nodes can belong to the same data center (DC), which helps improve communication efficiency. In another example, multiple training nodes can belong to the same availability zone (AZ), which helps improve communication efficiency. In yet another example, multiple training nodes can belong to the same region, which helps improve communication efficiency.

[0087] It should be noted that this application does not impose any restrictions on the locational relationship of multiple training nodes; the above is merely an illustrative example. For instance, multiple training nodes may belong to different data centers, different Availability Zones (AZs), or different regions.

[0088] In one example, the distributed training system includes multiple training nodes and a distributed software system deployed on the training nodes. Each training node can start at least one training process through the distributed software system. In other words, the distributed training system is a combination of hardware and software.

[0089] In another example, a distributed training system is a distributed software system used to train neural network models.

[0090] It should be noted that this application does not limit the form of the distributed training system; the above is merely an illustrative example. The following description uses a distributed training system as an example of a distributed software system to illustrate this application.

[0091] In one example, multiple training nodes are the user's training nodes. After purchasing the distributed training system provided in this application, the user deploys the distributed training system on multiple training nodes, thereby enabling the distributed training system to implement the erasure coding-based checkpoint management method provided in this application through multiple training nodes. Alternatively, after purchasing the checkpointing module provided in this application, the user deploys the checkpointing model on multiple training nodes, thereby enabling the training process on each training node to implement the erasure coding-based checkpoint management method provided in this application by calling the checkpointing module.

[0092] In another example, multiple training nodes are cloud training nodes rented by a tenant. After renting multiple training nodes, the tenant can apply to use the management service corresponding to the checkpoint management method provided in this application, thereby realizing the use of the erasure coding-based checkpoint management method provided in this application.

[0093] In this application, each training node may include a computing unit, a central processing unit (CPU), and a first memory of the CPU, etc., and at least one computing unit includes a second memory. Each computing unit can communicate with the CPU.

[0094] For example, as shown in Figure 2, training node 0 includes computing unit 0, CPU 0, and memory 01 of CPU 0. The computing unit 0 includes memory 02 and can communicate with CPU 0.

[0095] It should be noted that the reason for defining "first memory" and "second memory" in this application is to distinguish the memory of different devices. Specifically, the memory of the CPU is named "first memory," and the memory of the computing unit is named "second memory."

[0096] For example, the computing unit can be a computing unit with computing capabilities, such as a graphics processing unit (GPU), a data processing unit (DPU), a neural processing unit (NPU), or a tensor processing unit (TPU).

[0097] It should be noted that this application does not limit the type of computing unit; the above is merely an illustrative example. The following description uses a GPU as an example of a computing unit. The GPU's memory can also be referred to as video memory.

[0098] In one example, the number of training processes launched on a training node is the same as the number of computing units on the training node. For instance, if a training node includes two computing units, then two training processes are launched on the computing node. In this way, each training process can be allocated one computing unit, meaning that the resources of one computing unit can be used by one training process, thereby helping to improve the processing efficiency of each training process.

[0099] In another example, the number of training processes launched on a training node can be greater than the number of computing units on the training node. This helps to improve the resource utilization of each computing unit.

[0100] In yet another example, the number of training processes launched on a training node can be less than the number of computing units on the training node. This helps to increase the diversity of deployment methods.

[0101] It should be noted that this application does not limit the relationship between the number of training processes and the number of computing units on a training node; the above is merely an illustrative example. The following description assumes that the number of training processes and the number of computing units on a training node are equal.

[0102] It should be noted that this application does not limit the type of computing unit; the above is merely an illustrative example.

[0103] Optionally, the distributed training system can also communicate with a client on a user's electronic device. For example, a user can send configuration information to the distributed training system via a client on their electronic device, indicating the maximum number of faulty nodes supported by the distributed training system, i.e., the value of M.

[0104] Alternatively, the electronic device may be a terminal device or a network device.

[0105] For example, the terminal device can be a mobile phone, tablet computer, handheld computer, personal computer (PC), personal digital assistant (PDA), ultra-mobile personal computer (UMPC), laptop computer, netbook, desktop computer or all-in-one computer, etc.

[0106] It should be noted that this application does not impose any restrictions on the device form of the terminal equipment; the above is merely an illustrative example.

[0107] For example, network devices can be servers, bare metal servers, etc. A server can be a single physical server, or it can be two or more physical servers that share different responsibilities and work together to achieve the various functions of the server. For example, a server can be a blade server, a high-density server, a rack server, or a tower server, etc.

[0108] It should be noted that this application does not limit the form factor of the network device; the above is merely an illustrative example.

[0109] It should be noted that the system architecture shown in Figures 1 and 2 does not constitute a limitation on the system architecture for implementing the checkpoint management method provided in this application.

[0110] It should be noted that the system architecture and application scenarios described in this application are for the purpose of more clearly illustrating the technical solutions of this application, and do not constitute a limitation on the technical solutions provided in this application. As those skilled in the art will know, with the evolution of system architecture and the emergence of new application scenarios, the technical solutions provided in this application are also applicable to similar technical problems.

[0111] For ease of understanding, the checkpoint management method provided in this application will be described below with reference to the above system architecture and accompanying drawings.

[0112] Figure 3 is a flowchart of a checkpoint management method based on erasure coding provided in this application. For example, the checkpoint management method may include steps 301-303. Here, "step" in this application can be abbreviated as "S", and will not be described further hereafter.

[0113] For example, the number of training nodes is greater than or equal to N, that is, Z is greater than or equal to N. Here, N is equal to the sum of K and M.

[0114] In one example, as shown in Figure 4, multiple training nodes include N training nodes, where Z equals N. These N training nodes include training node 0, ..., training node N-1. Training node 0 includes GPU0, CPU0, and the CPU's first memory; GPU0 includes its second memory, etc. In another example, multiple training nodes include N+V training nodes, where Z is greater than N, and V is a positive integer greater than 1.

[0115] It should be noted that this application does not limit the number of GPUs included in the training node; Figure 4 only shows one GPU in the training node.

[0116] S301: Write the first training state data from the second memory on each training node into the first memory of each training node. The first training state data is obtained by training the neural network using the distributed training system.

[0117] This application analyzes the data types in the state dictionary, as shown in Figure 5. The data in the state dictionary is decomposed into tensor key-value data and training metadata. The tensor key-value data includes both keys and values. The values ​​include model parameters and optimizer states, while the keys include data loader states, etc. During the training of the neural network model in each training process of the distributed training system, most of the tensor key-value (KV) data values ​​are stored in the GPU's secondary memory, while a small portion of the tensor key-value data values, training metadata, and tensor key-value data keys are stored in the CPU's primary memory.

[0118] For ease of description, the values ​​of most of the tensor key-value data are referred to as the first training state data, the values ​​of a small portion of the tensor key-value data are referred to as the fourth training state data, and the keys of the tensor key-value data are referred to as the fifth data.

[0119] It should be noted that training metadata is used to indicate data in the state dictionary other than tensor key-value pairs. Therefore, training metadata can also be called non-tensor key-value pairs. For ease of description, the keys of tensor key-value data can be simply referred to as tensor keys, and the values ​​of tensor key-value data can be simply referred to as tensor data.

[0120] For example, training metadata is stored in the form of a dictionary, the keys of tensor key-value data are stored in the form of a list, and the values ​​of tensor key-value data are stored in the form of a list.

[0121] It should be noted that this application does not limit the storage format of training metadata, the keys of tensor key-value data, or the values ​​of tensor key-value data; the above is merely an illustrative example.

[0122] For example, the plurality of training nodes includes a first training node, which is any one of the plurality of training nodes. The following description uses the first training node and the first training process (worker) on the first training node as examples to illustrate this application.

[0123] For ease of description, the state dictionary slice maintained by the first training process will be referred to as the first state dictionary slice, and will not be described again hereafter.

[0124] For example, during the training of the neural network model in the first training process, a first state dictionary slice is obtained. The first training process stores the first training state data in the first state dictionary slice into the second memory of the GPU of the first training node, and stores the fourth training state data, the fifth data, and the training metadata in the first state dictionary slice into the first memory of the CPU of the first training node.

[0125] When checkpoints need to be stored, the first training process offloads the first training state data from the second memory of the first training node to the first memory, that is, writes the first training state data into the first memory. For example, as shown in "① offloading" in Figure 4, the first training state data in the second memory of GPU 0 is written into the first memory of CPU 0.

[0126] It should be noted that the operations performed by other training processes besides the first training process can be referenced from the operations performed by the first training process, and will not be repeated in this application.

[0127] Optionally, S301 includes: when each training process pauses training the neural network model, writing the first training state data of each training process into the first memory of the training node where each training process resides. The first training state data of each training process is used to indicate the first training state data obtained by each training process in training the neural network model. For example, the first training state data of the first training process is used to indicate the first training state data in the first state dictionary slice.

[0128] For example, during the training of the neural network model in the first training process, when it is necessary to save the checkpoint, the first training process pauses the training of the neural network model and writes the first training state data previously stored in the second memory into the first memory of the first training node.

[0129] In this embodiment, by setting each training process to pause training of the neural network model, the first training state data of each training process is written into the first memory. This ensures that the first training state data of each training process is data from the same moment, thereby helping to ensure the temporal consistency of the training state of each training process.

[0130] Optionally, the checkpoint management method provided in this application further includes: after writing the first training state data of each training process into the first memory of the training node where each training process is located, each training process continues to train the neural network model.

[0131] For example, after the first training process writes the first training state data previously stored in the second memory into the first memory of the first training node, the first training process cancels the operation of pausing the training of the neural network model and continues to train the neural network model.

[0132] In this embodiment, after writing the first training state data stored in the second memory into the first memory of the first training node, each training process continues to train the neural network model. This helps to shorten the time spent training the neural network model during the checkpoint saving process, thereby helping to ensure the throughput during the neural network model training process, and thus helping to ensure the performance of the neural network model training. Shortening the time spent training the neural network model during the checkpoint saving process helps to increase the checkpoint saving frequency while ensuring training throughput, thereby enabling the distributed training system to be restored to a more recent training state in the event of a failure.

[0133] Optionally, S301 includes: the distributed training system writing the first training state data in the second memory of each training node into the first memory of each training node according to the first cycle.

[0134] For example, a distributed training system can write the first training state data in the second memory of each training node into the first memory of each training node according to a pre-set first cycle, thereby realizing the saving of checkpoints according to the first cycle.

[0135] It should be noted that this application does not impose any restrictions on the specific value of the first cycle. For example, the first cycle can be set in conjunction with the training throughput requirements to ensure that the throughput (i.e., training throughput) of the distributed training system when training the neural network model can meet the target.

[0136] In this embodiment, by setting the first training state data in the second memory to be written into the first memory according to the first cycle, checkpoints can be saved according to the first cycle, which helps to restore the training state to a time point closer to the time of the failure when a failure occurs.

[0137] S302: Perform erasure coding on the second training state data in the first memory of each training node to obtain N checkpoint data.

[0138] The second training state data is the data obtained by training the neural network in the distributed training system, and it includes the first training state data. The N checkpoint data includes K raw data and M verification data. The K raw data includes the second training state data in the first memory of each training node. Any K checkpoint data from the N checkpoint data are used to obtain the second training state data in the first memory of each training node.

[0139] M is used to characterize the fault tolerance capability of the distributed training system. The larger the value of M, the stronger the fault tolerance capability of the distributed training system, and the more training nodes that can fail at the same time.

[0140] In one example, all training nodes across multiple training nodes are used to store N checkpoint data. The distributed training system receives initial setup information from the user, which includes the value of M. K equals NM.

[0141] In another example, a subset of the training nodes from a plurality of training nodes are used to store N checkpoint data. The distributed training system receives second setting information from the user, wherein the second setting information includes the values ​​of M and K.

[0142] For example, each training process performs erasure coding on the second training state data of each training process to obtain N checkpoint data.

[0143] Optionally, the second training state data includes various cases, which are illustrated below using cases 1 and 2.

[0144] In scenario 1, the second training state data includes the values ​​of the tensor key-value data. That is, the second training state data includes the first and fourth training state data mentioned above.

[0145] For example, as shown in Figure 5, an encoding operation is performed on the value of the tensor key-value data to obtain K original data and M check data.

[0146] This application analyzes the storage status of various data types in the state dictionary, revealing that the state dictionary is not stored in contiguous memory. Instead, most data (such as model parameters and optimizer states) is stored in GPU memory, while a smaller portion (such as data loader states and training metadata) is stored in CPU memory. Based on this, according to the encoding principles of related technologies, the state dictionary needs to be serialized into a continuous byte sequence before erasure coding is performed on this sequence. This serialization process incurs significant overhead, and because it cannot be piped with subsequent erasure coding operations (such as encoding and XOR operations), erasure coding can only be performed after serialization, increasing the time spent on storage checkpoints. However, this application, through analysis of the storage status of various data types in the state dictionary, finds that the first training state data (i.e., the values ​​of most tensor key-value pairs) is stored contiguously in the second memory of the GPU, and the fourth training state data (i.e., the values ​​of a small portion of tensor key-value pairs) is also stored contiguously in the first memory of the CPU. Based on this, this application can perform erasure coding operations only on the values ​​of tensor key-value data. Since the values ​​of tensor key-value data are stored contiguously, this application does not need to perform serialization operations on the first and fourth training state data before performing erasure coding operations on the values ​​of tensor key-value data. This not only helps to eliminate the resources consumed by serialization operations, but also helps to perform erasure coding operations in a pipelined manner, thereby helping to shorten the time spent storing checkpoints, and further helping to reduce the total time overhead when storing checkpoints and increase the storage frequency of checkpoints.

[0147] In addition, since tensor key-value data accounts for a large proportion of the total data volume of the state dictionary, for example, in the GPT2 345M model, the tensor key-value data exceeds 6.5GB and accounts for 99.999% of the total data volume of the state dictionary, performing erasure coding on the tensor key-value data generation can ensure the fault tolerance of most of the data in the state dictionary.

[0148] In this scenario, by performing erasure coding only on the values ​​of the tensor key data, a non-serialization coding protocol is used to achieve non-serialization encoding, which helps reduce the overhead of storing checkpoints, increase the frequency of storing checkpoints, and improve the efficiency of restoring the training state. Furthermore, when subsequently restoring the training state of the neural network model through checkpoints, a non-serialization decoding protocol can be used to achieve non-serialization decoding, further reducing the overhead of restoring the training state and improving its efficiency.

[0149] Case 2: In addition to the values ​​of the tensor key-value data, the second training state data also includes at least one of the keys of the tensor key-value data or training metadata.

[0150] In this case, it helps to ensure the integrity of the generated verification data by setting the values ​​of the tensor key-value data and performing erasure coding operations on at least one of the keys of the tensor key-value data or training metadata.

[0151] It should be noted that this application does not limit the range of data included in the second training state data; the above is merely an illustrative example.

[0152] Optionally, the checkpoint management method provided in this application further includes: storing the third training state data in the first memory of the first training node into the first memory of the second training node.

[0153] The second training node includes at least some of the training nodes other than the first training node, and the third training state data includes the training metadata and tensor key-value data obtained by the distributed training system from training the neural network.

[0154] In this application, at least part may be all, or it may be a part.

[0155] For example, when there are N training nodes, the second training node may include all training nodes except the first training node (i.e., N-1 training nodes), or the second training node may include some training nodes except the first training node (i.e., Np training nodes), where p is a positive integer greater than 1 and less than N.

[0156] For example, the first training process performs a serialization operation on its third training state data to obtain a first serialization result. Then, the first training process performs a broadcast operation to send the first serialization result to the second training node, thereby storing the third training state data in the first memory of the first training process into the first memory of the second training node.

[0157] For example, as shown in Figure 5, serialization and broadcast operations are performed on the keys of training metadata and tensor key-value data to distribute the keys of training metadata and tensor key-value data of each training process across various training nodes.

[0158] In this embodiment, the third training state data in the first memory of the first training node is stored in the first memory of the second training node, thereby achieving multiple backups of the third training state data in the first memory of the first training node. In this way, as long as some training nodes do not fail, the third training state data of each training node can be obtained, thus improving the fault tolerance of the third training state data. Furthermore, since the keys of tensor key-value data and training metadata account for a small proportion of the total data volume of the state dictionary—for example, in the GPT2345M model, the proportion of training metadata and tensor key-value data keys is 52KB (less than 0.001% of the total data volume)—storing the keys of tensor key-value data and training metadata from each training node to other training nodes for backup incurs low communication and computational resource overhead, thus helping to avoid affecting the training of the neural network model.

[0159] Optionally, S302 includes: during the process of training a neural network model through multiple training nodes in a distributed training system, performing erasure coding on the second training state data in the first memory of each training node using the CPU resources of each training node to obtain N checkpoint data.

[0160] Since the second training state data is stored in the CPU's memory, each training process can utilize the CPU resources on its respective training node to perform erasure coding on the second training state data. For example, during the process of the first training process using the GPU resources on the first training node to train the neural network model, it can utilize the CPU resources on the first training node to perform erasure coding on the second training state data in the first memory. For instance, as shown in "②encoding" in Figure 4, the training process on training node 0 utilizes the resources of CPU 0 to perform erasure coding on the second training state data in the first memory of CPU 0.

[0161] In this embodiment, by setting up the call to CPU resources to perform erasure coding operations on the data in the CPU's first memory, the erasure coding operations can be performed in parallel while calling GPU resources to train the neural network model. This helps to reduce the time spent training the neural network model when storing checkpoints, thereby helping to reduce the impact of the storage checkpoint process on training throughput.

[0162] The following describes one implementation of erasure coding operation through S302a-S302c.

[0163] S302a: Based on the encoding matrix of each training process in multiple training processes, perform encoding operations on the second training state data of each training process to obtain M encoding blocks for each training process.

[0164] For example, the first training process performs an encoding operation on the second training state data of the first training process according to the encoding matrix of the first training process, and obtains M encoding blocks of the first training process.

[0165] Example 1a, as shown in Figure 6, includes multiple training nodes: training node 0, training node 1, training node 2, and training node 3. Training process 0 is deployed on training node 0, training process 1 is deployed on training node 1, training process 2 is deployed on training node 2, and training process 3 is deployed on training node 3. During the training of the neural network model, training process 0 obtains second training state data d0, ..., and training process 3 obtains second training state data d3. With M equal to 2, encoding operations are performed on the second training state data of each training process to obtain two encoded blocks for each training process. Training process 0 performs encoding operations on d0 according to its encoding matrix to obtain e. 11 d0 and e 21 d0, training process 1 performs an encoding operation on d1 based on the encoding matrix of training process 1 to obtain e. 11 d1 and e 21 d1, training process 2 performs encoding operations on d2 based on the encoding matrix of training process 2 to obtain e. 12 d2 and e 22 d2, training process 3 performs encoding operations on d3 based on the encoding matrix of training process 3 to obtain e. 12 d3 and e 22 d3.

[0166] It should be noted that the second training state data of a training process can be considered as a data segment of all the second training state data in the state dictionary.

[0167] Example 1b, as shown in Figure 7, includes multiple training nodes: training node 0, training node 1, and training node 2. Training node 0 hosts training process 0 (referred to as process 0) and training process 1 (referred to as process 1), training node 2 hosts training process 2 (referred to as process 2) and training process 3 (referred to as process 3), and training node 3 hosts training process 4 (referred to as process 4) and training process 5 (referred to as process 5). During the training of the neural network model, training process 0 obtains second training state data d0, ..., and training process 5 obtains second training state data d5. For example, M equals 1; therefore, an encoding operation is performed on the second training state data of each training process to obtain one encoding block for each training process. Training process 0 performs an encoding operation on d0 based on its encoding matrix to obtain e. 11 d0, training process 1 performs an encoding operation on d1 based on the encoding matrix of training process 1 to obtain e. 11 d1, training process 2 performs encoding operations on d2 based on the encoding matrix of training process 2 to obtain e. 11 d2, training process 3 performs encoding operations on d3 based on the encoding matrix of training process 3 to obtain e. 12 d3, training process 4 performs encoding operations on d4 based on the encoding matrix of training process 4 to obtain e. 12 d4, training process 5 performs an encoding operation on d5 based on the encoding matrix of training process 5 to obtain e. 12 d5.

[0168] Optionally, S302a includes: performing encoding operations in parallel on the second training state data of each training process according to multiple threads of each training process, to obtain M encoding blocks for each training process.

[0169] For example, the first training process acquires multiple threads. Then, the multiple threads perform encoding operations in parallel on the second training state data of the first training process to obtain M encoded blocks of the first training process.

[0170] In one example, a thread pool for the first training process is deployed on the first training node, and the thread pool includes multiple threads of the first training process. When the first training process needs to perform encoding operations, it can obtain multiple threads from the thread pool. This helps improve the convenience and speed of the first training process obtaining multiple threads.

[0171] In another example, when the first training process needs to perform encoding operations, it creates multiple threads in real time. This helps avoid consuming excessive computing resources.

[0172] In this embodiment, encoding operations are performed in parallel by multiple threads, which helps to improve encoding efficiency and thus helps to reduce the time overhead of encoding operations.

[0173] The following section provides an exemplary description of the process for determining the encoding matrix for each training process, using steps S1-S3.

[0174] S1: Divide multiple logical blocks of multiple training processes into K logical segments. One logical block is used to indicate the second training state data obtained by a training process when training a neural network model.

[0175] For example, the multiple logic blocks include a first logic block, which is used to indicate the second training state data obtained by the first training process in training the neural network model. The first training process divides the multiple logic blocks of the multiple training processes into K logic segments.

[0176] Example 2a, referring to Figure 6, shows four logical blocks for the four training processes, d0', ..., d3', where d0' indicates d0, ..., d3' indicates d3. For example, when K equals 2, training process 0 divides the four logical blocks into two logical segments, D0' and D1', where D0' = [d′0 d′1] and D1' = [d′2 d′3].

[0177] Example 2b, referring to Figure 7, the six logical blocks of the six training processes are d0', ..., d5', where d0' is used to indicate d0, ..., d5' is used to indicate d5. For example, K equals 2, training process 0 divides the six logical blocks into two logical segments, namely D0' and D1', where D0' = [d′0 d′1 d′2] = , and D1' = [d′3 d′4 d′5].

[0178] S2: Based on the erasure coding matrix, perform encoding operations on K logical segments to obtain M logical check data, where the M logical check data are used to indicate the M check data.

[0179] For example, the first training process generates an erasure coding matrix based on the values ​​of K and N. This coding matrix is ​​an N*K matrix, where any K rows are linearly independent. Then, the first training process performs encoding operations on K logical segments based on the erasure coding matrix; that is, it multiplies the K logical segments by the coding matrix to obtain an encoding result matrix. This encoding result matrix includes N logical data points, where the N logical data points include K original logical data points and M parity logical data points. The K original logical data points indicate the K original logical data points, and the M parity logical data points indicate the M parity logical data points.

[0180] Example 3a, referring to Figure 6, the encoded result matrix satisfies the following relationship (1):

[0181] In this context, RE1 represents the encoding result matrix, and E1 represents the encoding matrix. D0' represents the original logical data, and D1' represents the original logical data. P0 represents the logical check data, and P1 represents the logical check data.

[0182] Example 3b, referring to Figure 7, the encoded result matrix satisfies the following relationship (2):

[0183] It should be noted that RE2 is used to represent the encoding result matrix, and E2 is used to represent the encoding matrix. D0' is used to represent the original logical data, and D1' is used to represent the original logical data. P0 is used to represent the logical verification data.

[0184] S3: Based on M logical verification data, determine the encoding matrix of each training process in multiple training processes.

[0185] For example, each training process determines its encoding matrix based on M logical check data. For instance, the first training process determines its encoding matrix based on the M logical check data.

[0186] Example 4a, referring to Figure 6, training process 0 determines the encoding matrix based on P0 and P1. Based on the same principle, the encoding matrix for training process 1 is: The encoding matrix for training process 2 is: The encoding matrix for training process 3 is:

[0187] Example 4b, referring to Figure 7, based on P0, the encoding matrix of training process 0 is determined to be [e 11 Based on the same principle, the encoding matrix [e] of training process 1 11 The encoding matrix of training process 2 [e] 11 The encoding matrix of training process 3 [e] 12 The encoding matrix of training process 4 [e] 12 The encoding matrix of training process 5 [e] 12 ].

[0188] In the above embodiments, determining the encoding matrix of each training process through the global encoding result matrix helps improve the accuracy and reliability of the final M check data. Furthermore, representing the second training state data of the training process through logical blocks, compared to actually obtaining the second training state data of other training processes, not only helps reduce the communication resources consumed when performing erasure coding operations, but also helps improve the execution efficiency of erasure coding operations, thereby helping to reduce the time spent storing checkpoints.

[0189] It should be noted that this application does not limit the method for determining the encoding matrix of each training process; the above is merely an illustrative example. For instance, the encoding matrix of each training process can also be determined using actual second training state data.

[0190] The following section uses the training process in the first XOR group as an example to illustrate one implementation method through S302a1-S302a2.

[0191] S302a1: Write the second training state data of each training process in the first XOR group into the data buffer of each training process in the first XOR group to obtain each data block (data packet) of each training process in the first XOR group.

[0192] For example, the first XOR group includes a first training process. The first training process requests data buffers from the CPU's first memory. These data buffers, referred to as the first training process's data buffers, are used to store each data block to be encoded. Then, the first training process writes its second training state data into the data buffers. After the first training process's data buffers are full, one data block of the first training process is obtained. By sequentially writing the second training state data of the first training process into the first training process's data buffers, each data block to be encoded by the first training process is obtained.

[0193] Example A1 can be obtained by writing the second training state data of the first training process into a data buffer, splitting the second training state data of the first training process into M data blocks, and then performing encoding operations on the M data blocks to obtain M encoded blocks.

[0194] Example A2 can be achieved by writing the second training state data of the first training process into a data buffer, splitting the second training state data of the first training process into G*M data blocks, where G is a positive integer greater than 1, and then performing encoding operations on the G*M data blocks to obtain G*M encoded sub-blocks, where G encoded sub-blocks constitute 1 encoded block.

[0195] The following description uses Example A1 as an example to illustrate this application.

[0196] Optionally, the first training process has at least one data buffer. This helps to increase the flexibility in the number of data buffers.

[0197] In one example, the first training process has a data buffer. The first training process can write the second training state data into a data buffer. This helps reduce the storage space occupied by the first memory.

[0198] In another example, the second training process has multiple data buffers. That is, the first training process can write the second training state data into multiple data buffers. This helps improve the efficiency of splitting the second training state data into blocks to be encoded, thereby helping to reduce the time consumed by performing the encoding operation.

[0199] S302a2: Based on the encoding matrix of each training process in the first XOR group, perform encoding operations on each data block of each training process in the first XOR group to obtain each encoding block of each training process in the first XOR group.

[0200] For example, the first training process performs encoding operations on each data block of the first training process according to the encoding matrix of the first training process, thereby obtaining each encoding block of the first training process.

[0201] Example 5a, referring to Figure 6, shows that the two encoded blocks obtained from training process 0 include e 11 d0 and e 21 d0. The two encoded blocks obtained from training process 1 include e 11 d1 and e 21 d1. The two encoded blocks obtained from training process 2 include e 12 d2 and e 22 d2. The two encoded blocks obtained from training process 3 include e 12 d3 and e 22 d3.

[0202] Example 5b, referring to Figure 7, shows that the encoding block obtained by training process 0 includes e. 11 d0. The encoded block obtained from training process 1 includes e 11 d1. The encoded block obtained from training process 2 includes e 11 d2. The encoded block obtained from training process 3 includes e 12 d3. The encoding block obtained from training process 4 includes e 12 d4. The encoded block obtained from training process 5 includes e 12 d5.

[0203] In this embodiment, by writing the second training state data of each training process into a data buffer, multiple data blocks are obtained for each training process. An encoding operation is then performed on each of these data blocks, thereby decomposing the encoding task of each training process into multiple sub-encoding tasks. Each sub-encoding task corresponds to one data block in the data buffer. Thus, by executing multiple sub-encoding tasks, M encoding blocks are obtained. After obtaining an encoding block, an XOR operation can be performed on it, allowing for pipelined execution of encoding and XOR operations. This helps improve the execution efficiency of erasure coding operations and reduce their time overhead.

[0204] Furthermore, by writing the second training state data into the data buffer, the second training state data can be split into multiple data blocks of the same size, thereby enabling different training processes to perform encoding operations on data blocks of the same size to obtain encoded blocks of the same size, and thus enabling different training processes to perform XOR operations on encoded blocks of the same size.

[0205] Optionally, S302a2 includes the following S4-S5. Performing the encoding operation through the scheme of S4-S5 helps to reduce the time overhead of the encoding operation.

[0206] S4: Get multiple threads for each training process in the first XOR group.

[0207] For example, the first training process acquires multiple threads to enable the encoding operations to be performed in parallel by multiple threads.

[0208] It should be noted that the relevant explanations of the multiple threads in S4 can be found in the descriptions in the above embodiments, and will not be repeated here.

[0209] S5: Encode each data block of each training process in the first XOR group in parallel using multiple threads of each training process in the first XOR group.

[0210] For example, after the first training process obtains multiple threads, it performs encoding operations on each data block of the first training process in parallel through the multiple threads, thereby obtaining each encoded block of the first training process.

[0211] In this embodiment, by assigning each sub-encoding task to multiple threads for simultaneous execution, the multiple cores of the CPU can be fully utilized to accelerate the encoding operation, thereby helping to reduce the time overhead of the encoding operation.

[0212] S302b: Determine R XOR groups for multiple training processes, where each XOR group includes K training processes.

[0213] In one example, each training process determines XOR groups based on M logical parity data in the encoding result matrix, resulting in R XOR groups. Erasure coding is then performed on a per-XOR-group basis. For instance, the first training process encodes the M logical parity data in the encoding result matrix to determine the XOR group to which the first XOR group belongs. Alternatively, different training processes can be assigned the same XOR group based on an element in the encoding result matrix, resulting in R XOR groups for M coded blocks.

[0214] For example, if multiple training processes include W training processes, then R equals W / K.

[0215] Example 6a, referring to Figure 6, based on e in RE1 11 d'0+e 12 d'2, determine the e of training process 0 11 d0 and training process 2's e 12 d2 is the XOR group 0. Based on e in RE1 21 d'0+e 22 d'2, determine the e of training process 0 21 d0 and training process 2's e 22 d2 is XOR group 1, based on e in RE1 11 d'1+e 12 d'3, determine e of training process 1 11 d1 and training process 3 e 12 d3 is XOR group 2. Based on e in RE1 21 d'1+e 22 d'3, determine e of training process 1 21 d1 and training process 3 e 22 d3 is the XOR group 3.

[0216] Example 6b, in conjunction with Figure 7, based on e in RE2 11 d'0+e 12 d'3, determine the training process 0 of e 11 d0 and training process 3's e 12 d3 is the XOR group 0. Based on e in RE2 11 d'1+e 12 d'4, determine e of training process 1 11 d1 and training process 4 of e 11 d4 is the XOR group 1, based on e in RE2. 11 d'2+e 12 d'5, determine e of training process 2 11 d2 and training process 5 of e 12 d5 is the XOR group 2.

[0217] In another example, the first training process determines multiple data interval groups for training processes, each data interval group comprising K data sub-interval groups. Then, the first training process groups training processes with the same index across different data sub-interval groups into the same XOR group. The index refers to the training process's subscript within the group, resulting in multiple XOR groups, and ultimately, the XOR group to which the first training process belongs.

[0218] For example, referring to Figure 6, training process 0 determines the data interval groups of the four training processes as [[0,1], [2,3]], where "0" represents training processes 0, ..., and "3" represents training process 3. [0,1] and [2,3] are data sub-interval groups. The operations performed by training processes 1, ..., and training process 3 are the same as those performed by training process 0, and will not be described again. Training process 0 divides the training processes with the same index in "[0,1]" and "[2,3]" into the same XOR group, resulting in two XOR groups: [0,2] and [1,3]. Training processes 0 and 2 belong to the same XOR group.

[0219] For example, referring to Figure 7, training process 0 determines the data interval groups of the six training processes as [[0,1,2], [3,4,5]], where "0" represents training processes 0, ..., and "5" represents training process 5. "[0,1,2]" and "[3,4,5]" are data sub-interval groups. The operations performed by training processes 1, ..., and training process 5 are the same as those performed by training process 0, and will not be elaborated further. Training process 0 divides the training processes with the same index in "[0,1,2]" and "[3,4,5]" into the same XOR group, resulting in three XOR groups: [0,3], [1,4], and [2,5]. Training processes 0 and 3 belong to the same XOR group.

[0220] It should be noted that this application does not limit the method of determining the XOR group; the above is merely an illustrative example.

[0221] It should be noted that this application does not restrict the execution order of S302a and S302b; the above is merely an illustrative example.

[0222] S302c: Perform an XOR operation on the encoding blocks of different training processes in each of the multiple XOR groups to obtain R*M parity blocks.

[0223] Among them, R*M check blocks contain M check data.

[0224] For example, the XOR operation is a type of reduction operation; therefore, the XOR operation can also be called the XOR reduction operation.

[0225] For example, the XOR operation can be performed on the encoding blocks of different training processes in each XOR group through communication between different training processes.

[0226] For example, the encoding blocks of different training processes in each XOR group undergo M XOR operations, where each XOR operation targets one encoding block from each training process in each XOR group, thereby obtaining M check blocks for each XOR group. After the XOR operation is completed, the resulting check blocks are stored in the first memory of the training node where the training process that performed the XOR operation is located.

[0227] Since each XOR group requires M XOR operations, and there are R XOR groups (i.e., W / K), saving a checkpoint requires R*M XOR operations. For example, Q training processes are deployed on each training node, where W = Q*N.

[0228] In one example, referring to Figure 6, training process 0's e to training process 0 11 d0 and training process 2's e 12 d2 performs an XOR operation to obtain parity block p0. Parity block p0 is stored in the first memory of training node 0 where training process 0 resides. Training process 2 performs an XOR operation on training process 0. 21 d0 and training process 2's e 22 d2 performs an XOR operation to obtain parity block p2. Parity block p2 is stored in the first memory of training node 2 where training process 2 resides. Training process 1 performs an XOR operation on training process 1. 11 d1 and training process 3 e 12 d3 performs an XOR operation to obtain parity block p1. Parity block p1 is stored in the first memory of training node 1, where training process 1 resides. Training process 3 performs an XOR operation on training process 1. 21 d1 and training process 3 e 22 d3 performs an XOR operation to obtain the parity block p3. The parity block p3 is stored in the first memory of training node 3, where training process 3 resides.

[0229] In another example, referring to Figure 7, training process 3 corresponds to the e of training process 0. 11 d0 and training process 3's e 12 d3 performs an XOR operation to obtain p0. p0 is stored in the first memory of training node 1 where training process 3 resides. Training process 1 performs an XOR operation on training process 1. 11 d1 and training process 4 of e 12 d4 performs an XOR operation to obtain p1. p1 is stored in the first memory of training node 0 where training process 1 resides. Training process 2 performs an XOR operation on training process 2. 11 d2 and training process 5 of e 12d5 performs an XOR operation to obtain p2. p2 is stored in the first memory of training node 1 where training process 2 resides.

[0230] In the above embodiments, M verification data are obtained through encoding and XOR operations, which optimizes the calculation of erasure coding operations, thereby simplifying the complexity of erasure coding operations, improving the efficiency of erasure coding operations, and thus achieving efficient acquisition of M verification data.

[0231] Optionally, the checkpoint management method provided in this application further includes: determining the XOR target process for each XOR group in a plurality of XOR groups.

[0232] In this context, the XOR target process of each XOR group is used to perform XOR operations on the encoding blocks of different training processes in each XOR group, and the training node where the XOR target process is located is used to store the verification block obtained by the XOR target process performing the XOR operation.

[0233] For example, the multiple XOR groups include a first XOR group, which includes a first training process. The first training process determines an XOR target process from the multiple training processes in the first XOR group, so that XOR operations can be performed on different training processes in the first XOR group through the XOR target process, thereby obtaining M parity blocks of the first XOR group. Furthermore, the parity blocks obtained by the XOR target process can be stored in the first memory of the training node where the XOR target process is located, which helps to reduce the communication resources occupied during the XOR operation.

[0234] In this embodiment, by setting the verification block to be stored in the first memory of the training node where the target XOR process is located, it helps to avoid the communication resources occupied by the process of storing the verification block, thereby helping to reduce the total communication resources occupied when performing the XOR operation.

[0235] In this application, the XOR target process includes several cases, which are illustrated below through cases A and B.

[0236] Case A: When the M parity blocks of each XOR group are stored on the training node where the XOR target process is located, the number of communications from the M parity data storage to the M parity nodes satisfies the target condition.

[0237] In this scenario, setting appropriate target conditions helps reduce the number of communications required for subsequent data transmission operations. This not only reduces the communication resources consumed by these operations but also decreases their latency, ultimately reducing the time required for storage checkpointing. The data transmission operation involves storing M parity data into M parity nodes.

[0238] Optionally, the communication count of M verification data stored on M verification nodes satisfies the target condition as follows: when the M verification blocks of each XOR group are stored on the training node where the XOR target process resides, the communication count of M verification data stored on M verification nodes is less than the communication count of M verification data stored on M verification nodes when the M verification blocks of each XOR group are stored on the training node where the fourth training process in each XOR group resides. The fourth training process includes every training process in each XOR group except the XOR target process. In other words, storing the M verification blocks of each XOR group on the training node where the XOR target process in each XOR group minimizes the communication count of M verification data stored on M verification nodes.

[0239] In this embodiment, by setting each XOR group to perform the XOR operation through the XOR target process, the number of communication times for storing M parity data to M parity nodes is minimized, thereby helping to reduce the number of communication times for subsequent data transmission operations. This not only helps to reduce the communication resources occupied by subsequent data transmission operations, but also helps to reduce the time consumed by subsequent data transmission operations.

[0240] Optionally, the target condition for the number of communications between M verification data stored in M ​​verification nodes to meet the target condition includes: the number of communications between M verification data stored in M ​​verification nodes is less than or equal to the number threshold.

[0241] It should be noted that this application does not impose a specific limit on the threshold value for the number of attempts. For example, it can be dynamically set according to actual scenarios.

[0242] It should be noted that when the A training processes (A is a positive integer greater than 1) in the first XOR group are used as the XOR target process, the number of communications is less than or equal to the number threshold. Therefore, the XOR target process can be any one of the A training processes.

[0243] In this embodiment, the target condition is met when the number of communications is less than or equal to a threshold number. By setting an appropriate threshold number, it is not only helpful to control the time consumption of subsequent data transmission processes, but also to improve the diversity of XOR target process selection.

[0244] In this application, the XOR target process that satisfies the target conditions includes multiple cases. Hereinafter, the first XOR group is used as an example, and cases A1 and A2 are introduced.

[0245] In case A1, where the first XOR group includes training processes on M verification nodes, the target XOR process also includes training processes on M verification nodes. For ease of distinction, the training processes on the M verification nodes will be referred to as verification processes below.

[0246] In one example, referring to Figure 7, for the XOR group [2,5], that is, the XOR group consisting of training process 2 and training process 5, when performing the XOR operation on d2 of training process 2 and d5 of training process 5, when training process 2 is determined to be the target XOR process, and the obtained verification block p2 is stored in training node 1 where training process 2 is located, the number of subsequent data transmissions is 3.

[0247] In another example, referring to Figure 8, when training process 5 is determined to be the XOR target process, and the resulting verification block p2 is stored in training node 3 where training process 5 resides, the subsequent data transmission communication will be 4 times. Therefore, by selecting the verification process as the XOR target process, the number of subsequent data transmission communications can be reduced. Other relevant explanations in Figure 8 can be found in Figure 7, and will not be repeated here.

[0248] For example, in a scenario where there are K training nodes among multiple training nodes, each training node includes Q training processes, and the multiple training nodes together include W training processes, there are a total of Q XOR groups containing the verification process. These Q XOR groups can select the verification process as the XOR target process, thereby avoiding additional communication operations during data transmission.

[0249] Example 1: The first XOR group includes M verification processes on M verification nodes, where the XOR target process includes M verification processes. For example, as shown in Figure 7, for the XOR group [2,5], when training node 1 is a verification node, since training process 2 is a verification process on the verification node and training process 5 is a training process on the data node, training process 2 is determined to be the XOR target process of XOR group [2,5].

[0250] In this example, each XOR operation of the first XOR group can be executed by different XOR target processes, which not only helps to balance the computing and communication resources consumed by each XOR target process, but also helps to balance the storage pressure of each training node storing the verification block.

[0251] It should be noted that when the first XOR group includes M verification processes, the XOR target process may also include only some of the M verification processes. This helps to improve the diversity of the selection method for the XOR target process.

[0252] Example 2: The first XOR group includes K verification processes on M verification nodes, where the XOR target processes include M of the K verification processes. In this way, each XOR operation in the first XOR group can be executed through different XOR target processes, which not only helps to balance the computing and communication resources consumed by each XOR target process, but also helps to balance the storage pressure on each training node storing verification blocks.

[0253] It should be noted that other relevant explanations for Example 2 can be found in the explanations for Example 1 above, and will not be repeated here.

[0254] Example 3: The first XOR group includes Y verification processes on M verification nodes, where Y is a positive integer less than M. The target XOR process includes Y verification processes. This helps to make full use of each verification process.

[0255] It should be noted that other relevant explanations for Example 3 can be found in the explanations of Example 1 and Example 2 above, and will not be repeated here.

[0256] In this scenario, by selecting the training processes on M verification nodes as the target XOR process, the verification block can be directly stored on the verification nodes, which helps reduce the number of subsequent data transmissions and thus reduces the communication resources occupied by subsequent data transmissions.

[0257] Case A2: The first XOR group does not include the training nodes on the verification node. The XOR target process includes the training processes on K data nodes, wherein the difference in the number of times the training processes on different data nodes are XOR target processes is less than or equal to a difference threshold.

[0258] It should be noted that this application does not impose any restrictions on the specific value of the difference threshold, which can be dynamically set according to the actual scenario.

[0259] For ease of distinction, the training process on K data nodes will be referred to as the data process below.

[0260] For example, for W / KG XOR groups that do not contain a verification process, the XOR target process for each XOR group can be selected from the K data processes of each XOR group.

[0261] For example, when determining the target XOR process from K data processes, the difference in the number of times the training processes on different data nodes are selected as the target XOR process is less than or equal to a difference threshold, thereby balancing the workload of XOR operations performed on each training node and the storage space occupied by the storage verification block. For instance, the target XOR process can be selected based on the relationship between K and M.

[0262] Example 4: When K equals M, each of the K data processes can be selected as the XOR target process. This way, each XOR operation can be executed through a different XOR target process, which not only helps to balance the computational and communication resources consumed by each XOR target process, but also helps to balance the storage pressure on the verification blocks of each training node.

[0263] Example 5: If K is greater than M, M data processes out of K data processes can be selected as the target XOR processes. Not all KM data processes need to be the target XOR processes. For example, as shown in Figure 6, for the XOR group [1,4], training process 1 and training process 4 are both data processes on data nodes, and K is 2 and M is 1. Therefore, either training process 1 or training process 4 can be selected as the target XOR process, such as selecting training process 1 as the target XOR process.

[0264] In this example, each XOR operation can be performed through different XOR target processes, which not only helps to balance the computing and communication resources consumed by each XOR target process, but also helps to balance the storage pressure of storing verification blocks on each training node.

[0265] For example, the target XOR process can be determined for K data processes in a K / M interval order. That is, M parity blocks are allocated to the M data processes out of the K data processes in a K / M interval order. This helps to select data processes on different data nodes as target XOR processes, thereby not only helping to balance the computing and communication resources consumed by each XOR target process, but also helping to balance the storage pressure of storing parity blocks on each training node.

[0266] It should be noted that other related explanations for Example 5 can be found in the explanations for Example 4 above, and will not be repeated here.

[0267] Example 6: If K is less than M, K data processes can be selected as the target XOR processes. At least some of these K data processes (e.g., MK processes) can be selected as the target XOR processes multiple times. This helps to make full use of each data process.

[0268] For example, the target XOR process for each XOR operation can be determined by cyclically among K data processes. That is, M parity blocks are allocated to the K data processes in a cyclic manner.

[0269] It should be noted that other related explanations for Example 6 can be found in the explanations of Examples 4 and 5 above, and will not be repeated here.

[0270] It should be noted that other relevant explanations for situation A2 can be found in the explanations for situation A1 above, and will not be repeated here.

[0271] In case B, the target process for each XOR operation can be any training process in the first XOR group. The target process is the training process that performs the XOR operation.

[0272] In one example, the target processes for different XOR operations can be the same. In another example, the target processes for different XOR operations can be different.

[0273] In this approach, any training process can be set as the XOR target process, which helps to improve the flexibility and diversity of selecting the XOR target process.

[0274] It should be noted that other relevant explanations for situation B can be found in the explanations for situation A above, and will not be repeated here.

[0275] It should be noted that this application does not limit the implementation method of erasure coding operations; the above is merely an illustrative example. For example, it can also be implemented using Reed-Solomon coding, low-density parity-checking (LDPC), or other methods.

[0276] Optionally, S302c includes the following S302c1-S302c2.

[0277] S302c1: Write each encoding block of each training process in the first XOR group into the encoding buffer of each training process in the first XOR group.

[0278] For example, the first training process requests encoding buffers from the CPU's first memory. These encoding buffers can be referred to as the encoding buffers of the first training process, and are used to store the encoded blocks of the first training process. After the first training process performs the encoding operation on each data block and obtains each encoded block of the first training process, it writes each encoded block of the first training process into the encoding buffer of the first training process.

[0279] S302c2: Perform an XOR operation on the encoding blocks in the encoding buffers of different training processes in the first XOR group.

[0280] For example, the second training state data of each training process in the first XOR group is converted into encoded blocks through an encoding operation and stored in the encoding buffer of each training process. Then, the target XOR process of the first XOR group performs an XOR operation on the encoded blocks of the encoding buffers of different training processes. For instance, after the encoding buffer of each training process in the first XOR group is full, resulting in one encoded block for each training process, the encoded blocks of different training processes in the first XOR group are then XORed.

[0281] In the above embodiments, by writing the encoded block into the encoding buffer and performing an XOR operation on the encoded block in the encoding buffer, the encoding operation of the data block and the XOR operation of the encoded block can be executed in a pipeline manner, which helps to improve the speed of erasure coding operation and reduce the time consumption of erasure coding operation.

[0282] Optionally, the number of encoding buffers in each training process is M*F. Here, F represents the number of data buffers in each training process. For example, if the first training process has one data buffer, then the first training process has M encoding buffers. Thus, after obtaining each encoding block of the first training process, by writing each encoding block into the M encoding buffers, M encoding blocks of the second training state of the first training process can be obtained.

[0283] In this embodiment, by setting the number of encoding buffers for each training process to M*F, it helps to improve the convenience and accuracy of obtaining M encoding blocks for each training process.

[0284] Optionally, S302c includes: when different training nodes of different training processes are in a network idle state in the first XOR group of multiple XOR groups, performing an XOR operation on the encoding blocks of different training processes in the first XOR group.

[0285] The "network idle state" between different training nodes can refer to the communication gap between them. This communication gap refers to the communication gap during the training process of the neural network model on different training nodes. For example, the communication gap between different training nodes could be the communication gap between different GPUs on different training nodes. Alternatively, it could be the communication gap between different training processes on different training nodes.

[0286] For example, as shown in Figure 7, for the XOR group [2,5], training process 2 is located at training node 1, and training process 5 is located at training node 2. Therefore, when training nodes 1 and 2 are in a network idle state, an XOR operation is performed on d2 of training process 2 and d5 of training process 5. For example, training processes 2 and 5 communicate during the simultaneous training of the neural network model. Based on this, when training processes 2 and 5 are in a network idle state, that is, when training processes 2 and 5 are not communicating about the operation of training the neural network model, an XOR operation is performed on d2 of training process 2 and d5 of training process 5. For example, by analyzing the first preset number of iterations (e.g., 50 times) in the process of training the neural network model, the network idle time period (i.e., the time period in a network idle state) in the process of training the neural network model can be determined, thereby realizing that the XOR operation is only performed during the network idle time period.

[0287] In this embodiment, by setting the XOR operation to be performed when different training nodes are in a network idle state, it helps to avoid the XOR operation from competing for network resources with the operation of training the neural network model, thereby helping to avoid affecting the communication during the training process of the neural network model, and thus helping to avoid affecting the training throughput.

[0288] In the above embodiment, each training process (worker) first encodes its own data packet to obtain m encoded packets. Then, in subsets (i.e., XOR groups) of different training processes, different encoded packets are XORed to obtain parity packets. Finally, the data packets and parity packets are communicated point-to-point (P2P) so that each training node stores its own original data or parity data.

[0289] The above examples illustrate the process of obtaining M verification data points from N checkpoint data points. The following describes the process of obtaining K raw data points from N checkpoint data points.

[0290] For example, the first training process determines K raw data based on K logical raw data. For instance, the first training data determines the second training state data indicated by each logical raw data as one raw data. Referring to Figure 6, based on d0' and d1' being one logical raw data, training process 0 determines d0 and d1 as one raw data. Based on d2' and d3' being one logical raw data, training process 0 determines d2 and d3 as one raw data, thus obtaining two raw data.

[0291] S303: Store N checkpoint data in the first memory of multiple training nodes, wherein different checkpoint data are stored in the first memory of different training nodes.

[0292] For example, as shown in “③ communication” in Figure 4, multiple training nodes include N training nodes. The distributed training system stores N checkpoint data in the first memory of different training nodes, such as the memory of CPU 0, ..., the memory of CPU N-1, by calling the CPU resources of the N training nodes, for example, calling the resources of CPU 0, ..., the resources of CPU N-1.

[0293] For example, during the process of storing N checkpoint data in the first memory of multiple training nodes, different training nodes can perform data transmission operations through point-to-point communication (P2P communication).

[0294] For ease of description, the operation of "storing N checkpoint data in the first memory of multiple training nodes" will be referred to as a data transfer operation.

[0295] After the erasure coding and data transmission operations are completed, the first memory of each training node's CPU stores the training metadata of the state dictionary, the keys of the tensor key-value data, and a portion of the values ​​of the tensor key-value data or at least a portion of the verification data from the M verification data.

[0296] Optionally, S303 includes the following S303a and S303b.

[0297] S303a: Determine K third training nodes from multiple training nodes as K data nodes, and determine M verification nodes from the training nodes other than the K third training nodes from multiple training nodes.

[0298] Specifically, when the third training node acts as a data node, the number of communication operations for storing K raw data to K data nodes and M parity data to M parity nodes is less than the number of communication operations for storing K raw data to K data nodes and M parity data to M parity nodes when the third training node acts as a parity node.

[0299] The following section provides an example of one implementation method for determining data nodes, using steps S6-S9.

[0300] S6: Obtain the physical interval group of multiple training processes. The physical interval group includes multiple physical sub-interval groups. The training processes in a physical sub-interval group are located on the same training node.

[0301] For example, the first training process obtains multiple physical sub-interval groups for multiple training processes based on the training node to which each training process belongs, thereby obtaining multiple physical interval groups for training processes. The number of multiple physical sub-interval groups is equal to the number of training nodes. The physical interval groups represent physical groups between training processes across nodes in a distributed training environment; therefore, the physical interval groups can be considered as physical groups of training processes. For example, as shown in Figure 7, the physical interval groups are [[0,1], [2,3], [4,5]], where the multiple physical sub-interval groups include [0,1], [2,3], and [4,5], and [0,1] indicates all training processes on training node 0.

[0302] S7: Obtain data interval groups for multiple training processes, where each data interval group includes K data sub-interval groups.

[0303] The data interval group is a logical grouping of all training processes based on K. Each data sub-interval group is used to indicate the training process to which the original data to be stored for each data node belongs, that is, to indicate the range of data to be stored for each data node. For example, as shown in Figure 7, the data interval group includes [[0,1,2], [3,4,5]].

[0304] S8: Determine the first physical sub-margin group from multiple physical sub-margin groups. The number of identical training processes between the first physical sub-margin group and the target data sub-margin group in the K data sub-margin groups is greater than the number of identical training processes between the second physical sub-margin group and the target data sub-margin group.

[0305] The second physical sub-spacer group includes each physical sub-spacer group other than the first physical sub-spacer group.

[0306] For example, the K data sub-interval groups include a target data sub-interval group, wherein the target data sub-interval group is any one of the K data sub-interval groups. The following description uses the target data sub-interval group as an example to illustrate S8.

[0307] For example, the first training process determines the number of training processes that the target data sub-interval group shares with each physical sub-interval group. Then, based on the number corresponding to each physical sub-interval, the first physical sub-interval group is determined from multiple physical sub-interval groups. For example, as shown in 7, the number of training processes that the data sub-interval group [0,1,2] shares with the physical sub-interval group [0,1] is 2, the number of training processes that the data sub-interval group [0,1,2] shares with the physical sub-interval group [2,3] is 1, and the number of training processes that the data sub-interval group [0,1,2] shares with the physical sub-interval group [4,5] is 0. Therefore, the physical sub-interval group [0,1] has a number of 2, which is the physical sub-interval group with the most training processes that it shares with the data sub-interval group [0,1,2] (i.e., it has the maximum overlap). Therefore, the physical sub-interval group [0,1] is the first physical sub-interval group of [0,1,2]. Since the index (i.e., subscript) of [0,1] in the physical interval group is 0, training node 0 is selected as a data node.

[0308] The number of training processes that data sub-margin [3,4,5] shares with physical sub-margin [0,1] is 0. The number of training processes that data sub-margin [3,4,5] shares with physical sub-margin [2,3] is 1. The number of training processes that data sub-margin [3,4,5] shares with physical sub-margin [4,5] is 2. Therefore, the physical sub-margin [4,5] has 2, which is the physical sub-margin that shares the most training processes with data sub-margin [3,4,5] (i.e., has the maximum overlap). Therefore, physical sub-margin [4,5] is the first physical sub-margin of [3,4,5].

[0309] S9: Determine the training node where the training process in the first physical sub-interval group is located as the data node.

[0310] Among them, the training node where the training process in the first physical sub-interval group is located is the third training node.

[0311] For example, if there are N training nodes, where N equals K+M, then the first training process determines K data nodes from the multiple training nodes according to S7-S9, and then determines the remaining M training nodes as M verification nodes.

[0312] For example, as shown in Figure 7, since the index (i.e., subscript) of [0,1] in the physical interval group is 0, training node 0 is selected as a data node. Since the index (i.e., subscript) of [4,5] in the physical interval group is 2, training node 2 is selected as a data node. Selecting training node 0 and training node 2 as data nodes, rather than selecting training node 1, helps to minimize the communication overhead during data transmission.

[0313] In the above embodiments, the data node selection problem is modeled as a maximum overlap interval pairing problem. Combining physical interval groups and data interval groups, for each data sub-interval group in the data interval group, the physical sub-interval group with the maximum overlap is determined among multiple physical sub-interval groups, thereby reducing communication overhead. Maximum overlap refers to having the largest number of identical training processes, and the index of the physical sub-interval group with the maximum overlap is determined as the identifier of the data node. After determining the physical sub-interval group with the maximum overlap for K data sub-interval groups in the data interval group, K data nodes are obtained.

[0314] Optionally, the first training process can use a scanning algorithm to determine the first physical sub-space group from multiple physical sub-space groups, that is, to solve the maximum overlapping space pairing problem.

[0315] For example, the scanning algorithm can use a scan line to traverse the endpoints of the physical interval group and the data interval group from left to right. By maintaining an active set of physical interval groups, whenever a data sub-interval group is encountered in the data interval group, that data sub-interval group is compared with the physical interval groups in the active set to determine the physical sub-interval group with the maximum overlap. Specifically, when determining the first physical sub-interval group of each data sub-interval group using the scanning method, the complexity is O((l1+l2)log(l1+l2)), where l1 represents the length of the physical interval group and l2 represents the length of the data interval group. Furthermore, when K equals M, when determining data nodes using the scanning method, nodes with odd indices are selected as data nodes, and nodes with even indices are selected as check nodes.

[0316] For example, as shown in Figures 7 and 8, N equals 3, K equals 2, and M equals 1. The distributed training system can tolerate the failure of one training node. The data size of d0, ..., d5 is all one unit (i.e., the data size of d0, ..., d5 is equal). Since M equals 1, each training process needs to obtain one encoding block. The six training processes are divided into three XOR groups. The three XOR groups are XORed to obtain three check blocks. Therefore, the final check data includes three check blocks. Since K equals 2, two original data sets will be obtained. Since there are a total of six second training state data sets, each original data set includes three second training state data sets.

[0317] As shown in Figure 7, when training node 1 is selected as the verification node, the process of storing the checkpoint requires 3 XOR operations and 3 data transmission operations, for a total of 6 communication operations, or 6 units of communication overhead. As shown in Figure 8, when training node 2 is selected as the verification node, the process of storing the checkpoint requires 3 XOR operations and 4 data transmission operations, for a total of 7 units of communication overhead. Therefore, the scheme shown in Figure 7 has lower communication overhead.

[0318] As shown in Figure 7, when training node 2 is selected as a data node, the second training state data that training node 2 needs to store includes d3, d4, and d5. Since training node 2 already stores d4 and d5, it only needs to receive one additional piece of second training state data. As shown in Figure 8, when training node 1 is selected as a data node, training node 1 needs to store the second training state data including d3, d4, and d5. Since training node 1 already stores d3, it needs to receive two additional pieces of second training state data. Therefore, the communication overhead of the scheme in Figure 7 is higher.

[0319] It should be noted that this application does not limit the method of determining data nodes; the above is merely an illustrative example. For instance, K data nodes can also be determined by counting the number of communication operations for data transmission when each training node acts as a data node, thereby identifying the communication frequency when different training nodes act as data nodes.

[0320] S303b: Store K raw data in the first memory of K data nodes, and store M parity data in the first memory of M parity nodes.

[0321] For example, referring to Figure 6, training nodes 0 and 2 are data nodes, and training nodes 1 and 2 are verification nodes. Training process 0 stores the verification block p0 obtained from the M verification data of the encoding result matrix into training node 1, thus making verification block p0 and verification block p1 obtained from training process 1 constitute one verification data. Training process 2 stores the verification block p2 obtained from the M verification data of the encoding result matrix into training node 3, thus making verification block p2 and verification block p3 obtained from training process 1 constitute one verification data, resulting in two verification data. Furthermore, training node 1 stores d1 into training node 0, thus making d1 and d0 of training process 0 constitute one original data. Training node 3 stores d3 into training node 2, thus making d3 and d2 of training process 2 constitute one original data, resulting in two original data.

[0322] For example, after determining K data nodes and M verification nodes, the second training state data (i.e., K original data) of multiple training processes are stored in the first memory of the K data nodes, and the M verification data are stored in the first memory of the M verification nodes. For instance, after the first training process determines K data nodes and M verification nodes from multiple training nodes, if the training node where the first training process is located is a data node, the first training process stores its verification block in the verification node. As shown in Figure 7, training node 0 where training process 1 is located is a data node, and training process 1 stores the verification block p1 of training process 2 in the verification node (i.e., training node 1).

[0323] For example, when the training node of the first training process is also the verification node, the first training process stores its second training state data according to the index indicated by the data sub-interval group to which it belongs. As shown in Figure 7, training process 2 belongs to the data sub-interval group with index 0. Therefore, training process 2 stores its d2 in training node 0, and training process 3 stores its d3 in training node 2. For example, training process 2 can send d2 to training process 0 so that d2 can be stored in the first memory of training node 0 through training process 0. Training node 3 can send d3 to training process 4 so that d3 can be stored in the first memory of training node 2 through training process 4.

[0324] In the above embodiments, by determining K third training nodes as K data nodes and storing K original data on K data nodes, the number of communication operations for data transmission can be reduced, thereby reducing the time overhead of data transmission operations.

[0325] The erasure coding-based checkpointing (ECCheck) provided in this application can determine and store checkpoints through encoding operations, XOR reduction operations, and peer-to-peer communication. The communication operations include XOR operations and data transmission operations. Furthermore, this application selects the target XOR process, target data nodes (i.e., K third training nodes), and verification nodes, and utilizes network idle time periods to perform communication tasks. The encoding, XOR, and data transmission operations are pipelined, achieving overlap between computation and communication. This helps reduce the time spent storing checkpoints, thereby increasing the frequency of storing checkpoints.

[0326] Optionally, S303 includes: when the fourth training node and the fifth training node among the multiple training nodes are in a network idle state, storing the original data or verification data on the fourth training node to the fifth training node.

[0327] It should be noted that other relevant explanations regarding network idle status can be found in the descriptions in the above embodiments, and will not be repeated here.

[0328] The following example uses the fourth training process on the fourth training node and the fifth training process on the fifth training process as examples for illustrative description.

[0329] In one example, the fourth training node is a data node and the fifth training node is a verification node. When both the fourth and fifth training nodes are in a network idle state, the fourth training process stores its verification block in the second memory of the fifth training node, thus storing the verification data in the second memory of the fifth training node. In another example, the fourth training node is a verification node and the fifth training node is a data node. When both the fourth and fifth training nodes are in a network idle state, the fourth training process stores its second training state data in the fifth training node, thus storing the original data in the fifth training node.

[0330] In another example, both the fourth and fifth training nodes are data nodes. When the fourth and fifth training nodes are in a network idle state, the fourth training process stores the second training state data of the fourth training process in the fifth training node, thereby storing the original data in the fifth training node.

[0331] In another example, both the fourth and fifth training nodes are verification nodes. When the fourth and fifth training nodes are in a network idle state, the fourth training process stores the verification block of the fourth training process to the fifth training node, thereby storing the verification data to the fifth training node.

[0332] In this embodiment, by performing data transmission operations between the two training nodes when the network is idle, it helps to avoid network resource competition between data transmission operations and the training of the neural network model, thereby helping to avoid affecting communication during the training of the neural network model and thus helping to avoid affecting training throughput.

[0333] In this application, since the checkpoints are stored in the CPU's memory, they can also be called in-memory checkpoints. Compared to related technologies that store checkpoints in a remote storage system, the solution in this application can utilize the high bandwidth of memory to improve the read and write efficiency of checkpoints. On the one hand, due to the higher write rate, the distributed training system can save more checkpoints over a period of time, thereby helping to increase the checkpoint saving frequency and thus helping to restore the training state to a more recent point in time in the event of a failure. On the other hand, due to the higher read rate, after a failure in the distributed training system, each training node can quickly obtain the checkpoints, thereby helping to improve the speed of training recovery.

[0334] Optionally, S303 includes: during the training of the neural network model in each training process, storing K raw data in the first memory of K data nodes, and storing M verification data in the first memory of M verification nodes.

[0335] For example, when the first training node is a verification node, during the process of the first training process calling the GPU resources of the first training node to train the neural network model, the CPU resources of the first training node are called to store the second training state data on the first training node into the first memory of K data nodes. For example, as shown in Figure 7, training node 1 is a verification node, and training node 0 and training node 2 are data nodes. Training node 1 stores d2 on its own training node into training node 0, and stores d3 on its own training node into training node 2.

[0336] When the first training node is a data node, during the process of the first training process calling the GPU resources of the first training node to train the neural network model, the CPU resources of the first training node are called to store the verification data on the first training node into the first memory of M verification nodes. For example, as shown in Figure 7, training node 0 stores p1 on this training node into training node 1.

[0337] In this embodiment, since the data in the CPU's first memory is transferred to other training nodes, CPU resources can be called to perform data transfer operations. In this way, while the GPU resources are being used to train the neural network model, the CPU resources can be called to perform data transfer operations, thus achieving parallel execution of the neural network model training operation and data transfer operation. This helps to reduce the time occupied by the neural network model training during storage checkpoints, and further helps to reduce the impact of the storage checkpoint process on training throughput.

[0338] Optionally, the first training process executes encoding operations, XOR operations, and data communication operations in a pipeline manner.

[0339] For example, the first training process executes encoding operations, XOR operations, and data communication operations through different threads. For instance, the first training process includes an encoding thread, an XOR thread, and a communication thread. The encoding thread performs encoding operations, the XOR thread performs XOR operations, and the communication thread performs communication operations. Thus, after a data block is encoded by the encoding thread, the XOR thread takes over and performs the XOR operation. Simultaneously, the encoding thread intermittently continues processing the next data block. After an XOR operation is completed, the communication thread performs data transmission operations on the corresponding check block or data block.

[0340] The encoding thread can continuously execute encoding operations. The XOR thread first places XOR operation tasks into a buffer queue so that the XOR operation can be executed during network idle periods. The communication process places data transmission operation tasks into a buffer queue so that data transmission operations can be executed during network idle periods. This helps to avoid interfering with communication during the training of the neural network model.

[0341] In this embodiment, by performing encoding operations, XOR operations, and data communication operations in a pipelined manner, computation and communication can be overlapped, and multithreading technology can be used to further reduce the time overhead of storage checkpoints.

[0342] Optionally, the checkpoint management method provided in this application further includes: storing the second training state data and the third training state data into a storage system according to a second cycle. The second cycle is shorter than the first cycle.

[0343] For example, the first training node stores the second training state data and the third training state data from its first memory to the storage system according to the second cycle. For instance, the first training process stores its second training state data and the second training state data to the storage system.

[0344] For example, as shown in "④ Low-frequency persisting" in Figure 4, training node 0 uses its CPU resources to store the second and third training state data from its CPU memory to a remote storage system. The operations performed by training nodes 1, ..., N-1 can be referenced from those performed by training node 0 and will not be described again.

[0345] It should be noted that this application does not impose any limitation on the specific duration of the second cycle. For example, the second cycle can be dynamically set according to the actual scenario.

[0346] In this embodiment, by backing up the second and third training state data to a remote storage system, when the number of failed training nodes exceeds M, the distributed training system can retrieve the previously backed-up checkpoints from the remote storage system, thereby helping to ensure the reliability of the neural model training process.

[0347] The above embodiments have detailed the scheme for storing checkpoints. The following provides an exemplary description of a scheme for recovering the training state of a neural network model through checkpoints.

[0348] For ease of description, training nodes that are alive will be referred to as live nodes, data nodes that are alive will be referred to as live data nodes, and verification nodes that are alive will be referred to as live verification nodes. Training nodes that fail will be referred to as fail nodes, data nodes that fail will be referred to as fail data nodes, and verification nodes that fail will be referred to as fail verification nodes. Training nodes used to replace fail nodes will be referred to as replacement nodes, training nodes that replace fail data nodes will be referred to as replacement data nodes, and training nodes that replace fail verification nodes will be referred to as replacement data nodes.

[0349] It should be noted that a training node that is alive can be considered as a training node that has not experienced a failure.

[0350] For example, a scenario where each training node storing K original data points is in a live state, while the training node storing validation data fails, is called a Type I failure (or Type I error). A scenario where the training node storing the original data fails is called a Type II failure (or Type II error).

[0351] Optionally, the checkpoint management method provided in this application further includes: when a training node storing verification data fails, and each training node storing K original data is in a live state, restoring the training state of the neural network model through the K original data.

[0352] Example 1: If every training node storing K original data points is alive, this indicates a Type I failure, and the K original data points have not been lost. Based on this, data fragments belonging to the failed node from the K original data points can be sent to the replacement node of the failed node. This allows the training process on the replacement node to recover the state dictionary fragments of the training process on the failed node using the data fragments belonging to the failed node.

[0353] Optionally, the checkpoint management method provided in this application further includes: the first surviving node among multiple surviving nodes sends the second training state data of the failed node stored in the first surviving node to the replacement node of the failed node. The first surviving node can be any one of the multiple surviving nodes.

[0354] For example, the surviving node sends the third training state data belonging to the failed node to the replacement node of the failed node, so that the training process on the replacement node can recover the complete state dictionary fragments. Based on these complete state dictionary fragments, the training state of the neural network model can be recovered more accurately. After the training processes on the surviving node and the replacement node each obtain their own state dictionary fragments, each training process can recover the training state of the neural network model using its respective state dictionary fragments, thereby resuming the training of the neural network model.

[0355] In this embodiment, by sending the second training state data of the faulty node to the replacement node of the faulty node, the replacement node can obtain the complete state dictionary fragments of the faulty node, which helps the replacement node to recover the accuracy of the training state.

[0356] For example, referring to Figure 6, when training node 1 and training node 3 fail simultaneously, training node 0 and training node 2 remain alive. The training state of the neural network model is restored using the original data from training node 0 and training node 2. Specifically, training process 0 can send data segment d1 belonging to training node 1 and the third training state data 1 of training node 1 to the replacement node 1, allowing the training process on the replacement node 1 to restore the state dictionary fragment 1 of training process 1 on training node 1 based on data segment d1 and the third training state data 1. Training node 2 sends data segment d3 belonging to training node 3 and the third training state data 3 of training node 3 to the replacement node 3, allowing the training process on the replacement node 3 to restore the state dictionary fragment of training process 3 on training node 3 based on data segment d3 and the third training state data 3. Training process 0 can restore the state dictionary fragment of training process 0 based on the locally stored data segment d0 and the third training state data 0 of training process 0. Training process 2 can restore the state dictionary fragments of training process 2 based on the locally stored data fragment d2 and the third training state data 2 of training process 2.

[0357] It should be noted that this application does not limit the surviving nodes that provide third training state data for replacement nodes; the above is merely an illustrative example.

[0358] Example 2: When at least one of the M verification nodes fails and K data nodes are alive, the training state of the neural network model is restored using K original data from the K data nodes.

[0359] For example, if K data nodes are alive, meaning they have not failed, it indicates that the K original data points have not been lost. Based on this, data fragments belonging to the failed nodes from the K alive data nodes can be sent to the replacement verification node of the fault verification node. This allows the training process on the replacement verification node to recover the state dictionary fragments of the training process on the fault verification node using the data fragments belonging to the fault verification node.

[0360] It should be noted that other relevant explanations for Example 2 can be found in the explanations for Example 1 above, and will not be repeated here.

[0361] In this embodiment, the training state of the neural network model is restored using K original data points in the second memory of the surviving node. This not only helps improve the speed of restoring the training state and thus reduces the time overhead, but also helps reduce the communication overhead. For example, after restoring the training state of the neural network model, the surviving node and the replacement node can reacquire M verification data points by performing the operation in S302 described above, and store the reacquired M verification data points on the surviving verification node and / or the replacement verification node to restore the fault tolerance capability of the distributed training system. The process of storing the M verification data points on the surviving verification node and / or the replacement verification node can be referred to the description in S303 above, and will not be repeated here.

[0362] Optionally, the checkpoint management method provided in this application further includes: when the training node storing the original data fails, and there are more than or equal to M training nodes in the surviving state among the N training nodes storing N checkpoint data, the training state of the neural network model is restored through the K checkpoint data on the surviving training nodes.

[0363] Example 3: When the training node storing the original data fails, it indicates a Type II failure, and K pieces of original data are incomplete. Based on this, the training state of the neural network model can be restored using the K checkpoint data from the surviving training nodes. For example, the lost original and checkpoint data can be reconstructed through distributed decoding operations, allowing the training processes on surviving nodes and replacement nodes (e.g., replacement nodes for data failure nodes, replacement nodes for checkpoint failure nodes, etc.) to obtain their respective state dictionary fragments, thereby restoring the training state of the neural network model.

[0364] For example, the process of restoring the training state of a neural network model can be viewed as the inverse process of erasure coding. For instance, a decoding operation can be performed first to allow each training process to obtain the original tensor key-value data values. Then, the training metadata, the keys of the tensor key-value data, and the values ​​of the tensor key-value data can be combined to form a state dictionary, thereby restoring the training state of the neural network model.

[0365] For example, referring to Figure 9, based on M equal to 2, the distributed training system can tolerate a maximum of two concurrent failures of training nodes. If training node 1 and training node 2 fail simultaneously, the test data P0 and the original data D1 will be lost. Where D0 = [d0d1d2], D1 = [d3d4d5]. Based on this, the original data can be recovered through the following relationship (3), thereby obtaining all the fragments of the state dictionary, and then recovering the training state of the neural network model.

[0366] As shown in Figure 9, replacement node 1 is the replacement node for the failed training node 1, and replacement node 2 is the replacement node for the failed training node 2. Since training node 1 is a data node, the current failure scenario belongs to the second type of failure, i.e., K original data are incomplete. Based on this, training process 0 sends d1 to replacement node 1, and training process 3 sends a verification block p2 to replacement node 3, thereby distributing the workload of the decoding operation to the replacement nodes, which helps to improve the speed of restoring the training state. Among them, training node 0 still retains d1, and training node 3 still retains the verification block p2. Afterwards, training process 0 obtains the decoding block e′ through the decoding matrix E'. 11 d0 and e′ 21 d0, Training process 1 obtains decoded block e′ through decoded matrix E'. 11 d1 and e′ 21 d1, Training process 2 obtains the decoded block e′ through the decoded matrix E'. 12 p2 and e′ 22 p2, Training process 3 obtains the decoded block e′ through the decoded matrix E'. 12 p3 and e′ 22 p3, where the decoding matrix is ​​obtained by inverting a submatrix of the encoding matrix. Subsequently, the decoding blocks from different training processes within the same XOR group undergo an XOR operation.

[0367] Training process 0 for e′ 21 d0 and e′ 22 p2 performs an XOR operation to obtain the recovered parity block p0. The recovered parity block p0 is stored in the first memory of training node 0 where training process 0 resides. Training process 2' performs an XOR operation on e'. 11 d0 and e′ 12p2 performs an XOR operation to obtain the recovered d2. The recovered d2 is stored in the first memory of training node 2, where training process 2 resides. Training process 1' performs an XOR operation on e'. 21 d1 and e′ 22 p3 performs an XOR operation to obtain the recovered parity block p1. The recovered parity block p1 is stored in the first memory of replacement node 1 where training process 1' resides. Training process 3 performs an XOR operation on e'. 11 d1 and e′ 12 p3 performs an XOR operation to obtain the recovered d3. The recovered d3 is stored in the first memory of training node 3, where training process 3 resides. Afterwards, training process 0 can recover its state dictionary fragment using d0. Training process 1' on replacement node 1 can recover its state dictionary fragment using d1. Training process 2' on replacement node 2 can recover its state dictionary fragment using d2. Training process 3 can recover its state dictionary fragment using d3.

[0368] For example, after restoring the training state of the neural network model, K original data points and M verification data points are stored on the surviving nodes and replacement nodes to restore the fault tolerance of the distributed training system. For instance, the fault tolerance of the distributed training system can be restored by performing data transfer operations to store the restored data fragments on the replacement nodes of the failed data nodes and the restored verification blocks on the replacement nodes of the failed verification nodes. As shown in Figure 9, training node 0 stores the restored verification block p0 on replacement node 1, and training process 3 stores the restored d3 on replacement node 2.

[0369] It should be noted that the process of performing the XOR operation during the recovery phase can be referred to the process of performing the XOR operation during the storage checkpoint, and will not be repeated here.

[0370] It should be noted that other relevant explanations for Example 3 can be found in Examples 1 and 2 above, and will not be repeated here.

[0371] Example 4: When at least one of the K data nodes fails, and the number of training nodes in a surviving state is greater than or equal to K, the training state of the neural network model is restored using the K checkpoint data from the surviving training nodes.

[0372] It should be noted that other relevant explanations for Example 4 can be found in Examples 1 to 3 above, and will not be repeated here.

[0373] In the above embodiments, when the original data is incomplete, the training state of the neural network model is restored by using K checkpoint data, thereby improving the fault tolerance of the distributed training system and thus improving the training performance of the distributed training system.

[0374] The foregoing mainly describes the solutions provided by the embodiments of this application from a methodological perspective. To achieve the above functions, the distributed training system includes corresponding hardware structures and / or software modules for executing each function. Those skilled in the art should readily recognize that, based on the units and algorithm steps of the examples described in conjunction with the embodiments disclosed herein, this application can be implemented in hardware or a combination of hardware and computer software. Whether a function is executed in hardware or by computer software driving hardware depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.

[0375] This application embodiment can, based on the above method, exemplarily divide a distributed training system into functional modules. For example, the distributed training system may include functional modules corresponding to each functional division, or two or more functions may be integrated into one processing module. The integrated module can be implemented in hardware or as a software functional module. It should be noted that the module division in this application embodiment is illustrative and only represents one logical functional division; in actual implementation, there may be other division methods.

[0376] For example, Figure 10 shows a possible schematic diagram of the distributed training system (i.e., distributed training system 1000) involved in the above embodiments. The actions performed by the distributed training system 1000 can be executed by a computing device or implemented by executing corresponding software through a computing device. The distributed training system 1000 may include: a writing module 1001, an erasure coding module 1002, and a storage module 1003. The writing module 1001 is used to write the first training state data in the second memory of each training node to the first memory of each training node, wherein the first training state data is obtained by the distributed training system training the neural network. For example, as shown in S301 of Figure 3. The erasure coding module 1002 is used to perform erasure coding on the second training state data in the first memory of each training node to obtain N checkpoint data. The second training state data is the data obtained by the distributed training system training the neural network, and includes the first training state data. The N checkpoint data includes K original data and M verification data. The K original data includes the second training state data of each training node. Any K checkpoint data from the N checkpoint data are used to obtain the second training state data of each training node. For example, as shown in S302 of Figure 3. The storage module 1003 is used to store the N checkpoint data in the first memory of multiple training nodes, wherein different checkpoint data from the N checkpoint data are stored in the first memory of different training nodes. For example, as shown in S303 of Figure 3.

[0377] Optionally, the second training state data includes the values ​​of the tensor key-value data.

[0378] Optionally, the plurality of training nodes include a first training node, and the storage module 1003 is further configured to: store the third training state data in the first memory of the first training node into the first memory of the second training node, wherein the second training node includes at least some of the training nodes other than the first training node, and the third training state data includes the training metadata and tensor key-value data obtained by the distributed training system from training the neural network.

[0379] Optionally, the storage module 1003 is specifically used to: determine K third training nodes from a plurality of training nodes as K data nodes, and determine M verification nodes from the training nodes other than the K third training nodes from a plurality of training nodes, wherein when the third training node is a data node, the number of communication times for storing K raw data to the K data nodes and M verification data to the M verification nodes is less than the number of communication times for storing K raw data to the K data nodes and M verification data to the M verification nodes when the third training node is a verification node; store the K raw data in the first memory of the K data nodes, and store the M verification data in the first memory of the M verification nodes.

[0380] Optionally, the distributed training system includes multiple training processes deployed on multiple training nodes for training neural network models. The storage module 1003 is further configured to: obtain physical interval groups of the multiple training processes, wherein each physical interval group includes multiple physical sub-interval groups, and training processes in a physical sub-interval group are located on the same training node; obtain data interval groups of the multiple training processes, wherein each data interval group includes K data sub-interval groups; determine a first physical sub-interval group from the multiple physical sub-interval groups, wherein the number of identical training processes in the first physical sub-interval group and the target data sub-interval group in the K data sub-interval groups is greater than the number of identical training processes in the second physical sub-interval group and the target data sub-interval group, wherein the second physical sub-interval group includes at least some of the physical sub-interval groups other than the first physical sub-interval group; and determine the training node where the training processes in the first physical sub-interval group are located as a data node, wherein the training node where the training processes in the first physical sub-interval group are located is a third training node.

[0381] Optionally, the erasure coding module 1002 is specifically used for: performing an encoding operation on the second training state data of each training process according to the encoding matrix of each training process in multiple training processes, to obtain M encoding blocks of each training process; determining R XOR groups of multiple training processes, wherein one XOR group includes K training processes; performing an XOR operation on the encoding blocks of different training processes in each XOR group to obtain R*M check blocks, wherein the R*M check blocks are M check data.

[0382] Optionally, the erasure coding module 1002 is further configured to: determine the XOR target process of each XOR group in multiple XOR groups, wherein the XOR target process of each XOR group is used to perform XOR operation on the encoding blocks of different training processes of each XOR group, the training node where the XOR target process is located is used to store the check block obtained by the XOR target process performing the XOR operation, and when M check blocks of each XOR group are stored in the training node where the XOR target process of each XOR group is located, the number of communication times of storing M check blocks in M ​​check nodes satisfies the target condition.

[0383] Optionally, when the first XOR group in a plurality of XOR groups includes training processes on M verification nodes, the target XOR process of the first XOR group is the training process on the M verification nodes.

[0384] Optionally, the erasure coding module 1002 is specifically used to: write the second training state data of each training process in the first XOR group into the data buffer of each training process in the first XOR group to obtain each data block of each training process in the first XOR group; and perform encoding operation on each data block of each training process in the first XOR group according to the encoding matrix of each training process in the first XOR group to obtain each encoding block of each training process in the first XOR group.

[0385] Optionally, the erasure coding module 1002 is specifically used to: obtain multiple threads of each training process in the first XOR group; and perform encoding operations in parallel on each data block of each training process in the first XOR group through the multiple threads of each training process in the first XOR group.

[0386] Optionally, the erasure coding module 1002 is specifically used to: write each coding block of each training process in the first XOR group into the coding buffer of each training process in the first XOR group; and perform an XOR operation on the coding blocks in the coding buffers of different training processes in the first XOR group.

[0387] Optionally, the erasure coding module 1002 is further configured to: divide multiple logical blocks of multiple training processes into K logical segments, wherein one logical block is used to indicate the second training state data obtained by a training process in training a neural network model; perform encoding operations on the K logical segments according to the erasure coding matrix to obtain M logical verification data, wherein the M logical verification data are used to indicate M verification data; and determine the encoding matrix of each training process in multiple training processes based on the M logical verification data.

[0388] Optionally, the distributed training system 1000 also includes a recovery module 1004, which is used to: when a training node storing verification data fails and each training node storing K original data is alive, restore the training state of the neural network model using the K original data.

[0389] Optionally, the recovery module 1004 is further configured to: when a training node storing the original data fails, and there are more than or equal to M training nodes in a surviving state among the N training nodes storing N checkpoint data, restore the training state of the neural network model using the K checkpoint data on the surviving training nodes.

[0390] Optionally, the erasure coding module 1002 is specifically used to: in the process of training a neural network model through multiple training nodes in a distributed training system, perform erasure coding operations on the second training state data in the first memory of each training node using the CPU resources of each training node.

[0391] Optionally, the storage module 1003 is specifically used to: during the process of training a neural network model through multiple training nodes in a distributed training system, store N checkpoint data in the first memory of multiple training nodes using the CPU resources of each training node.

[0392] Optionally, the erasure coding module 1002 is specifically used to: when different training nodes of different training processes are in a network idle state in the first XOR group of multiple XOR groups, perform an XOR operation on the coding blocks of different training processes in the first XOR group.

[0393] Optionally, the storage module 1003 is specifically used to: when the fourth training node and the fifth training node among multiple training nodes are in a network idle state, store the original data or verification data on the fourth training node to the fifth training node.

[0394] For a detailed description of the above-mentioned optional methods, please refer to the foregoing method embodiments, which will not be repeated here. Furthermore, the explanation of any of the distributed training systems 1000 provided above, as well as the description of its beneficial effects, can be found in the corresponding method embodiments described above, and will not be repeated here.

[0395] In this application, the writing module 1001, erasure coding module 1002, storage module 1003, and recovery module 1004 can all be implemented in software or in hardware. For example, the implementation of the writing module 1001 will be described below. Similarly, the implementation of the erasure coding module 1002, storage module 1003, and recovery module 1004 can refer to the implementation of the writing module 1001.

[0396] As an example of a software functional unit, the writing module 1001 may include code running on a computing instance. The computing instance may include at least one of a physical host (computing device), a virtual machine, or a container. Further, the aforementioned computing instance may be one or more. For example, the writing module 1001 may include code running on multiple hosts / virtual machines / containers. It should be noted that the multiple hosts / virtual machines / containers used to run the code may be distributed within the same region or in different regions. Further, the multiple hosts / virtual machines / containers used to run the code may be distributed within the same availability zone (AZ) or in different AZs, each AZ including one or more geographically proximate data centers. Typically, a region may include multiple AZs.

[0397] Similarly, multiple hosts / virtual machines / containers used to run this code can be distributed within the same Virtual Private Cloud (VPC) or across multiple VPCs. Typically, a VPC is set up within a region. Communication between two VPCs within the same region, as well as between VPCs in different regions, requires a communication gateway to be set up within each VPC to enable interconnection between VPCs.

[0398] As an example of a hardware functional unit, the write module 1001 may include at least one computing device, such as a server. Alternatively, the write module 1001 may also be a device implemented using an application-specific integrated circuit (ASIC) or a programmable logic device (PLD). The PLD may be implemented using a complex programmable logical device (CPLD), a field-programmable gate array (FPGA), generic array logic (GAL), or any combination thereof.

[0399] The multiple computing devices included in the write module 1001 can be distributed in the same region or in different regions. Similarly, the multiple computing devices included in the write module 1001 can be distributed in the same Availability Zone (AZ) or in different AZs. Likewise, the multiple computing devices included in the write module 1001 can be distributed in the same Virtual Private Cloud (VPC) or in multiple VPCs. These multiple computing devices can be any combination of computing devices such as servers, ASICs, PLDs, CPLDs, FPGAs, and GALs.

[0400] It should be noted that, in other embodiments, the writing module 1001, the erasure coding module 1002, the storage module 1003, and the recovery module 1004 can all execute any step in the checkpoint management method. The steps implemented by the writing module 1001, erasure coding module 1002, storage module 1003, and recovery module 1004 can be specified as needed. By implementing different steps in the checkpoint management method through the writing module 1001, erasure coding module 1002, storage module 1003, and recovery module 1004, all functions of the data storage device can be achieved.

[0401] This application also provides a computing device 1100. As shown in FIG11, the computing device 1100 includes: a bus 1102, a processor 1104, a memory 1106, and a communication interface 1108. The processor 1104, the memory 1106, and the communication interface 1108 communicate with each other via the bus 1102. The computing device 1100 may be a server or a terminal device. It should be understood that this application does not limit the number of processors and memories in the computing device 1100.

[0402] Bus 1102 can be a Peripheral Component Interconnect (PCI) bus or an Extended Industry Standard Architecture (EISA) bus, etc. Buses can be categorized as address buses, data buses, control buses, etc. For ease of illustration, only one line is used in Figure 11, but this does not imply that there is only one bus or one type of bus. Bus 1102 can include pathways for transmitting information between various components of computing device 1100 (e.g., memory 1106, processor 1104, communication interface 1108).

[0403] The processor 1104 may include any one or more processors such as a central processing unit (CPU), a graphics processing unit (GPU), a microprocessor (MP), or a digital signal processor (DSP).

[0404] The memory 1106 may include volatile memory, such as random access memory (RAM). The processor 1104 may also include non-volatile memory, such as read-only memory (ROM), flash memory, hard disk drive (HDD), or solid state drive (SSD).

[0405] The memory 1106 stores executable program code, and the processor 1104 executes the executable program code to implement the functions of the aforementioned write module 1001, erasure coding module 1002, storage module 1003, and recovery module 1004, thereby realizing the checkpoint management method. That is, the memory 1106 stores instructions for executing the checkpoint management method.

[0406] The communication interface 1108 uses transceiver modules such as, but not limited to, network interface cards and transceivers to enable communication between the computing device 1100 and other devices or communication networks.

[0407] For example, the computing device 1100 described above may be the training node shown in Figure 1.

[0408] This application also provides a computing device cluster. The computing device cluster includes at least one computing device. The computing device can be a server, such as a central server, an edge server, or a local server in a local data center. In some embodiments, the computing device can also be a terminal device such as a desktop computer, a laptop computer, or a smartphone.

[0409] As shown in Figure 12, the computing device cluster 1200 includes at least one computing device 1100. The memory 1106 of one or more computing devices 1100 in the computing device cluster may store the same instructions for executing the checkpoint management method.

[0410] In some possible implementations, the memory 1106 of one or more computing devices 1100 in the computing device cluster may also store partial instructions for executing the checkpoint management method. In other words, a combination of one or more computing devices 1100 can jointly execute instructions for executing the checkpoint management method.

[0411] It should be noted that the memory 1106 in different computing devices 1100 within the computing device cluster can store different instructions, which are used to execute some functions of the resource scheduling device. That is, the instructions stored in the memory 1106 of different computing devices 1100 can implement the functions of one or more modules among the write module 1001, erasure coding module 1002, storage module 1003, and recovery module 1004.

[0412] In some possible implementations, one or more computing devices in a computing device cluster can be connected via a network. This network can be a wide area network (WAN) or a local area network (LAN), etc. Figure 13 illustrates one possible implementation. As shown in Figure 13, two computing devices 1100A and 1100B are connected via a network. Specifically, they are connected to the network through communication interfaces in each computing device.

[0413] For example, computing device 1100A can be used to perform functions executed by computing units on training nodes, and computing device 1100B can be used to perform functions executed by central processing units on training nodes.

[0414] In this type of possible implementation, the memory 1106 in computing device 1100A stores instructions for performing the functions of the write module 1001 and the restore module 1004. Meanwhile, the memory 1106 in computing device 1100B stores instructions for performing the functions of the erasure coding module 1002 and the storage module 1003.

[0415] The connection method between the computing device clusters shown in Figure 13 can be considered as follows: taking into account that the checkpoint management method provided in this application needs to perform a large amount of computation, the functions of the erasure coding module 1002 and the storage module 1003 are considered to be performed by the computing device 1100B.

[0416] It should be understood that the functions of computing device 1100A shown in Figure 13 can also be performed by multiple computing devices 1100. Similarly, the functions of computing device 1100B can also be performed by multiple computing devices 1100.

[0417] This application also provides another computing device cluster. The connection relationship between the computing devices in this computing device cluster can be similarly referred to the connection method of the computing device cluster shown in Figures 12 and 13. The difference is that the memory 1106 of one or more computing devices 1100 in this computing device cluster can store the same instructions for executing the checkpoint management method.

[0418] In some possible implementations, the memory 1106 of one or more computing devices 1100 in the computing device cluster may also store partial instructions for executing the checkpoint management method. In other words, a combination of one or more computing devices 1100 can jointly execute instructions for executing the checkpoint management method.

[0419] This application also provides a processor that can be used to perform the above-described methods.

[0420] This application also provides a chip, including: a processor and a power supply circuit; the power supply circuit can be used to supply power to the processor; the processor can be used to perform the above-described methods.

[0421] This application also provides a computing device, which may include a processor, a memory, and a computer program / instructions stored in the memory; the processor executes the computer program / instructions to enable the computing device to implement the above-described method.

[0422] This application also provides a computing device cluster, which includes at least one computing device; each computing device includes a processor, a memory, and computer programs / instructions stored in the memory, wherein the processor of each computing device executes the computer programs / instructions stored in the memory of each computing device to enable each computing device to implement the above-described method.

[0423] This application also provides a computer program product. The computer program product includes a computer program / instructions, which are software or program products capable of running on a computing device or stored on any usable medium. When the computer program / instructions are executed on at least one computing device, the at least one computing device can perform the methods described above.

[0424] This application also provides a computer-readable storage medium, which can be any available medium capable of being stored by a computing device or a data storage device such as a data center containing one or more available media. The computer-readable storage medium stores a computer program / instructions that, when executed on at least one computing device, enable the at least one computing device to perform the described method.

[0425] For example, the available media may be magnetic media (e.g., floppy disks, magnetic disks, magnetic tapes), optical media (e.g., digital video discs (DVDs)), or semiconductor media (e.g., solid-state drives (SSDs)).

[0426] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, and not to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features; and these modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the protection scope of the technical solutions of the embodiments of the present invention.

Claims

1. A method for checkpoint management based on erasure code, characterized by, The method is applied to a distributed training system for training a neural network model using multiple training nodes. Each training node includes a computing unit, a central processing unit (CPU), and a first memory of the CPU. The computing unit includes a second memory. Write the first training state data in the second memory of each training node into the first memory of each training node, wherein the first training state data is obtained by the distributed training system training the neural network; An erasure coding operation is performed on the second training state data in the first memory of each training node to obtain N checkpoint data. The second training state data is the data obtained by the distributed training system in training the neural network. The second training state data includes the first training state data. The N checkpoint data includes K original data and M verification data. The K original data includes the second training state data of each training node. Any K checkpoint data from the N checkpoint data is used to obtain the second training state data of each training node. The N checkpoint data are stored in the first memory of the plurality of training nodes, wherein different checkpoint data in the N checkpoint data are stored in the first memory of different training nodes in the plurality of training nodes.

2. The method of claim 1, wherein, The second training state data includes the values ​​of tensor key-value data.

3. The method of claim 2, wherein, The plurality of training nodes includes a first training node, and the method further includes: The third training state data in the first memory of the first training node is stored in the first memory of the second training node, wherein the second training node includes at least some of the training nodes other than the first training node, and the third training state data includes training metadata obtained by the distributed training system from training the neural network and the key of the tensor key value data.

4. The method according to any one of claims 1 to 3, characterized in that, The step of storing the N checkpoint data in the first memory of the plurality of training nodes includes: K third training nodes among the plurality of training nodes are determined as K data nodes, and M verification nodes are determined from the training nodes other than the K third training nodes among the plurality of training nodes. Wherein, when the third training node is the data node, the number of communication times for storing the K original data to the K data nodes and the number of communication times for storing the M verification data to the M verification nodes is less than when the third training node is the verification node. The K original data are stored in the first memory of the K data nodes, and the M verification data are stored in the first memory of the M verification nodes.

5. The method of claim 4, wherein, The distributed training system includes multiple training processes deployed on the multiple training nodes for training neural network models. The step of determining K third training nodes among the multiple training nodes as K data nodes includes: Obtain the physical interval group of the multiple training processes, wherein the physical interval group includes multiple physical sub-interval groups, and the training processes in a physical sub-interval group are located on the same training node; Obtain the data interval group of the plurality of training processes, wherein the data interval group includes K data sub-interval groups; A first physical sub-interval group is determined from the plurality of physical sub-interval groups, wherein the number of identical training processes of the first physical sub-interval group and the target data sub-interval group in the K data sub-interval groups is greater than the number of identical training processes of the second physical sub-interval group and the target data sub-interval group, and the second physical sub-interval includes at least some physical sub-interval groups other than the first physical sub-interval group in the plurality of physical sub-interval groups. The training node where the training process in the first physical sub-interval group is located is determined as the data node, wherein the training node where the training process in the first physical sub-interval group is located is the third training node.

6. The method of claim 5, wherein, The step of performing erasure coding on the second training state data in the first memory of each training node includes: Based on the encoding matrix of each of the plurality of training processes, an encoding operation is performed on the second training state data of each training process to obtain M encoding blocks of each training process; Determine R XOR groups for the plurality of training processes, wherein one XOR group includes K training processes; An XOR operation is performed on the encoding blocks of different training processes in each of the plurality of XOR groups to obtain R*M check blocks, wherein the R*M check blocks are the M check data.

7. The method of claim 6, wherein, The method further includes: The XOR target process of each XOR group is determined, wherein the XOR target process of each XOR group is used to perform XOR operation on the encoding blocks of different training processes of each XOR group, and the training node where the XOR target process is located is used to store the check block obtained by the XOR target process performing the XOR operation. When M check blocks of each XOR group are stored in the training node where the XOR target process of each XOR group is located, the number of communication times of storing the M check blocks in the M check nodes satisfies the target condition.

8. The method according to claim 7, characterized in that, When the first XOR group in the plurality of XOR groups includes the training process on the M verification nodes, the target XOR process of the first XOR group is the training process on the M verification nodes.

9. The method according to any one of claims 6-8, characterized in that, The step of performing an encoding operation on the second training state data of each training process based on the encoding matrix of each of the plurality of training processes to obtain M encoding blocks for each training process includes: Write the second training state data of each training process in the first XOR group of the plurality of XOR groups into the data buffer of each training process in the first XOR group to obtain each data block of each training process in the first XOR group. Based on the encoding matrix of each training process in the first XOR group, an encoding operation is performed on each data block of each training process in the first XOR group to obtain each encoding block of each training process in the first XOR group.

10. The method of claim 9, wherein, The step of performing encoding operations on each data block of each training process in the first XOR group includes: Obtain multiple threads for each training process in the first XOR group; Encoding operations are performed in parallel on each data block of each training process in the first XOR group through multiple threads of each training process in the first XOR group.

11. The method according to claim 9 or 10, characterized in that, The step of performing encoding operations on each data block of each training process in the first XOR group includes: Write each encoding block of each training process in the first XOR group into the encoding buffer of each training process in the first XOR group; Perform an XOR operation on the encoding blocks in the encoding buffers of different training processes in the first XOR group.

12. The method according to any one of claims 6-11, characterized in that, The step of performing erasure coding on the second training state data in the first memory of each training node further includes: The multiple logical blocks of the multiple training processes are divided into K logical segments, wherein one logical block is used to indicate the second training state data of a training process; Based on the erasure coding matrix, an encoding operation is performed on the K logical segments to obtain M logical check data, wherein the M logical check data are used to indicate the M check data; Based on the M logical verification data, determine the encoding matrix of each training process in the plurality of training processes.

13. The method according to any one of claims 1-12, characterized in that, The method further includes: When the training node storing the verification data fails, and each training node storing the K original data is alive, the training state of the neural network model is restored using the K original data; or, When a training node storing the original data fails, and M or more of the N training nodes storing the N checkpoint data are in a surviving state, the training state of the neural network model is restored using the K checkpoint data from the surviving training nodes.

14. A distributed training system, comprising: The distributed training system is used to train a neural network model through multiple training nodes. Each training node includes a computing unit, a central processing unit (CPU), and a first memory of the CPU. The computing unit includes a second memory. The distributed training system includes: The writing module is used to write the first training state data in the second memory of each training node into the first memory of each training node, wherein the first training state data is obtained by the distributed training system training the neural network; The erasure coding module is used to perform erasure coding on the second training state data in the first memory of each training node to obtain N checkpoint data. The second training state data is the data obtained by the distributed training system in training the neural network. The second training state data includes the first training state data. The N checkpoint data includes K original data and M verification data. The K original data includes the second training state data of each training node. Any K checkpoint data from the N checkpoint data is used to obtain the second training state data of each training node. A storage module is used to store the N checkpoint data in the first memory of the plurality of training nodes, wherein different checkpoint data in the N checkpoint data are stored in the first memory of different training nodes in the plurality of training nodes.

15. The distributed training system of claim 14, wherein, The second training state data includes the values ​​of tensor key-value data.

16. The distributed training system of claim 15, wherein, The plurality of training nodes includes a first training node, and the storage module is further configured to: The third training state data in the first memory of the first training node is stored in the first memory of the second training node, wherein the second training node includes at least some of the training nodes other than the first training node, and the third training state data includes training metadata obtained by the distributed training system from training the neural network and the key of the tensor key value data.

17. The distributed training system of any one of claims 14-16, wherein, The storage module is specifically used to: determine K third training nodes among the plurality of training nodes as K data nodes, and determine M verification nodes from the training nodes other than the K third training nodes among the plurality of training nodes, wherein, when the third training node is used as the data node, the number of communication times for storing the K original data to the K data nodes and the number of communication times for storing the M verification data to the M verification nodes is less than when the third training node is used as the verification node; The K original data are stored in the first memory of the K data nodes, and the M verification data are stored in the first memory of the M verification nodes.

18. The distributed training system of claim 17, wherein, The distributed training system includes multiple training processes deployed on the multiple training nodes for training neural network models, and the storage module is further used for: Obtain the physical interval group of the multiple training processes, wherein the physical interval group includes multiple physical sub-interval groups, and the training processes in a physical sub-interval group are located on the same training node; Obtain the data interval group of the plurality of training processes, wherein the data interval group includes K data sub-interval groups; A first physical sub-interval group is determined from the plurality of physical sub-interval groups, wherein the number of identical training processes of the first physical sub-interval group and the target data sub-interval group in the K data sub-interval groups is greater than the number of identical training processes of the second physical sub-interval group and the target data sub-interval group, and the second physical sub-interval includes at least some physical sub-interval groups other than the first physical sub-interval group in the plurality of physical sub-interval groups. The training node where the training process in the first physical sub-interval group is located is determined as the data node, wherein the training node where the training process in the first physical sub-interval group is located is the third training node.

19. The distributed training system of claim 18, wherein, The erasure coding module is specifically used for: Based on the encoding matrix of each of the plurality of training processes, an encoding operation is performed on the second training state data of each training process to obtain M encoding blocks of each training process; Determine R XOR groups for the plurality of training processes, wherein one XOR group includes K training processes; An XOR operation is performed on the encoding blocks of different training processes in each of the plurality of XOR groups to obtain R*M check blocks, wherein the R*M check blocks are the M check data.

20. The distributed training system of claim 19, wherein, The erasure coding module is also used for: The XOR target process of each XOR group is determined, wherein the XOR target process of each XOR group is used to perform XOR operation on the encoding blocks of different training processes of each XOR group, and the training node where the XOR target process is located is used to store the check block obtained by the XOR target process performing the XOR operation. When M check blocks of each XOR group are stored in the training node where the XOR target process of each XOR group is located, the number of communication times of storing the M check blocks in the M check nodes satisfies the target condition.

21. The distributed training system according to claim 19, characterized in that, When the first XOR group in the plurality of XOR groups includes the training process on the M verification nodes, the target XOR process of the first XOR group is the training process on the M verification nodes.

22. The distributed training system of any one of claims 19-21, wherein, The erasure coding module is specifically used for: Write the second training state data of each training process in the first XOR group of the plurality of XOR groups into the data buffer of each training process in the first XOR group to obtain each data block of each training process in the first XOR group. Based on the encoding matrix of each training process in the first XOR group, an encoding operation is performed on each data block of each training process in the first XOR group to obtain each encoding block of each training process in the first XOR group.

23. The distributed training system of claim 22, wherein, The erasure coding module is specifically used for: Obtain multiple threads for each training process in the first XOR group; Encoding operations are performed in parallel on each data block of each training process in the first XOR group through multiple threads of each training process in the first XOR group.

24. The distributed training system of claim 22 or 23, wherein, The erasure coding module is specifically used for: Write each encoding block of each training process in the first XOR group into the encoding buffer of each training process in the first XOR group; Perform an XOR operation on the encoding blocks in the encoding buffers of different training processes in the first XOR group.

25. The distributed training system of any one of claims 19-24, wherein, The erasure coding module is also used for: The multiple logical blocks of the multiple training processes are divided into K logical segments, wherein one logical block is used to indicate the second training state data of a training process; Based on the erasure coding matrix, an encoding operation is performed on the K logical segments to obtain M logical check data, wherein the M logical check data are used to indicate the M check data; Based on the M logical verification data, determine the encoding matrix of each training process in the plurality of training processes.

26. The distributed training system of any one of claims 14-25, wherein, The distributed training system further includes a recovery module, which is used for: When the training node storing the verification data fails, and each training node storing the K original data is alive, the training state of the neural network model is restored using the K original data. or, When a training node storing the original data fails, and M or more of the N training nodes storing the N checkpoint data are in a surviving state, the training state of the neural network model is restored using the K checkpoint data from the surviving training nodes.

27. A cluster of computing devices, characterized in that, The computing device cluster includes at least one computing device; Each of the at least one computing device includes a processor, a memory, and a computer program / instructions stored in the memory; The processor of each computing device executes computer programs / instructions stored in the memory of each computing device to enable each computing device to implement the method as described in any one of claims 1-13.

28. A computer program product, characterized in that, The computer program product includes a computer program / instruction, which, when executed by a computing device, implements the method as described in any one of claims 1-13.

29. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program / instruction, which, when executed by a computing device, implements the method as described in any one of claims 1-13.