Checkpoint recovery method of distributed training cluster, cluster and medium

CN122816773APending Publication Date: 2026-09-25SHANGHAI BIREN TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202611242741.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-08-17
Publication Date
2026-09-25

AI Technical Summary

Technical Problem

当某个计算节点发生故障并被热备节点替代时,热备节点会被分配新的进程标识,原有的固定备份关系映射因此失效,导致热备节点无法通过预配置的备份关系找到正确的检查点数据,系统频繁回退至从持久化存储加载,恢复时间从预期的秒级退化为分钟级,热备节点的检查点恢复成功率极低

Benefits of technology

[0016]通过本公开的实施例,在目标计算节点故障后启动分配全新进程标识的热备进程替代故障进程,通过汇总全部存活节点检查点元数据构建包含原始进程标识、当前进程标识、训练迭代步数的全局检查点可用性对象,基于该对象匹配原本归属故障进程的目标检查点并加载至热备进程,借助原始进程标识解耦热备新进程标识与检查点数据归属的绑定关系,解决传统静态备份映射因进程标识重分配失效的问题,依托集群全局检查点可用性对象实现跨节点内存分片快速恢复,无需依赖低速持久化存储,能够在集群节点故障、拓扑动态变更场景下自动定位有效备份,兼顾故障恢复速度与分布式训练状态一致性,有效提升分布式训练集群的故障容错能力与运行连续性。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122816773A_ABST
    Figure CN122816773A_ABST
Patent Text Reader

Abstract

The present disclosure relates to the technical field of distributed machine learning training cluster, and provides a checkpoint recovery method of a distributed training cluster, a cluster and a medium. The method comprises: in the case where a target computing node fails, starting a hot backup process to replace a failed process of the target computing node, and assigning a hot backup process identifier to the hot backup process; obtaining metadata of a checkpoint stored by each surviving computing node, and obtaining a global checkpoint availability object according to the metadata; determining, according to the global checkpoint availability object, a target checkpoint in which an original process identifier matches the failed process among the checkpoints stored by each surviving process, and determining a sending process corresponding to the target checkpoint; and loading the target checkpoint from the sending process to the hot backup process. The problem that a traditional static backup mapping is invalid due to process identifier reassignment can be solved, and cross-node memory shard fast recovery is realized relying on the global checkpoint availability object of the cluster.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This disclosure relates to the field of distributed machine learning training cluster technology, and more specifically, to a checkpoint recovery method for a distributed training cluster, a distributed training cluster, and a computer-readable storage medium. Background Technology

[0002] In large-scale deep learning training, distributed training clusters typically consist of dozens or even hundreds of computing nodes. Each computing node runs multiple training processes, and these processes identify and communicate with each other through unique process identifiers. To cope with abnormal situations such as computing node hardware failures and network interruptions, the system usually employs a checkpoint mechanism to periodically save the training state and resume training from the checkpoint after a failure occurs.

[0003] In related technologies, to achieve a balance between local memory recovery and persistent storage, some systems introduce a neighbor node backup mechanism: each process periodically pushes checkpoint data to the pre-configured backup shared memory of neighbor nodes. When a failure occurs, the hot standby node can quickly obtain backup data from the neighbor node via network communication. However, this mechanism typically uses a fixed one-to-one backup relationship mapping, which is established based on a fixed process identifier relationship at the start of training and remains unchanged throughout the training process. When a compute node fails and is replaced by a hot standby node, the hot standby node is assigned a new process identifier, and the original fixed backup relationship mapping becomes invalid. This causes the hot standby node to be unable to find the correct checkpoint data through the pre-configured backup relationship, resulting in frequent system rollbacks to loading from persistent storage. The recovery time degrades from the expected seconds to minutes, and the checkpoint recovery success rate of the hot standby node is extremely low. Summary of the Invention

[0004] One objective of this disclosure is to provide a new technical solution that enables checkpoint recovery in a distributed training cluster, thereby improving the fault recovery capability of the distributed training cluster.

[0005] According to a first aspect of this disclosure, a checkpoint recovery method for a distributed training cluster is provided, the distributed training cluster comprising multiple computing nodes, each computing node comprising at least one process, each process having a unique process identifier; the method comprising: In the event of a failure in the target computing node, a hot standby process is initiated to replace the failed process of the target computing node, and a hot standby process identifier is assigned to the hot standby process. Obtain the metadata of the checkpoints stored in each surviving compute node, and obtain a global checkpoint availability object based on the metadata; wherein, the global checkpoint availability object includes the original process identifier of the process to which all checkpoints stored in each surviving process originally belonged, and the current process identifier of the process to which they currently belong; Based on the global checkpoint availability object, determine the target checkpoint in the checkpoints stored by each surviving process that matches the original process identifier of the faulty process, and determine the sending process corresponding to the target checkpoint; The target checkpoint is loaded from the sending process to the hot standby process.

[0006] Optionally, the global checkpoint availability object further includes the training iteration steps of all checkpoints stored by each surviving process; the step of determining the target checkpoint among the checkpoints stored by each surviving process that matches the original process identifier of the faulty process, and determining the surviving process storing the target checkpoint as the sending process, includes: determining the target training iteration steps according to the global checkpoint availability object; wherein, the target training iteration steps represent the maximum training iteration steps for which each process identifier has a corresponding checkpoint in the global checkpoint availability object; and determining, according to the global checkpoint availability object, the checkpoint whose original process identifier is the process identifier of the faulty process and whose training iteration steps are the target training iteration steps, as the target checkpoint.

[0007] Optionally, determining the target training iteration steps based on the global checkpoint availability object includes: obtaining the training iteration steps of all checkpoints based on the global checkpoint availability object and performing deduplication; traversing the deduplicated training iteration steps in ascending order, verifying that each process identifier has a checkpoint in the global checkpoint availability object with a training iteration step number equal to the currently traversed training iteration step number; if so, stopping the traversal and using the currently traversed training iteration step number as the target training iteration step number; otherwise, continuing the traversal.

[0008] Optionally, determining the sending process corresponding to the target checkpoint includes: determining the candidate live processes to which the target checkpoint currently belongs based on the global checkpoint availability object; and selecting one of the candidate live processes as the sending process.

[0009] Optionally, selecting one candidate surviving process from the candidate surviving processes as the sending process includes: determining the priority of the candidate surviving process based on the network distance between the candidate surviving process and the hot standby process; and selecting the candidate surviving process with the highest priority as the sending process.

[0010] Optionally, obtaining the metadata of checkpoints stored by each surviving computing node and obtaining a global checkpoint availability object based on the metadata includes: obtaining the configuration information of checkpoints stored in the local shared memory of each surviving computing node and the configuration information of checkpoints stored in the backup shared memory; obtaining the metadata of each checkpoint based on the configuration information; and assembling the metadata of each checkpoint to obtain the global checkpoint availability object.

[0011] Optionally, the global checkpoint availability object also includes the checkpoint types of all checkpoints stored by each surviving process, where the checkpoint type indicates whether the checkpoint is stored in local shared memory or backup shared memory in the corresponding surviving compute node; loading the target checkpoint into the hot standby process includes: the sending process reading the target checkpoint from local shared memory or backup shared memory according to the checkpoint type of the target checkpoint, and sending the target checkpoint to the hot standby process according to the hot standby process identifier; the hot standby process receiving the target checkpoint.

[0012] Optionally, the method further includes: obtaining a first number of parallel data copies at the time of checkpoint storage and a second number of parallel data copies for the current training task; comparing whether the first number of parallel data copies and the second number of parallel data copies are consistent; and if the first number of parallel data copies and the second number of parallel data copies are consistent, performing the step of determining the target checkpoint.

[0013] Optionally, the method further includes: if the first number of parallel data copies is inconsistent with the second number of parallel data copies, the hot standby process loads a checkpoint from the persistent storage system.

[0014] According to a second aspect of this disclosure, a distributed training cluster is provided, including a memory and a processor, the memory for storing a computer program; the processor for executing the computer program to implement the method described in the first aspect of this disclosure.

[0015] According to a third aspect of this disclosure, a computer-readable storage medium is provided storing a computer program that, when executed by a processor, implements the method described in the first aspect of this disclosure.

[0016] Through the embodiments of this disclosure, after the target computing node fails, a hot standby process with a newly assigned process identifier is started to replace the failed process. By aggregating the checkpoint metadata of all surviving nodes, a global checkpoint availability object is constructed, which includes the original process identifier, the current process identifier, and the number of training iteration steps. Based on this object, the target checkpoint originally belonging to the failed process is matched and loaded into the hot standby process. The binding relationship between the new hot standby process identifier and the checkpoint data ownership is decoupled by the original process identifier, which solves the problem of traditional static backup mapping failing due to process identifier reallocation. Relying on the cluster global checkpoint availability object, cross-node memory sharding fast recovery is achieved without relying on low-speed persistent storage. It can automatically locate effective backups in scenarios of cluster node failure and dynamic topology changes, taking into account both fault recovery speed and distributed training state consistency, effectively improving the fault tolerance capability and operational continuity of the distributed training cluster.

[0017] The features and advantages of the embodiments of this specification will become clear from the following detailed description of exemplary embodiments with reference to the accompanying drawings. Attached Figure Description

[0018] The accompanying drawings, which are incorporated in and form a part of this specification, illustrate embodiments of this specification and, together with their description, serve to explain the principles of these embodiments.

[0019] Figure 1 This is a schematic diagram of a distributed training cluster according to some embodiments; Figure 2 This is a flowchart of checkpoint recovery for a distributed training cluster according to some embodiments; Figure 3 It is a timing diagram of checkpoint recovery of a distributed training cluster according to some embodiments; Figure 4 This is a schematic diagram of a distributed training cluster according to some embodiments. Detailed Implementation

[0020] Various exemplary embodiments of this specification will now be described in detail with reference to the accompanying drawings.

[0021] The following description of at least one exemplary embodiment is merely illustrative and is in no way intended to limit the embodiments of this specification or their application or use.

[0022] It should be noted that similar labels and letters in the following figures indicate similar items; therefore, once an item is defined in one figure, it does not need to be discussed further in subsequent figures.

[0023] This disclosure relates to a checkpoint recovery scheme for a distributed training cluster. Figure 1This is a schematic diagram of the structure of a distributed training cluster that can apply the checkpoint recovery method of the distributed training cluster according to the embodiments of this disclosure. Figure 1 As shown, the distributed training cluster 100 includes multiple compute nodes 110 and at least one hot standby node 120. Each compute node runs several training processes, and each process has a unique process identifier within the cluster. Each compute node 110 maintains local shared memory 111 and backup shared memory 112, used to store checkpoints of its own running processes and backup checkpoints pushed by the running processes of neighboring compute nodes, respectively. The hot standby node 120 may maintain local shared memory 121.

[0024] In this embodiment, a computing node can be provided by a server.

[0025] During normal training in a distributed training cluster, each training process writes the current model parameters, optimizer state, and random number state to local shared memory after each iteration, generates a local checkpoint, and constructs corresponding metadata. The checkpoint type in the metadata is marked as local, and the original process identifier and the current process identifier are both assigned to the process identifier of the current process itself.

[0026] In addition, each process pushes a copy of the local checkpoint to the backup shared memory of the preset neighbor node. When the neighbor node stores the checkpoint, it constructs the corresponding metadata. The checkpoint type in the metadata is marked as backup, the original process identifier is assigned the process identifier of the process that pushed the checkpoint, and the current process identifier is assigned the process identifier of the current process itself.

[0027] The distributed training cluster 100 can also be configured with a persistent storage system (not shown in the figure) as a final fallback recovery resource. The checkpoint recovery process of the distributed training cluster consists of four main stages: checkpoint information collection stage, global checkpoint availability object construction stage, dynamic matching calculation stage, and data transmission execution stage.

[0028] In related technologies, when a computing node fails and is taken over by a hot standby node, the hot standby node is assigned a hot standby process identifier, thus invalidating the original fixed one-to-one backup relationship mapping. For example, in training scenarios employing data parallelism, pipelined parallelism, tensor parallelism, and model parallelism, if one of the computing nodes fails and is replaced by a hot standby node, when the cluster queries backup data based on the pre-configured backup relationship, the process identifier remapping will return the wrong backup node. This causes the hot standby node to be unable to correctly load checkpoint data from neighboring nodes, frequently rolling back to loading from persistent storage, and the recovery time degrades from the expected seconds to minutes.

[0029] To this end, this embodiment of the disclosure introduces an original process identifier field into the checkpoint metadata, constructs a global checkpoint availability object through set communication, and adopts a multi-objective optimized dynamic matching algorithm, which completely decouples the process identifier mapping from the data storage location, significantly improves the success rate of hot standby node checkpoint recovery, and compresses the recovery time from minutes to seconds.

[0030] <First Embodiment> This disclosure provides a checkpoint recovery method for a distributed training cluster. This method can be implemented by the distributed training cluster, specifically, it can be implemented by, for example... Figure 1 The distributed training cluster 100 shown is implemented. For example... Figure 2 As shown, the checkpoint recovery method for the distributed training cluster in this embodiment may include the following steps S210 to S240.

[0031] Step S210: In the event of a failure in the target computing node, start a hot standby process to replace the failed process of the target computing node, and assign a hot standby process identifier to the hot standby process.

[0032] In this embodiment, each compute node deploys multiple training processes, each process is assigned a unique process identifier; each compute node is configured with two independent shared memory blocks, including local shared memory and backup shared memory. The local shared memory stores checkpoints of the processes running on that compute node, while the backup shared memory stores backup checkpoints pushed from other remote processes.

[0033] Checkpoints are periodic snapshots of the training state saved during the training process, including model parameters (including weights, biases, etc.), optimizer state (including momentum, learning rate, etc.), random number generator state (to ensure the reproducibility of training), and the current number of training iterations.

[0034] Shared memory is a memory-sharing mechanism provided at the operating system level, allowing multiple processes on the same computing node to access the same memory region. In distributed training, the training process and the checkpoint saving process can achieve efficient data exchange through shared memory, avoiding the overhead of data copying.

[0035] A hot standby node is a pre-configured backup node in a distributed training cluster that is not involved in training. When an active node fails, a hot standby node can immediately take over its workload, reducing business downtime. Hot standby nodes typically come pre-installed with the necessary training environment and dependencies, allowing for rapid startup of the training process.

[0036] A process identifier is a unique numerical identifier assigned to each training process in distributed training. Process identifiers can start from 0 and are used to identify the process's position within the global process group. The process identifier allows for precise location of each checkpoint slice stored in shared memory.

[0037] Checkpoint shards are portions of the model's parameters or optimizer state held by each process during training using data parallelism or model parallelism. Therefore, checkpoints are also stored in shards, with each process storing its own shard. Restoring a checkpoint requires all processes to collaboratively load the complete set of shards to reconstruct the full training state.

[0038] When a distributed training cluster detects a hardware failure or network disconnection on a certain computing node (i.e., the target computing node), all training processes under that target computing node fail, and training is interrupted. The distributed training cluster starts a hot standby process on a hot standby node to replace all failed processes; and assigns a hot standby process identifier to each hot standby process. The hot standby process identifier can be consistent with the original process identifier of the failed process or be reassigned, and the process topology of the distributed training cluster changes.

[0039] Since the hot standby process has no historical local checkpoints, the distributed training cluster automatically triggers the checkpoint recovery process in this embodiment.

[0040] Step S220: Obtain the metadata of the checkpoints stored by each surviving computing node, and obtain the global checkpoint availability object based on the metadata; wherein, the global checkpoint availability object includes the original process identifier of the process to which all checkpoints originally belonged and the current process identifier of the process to which they currently belong, which are stored by each surviving process.

[0041] All surviving, fault-free computing nodes in the distributed training cluster collect complete metadata corresponding to the two types of checkpoints stored in their local shared memory and backup shared memory in parallel. With the help of distributed set communication, the metadata of all nodes is interconnected and summarized, and a global checkpoint availability object is generated. This global checkpoint availability object fully contains the original process identifier and current process identifier of each checkpoint held by all surviving processes in the distributed training cluster, providing a globally unified query data source for subsequent matching of target checkpoints.

[0042] The distributed training cluster can collect the heartbeat and network connectivity status of each training process in real time. Computing nodes with interrupted heartbeats or broken links are marked as target computing nodes, and the remaining nodes that can read and write shared memory and perform distributed communication are determined to be live computing nodes.

[0043] Metadata is auxiliary information used to describe checkpoint data, including the number of training iterations for the checkpoint, the original process identifier, the current process identifier, the checkpoint type (local / backup), and the checkpoint's sharding information (tensor shape, data type, etc.).

[0044] In this embodiment, all surviving compute nodes within the distributed training cluster can perform memory reads in parallel: reading the metadata of checkpoints stored in local shared memory and backup shared memory respectively. Then, the metadata of checkpoints obtained by all surviving compute nodes is aggregated to obtain a global checkpoint availability object.

[0045] All surviving nodes report all checkpoint metadata on their own machines through distributed aggregated communication, which aggregates all metadata entries of the distributed training cluster, organizes them by storage process dimension, and uniformly collects the original process identifiers and current process identifiers of all checkpoints to form a structured and searchable global checkpoint availability object; the matrix provides the ability to query checkpoints by original process identifier and training iteration steps.

[0046] In some embodiments, the global checkpoint availability object may also include the number of training iterations and / or checkpoint types for all checkpoints stored by each surviving process.

[0047] In some embodiments, obtaining the metadata of checkpoints stored by each surviving compute node and obtaining a global checkpoint availability object based on the metadata includes: obtaining the configuration information of checkpoints stored in the local shared memory of each surviving compute node and the configuration information of checkpoints stored in the backup shared memory; obtaining the metadata of each checkpoint based on the configuration information; and assembling the metadata of each checkpoint to obtain a global checkpoint availability object.

[0048] Specifically, each surviving compute node can read its local shared memory in parallel to extract the configuration information of its own checkpoints, including the number of training iterations and the current process identifier of the currently stored process; each surviving compute node can also read its local backup shared memory in parallel to extract the configuration information of the managed backup checkpoints, including the number of training iterations and the stored_rank field, where the stored_rank field represents the process identifier of the process to which the corresponding checkpoint originally belonged.

[0049] For local checkpoints, the original process identifier is equal to the current process identifier; for backup checkpoints, the original process identifier is equal to the source process identifier that pushed the backup, while the current process identifier is the process identifier that actually stores the backup. Through the original process identifier field, the cluster completely decouples the process identifier mapping from the data storage location, ensuring that even after the hot standby node's process identifier is remapped, the cluster can still accurately locate the checkpoint data source using the original process identifier.

[0050] For locally owned checkpoints, the metadata for each checkpoint is obtained based on the configuration information, which can be such that its original process identifier is the current process identifier. For managed backup checkpoints, the metadata for each checkpoint is obtained based on the configuration information, which can be such that its original process identifier is the process identifier represented by the stored_rank field.

[0051] In some examples, a checkpoint information structure can be defined as follows to describe the metadata of a checkpoint: dataclass CheckpointInfo: - rank: int# Stores the currently stored process at this checkpoint. - step: int# Number of training iterations for checkpoints - type: str# Checkpoint type: 'local' or 'backup' - original_rank: int# The process to which this checkpoint originally belonged - sharded: bool# Whether it is a distributed checkpoint (DP / PP / TP sharding) This structure enables the cluster to accurately identify the source, type, and affiliation of each checkpoint, providing the necessary information for dynamic matching.

[0052] In some examples, each surviving compute node synchronizes its acquired metadata to all surviving compute nodes in the distributed cluster via distributed aggregate communication, enabling each surviving compute node to obtain the complete metadata set of the cluster. Specifically, each surviving compute node broadcasts its stored checkpoint metadata to all surviving compute nodes via aggregate communication (all_gather), and assembles the collected metadata of all processes into a global checkpoint availability object.

[0053] The global checkpoint availability object, indexed by the current process identifier, records metadata for all checkpoints stored by each surviving process in the cluster, including the original process identifier, the current process identifier, the number of training iterations, and the optional checkpoint type. In this way, each surviving process obtains a complete view of all available checkpoints in the cluster, providing a sufficient data foundation for subsequent dynamic matching algorithms.

[0054] In other examples, each surviving compute node can synchronize its acquired metadata to the hot standby node via distributed collection communication, allowing the hot standby node to assemble the collected metadata of all processes into a global checkpoint availability object.

[0055] The metadata of each checkpoint is assembled to obtain a global checkpoint availability object, which may include: an initial matrix storage structure with the current process identifier as the grouping dimension, which classifies and stores the metadata of all checkpoints corresponding to the same live process; and a built-in search interface that supports inputting process identifier and training iteration steps to quickly filter matching checkpoints.

[0056] In these examples, by distinguishing between local shared memory and backup shared memory and collecting two types of metadata in parallel through dual channels, and standardizing and filling the original process identifier field, and then using global set communication to aggregate all information to construct a global checkpoint availability object, it is possible to completely collect all checkpoint resources of both cluster own and backup. This breaks the limitation of traditional solutions that only read a single pre-configured backup and lack a global checkpoint view, and provides a complete and accurate global retrieval data source for dynamic matching logic.

[0057] Step S230: Based on the global checkpoint availability object, determine the target checkpoint in the checkpoints stored by each surviving process that matches the original process identifier of the faulty process, and determine the sending process corresponding to the target checkpoint.

[0058] Based on the global checkpoint availability object generated in the preceding steps, and using the process identifier of the faulty process as the matching benchmark, the checkpoints whose original process identifiers match the faulty process are selected from all checkpoints held by all surviving processes in the cluster as the target checkpoints to be transmitted and restored.

[0059] In this embodiment, the target checkpoint that matches the original process identifier with the faulty process can be a checkpoint trained by the faulty process that was originally replaced by the hot standby process.

[0060] In some examples, when multiple hot standby processes are started, the integrity verification phase can use multi-threading to determine the target checkpoint that matches the faulty process replaced by each hot standby process, thereby reducing the time consumed by single-step dynamic matching.

[0061] In this embodiment, the sending process is a live process that stores the target checkpoint and is the live process that needs to send the stored target checkpoint to the hot standby process.

[0062] In some embodiments, determining the target checkpoint that matches the hot standby process from the checkpoints stored by each surviving process based on the global checkpoint availability object includes the following steps S231 to S232: Step S231: Determine the target training iteration steps based on the global checkpoint availability object; wherein, the target training iteration steps represent the maximum number of training iteration steps for each process identifier to have a corresponding checkpoint in the global checkpoint availability object.

[0063] Based on the iteration step information of all checkpoints within the global checkpoint availability object, training iteration versions with complete backups for all hot standby processes to be recovered in the cluster are selected. The iteration step with the largest value, representing the latest training progress of all computing nodes, is selected as the target training iteration step. This target training iteration step is the version benchmark for all subsequent hot standby processes to load checkpoints uniformly, ensuring that the training status is synchronized after the entire cluster is restored.

[0064] The training iteration count is the iteration count recorded when a checkpoint is generated after each complete round of data iteration in the training process. It is used to characterize the training progress; a larger value indicates a more recent training snapshot. This step uses the training iteration count as the filtering dimension, prioritizing the latest snapshot to minimize training rollbacks caused by failures and reduce wasted computing power.

[0065] In some embodiments, determining the target training iteration steps based on the global checkpoint availability object includes: obtaining the training iteration steps of all checkpoints based on the global checkpoint availability object and performing deduplication; traversing the deduplicated training iteration steps in ascending order, verifying that each process identifier has a checkpoint in the global checkpoint availability object with a training iteration step number equal to the currently traversed training iteration step number; if so, stopping the traversal and using the currently traversed training iteration step number as the target training iteration step number; otherwise, continuing the traversal.

[0066] In this embodiment, the training iteration steps of all checkpoints in the global checkpoint availability object can be extracted and stored in a set to remove duplicates, resulting in a set of training iteration steps without repetition. The deduplicated training iteration step set is then sorted in descending order of numerical value, generating a traversal queue of iteration steps from newest to oldest. Individual training iteration steps are sequentially retrieved from the traversal queue, and integrity checks are performed: for the currently traversed training iteration step, the global checkpoint availability object is searched to determine whether each process identifier in the distributed training cluster (including the process identifiers of surviving processes and failed processes) has a checkpoint in the global checkpoint availability object with the training iteration step number of the currently traversed step. If each process identifier has a checkpoint in the global checkpoint availability object with the training iteration step number of the currently traversed step, the traversal stops, and the current iteration step is determined as the target training iteration step. If no process identifier has a checkpoint in the global checkpoint availability object with the training iteration step number of the currently traversed step, the next training iteration step is retrieved from the queue, and the integrity check is repeated.

[0067] If a process identifier has a checkpoint in the global checkpoint availability object where the original process identifier is the same as the process identifier and / or the current process identifier is the same as the process identifier, and the number of training iterations is the number of training iterations in the current traversal, then it is determined that the process identifier has a checkpoint in the global checkpoint availability object where the number of training iterations is the number of training iterations in the current traversal. If a process identifier does not have a checkpoint in the global checkpoint availability object where the original process identifier or the current process identifier is the same as the process identifier, and the number of training iterations is the number of training iterations in the current traversal, then it is determined that the process identifier does not have a checkpoint in the global checkpoint availability object where the number of training iterations is the number of training iterations in the current traversal.

[0068] In these examples, by extracting and deduplicating all iteration steps, prioritizing the latest training snapshot in descending order, and performing a step-by-step global backup integrity check on each hot standby process, the maximum number of iteration steps for which all hot standby processes have complete backups can be accurately selected. Compared to implementations that randomly select iteration steps or only check the backup of a single process, this effectively avoids training restart failures caused by inconsistent shard versions, maximizes the preservation of the latest training progress, and reduces computational cost.

[0069] In other embodiments, determining the target training iteration steps based on the global checkpoint availability object includes: using each process identifier and the current training iteration step as keywords, calling the built-in query interface of the global checkpoint availability object to perform a query; if the query result returns a non-empty checkpoint list, it is determined that the training iteration step has passed the integrity check; if the query result is empty, it is determined that the current training iteration step has not passed the integrity check, and it can automatically roll back to the next training iteration step and query again, ensuring that the checkpoint steps loaded by all processes are completely consistent, satisfying the consistency constraints of distributed training.

[0070] Step S232: Based on the global checkpoint availability object, determine the checkpoint whose original process identifier is the faulty process and whose training iteration step number is the target training iteration step number, and use it as the target checkpoint.

[0071] Based on the generated global checkpoint availability objects and the determined target training iteration steps, firstly, select all available backup checkpoints corresponding to the faulty processes replaced by the hot standby process and whose training iteration steps match the target training iteration steps. Then, select the optimal node from the above-mentioned surviving processes as the sending process according to the preset selection rules. The sending process locally retains the target checkpoint fragment to be transmitted to the hot standby process.

[0072] The target checkpoint is the checkpoint where the original process identifier is the process identifier of the faulty process replaced by the hot standby process, and the training iteration step number is the target training iteration step number.

[0073] In this embodiment, the process identifier of the faulty process replaced by the hot standby process and the target training iteration steps can be used as matching keywords to retrieve global checkpoint availability objects, and checkpoints whose original process identifier is equal to the process identifier of the faulty process replaced by the hot standby process and whose training iteration steps are equal to the target training iteration steps can be selected as target checkpoints.

[0074] In these examples, the maximum target training iteration steps for all failed processes to have complete backups are first filtered using a global checkpoint availability object. This ensures a consistent training version after cluster recovery and reduces computational rollback losses. The mapping failure issue caused by the new hot standby process identifier is decoupled from the original process identifier, ensuring the integrity of fault recovery and the consistency of training state. In some embodiments, determining the sending process corresponding to the target checkpoint includes: determining the candidate live processes currently belonging to the target checkpoint based on the global checkpoint availability object; and selecting one of the candidate live processes as the sending process.

[0075] The candidate live process to which the target checkpoint currently belongs can be the live process that actually stores the candidate checkpoint, specifically the live process represented by the current process identifier corresponding to the target checkpoint.

[0076] In these examples, by matching faulty process identifiers with original process identifiers and superimposing the target training iteration steps as a dual constraint to screen candidate checkpoints of the same source and version, and by summarizing all candidate live processes holding backups, a single sending process is selected based on the best one. This can decouple data ownership from dynamically changing process identifiers based on the original process identifiers, and solve the problem of backup matching failure after the hot standby process is assigned a new identifier. At the same time, it can collect multiple replica backups of the cluster to form a candidate pool, which has the ability to redundancy for multiple node failures. Furthermore, it selects only a single optimal node to transmit data, avoiding network congestion caused by concurrent transmission of multiple nodes, which greatly improves the success rate of hot standby node failure recovery and shortens the overall recovery time.

[0077] In some embodiments, a process may be randomly selected from the candidate surviving processes as the sending process.

[0078] In some embodiments, selecting one candidate surviving process as the sending process includes: determining the priority of the candidate surviving process based on the network distance between the candidate surviving process and the hot standby process; and selecting the candidate surviving process with the highest priority as the sending process.

[0079] This step uses the network communication distance between the candidate live process and the hot standby process as the evaluation criterion to prioritize all candidate live processes with valid candidate checkpoints. The closer the communication distance and the smaller the transmission overhead, the higher the priority. Finally, the candidate live process with the highest priority is selected as the process that sends the checkpoint to the hot standby process.

[0080] Network distance is determined based on the physical deployment location of processes. Processes within the same server have the shortest network distance; processes on different servers within the same data center have the next shortest; and processes on servers in different data centers have the longest network distance. Network distance is used to quantify the latency and bandwidth loss of data transmission between processes, and is a core indicator for measuring transmission costs, providing an objective basis for prioritization.

[0081] The closer the network distance between the candidate live process and the hot standby process, the higher the priority of the candidate live process; the farther the network distance between the candidate live process and the hot standby process, the lower the priority of the candidate live process.

[0082] Specifically, the process can be as follows: first, mark the candidate surviving processes deployed on the same server as the hot standby process as first-level priority; mark the candidate surviving processes on different servers but in the same data center as the hot standby process as second-level priority; mark the candidate surviving processes in different data centers as third-level priority; sort all candidate surviving processes in the order of first-level priority > second-level priority > third-level priority; and determine the candidate surviving process ranked first in the sorting results as the sending process.

[0083] In these examples, by differentiating network distances by server and data center level and dividing them into multiple priority levels, nodes with local high-speed memory communication or low-latency links within the same data center are automatically selected as the sending process. This can significantly reduce network bandwidth consumption and data transmission latency caused by large-volume cross-data center transmissions, and effectively compress the overall fault recovery time of hot standby nodes.

[0084] In these examples, by filtering candidate checkpoints of the same source and version through dual-condition filtering, summarizing all backup storage processes, and selecting the sending process based on network distance, the system can automatically select the data source with the lowest transmission latency in scenarios with multiple backups coexisting and dynamic changes in cluster topology, thereby improving the recovery speed of hot standby nodes and making full use of cluster multi-replica redundancy to improve recovery reliability.

[0085] In other examples, network distance can be differentiated into only two levels: processes on the same local server are given high priority, while all processes across servers are given low priority, simplifying the priority determination logic and reducing the computational cost of the matching phase.

[0086] In some other examples, when prioritizing, the real-time bandwidth load index of the computing node can be superimposed. Under the same network distance, candidate surviving processes with higher bandwidth idle rate are selected first to avoid single-node network congestion slowing down the transmission speed.

[0087] In some other examples, if there are multiple highest-priority candidate surviving processes, a round-robin strategy can be used to select the sending process in turn, thereby balancing the transmission task pressure of each computing node in the distributed training cluster.

[0088] In some embodiments, the method further includes: obtaining a first number of parallel data copies at the time of checkpoint storage and a second number of parallel data copies for the current training task; comparing whether the first number of parallel data copies and the second number of parallel data copies are consistent; and if the first number of parallel data copies and the second number of parallel data copies are consistent, performing the step of determining the target checkpoint.

[0089] The system reads and compares the first number of parallel data copies recorded when the checkpoint is generated with the second number of parallel data copies of the currently running hot standby task. Only when the number of parallel data copies is the same will the system be allowed to enter the memory recovery process based on the global checkpoint availability object to match the target checkpoint. If the two values ​​do not match, the memory shard matching logic will not be executed.

[0090] The first number of parallel data copies can be obtained by parsing historical checkpoint metadata and reading the data parallel configuration value of the training task at the time of the checkpoint. The first number of parallel data copies represents the parallel sharding rules when generating this batch of shard checkpoints. The sharding logic is strongly bound to this number of copies and is the core criterion for determining whether shards can be directly spliced ​​and restored.

[0091] The second data parallelism count can be obtained by reading the parameters of the training tasks currently running in the cluster and those awaiting hot standby process access. The second data parallelism count represents the sharding rule of the current distributed training and serves as a benchmark for compatibility verification.

[0092] The number of data parallel copies is a key parameter affecting the checkpoint sharding method. If the second number of data parallel copies for the current training task is inconsistent with the first number of data parallel copies configured when the checkpoint is stored, the optimizer state sharding method for the checkpoint changes, making it impossible to load directly and requiring special handling. When the first number of data parallel copies is consistent with the second number of data parallel copies, the cluster executes steps S230 and S240 to load the checkpoint from the neighbor node's memory backup.

[0093] By comparing the number of parallel data copies stored at the checkpoint with the number of parallel data copies in the current task, the target checkpoint matching process is only executed when the two are consistent. This can identify data parallel configuration change scenarios in advance, avoid parameter loading anomalies and training task interruptions caused by incompatibility between the old and new checkpoint sharding dimensions, and enable low-latency memory backup and recovery mechanisms only for compatible scenarios while ensuring the integrity of distributed training data and the correctness of calculations, thus balancing fault recovery efficiency and training operation stability.

[0094] In some embodiments, the method further includes: if the first number of parallel data copies is inconsistent with the second number of parallel data copies, the hot standby process loads a checkpoint from the persistent storage system.

[0095] If the number of parallel copies of the first data stored at the checkpoint is different from the number of parallel copies of the second data in the current training task, it is determined that the memory sharding structure cannot adapt to the current parallel topology. The cross-node memory backup matching and transmission logic is no longer executed, and the degradation scheme is switched. The hot standby process can directly read the complete unsharded global checkpoint from the persistent storage system to complete the recovery.

[0096] The persistent storage system is a clustered, externally distributed persistent storage system, which can be HDFS, object storage, shared file storage, etc. During the training cycle, it periodically outputs complete global checkpoints, independent of memory sharding on each node. As a fallback recovery medium in scenarios where the number of data parallel copies changes, the persistent storage system stores unsharded, full, and complete model parameters, unaffected by changes in the number of data parallel copies, ensuring the normal completion of the fault recovery process.

[0097] The hot standby process loads checkpoints from the persistent storage system. This can be achieved by the hot standby process initiating a persistent storage read request, pulling the complete global checkpoint, and completing parameter initialization.

[0098] In other examples, after the hot standby process reads the complete checkpoint from the persistent storage, it re-shards the complete checkpoint according to the current number of parallel copies of the second data, adapting it to the current cluster parallel topology.

[0099] If the first number of parallel data copies is inconsistent with the second number of parallel data copies, the hot standby process loads the checkpoint from the persistent storage system and performs the corresponding resharding operation to accommodate the second number of parallel data copies.

[0100] When the number of parallel data copies stored at the checkpoint is inconsistent with the number of parallel data copies in the current task, the system switches to persistent storage to load the complete global checkpoint. This avoids model loading anomalies caused by conflicts between different parallel sharding dimensions, provides a reliable fallback recovery path for scenarios involving scaling up or down the number of parallel data copies in the cluster, ensures the continuous and executable nature of the fault recovery process, and balances the efficiency of fast memory recovery with the robustness of recovery under parallel configuration change scenarios.

[0101] In some examples, when multiple nodes fail simultaneously, the system can automatically find the corresponding target checkpoint and sending node for each hot standby process without manual intervention; the cluster automatically completes the globally optimal allocation. This approach supports multi-node concurrent failure scenarios, significantly improving the cluster's robustness.

[0102] Step S240: Load the target checkpoint from the sending process to the hot standby process.

[0103] Specifically, the selected sending process reads the target checkpoint fragments from local storage and transmits them to the hot standby process via cluster communication. The hot standby process receives the fragment data and writes it into its own memory, completing the training state reconstruction and enabling the hot standby process to continue executing distributed training from the target training iteration steps.

[0104] In some examples, the sending process may send both the target checkpoint's metadata and checkpoint fragment data to the hot standby process.

[0105] After reading the metadata and data buffer, the sending process sends the metadata and data to the hot standby process using distributed communication primitives, based on the hot standby process identifier. Upon receiving the metadata, the hot standby process initializes its local shared memory according to the buffer size, sets the metadata, and then receives the data into its local shared memory, completing the checkpoint load.

[0106] Through the embodiments of this disclosure, after the target computing node fails, a hot standby process with a newly assigned process identifier is started to replace the failed process. By aggregating the checkpoint metadata of all surviving nodes, a global checkpoint availability object containing the original process identifier, the current process identifier, and the number of training iteration steps is constructed. Based on this matrix, a target checkpoint belonging to the failed process and with the same version is matched and loaded into the hot standby process. The binding relationship between the new hot standby process identifier and the checkpoint data ownership is decoupled by the original process identifier, solving the problem of traditional static backup mapping failing due to process identifier reallocation. Relying on the cluster global checkpoint availability object, cross-node memory sharding fast recovery is achieved without relying on low-speed persistent storage. It can automatically locate effective backups in scenarios of cluster node failure and dynamic topology changes, taking into account both fault recovery speed and distributed training state consistency, effectively improving the fault tolerance capability and operational continuity of the distributed training cluster.

[0107] In some embodiments, the global checkpoint availability object also includes the checkpoint type of all checkpoints stored by each live process, where the checkpoint type indicates whether the checkpoint is stored in local shared memory or backup shared memory in the corresponding live compute node; loading the target checkpoint into the hot standby process includes: the sending process reading the target checkpoint from local shared memory or backup shared memory according to the checkpoint type of the target checkpoint, and sending the target checkpoint to the hot standby process according to the hot standby process identifier; the hot standby process receiving the target checkpoint.

[0108] In this embodiment, the global checkpoint availability object additionally stores the type identifier corresponding to each checkpoint, which is used to distinguish whether the checkpoint is stored in the node's local shared memory or the backup shared memory. When executing the target checkpoint loading process, the sending process directly reads the checkpoint type recorded in the matrix, selects the corresponding shared memory to read the fragment data, and then transmits the checkpoint fragments according to the hot standby process identifier. Finally, the hot standby process completes the data reception.

[0109] Specifically, when the target checkpoint type is local, the sending process can read the target checkpoint from local shared memory; when the target checkpoint type is backup, the sending process can read the target checkpoint from backup shared memory.

[0110] In this embodiment, the sending process uses the newly allocated process identifier of the hot standby process as the communication target, which can accurately push fragmented data, avoid the waste of cluster bandwidth caused by broadcast transmission, and ensure that data is only transmitted to the hot standby process to be recovered.

[0111] The hot standby process creates a memory buffer, receives the sharded data pushed by the sending process, caches it in local shared memory, completes the local storage of sharded data, reconstructs the complete training context, and supports subsequent training.

[0112] In these examples, by pre-delegate the checkpoint types in the global checkpoint availability object to distinguish between local and backup storage locations, the sending process can directly read the target checkpoint in the corresponding shared memory based on the type, without having to repeatedly traverse the two shared memory locations, thus reducing memory retrieval overhead. Then, based on the hot standby process identifier, the checkpoint fragments of the target checkpoint are transmitted, avoiding the bandwidth waste caused by broadcast transmission. High-speed loading is achieved by relying on memory fragments. While ensuring that the fragmented data is accurately transmitted to the hot standby process, the overall time of fault recovery is further reduced, and the efficiency of distributed training fault recovery is improved.

[0113] In other embodiments, the sending process may determine whether the current process identifier and the original process identifier of the target checkpoint are the same. If they are the same, the target checkpoint is read from the local shared memory; if they are different, the target checkpoint is read from the backup shared memory.

[0114] In some examples, for live processes that need to send checkpoints to multiple target processes, the cluster can use a thread pool to perform parallel multi-path sending, improving transmission efficiency through asynchronous communication and further shortening the overall recovery time. This approach parallelizes the checkpoint recovery operations of multiple hot standby processes, reducing overall recovery time.

[0115] Figure 3 This is a timing diagram of checkpoint recovery for a distributed training cluster according to some embodiments. For example... Figure 3As shown, the cluster in this embodiment includes multiple compute nodes and one hot standby node. Each compute node carries a continuous training process, rank. During normal training, each node periodically writes its local checkpoint to its local shared memory and entrusts a copy of the checkpoint to the backup shared memory of its neighboring nodes. Each surviving node can aggregate checkpoint metadata and construct a global checkpoint availability object through global communication to achieve dynamic matching and recovery after a failure. At the 20th second of training, compute node Node2 in the cluster experiences a hardware failure, causing the training process Rank16 carried by Node2 to exit abnormally, resulting in the overall interruption of the distributed training task of the cluster.

[0116] At the 30th second, the cluster scheduling module detected a failure in Node2 and immediately started a preset hot standby node Node2' to replace the failed node Node2. A new process identifier Rank16 was assigned to the recovery process on the hot standby node to completely take over the training task of the original failed process Rank16. The hot standby process has no historical checkpoint data cpkt locally and enters the fault recovery process.

[0117] At the 31st second, all normally surviving computing nodes in the cluster started the checkpoint metadata collection process in parallel, reading all checkpoint metadata stored in the local shared memory and backup shared memory respectively. Each piece of metadata carries the original process identifier, the current process identifier, the number of training iterations, and the checkpoint type, and fully collects all self-owned checkpoints and managed backup checkpoint information in the cluster.

[0118] At the 33rd second, all surviving processes synchronized their metadata through the all_gather global collection communication, aggregated all checkpoint metadata of the cluster, and assembled a global checkpoint availability object containing the original process identifier, current process identifier, and training iteration number of all checkpoints stored by each surviving process.

[0119] At the 34th second, dynamic matching logic is executed based on the global checkpoint availability object. First, the maximum number of training iteration steps in which the process identifier of each surviving process has a corresponding checkpoint in the global checkpoint availability object is selected as the target training iteration step. Then, the hot standby process Rank16 is selected as the matching target, and candidate checkpoints with the original process identifier 16 and the iteration step number matching the target training iteration step number are selected, i.e., cpkt=Rank16. The optimal sending process is determined by combining the network distance priority, and the mapping relationship load_targets

[16] =17 is generated, i.e. the surviving process Rank17 is used as the sending process to provide the target checkpoint for the hot standby process Rank16.

[0120] Between the 35th and 40th seconds, the sending process Rank17 reads the target checkpoint fragment data from the corresponding local shared memory or backup shared memory according to the checkpoint type recorded in the global checkpoint availability object, and transmits the target checkpoint metadata and fragment data to the hot standby process Rank16 through distributed communication primitives according to the process identifier of the hot standby process Rank16.

[0121] At the 40-second mark, the hot standby process Rank16 fully received the target checkpoint data transmitted from the remote end, completed the local memory writing and training state initialization, and successfully loaded the target checkpoint adapted to its own task.

[0122] At the 45th second, all surviving training processes in the cluster completed training state synchronization, the cluster training topology was restored to complete, and the distributed training tasks resumed normal iteration from the target training iteration step, completing the entire process of node failure, hot standby replacement, and dynamic matching checkpoint recovery.

[0123] <Second Embodiment> This embodiment provides a distributed training cluster. Figure 4 A schematic diagram of the hardware structure of the electronic device is shown.

[0124] like Figure 4 As shown, the distributed training cluster 100 includes a processor 410 and a memory 420. The memory can be used to store computer programs, and the processor can be used to retrieve the computer programs from the memory to execute any method embodiment of this disclosure. The processor can be one or more, and these processors can execute instructions individually or jointly. Similarly, the memory can be one or more, and these memories can store the aforementioned computer programs individually or jointly.

[0125] This disclosure also provides a computer-readable storage medium storing a computer program thereon, which, when executed by a processor, implements the checkpoint recovery method for a distributed training cluster in any embodiment of this disclosure. Optionally, the computer-readable storage medium may be a non-transitory storage medium, but is not limited thereto; it may also be a temporary storage medium.

[0126] The various embodiments in this specification are described in a progressive manner, with each embodiment focusing on the differences from other embodiments. The same or similar parts between the various embodiments can be referred to each other.

[0127] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and apparatuses according to various embodiments of this specification. In this regard, each block in a flowchart or block diagram may represent a module, unit, or part of a circuit. In some alternative implementations, the functions marked in the blocks may occur in a different order than those marked in the drawings. For example, two consecutive blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, can be implemented in hardware that performs the specified function or action, or in a combination of dedicated hardware and computer instructions. It will be known to those skilled in the art that implementation in hardware, implementation in software, and implementation in a combination of software and hardware are equivalent.

[0128] Various embodiments of this specification have been described above. These descriptions are exemplary and not exhaustive, nor are they limited to the disclosed embodiments. Many modifications and variations will be apparent to those skilled in the art without departing from the scope of the described embodiments. The terminology used herein is chosen to best explain the principles, practical application, or improvement of the technology in the market, or to enable others skilled in the art to understand the embodiments disclosed herein.

Claims

1. A checkpoint recovery method for a distributed training cluster, characterized in that, The distributed training cluster includes multiple computing nodes, each computing node includes at least one process, and each process has a unique process identifier; the method includes: In the event of a failure in the target computing node, a hot standby process is initiated to replace the failed process of the target computing node, and a hot standby process identifier is assigned to the hot standby process. Obtain the metadata of the checkpoints stored in each surviving compute node, and obtain a global checkpoint availability object based on the metadata; wherein, the global checkpoint availability object includes the original process identifier of the process to which all checkpoints stored in each surviving process originally belonged, and the current process identifier of the process to which they currently belong; Based on the global checkpoint availability object, determine the target checkpoint in the checkpoints stored by each surviving process that matches the original process identifier of the faulty process, and determine the sending process corresponding to the target checkpoint; The target checkpoint is loaded from the sending process to the hot standby process.

2. The method according to claim 1, characterized in that, The global checkpoint availability object also includes the training iteration steps of all checkpoints stored by each surviving process; The step of determining the target checkpoint that matches the original process identifier of the failed process among the checkpoints stored by each surviving process based on the global checkpoint availability object, and determining the surviving process storing the target checkpoint as the sending process, includes: The target number of training iterations is determined based on the global checkpoint availability object; wherein, the target number of training iterations represents the maximum number of training iterations for each process identifier to have a corresponding checkpoint in the global checkpoint availability object; Based on the global checkpoint availability object, a checkpoint whose original process identifier is the process identifier of the faulty process and whose training iteration steps are the target training iteration steps is determined as the target checkpoint.

3. The method according to claim 2, characterized in that, Determining the target training iteration steps based on the global checkpoint availability object includes: Based on the global checkpoint availability object, obtain the training iteration steps of all checkpoints and perform deduplication. The deduplicated training iteration steps are traversed in order from newest to oldest. It is verified that each process identifier has a checkpoint in the global checkpoint availability object with the training iteration steps of the current iteration step. If so, the traversal is stopped and the current training iteration step is taken as the target training iteration step. If not, the traversal continues.

4. The method according to claim 2, characterized in that, The step of determining the sending process corresponding to the target checkpoint includes: Based on the global checkpoint availability object, determine the candidate live process to which the target checkpoint currently belongs; Select one of the candidate surviving processes as the sending process.

5. The method according to claim 4, characterized in that, Selecting one of the candidate surviving processes as the sending process includes: The priority of the candidate surviving process is determined based on the network distance between the candidate surviving process and the hot standby process; The candidate surviving process with the highest priority is selected as the sending process.

6. The method according to claim 1, characterized in that, The process of obtaining the metadata of checkpoints stored on each surviving compute node and obtaining a global checkpoint availability object based on the metadata includes: Obtain the configuration information of checkpoints stored in the local shared memory of each surviving computing node and the configuration information of checkpoints stored in the backup shared memory; Metadata for each checkpoint is obtained based on configuration information; The metadata of each checkpoint is assembled to obtain the global checkpoint availability object.

7. The method according to claim 1, characterized in that, The global checkpoint availability object also includes the checkpoint type of all checkpoints stored by each surviving process. The checkpoint type indicates whether the checkpoint is stored in local shared memory or backup shared memory in the corresponding surviving compute node. The step of loading the target checkpoint into the hot standby process includes: The sending process reads the target checkpoint from local shared memory or backup shared memory according to the checkpoint type of the target checkpoint, and sends the target checkpoint to the hot standby process according to the hot standby process identifier; The hot standby process receives the target checkpoint.

8. The method according to claim 1, characterized in that, The method further includes: Get the number of parallel copies of the first data at checkpoint storage and the number of parallel copies of the second data for the current training task; Compare whether the number of parallel data copies in the first data segment is consistent with the number of parallel data copies in the second data segment; If the number of parallel data copies is the same as the number of parallel data copies, then the step of determining the target checkpoint is performed.

9. The method according to claim 8, characterized in that, The method further includes: If the number of parallel data copies in the first case is inconsistent with the number of parallel data copies in the second case, the hot standby process loads a checkpoint from the persistent storage system.

10. A distributed training cluster, characterized in that, It includes a memory and a processor, the memory being used to store a computer program; the processor being used to execute the computer program to implement the method according to any one of claims 1-9.

11. A computer-readable storage medium, characterized in that, The system contains a computer program that, when executed by a processor, implements the steps of the checkpoint recovery method according to any one of claims 1-9.