Large model check point disaster recovery system based on network calculation and asynchronous check points

Through distributed training systems and asynchronous checkpoint technology, the interruption problem caused by GPU failures in large-scale deep learning model training is solved, efficient checkpoint preservation and rapid failure recovery are achieved, and the system reliability and training efficiency are improved.

CN120276894AActive Publication Date: 2025-07-08JIUWEI DIGITAL INTELLIGENCE (BEIJING) TECHNOLOGY CO LTD

Patent Information

Application Number
CN202510321034.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-18
Publication Date
2025-07-08
Estimated Expiration
2045-03-18

AI Technical Summary

Technical Problem

During the training process of large-scale deep learning model, training interruptions caused by GPU failures and high overhead problems caused by frequent storage of checkpoints affect training efficiency and reliability.

Method used

The disaster recovery system based on network computing and asynchronous checkpoints is adopted. Through distributed training systems and asynchronous programming methods, CPU memory and remote CPU memory save checkpoints, reduce checkpoint storage frequency blockage, use idle time to conduct checkpoint transmission, and combine hierarchical storage design and data parallel strategy to achieve rapid failure recovery.

Benefits of technology

It realizes high-frequency checkpoint preservation without increasing training throughput overhead, fast failure recovery, and reduces training and recovery time. It is suitable for existing distributed parallel training methods, improving the high performance and reliability of the system.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120276894A_ABST
    Figure CN120276894A_ABST
Patent Text Reader

Abstract

The invention discloses a large model check point disaster recovery system based on network computing and asynchronous check points, the disaster recovery system adopts a hierarchical storage design and comprises a plurality of groups, each group comprises at least two computing nodes, and the computing nodes in the groups mutually store check points of each other; the computing node at least comprises a CPU, a memory, an NIC and a plurality of GPUs; the last check point of each GPU is stored in a CPU memory of the current computing node, and the last check point is stored in CPU memories of adjacent computing nodes in the same group through a network; and the check points are stored in the storage nodes by using the characteristic of data parallelism and network calculation. According to the large-model disaster recovery system under the distributed system, rapid fault recovery can be provided for large-model training, high-frequency check points are achieved, time expenditure during training and fault recovery is reduced, and extra training throughput expenditure cannot be generated. Meanwhile, the system is suitable for the existing distributed parallel training method and training framework.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of deep learning large models, and in particular to a large model checkpoint disaster tolerance system based on network computing and asynchronous checkpoints. Background Art

[0002] Deep learning models have demonstrated excellent performance in multi-field tasks including computer vision and natural language processing, and are widely used in scenarios such as image recognition, autonomous driving, machine translation, text generation, medical imaging, financial services, intelligent assistants, and personalized recommendations. Since the American company Google proposed the Transformer architecture and introduced the self-attention mechanism in 2017, large language models based on the Transformer architecture have developed rapidly. Such models have unprecedented language understanding and generation capabilities and have received extensive attention in the academic and industrial fields. However, such models often have tens of billions of parameters, making the training process extremely costly, requiring a large number of GPUs and training durations of several weeks or even months. This large-scale computing system poses huge challenges in reliable training. More specifically, due to the large scale and high synchronization of GPU training tasks, hardware failures are common during the training process, and even a single GPU failure will interrupt the entire training process, resulting in the need to restart. Therefore, saving the current state during model training so that it can be restored when needed is the basis and core of the large model disaster tolerance strategy.

[0003] Model checkpoints are an important part of large model training, but checkpointing is an expensive process because each time the latest weight file is saved as a checkpoint, it blocks the training process. However, not performing checkpoints or reducing the checkpoint frequency will result in significant losses in training progress. At the same time, when software or hardware errors occur and the training process needs to be restarted, all training tasks must stop and restart from the last saved checkpoint. Summary of the Invention

[0004] In order to reduce the fault recovery overhead in existing deep learning training, this patent proposes a large model disaster tolerance system under a distributed training system. By utilizing network computing and storage-side computing power, it realizes fast checkpoint storage during the large model training process without explicitly determining the checkpoint saving frequency. On the other hand, to accelerate the recovery speed in case of faults, the high bandwidth of CPU memory (the memory of a computing platform represented by HGX, which is distinguished from GPU video memory and is thus called CPU memory) is utilized to save the latest checkpoints in local CPU memory and remote CPU memory.

[0005] The present invention discloses a disaster tolerance system for large model checkpoints based on network computing and asynchronous checkpoints. The disaster tolerance system adopts a hierarchical storage design. The disaster tolerance system includes multiple groups, and each group includes at least two computing nodes. The computing nodes within a group save the checkpoints of each other.

[0006] The computing node includes at least a CPU, a memory, a network transmission module, and multiple GPUs.

[0007] The most recent checkpoint of each GPU is saved in the CPU memory of the current computing node, and the most recent checkpoint is also saved in the CPU memory of adjacent computing nodes within the same group through the network transmission module.

[0008] Furthermore, when writing the checkpoint to the local CPU memory and adjacent CPU memories, after each iteration's weight update step, the write operation is called through an asynchronous programming method without blocking the training process.

[0009] Furthermore, when writing the checkpoint to adjacent CPU memories, by utilizing the idle time in the training process for checkpoint communication, the impact of checkpoint communication on training communication can be reduced.

[0010] The method for confirming idle time includes the following steps:

[0011] Step 1, during each iteration, set hooks at the entrance and exit of all collective communication operations for timestamp recording; record the start time and end time of each communication operation, and additionally collect the communication packet size, latency, and current network bandwidth utilization rate.

[0012] The communication operations include gradient synchronization, parameter acquisition, and RDMA calls.

[0013] Step 2, for all recorded communication intervals, construct a continuous timeline and mark the communication process as occupying a "busy interval".

[0014] Perform discretization processing on the entire timeline, divide the time into several tiny time slices, and mark whether there is a communication operation within each time slice; count the distribution and cumulative duration of the "busy intervals" in each iteration, and calculate the candidate idle time window.

[0015] Step 3, traverse the discretized timeline, merge adjacent time windows according to the "busy interval" or "idle" state to generate a list of continuous idle time periods.

[0016] And set an initial minimum idle threshold, and only the idle segments longer than this duration are considered as the truly available idle time periods for checkpoint transmission.

[0017] Step 4: Use a sliding window to statistically analyze the duration distribution of the idle time periods detected in the most recent N iterations, calculate their mean value, and use a simple exponential smoothing algorithm to predict the possible length of the idle time period in the next iteration:

[0018] S t = αX t +(1 - α)S t-1

[0019] where X t is the length of the idle time period detected in the current iteration, S t is the smoothed value at the current moment, S t-1 is the previous smoothed value, and α is the smoothing coefficient, which determines the weight distribution between the new observation value and the historical smoothed sequence; in the first round, there is no historical data, and the first observation value is used for initialization, setting S0 = X0;

[0020] Step 5: When the detected idle interval exceeds the current dynamic threshold, the checkpoint transfer operation is triggered; if the length of the idle interval is not sufficient to complete the entire transfer task, the data transfer is split into multiple times, and the "partial transfer - feedback - judge idle continuation" method is used for dynamic supplementary transfer; at the same time, the actual transfer delay and the network resources already used are monitored online. If it is found that the network utilization rate suddenly increases during the transfer process, the transfer rate is dynamically reduced, or the transfer is temporarily suspended and waits for the next idle window;

[0021] Step 6: Dynamically adjust the minimum idle threshold according to the distribution. When the idle time is generally short in the past several iterations, the threshold can be reduced; when the idle time is long, the threshold is increased to reserve enough time for the transfer of larger data blocks; when it is detected that the current network utilization rate decreases, the idle judgment threshold can be appropriately reduced to make the best use of the short idle time as much as possible.

[0022] Furthermore, the large model checkpoint disaster tolerance system based on network calculation and asynchronous checkpoints further includes storage nodes, and the storage nodes include multiple non - volatile storage devices and fewer computing power devices;

[0023]

[0024] FLOPs 前向传播+反向传播 represents the floating - point computation volume of the forward propagation and the backward propagation, T 前向传播+反向传播 represents the time required for the forward propagation and the backward propagation, T 参数更新 represents the time required for parameter update. Calculate P 存储 represents the computing performance required for the group where the storage node is located, and the total computing power within the group shall not be lower than the calculation result.

[0025] Furthermore, the storage process of the checkpoints on the storage nodes utilizes the characteristics of the data parallel strategy, including the following steps:

[0026] Step 21, perform initialization. The storage node prepares the checkpoint data required for this training in advance, and according to the data parallelism strategy adopted, the storage node distributes the complete model to each computing node according to a predetermined rule;

[0027] Step 22, each model parallel group calculates the parameter gradients locally;

[0028] Step 23, perform global gradient aggregation and synchronize it to all nodes;

[0029] Step 24, the model update stage.

[0030] Furthermore, in Step 21, it also includes:

[0031] Step 211, before the training starts, the storage node prepares the checkpoint data, gradient synchronization data, and parameter update data required for this training in advance;

[0032] The checkpoint data is sourced from the archive of the previous training or configured manually by the user; the checkpoint usually includes: the initial values of the model parameters, the optimizer state, and the global state to be restored;

[0033] The checkpoint data, gradient synchronization data, and parameter update data are classified according to the transmission priority. Gradient synchronization and parameter update must be synchronized in a timely manner as high-priority tasks, while checkpoint data and backups are low-priority tasks;

[0034] The distributed training process requires checkpoint saving and parallel execution with the CPU. Asynchronous programming is used to implement non-blocking operations, and a central coordinator is used to uniformly manage the scattered monitoring data, prediction results, and scheduling policies, and feedback this information to the running distributed training process, so that low-priority tasks are not triggered during the communication-intensive phase, and checkpoint transmission is quickly started when the idle window is sufficient;

[0035] Step 212, according to the data parallelism strategy adopted, the storage node distributes the complete model to each computing node according to a predetermined rule; each computing node loads a model copy in the memory of the local computing device at startup, and the structure of each copy is exactly the same.

[0036] Furthermore, in Step 22, it also includes:

[0037] Step 221, each model parallel group performs forward propagation on its allocated data subset, that is, the input data is processed layer by layer through the model to calculate the output;

[0038] In the forward propagation calculation, each device uses the locally loaded parameter copy, and there is no communication overhead between model parallel groups;

[0039] Step 222: Compared with the true label, the model outputs the local loss value calculated by the loss function, and the loss calculation result provides a numerical basis for backpropagation;

[0040] Step 223: Each computing node calculates the gradient on the local data subset through the backpropagation algorithm based on the local loss value.

[0041] Furthermore, in Step 23, it also includes:

[0042] Step 231: Each computing node first obtains the local gradient, and then needs to integrate the gradients of each node to obtain the global gradient;

[0043] Step 232: The aggregated global gradient will be broadcast to all computing nodes to ensure that each computing node has consistent gradient information;

[0044] Step 233: Synchronize the gradient to the storage node for subsequent checkpoint update operations, so that the storage node can obtain the latest global state and provide a basis for recovery and debugging.

[0045] Furthermore, in Step 231, an all-reduce operation is adopted:

[0046] All hosts submit communication data to their respective connected switches. After receiving the data, the leaf switches will use the built-in engine to calculate and process the data, and then submit the result data to the spine switch. The spine switch also uses its own engine to aggregate the result data received from several switches and submit it to the root switch. The root switch performs the final calculation and returns the result to all host nodes;

[0047] All nodes exchange data with each other to jointly complete gradient averaging or accumulation; participate in part of the calculation through network switches or dedicated high-performance network communication libraries to further accelerate the gradient aggregation speed.

[0048] Furthermore, in Step 24, it also includes:

[0049] Step 241: After each computing node receives the global gradient, it performs parameter update calculation according to the preset optimizer;

[0050] During the update process, the computing node calculates and adjusts the parameter value according to the current parameters, the global gradient, and the internal state of the optimizer to keep the model state of each node consistent;

[0051] Step 242: After the storage node receives the global gradient synchronization, it needs to complete a similar update operation;

[0052] The entire parameter matrix or optimizer state is divided into multiple blocks of a certain size, and the state update of each block is completed on the computing device in sequence. After each block is updated, it is immediately written back to the memory. Space is reserved in the memory to store the updated blocks. The write-back operation from the memory to the non-volatile storage device is asynchronous to reduce the blocking impact on parameter updates. By not waiting for the data to be written to the storage node and directly proceeding with the state update of the next block, an independent background thread or process asynchronously writes the data in the memory to the storage device in real time using idle resources, thereby reducing the latency caused by IO operations;

[0053] Step 243, after the update is completed, the checkpoint data on the storage node contains the latest model parameters, optimizer state, and global gradient;

[0054] Based on the existing storage nodes, it is expanded to a multi-copy checkpoint, and the checkpoint data is stored on multiple physical or logical nodes to ensure that if a certain storage node fails to read the data, other replicas can still ensure the smooth progress of the recovery operation;

[0055] Give priority to ensuring the network bandwidth of high-priority tasks; low-priority tasks use preemptive transmission or background transmission methods and allow retransmission with interruption when necessary.

[0056] The beneficial effects achieved by the present invention are:

[0057] The large model disaster tolerance system under the distributed system proposed by the present invention can provide fast fault recovery for large model training, achieve high-frequency checkpoints, reduce the time overhead during training and fault recovery, and do not generate additional training throughput overhead. At the same time, this system is applicable to existing distributed parallel training methods and training frameworks.

[0058] The idle time period detection method proposed by the present invention can utilize the idle network bandwidth more flexibly and efficiently, achieving a better balance between checkpoint transmission and training calculation. In addition, this solution uses non-intrusive scheduling, minimizing the interference to the training process. At the same time, through fault tolerance guarantee and resource reservation capabilities, it ensures the improvement of the reliability of checkpoint transmission. In addition, this method supports various collective communication operations and network protocols, is applicable to multiple distributed training frameworks, and ensures scalability and compatibility.

[0059] The present invention proposes a data parallel strategy that fully utilizes network resources in gradient aggregation to accelerate the data exchange speed. The training process and the process of writing checkpoints to storage nodes are executed in parallel, so that under the condition of paying a small amount of computing cost, the checkpoints can be written to storage nodes without increasing the training time overhead. In addition, the storage node adopts a block processing and asynchronous write-back method when updating checkpoints, optimizing the system performance. The above steps enable the system to maintain high performance, high reliability, and low latency in a large-scale distributed training environment, meeting the requirements of training complex models and massive data.

[0060] During the distributed training process of the present invention, a central coordinator is used to uniformly manage scattered monitoring data, prediction results, and scheduling strategies, and feedback this information to the running distributed training process, so that low-priority tasks are not triggered during the communication-intensive phase, and checkpoint transmission is quickly started when the idle window is sufficient. In addition, by using a notification mechanism and reserved bandwidth, the distributed training system itself senses the idle window and collaboratively notifies the background transmission task to start or pause. Brief Description of the Drawings

[0061] Figure 1 . Checkpoint replica placement strategy;

[0062] Figure 2 . Asynchronous storage of local CPU memory checkpoints;

[0063] Figure 3 . Communication of redundant CPU memory checkpoints;

[0064] Figure 4 . Remote storage checkpoint architecture. Detailed Embodiment

[0065] The present invention will be further described below in conjunction with specific embodiments, and the advantages and features of the present invention will become clearer as the description progresses. However, these embodiments are exemplary only and do not constitute any limitation to the scope of the present invention. Those skilled in the art should understand that the details and forms of the technical solutions of the present invention can be modified or replaced without departing from the spirit and scope of the present invention, but such modifications and replacements all fall within the protection scope of the present invention.

[0066] The large model checkpoint disaster tolerance system based on network computing and asynchronous checkpoints provided by the present invention adopts a hierarchical storage design. A computing node contains multiple GPUs, and the checkpoints of the GPUs in the computing node are saved in the CPU memory of the computing node, the CPU memory of adjacent computing nodes in the same group, and the non-volatile storage device of the remote storage node.

[0067] Only the most recent checkpoint is saved in the local CPU memory and the adjacent CPU memory, which is transparent to the user. The saving and restoration of the checkpoint are automatically completed by the system. The remote storage saves the historical checkpoints visible to the user for fault recovery, model debugging, model deployment, etc.

[0068] The checkpoints in the local CPU memory and the adjacent CPU memory are different in format from those in the remote storage, but basically the same in content. The checkpoint in the memory can be regarded as a copy of the GPU video memory, while the checkpoint in the remote storage is saved in a serialized format for convenient storage and compression. During the general training process, the GPU video memory stores (part of) the model parameters and the optimizer state, and the CPU memory stores some parameter settings of the current training. Therefore, after the CPU memory obtains a copy of the GPU video memory, the CPU obtains all the states that need to be saved for the current training, and the checkpoint in the remote storage is the serialized content of the corresponding data.

[0069] The writing of the checkpoint in the local CPU memory and the adjacent CPU memory is completed by the working agent of the local machine. The placement strategy is determined according to the computing node topology structure during the training initialization. During the runtime, the working agent transfers the checkpoint from the GPU video memory to the CPU memory according to the placement strategy.

[0070] The purpose of the checkpoint in the adjacent CPU memory is that when a failure occurs, the checkpoint in the local CPU memory may be unavailable. At this time, the system has to rely on the checkpoint in the remote storage, resulting in a significant cost for fault recovery. To increase the probability of successfully recovering from a failure in the CPU memory, this system adopts a redundant checkpoint mechanism to save a copy of the local CPU checkpoint in the adjacent CPU.

[0071] Saving a copy incurs significant overheads, including CPU memory occupancy and the overhead of transmitting the copy. Since failures occur frequently but not concentratedly during the training process, this system selects a relatively simple copy measure, that is, the grouping measure of computing nodes (herein referred to as combination to denote the computing node allocation method, which is different from the group in the following text). A computing node usually contains 8 GPUs and undertakes the main computing tasks in model training. The computing node also contains devices such as a CPU, RAM, and NIC. The computing nodes within a combination save each other's checkpoints. The minimum unit of the combination is 2. When there is an odd number of computing nodes, the last combination has three nodes. As Figure 1 shown, the CPU memory of computing node 01 saves the checkpoints of the GPUs of computing nodes 01 and 02, and the CPU memory of computing node 02 saves the checkpoints of the GPUs of computing nodes 01 and 02. The same applies to the three-node combination.

[0072] When writing to local CPU memory and adjacent CPU memory, to further improve the writing speed, the system adopts a fine-grained pipelining approach. Checkpoint writing refers to copying model parameters from GPU memory to CPU memory. When writing to local CPU memory, if the write operation is executed synchronously, it will significantly block the training process. During the training of large models, it is obvious that model parameters only change when the parameters are updated and do not change during forward and backward propagation. Therefore, the copying process can be processed asynchronously with the computing process.

[0073] As Figure 2 shown, after each iteration's weight update step, the write operation is called through asynchronous programming methods such as callback functions, Promises, and async / await, without blocking the training process.

[0074] When writing to adjacent CPU memory, the system needs to regulate the checkpoint traffic. The reason is that during the training of large models, due to the large number of model parameters, certain distributed training strategies are adopted, such as data parallelism, model parallelism, pipeline parallelism, and hybrid strategies, etc. These parallel strategies rely on collective communication operations for synchronization and have obvious requirements for network bandwidth. When transmitting checkpoints from the local CPU to adjacent CPUs, it will occupy network bandwidth, causing potential network resource competition, which may delay communication operations during training and hinder computing. To alleviate this problem, this system chooses to regulate the checkpoint traffic to minimize its impact on training.

[0075] By utilizing the idle time in the training process for checkpoint communication, the impact of checkpoint communication on training communication can be reduced. During the forward and backward propagation of the model, there are a large number of training communication operations, including gradient synchronization and parameter acquisition. There are idle time points in these operations, and there is no communication during parameter update. Therefore, checkpoint communication operations can be carried out at these idle time points. When the idle time is insufficient, the training process will be blocked. As Figure 3 shown.

[0076] The confirmation of the idle time period uses an adaptive method for detecting the idle time period, which mainly includes three steps: data collection and time series modeling, idle time period identification and adaptive threshold adjustment, and real-time scheduling feedback.

[0077] Step 1, Network Event Log Collection: In each iteration, set hooks at the entry and exit of all collective communication operations (such as gradient synchronization, parameter acquisition, RDMA calls) for timestamp recording. Record the start time t_start and end time t_end for each communication operation and save them to the log system. Additionally, collect auxiliary information such as the communication packet size, latency, and current network bandwidth utilization.

[0078] Step 2, Timing Data Construction: Based on all recorded communication intervals, construct a continuous timeline and mark the communication process as occupied "busy intervals". Discretize the entire timeline, divide time into several tiny time slices (e.g., a 1-millisecond time window), and mark whether there is a communication operation within each time slice. Statistically analyze the distribution and cumulative duration of "busy intervals" in each iteration, and calculate the candidate idle time windows.

[0079] Step 3, Initial Idle Interval Detection: Traverse the discretized timeline, merge adjacent time windows according to the "busy" or "idle" status to generate a list of continuous idle time periods. Set an initial minimum idle threshold (e.g., 50 ms), and only idle segments longer than this duration are considered as truly available idle time periods for checkpoint transfer.

[0080] Step 4, Time Prediction and Model Update: Use a sliding window to statistically analyze the duration distribution of the idle time periods detected in the most recent N (e.g., 10 or 20) iterations, calculate their mean value, and use a simple exponential smoothing algorithm to predict the possible length of the idle time period in the next iteration.

[0081] S t = αX t +(1 - α)S t-1

[0082] where X t is the length of the idle time period detected in the current iteration, S t is the smoothed value at the current moment, S t-1 is the previous smoothed value, and α is the smoothing coefficient, which determines the weight distribution between the new observation value and the historical smoothed sequence. There is no historical data in the first round, and S0 = X0 can be set, i.e., initialize with the first observation value.

[0083] Step 5, Idle Time Confirmation and Checkpoint Scheduling: When it is detected that the idle interval exceeds the current dynamic threshold, the checkpoint transfer operation is triggered. If the length of the idle interval is not sufficient to complete the entire transfer task, the data transfer is split into multiple times, and the "partial transfer - feedback - judge idle continuation" method is used for dynamic retransmission. At the same time, the actual transmission delay and the network resources already used are monitored online. If it is found that the network utilization rate surges during the transmission process, the transmission rate is dynamically reduced, or the transmission is suspended temporarily and waits for the next idle window.

[0084] Step 6, Dynamically adjust the minimum idle threshold according to the distribution. When the idle time is generally short in the past several iterations, the threshold can be reduced (but it is necessary to ensure that the transmission will not be severely interrupted), for example, adjusted to 0.8 times the median. When the idle time is long, the threshold is increased to reserve sufficient time for the transmission of larger data blocks. When it is detected that the current network utilization rate drops, the idle judgment threshold can be appropriately reduced to make the best use of the short idle time as much as possible.

[0085] Through the above solutions, the idle network bandwidth can be utilized more flexibly and efficiently, achieving a better balance between checkpoint transfer and training calculation. In addition, this solution uses non-intrusive scheduling, minimizing the interference to the training process. At the same time, through the fault tolerance guarantee and resource reservation capabilities, the reliability of checkpoint transfer is improved. In addition, this method supports various collective communication operations and network protocols, is applicable to a variety of distributed training frameworks, and ensures scalability and compatibility.

[0086] During the training process of large models, data parallelism is the most commonly used parallel training scheme. Because data parallelism can improve the model training speed while having a relatively small network overhead. Compared with other parallel strategies that are adopted passively due to large model parameters, data parallelism is more of an active choice. Common data parallel strategies include DP, DDP, FSDP, ZERO, etc. This system is designed based on the general process of data parallelism, that is, data segmentation, model replication, parallel computing, and gradient aggregation. Among them, gradient aggregation will aggregate the calculated gradients (all-reduce operation). Since the model is replicated on different model parallel groups (the model parallel strategy is implemented among model parallel groups), forward and backward calculations are performed based on the same model parameters and different data among model parallel groups to obtain gradients for aggregation among model parallel groups. After obtaining the aggregated gradients, local parameter updates are performed by the local GPU.

[0087] Such as Figure 4As shown in the figure, a group contains multiple computing nodes and leaf switches. The leaf switches are fully connected to the computing nodes and fully connected to the spine switches. The storage nodes also have corresponding leaf switches to connect to the spine switches. Taking a two-layer fat tree network topology as an example, the computing nodes usually contain 8 GPUs, which undertake the main computing tasks in model training. The computing nodes also contain devices such as CPUs, RAMs, and NICs. The number of leaf switches and spine switches is determined by the number of switch ports and the network topology.

[0088] The storage nodes contain multiple non-volatile storage devices and less computing power, usually several GPUs. They also contain CPUs, RAMs, NICs, etc. The number of storage nodes can be more than one. The number of storage nodes does not affect the algorithm process. Multiple storage nodes store the complete checkpoint information respectively and are redundant with each other. The computing power of the storage nodes does not participate in model training. Therefore, devices with lower computing power than other computing nodes are selected. The specific selection can be based on the formula:

[0089]

[0090] F: OPs 前向传播+反向传播 represents the floating-point computation volume of forward propagation and backward propagation, T 前向传播+反向传播 represents the time required for forward propagation and backward propagation, T 参数更新 represents the time required for parameter update, and calculating gives P 存储 represents the computing performance required for the group where the storage node is located. The total computing power within the group must not be lower than the calculation result. If the total computing power is lower than the requirement, it will not be able to meet the need to save checkpoints remotely during each iteration. At this time, the checkpoint frequency can be adjusted. For example, save checkpoints every other parameter update, which can reduce the computing power requirement by half. Considering that the computing power of the group where the storage node is located can be ignored compared with the overall training computing power, it is not recommended to reduce the computing power of the group where the storage node is located.

[0091] The storage process of the checkpoints on the storage nodes utilizes the characteristics of the data parallel strategy. The specific process is as follows:

[0092] Step 21, initialization phase

[0093] Step 211, checkpoint preparation. Before the training starts, the storage nodes prepare in advance the checkpoint data, gradient synchronization data, and parameter update data required for this training. The checkpoints can be sourced from the archives of the previous training or configured manually by the user. The checkpoints usually contain: the initial values of the model parameters, the optimizer state (such as momentum, learning rate decay parameters, etc.), and other global states that may need to be restored, such as the global training step, random seed, etc.

[0094] Checkpoint data, gradient synchronization data, and parameter update data are classified according to transmission priorities. Gradient synchronization and parameter updates must be synchronized in a timely manner and are regarded as high-priority tasks, while checkpoint data and backups are low-priority tasks.

[0095] The distributed training process requires checkpoint saving with the CPU and is executed in parallel. Non-blocking operations are achieved using asynchronous programming (such as async / await, Promise, callback functions, etc.). At this time, an integrator, namely the central coordinator, is used to uniformly manage the scattered monitoring data, prediction results, and scheduling strategies, and feedback this information to the running distributed training process, so that low-priority tasks are not triggered during the communication-intensive phase, and checkpoint transmission is quickly started when the idle window is sufficient.

[0096] Step 212, model distribution. According to the adopted data parallel strategy, the storage node distributes the complete model to each computing node according to a predetermined rule. Each node loads a copy of the model in the memory of the local computing device (such as GPU or TPU) when starting. Each copy has exactly the same structure but will process different data subsets next.

[0097] Step 22, parallel computing phase

[0098] Step 221, forward propagation calculation. Each parallel group of models performs forward propagation on its allocated data subset, that is, the input data is processed layer by layer through the model to calculate the output. In this process, each device uses the locally loaded parameter copy, and there is no communication overhead between parallel groups of models. There will be corresponding overhead introduced by other parallel strategies when using them.

[0099] Step 222, loss calculation. The model output is compared with the true label, and the local loss value is calculated through a loss function (such as cross-entropy, mean squared error, etc.). The loss calculation result provides a numerical basis for backpropagation.

[0100] Step 223, backpropagation to calculate gradients. Each computing node calculates the gradients on the local data subset through the backpropagation algorithm based on the local loss value. In this stage, only local parameter gradients are calculated and cross-node synchronization is not involved. Therefore, no additional network overhead needs to be introduced under the data parallel strategy.

[0101] Step 23, parameter synchronization phase

[0102] Step 231, Local-to-Global Gradient Aggregation. Each computing node first obtains the local gradient, and then the gradients of each node need to be integrated to obtain the global gradient. This process usually adopts the all-reduce operation: all hosts submit the communication data to their respective connected switches. After receiving the data, the leaf switches will use the built-in engine to calculate and process the data, and then submit the result data to the spine switch. The spine switch also uses its own engine to aggregate the result data received from several switches and submit it to the root switch. The root switch performs the final calculation and returns the result to all host nodes. All nodes exchange data with each other to jointly complete gradient averaging or accumulation. Network switches or dedicated high-performance network communication libraries (such as those based on RDMA) can participate in part of the calculation to further accelerate the gradient aggregation speed.

[0103] Step 232, Synchronize the Gradient to All Nodes. The aggregated global gradient will be broadcast to all computing nodes to ensure that each node has consistent gradient information, which is convenient for subsequent model parameter updates.

[0104] Step 233, Synchronize the Gradient to the Storage Node. In addition to being used for local updates, the global gradient also needs to be synchronized and stored in the storage node. This gradient data is used for subsequent checkpoint update operations, enabling the storage node to obtain the latest global state and providing a basis for recovery and debugging.

[0105] Step 24, Model Update Phase

[0106] Step 241, Parameter Update (on the Computing Node Side). After each computing node receives the global gradient, it performs parameter update calculations according to a preset optimizer (such as SGD, Adam, etc.). During the update process, the computing node calculates and adjusts the parameter values based on the current parameters, the global gradient, and the internal state of the optimizer (such as momentum, etc.) to keep the model state of each node consistent.

[0107] Step 242, Update Checkpoint (Storage Node Side) After receiving the global gradient synchronization, the storage node needs to complete a similar update operation. First, it loads the checkpoint data (including model parameters and optimizer state) from the persistent storage into the memory. Considering the large amount of model parameter data, during the update, it will be processed in chunks according to the device video memory capacity or system memory. The entire parameter matrix or optimizer state will be divided into multiple chunks of a certain size, and the state update of each chunk will be completed on the computing device (which may use a GPU or a dedicated accelerator) in turn. After each chunk is updated, it will be immediately written back to the memory. Space is reserved in the memory to save the updated chunks. The write-back operation from the memory to the non-volatile storage device is asynchronous to reduce the blocking effect on parameter updates, that is, without waiting for the data to be written to the storage node, directly proceed with the state update of the next chunk. An independent background thread or process uses idle resources to asynchronously write the memory data to the storage device in real time, thereby reducing the latency caused by IO operations.

[0108] Step 243, Overall Consistency and Storage Write-back After the update is completed, the checkpoint data on the storage node contains the latest model parameters, optimizer state, and global gradient. This ensures that when the system needs to be restored, it can be started from this checkpoint data and continue training without losing progress. Data consistency verification will also be performed at this stage to ensure that no errors occur during data storage and reading. Further, to ensure data reliability, based on the existing storage node, it can be extended to a multi-copy checkpoint, that is, the checkpoint data (including model parameters, optimizer state, global gradient, etc.) is stored on multiple physical or logical nodes. At this time, if a certain storage node fails to read the data due to a fault, other copies can still ensure the smooth progress of the recovery operation. At the same time, the network bandwidth for high-priority tasks is preferentially guaranteed; low-priority tasks use preemptive transmission or background transmission methods and are allowed to interrupt and retransmit when necessary.

[0109] The above steps make full use of network resources in gradient aggregation to accelerate the data exchange speed. The training process and the process of writing the checkpoint to the storage node are executed in parallel, so that the checkpoint can be written to the storage node without increasing the training time overhead under the condition of paying a small amount of computational cost. In addition, the storage node adopts chunk processing and asynchronous write-back methods when updating the checkpoint, which optimizes the system performance. The above steps enable the system to maintain high performance, high reliability, and low latency in a large-scale distributed training environment, meeting the needs of training complex models and massive data.

[0110] The fault recovery function of this system involves four components, a group of worker agents, a root agent, a distributed key-value store, and an operator.

[0111] The working agent is responsible for monitoring the health status of the machines it is in charge of, updating it to the distributed key-value storage system, and completing the writing of the local CPU memory and adjacent CPU memory of the checkpoint.

[0112] The root agent runs on a regular training machine or a dedicated machine. The root agent periodically checks the health status of each training machine from the distributed key-value storage. If the root agent detects a training machine failure, the root agent will take corresponding actions according to the type of failure. When the root machine fails, the root machine is selected through the leader election method in the distributed key-value storage.

[0113] The cloud operator manages the computing resources and replaces the failed machines with healthy machines when needed.

[0114] This system is designed based on hierarchical storage and takes different measures for different types of failures.

[0115] When a software failure occurs, the training process will be interrupted, but the hardware is still healthy, and the checkpoint stored in the local CPU memory can still be accessed, so the training can be directly resumed from the local checkpoint.

[0116] When a hardware failure occurs, it is necessary to locate the failed machine and replace it. The location of the failed machine depends on the hardware monitoring tool and the distributed key-value storage system. The former is provided by the hardware manufacturer, and the latter is a function of this system. Specifically, the working agent and the root agent will periodically send heartbeat signals to the distributed key-value storage. When the distributed key-value system cannot hear the heartbeat signal, the corresponding machine or network is regarded as having a failure. At this time, recovery needs to be carried out according to the grouping situation. When the checkpoint replicas within the group are available, recovery can be directly performed from within the group. When the checkpoint replicas within the group are not available, it is necessary to recover from the storage node.

[0117] When a hardware failure occurs, the cloud operator should immediately provide a healthy machine to replace the failed machine. However, this replacement operation causes the training to wait, so the training cluster can reserve some spare machines. When a certain machine has a hardware failure, the spare machine can be immediately activated, replace the failed machine for failure recovery, and allocate the tasks on the failed machine to the spare machine to ensure that the overall training is not interrupted. At the storage node level, when a storage node fails, the distributed file system within the cluster is used to ensure the high availability and fast recovery of data.

[0118] After that, the root agent returns the failed machine and requests another spare machine.

[0119] The above are only the specific steps of the present invention and do not constitute any limitation to the protection scope of the present invention; all technical solutions formed by equivalent transformation or equivalent substitution fall within the scope of the protection of the present invention; the parts not elaborated in detail in the present invention belong to the well-known technologies in the art.

Claims

1. A large model checkpoint disaster tolerance system based on network computing and asynchronous checkpoints, characterized in that, The disaster recovery system adopts a hierarchical storage design. The disaster recovery system includes multiple groups, and each group includes at least two computing nodes. The computing nodes within the group save the checkpoints of each other. The computing node includes at least a CPU, a memory, a network transmission module, and multiple GPUs. The most recent checkpoint of each GPU is saved in the CPU memory of the current computing node, and the most recent checkpoint is also saved in the CPU memory of the adjacent computing node within the same group through the network transmission module.

2. The large model checkpoint disaster tolerance system based on network computing and asynchronous checkpoints according to claim 1, characterized in that, When writing the checkpoint to the local CPU memory and the adjacent CPU memory, the write operation is called through an asynchronous programming method after the weight update step of each iteration, without blocking the training process.

3. The large model checkpoint disaster tolerance system based on network computing and asynchronous checkpoints according to claim 1, characterized in that, When writing the checkpoint to the adjacent CPU memory, by utilizing the idle time in the training process for checkpoint communication, the impact of checkpoint communication on training communication can be reduced. The method for confirming the idle time includes the following steps: Step 1, during each iteration, set hooks at the entrance and exit of all collective communication operations to record timestamps; record the start time and end time for each communication operation, and additionally collect the communication packet size, latency, and the current network bandwidth utilization rate. The communication operations include gradient synchronization, parameter acquisition, and RDMA calls. Step 2, for all the recorded communication intervals, construct a continuous timeline and mark the communication process as an occupied "busy interval". Discretize the entire timeline, divide the time into several tiny time slices, and mark whether there is a communication operation within each time slice; count the distribution and cumulative duration of the "busy intervals" in each iteration, and calculate the candidate window for idle time. Step 3, traverse the discretized timeline, merge adjacent time windows according to the "busy interval" or "idle" state to generate a list of continuous idle time periods. And set an initial minimum idle threshold, and only the idle segments longer than this duration are considered as the truly available idle time periods for checkpoint transmission. Step 4, use a sliding window to statistically analyze the duration distribution of the idle time periods detected in the most recent N iterations, calculate their mean value, and use a simple exponential smoothing algorithm to predict the possible length of the idle time period in the next iteration. S t = αX t + (1 - α)S t-1 Among them, X t is the length of the idle time period detected within the current iteration, S t is the smoothed value at the current moment, S t-1 is the previous smoothed value, and α is the smoothing coefficient, which determines the weight allocation between the new observation value and the historical smoothed sequence; there is no historical data in the first round, and the first observation value is used for initialization, setting S0 = X0; Step 5, when it is detected that the idle interval exceeds the current dynamic threshold, trigger the checkpoint transmission operation; if the length of the idle interval is not sufficient to complete the entire transmission task, split the data transmission into multiple times and use the method of "partial transmission - feedback - judging idle continuation" for dynamic supplementary transmission; at the same time, online monitor the actual transmission latency and the network resources already used. If it is found that the network utilization rate suddenly increases during the transmission process, dynamically reduce the transmission rate, or temporarily suspend the transmission and wait for the next idle window. Step 6, dynamically adjust the minimum idle threshold according to the distribution. When the idle time is generally short in the past several iterations, the threshold can be reduced; when the idle time is long, the threshold is increased to reserve enough time for transmitting larger data blocks; when it is detected that the current network utilization rate decreases, the idle judgment threshold can be appropriately reduced to make the best use of the short idle time as much as possible.

4. The large model checkpoint disaster tolerance system based on network computing and asynchronous checkpoints according to claim 1, wherein The large model checkpoint disaster tolerance system based on network computing and asynchronous checkpoints further includes storage nodes, which contain multiple non-volatile storage devices and fewer computing power devices; FLOPs 前向传播+反向传播 Represents the floating-point computation volume for forward and backward propagation, T 前向传播+反向传播 Represents the time required for forward and backward propagation, T 参数更新 Represents the time required for parameter update, calculated as P 存储 Represents the computing performance required for the group where the storage node is located. The total computing power within the group must not be lower than the calculation result.

5. The large model checkpoint disaster tolerance system based on network computing and asynchronous checkpointing according to claim 4, wherein The storage process of the checkpoints on the storage nodes makes use of the characteristics of the data parallel strategy, including the following steps: Step 21, perform initialization. The storage nodes prepare in advance the checkpoint data to be used in this training, and according to the adopted data parallel strategy, the storage nodes distribute the complete model to each computing node according to a predetermined rule; Step 22, each model parallel group calculates the parameter gradients locally; Step 23, perform global gradient aggregation and synchronize it to all nodes; Step 24, the model update phase.

6. The large model checkpoint disaster tolerance system based on network computing and asynchronous checkpoints according to claim 5, wherein In step 21, it also includes: Step 211, before the training starts, the storage nodes prepare in advance the checkpoint data, gradient synchronization data, and parameter update data to be used in this training; The checkpoint data is sourced from the archive of the previous training or configured manually by the user; the checkpoint usually contains: the initial values of the model parameters, the optimizer state, and the global state to be restored; The checkpoint data, gradient synchronization data, and parameter update data are classified according to the transmission priority. Gradient synchronization and parameter update must be synchronized in a timely manner and are regarded as high-priority tasks, while checkpoint data and backups are low-priority tasks; The distributed training process needs to save checkpoints with the CPU and execute in parallel. Use asynchronous programming to achieve non-blocking operations, and use a central coordinator to uniformly manage the scattered monitoring data, prediction results, and scheduling strategies, and feedback this information to the running distributed training process, so that low-priority tasks are not triggered during the communication-intensive phase, and checkpoint transmission is quickly started when the idle window is sufficient; Step 212, according to the adopted data parallel strategy, the storage nodes distribute the complete model to each computing node according to a predetermined rule; each computing node loads a model copy in the memory of the local computing device at startup, and each copy has exactly the same structure.

7. The large model checkpoint disaster tolerance system based on network computing and asynchronous checkpoints according to claim 5, characterized in that, In step 22, it also includes: Step 221, each model parallel group performs forward propagation on its allocated data subset, that is, the input data is processed layer by layer through the model to calculate the output; In the forward propagation calculation, each device uses the locally loaded parameter copy, and there is no communication overhead between model parallel groups; Step 222, the model output is compared with the true label, and the local loss value is calculated through the loss function. The loss calculation result provides a numerical basis for backpropagation; Step 223, each computing node calculates the gradient on the local data subset through the backpropagation algorithm based on the local loss value.

8. The large model checkpoint disaster tolerance system based on network computing and asynchronous checkpoints according to claim 5, characterized in that, In step 23, it also includes: Step 231, each computing node first obtains the local gradient, and then needs to integrate the gradients of each node to obtain the global gradient; Step 232, the aggregated global gradient will be broadcast to all computing nodes to ensure that each computing node has consistent gradient information; Step 233, synchronize the gradient to the storage node for subsequent checkpoint update operations, so that the storage node can obtain the latest global state and provide a basis for recovery and debugging.

9. The large model checkpoint disaster tolerance system based on network computing and asynchronous checkpoints according to claim 8, characterized in that, In step 231, a full reduction operation is adopted: All hosts submit communication data to their respective connected switches. After receiving the data, the leaf switches use the built-in engines to calculate and process the data, and then submit the result data to the spine switches. The spine switches also use their own engines to aggregate the result data received from several switches and submit it to the root switch. The root switch performs the final calculation and returns the result to all host nodes; All nodes exchange data with each other to jointly complete gradient averaging or accumulation; participate in part of the calculation through network switches or dedicated high-performance network communication libraries to further accelerate the gradient aggregation speed.

10. The large model checkpoint disaster tolerance system based on network computing and asynchronous checkpoint according to claim 5, wherein In step 24, it further includes: Step 241, after each computing node receives the global gradient, it performs parameter update calculation according to a preset optimizer; During the update process, the computing node calculates and adjusts the parameter values according to the current parameters, global gradient, and internal state of the optimizer to keep the model state of each node consistent; Step 242, after the storage node receives the global gradient synchronization, it needs to complete a similar update operation; The entire parameter matrix or optimizer state is divided into multiple blocks of a certain size, and the state update of each block is completed on the computing device in turn. After each block is updated, it is immediately written back to the memory. Space is reserved in the memory to save the updated blocks. The write-back operation from the memory to the non-volatile storage device is asynchronous to reduce the blocking impact on parameter updates. By not waiting for the data to be written to the storage node and directly performing the state update of the next block, a background independent thread or process uses idle resources to asynchronously write the memory data to the storage device in real time, thereby reducing the delay caused by IO operations; Step 243, after the update is completed, the checkpoint data on the storage node contains the latest model parameters, optimizer state, and global gradient; Based on the existing storage nodes, it is expanded to a multi-copy checkpoint, and the checkpoint data is stored on multiple physical or logical nodes to ensure that if a certain storage node fails to read the data, other copies can still ensure the smooth progress of the recovery operation; Give priority to ensuring the network bandwidth for high-priority tasks; low-priority tasks use preemptive transmission or background transmission methods and allow retransmission with interruption when necessary.

Citation Information

Patent Citations

  • Network load dynamic adaptive parameter adjusting method based on priorities

    CN104185298A

  • System and method for supporting efficient load-balancing in a high performance computing (hpc) environment

    CN106489255A

  • Parallel processing of reduction and broadcast operations on large datasets of non-scalar data

    CN108537341A

  • Large language model federation pre-training method based on trusted execution environment

    CN117648998A

  • Large model training fault recovery method and device based on distributed memory management

    CN119473732A

Cited By

  • Small sample cross-domain equipment fault diagnosis method based on mixed attention and meta learning

    CN120876451A

  • Data processing method and device applied to distributed training system and chip product

    CN121279482A