Large model distributed training fault processing method based on dynamic check point strategy
The dynamic checkpoint strategy enhances LLM training efficiency and reliability by using a four-layer collaborative architecture and predictive models to optimize checkpoint placement and frequency, addressing the inefficiencies in traditional methods.
Patent Information
- Application Number
- CN202510820987.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-19
- Publication Date
- 2025-07-15
- Estimated Expiration
- 2045-06-19
AI Technical Summary
In the existing large-scale distributed training technology, checkpoint recovery speed is slow, failure loss is large, and existing methods cannot effectively cope with dynamic changes in large-scale data and long-term iteration, resulting in insufficiency in training.
A four-layer collaborative checkpoint placement architecture driven by heterogeneous priority is adopted, combined with deep reinforcement learning model agents, dynamically decide the checkpoint storage location, and predict the training time change trend through graph neural networks and long-term memory networks to achieve adaptive optimization of checkpoint update frequency.
It improves the efficiency and reliability of checkpoint transmission, reduces training time loss, optimizes resource utilization, and improves the reliability and efficiency of large model training.
Smart Images

Figure CN120317318A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of large model distributed training, and particularly relates to a method for handling faults in large model distributed training based on a dynamic checkpoint strategy. Background Art
[0002] Large language models (LLMs) have currently attracted extensive attention in both the academic and industrial fields. However, the number of parameters of LLMs is continuously increasing. In 2018, GPT-1 had only 117 million parameters, while the number of parameters of GPT-3.5 released in 2023 has reached 175 billion. The increase in model scale has greatly increased the cost of training a model. Due to the large number of parameters of LLMs and the need to use massive amounts of data for training, the training process must be completed in a large-scale GPU cluster, and the entire process will take thousands of GPUs and training time measured in months. During this process, faults will inevitably occur frequently. For example, during the training of the OPT model, faults occur twice a day on average; Llama 3.1 had 466 faults during a 54-day training process. Training faults have become a bottleneck affecting training efficiency.
[0003] In the context of existing large model distributed training technologies, in a large-scale GPU cluster, existing model training frameworks such as Pytorch and DeepSpeed usually resume the training progress from model checkpoints to handle faults. Specifically, checkpoint information such as parameters and optimizers during the training process of deep learning models is saved at a location accessible to a certain training node according to certain rules, and the training process is restarted and the checkpoint is loaded after a fault occurs. However, for LLM training faults with large amounts of data and long single-iteration times, traditional checkpoint schemes have problems such as slow recovery speed and large fault losses.
[0004] In this context, the current state-of-the-art prior art solutions closest to the present invention are mainly reflected in the following aspects: 1. Placement location of large model checkpoints Most existing methods store the training state externally. Check-N-Run uses model compression technology to reduce the time and storage overhead for saving checkpoint of TB-level size to external storage. DeepFreeze studies the implementation of asynchronous checkpoint technology for large models in the scenario of external storage of checkpoints. Chronicles is a 176B multilingual LLM that uses external storage to save checkpoints during training. However, the large size of LLM checkpoint results in a very high communication latency between the external storage and the GPU training cluster, leading to a large time overhead in the retraining process when training failures occur. A small number of methods use the CPU memory of training nodes to store checkpoints. Unicorn stores a copy of the checkpoint in both the local memory and external storage of large model training nodes, but the local memory is highly unreliable and prone to losing copies, resulting in the need to still recover the training state from external storage. GEMINI stores checkpoints in distributed CPU memory. By grouping the CPUs of large model training nodes and storing them in a circular strategy within the same group, it achieves multi-copy fault tolerance under distributed storage. However, the lack of analysis of the physical topology relationship between training nodes leads to redundancy in copies and a large resource overhead.
[0005] 2. Method for analyzing large model training time: The effectiveness of the checkpoint acquisition frequency decision method depends on the time of iterative update of model parameters. Currently, Zeng et al. model the iterative time by analyzing the calculation and communication steps in training iterations and summarizing the overlap and interference among multiple worker nodes. Gao et al. build a graph neural network to predict the execution time and performance of deep learning models by modeling the model information, parameter transfer size, and node communication bandwidth information in the model topology. Zancato et al. approximately calculate the training loss and accuracy at any point during training by solving low-dimensional stochastic differential equations (SDEs) in the function space without any training. However, existing methods only consider the performance and time prediction for a single training batch at the model level and do not support the prediction of the model iteration time series and the analysis of change trends in the future for a period of time in the large model distributed training scenario, which results in poor timeliness of the checkpoint acquisition frequency decision method.
[0006] 3. Update frequency of large model checkpoints The update frequency determines how often checkpoint operations are performed. Most existing methods adopt a static update frequency. Methods such as ByteCheck, PCcheck, and DeepFreeze mainly focus on optimizing the efficiency of the checkpoint saving and recovery processes at the operating system layer and support statically specifying the checkpoint frequency during training. They usually adopt a relatively low frequency. Both GEMINI and Unicorn provide fault tolerance by performing a checkpoint operation after each iteration. However, a low frequency results in a large amount of training state loss in case of a failure, while a high frequency leads to redundancy in a large number of training state saving operations. Very few methods consider a dynamic checkpoint frequency. CheckFreq takes into account the huge time overhead of storing a checkpoint to external storage, which spans multiple iterations, so it dynamically adjusts the checkpoint frequency to approximately perform the next checkpoint after each storage. However, it is not applicable to efficient local memory distributed scenarios and still does not solve the performance problems at high frequencies.
[0007] During the distributed training process of existing large models, reasonable checkpoints are used to prevent operation optimization and fault recovery processes. However, random placement strategies lead to unreliable or high-latency checkpoint placement positions. Most existing methods store checkpoints in external storage, which results in a large delay in the checkpoint acquisition process and the fault recovery process. A few methods randomly select multiple nodes locally and store a copy of the checkpoint in the memory of each of these nodes, reducing the delay of fault recovery. However, they treat different storage nodes equally and ignore the differences in reliability and available resource amounts between local distributed storage locations, resulting in unreliable placement schemes or high latency in the checkpoint transmission process.
[0008] In addition to optimizing the checkpoint storage process, the update frequency of checkpoints is also an important issue. Checkpoint update frequency decisions only consider average training performance and are difficult to adapt to dynamic environmental changes, resulting in large redundancy or fault losses. Most existing methods adopt a static checkpoint update frequency, but a low frequency causes a large amount of training state loss due to failures, while a large number of redundant operations at high frequencies lead to serious resource waste, both of which affect the execution efficiency of the training process. A few methods dynamically adjust the update frequency during training. They obtain the single-iteration time of the model through calculation and prediction methods, combine the life cycle time overhead of a single checkpoint update operation, and then select the number of iterations acceptable between two checkpoint updates based on this time. However, in fact, due to the dynamic changes in computing, storage, and communication resources, the single-word iteration time of DL training tasks also changes dynamically as the training process progresses. This change makes the analysis strategy based only on a single average value may be difficult to match the updated iteration situation. For example, when the network environment deteriorates and the training slows down, larger fault losses may occur under the same update interval for the same number of iterations.
[0009] For the evaluation method of training time, existing methods only support analyzing the performance of a single iteration based on the mean of historical iteration times or on the calculation of basic information such as the model network architecture, distributed topology, and hyperparameters, and are unable to accurately capture the law of iteration time changes in a dynamically changing environment. The error in performance estimation exacerbates the unreasonable setting of checkpoint frequencies, making it difficult to match future situations, resulting in operational redundancy or large fault losses. Summary of the Invention
[0010] In view of the problems existing in the prior art, the present invention proposes a method for handling faults in large model distributed training based on a dynamic checkpoint strategy. The present invention designs a four-layer collaborative checkpoint placement architecture driven by heterogeneous priorities, constructs a proxy for the checkpoint location decision model at the cluster layer based on deep reinforcement learning, and through a multi-dimensional reliability perception mechanism, integrates real-time resource load dynamics, fault mode characteristics, and distributed storage reliability information to ensure the efficient and reliable transmission of checkpoints during fault recovery. A multi-step deep learning training time prediction method based on spatio-temporal joint modeling is proposed. The spatio-topological features and temporal dynamics of distributed training are captured by a graph neural network and a long short-term memory network respectively to predict the training time for multiple future steps. Furthermore, the change trend and amplitude of future loads are extracted as key indicators to drive the adaptive optimization of checkpoint update frequencies. A checkpoint dual-frequency update method for a multi-layer placement architecture is proposed. By establishing a continuous characterization model for fault losses, the cost function of the method under different update frequencies is quantitatively characterized. This method combines the identification of the change trend of future training time and the quantification of the amplitude to achieve the online dynamic adjustment of the checkpoint update frequency, reducing the checkpoint operation overhead and improving the training efficiency while ensuring fault tolerance.
[0011] Technical Solution A method for handling faults in large model distributed training based on a dynamic checkpoint strategy, comprising: (1) A checkpoint distributed access strategy for cluster topology and environment dynamic perception: Design a four-layer access topology checkpoint distributed access method to dynamically perceive the topology structure and resource status of the GPU cluster, analyze the multi-copy access locations of checkpoints during training, and determine the optimal access strategy.
[0012] (2) A method for predicting the iteration time of a large model with change trend perception: Implement time series prediction of large model training iterations, and analyze the future iteration time trend through historical data. Combining change trend analysis, improve the timeliness of checkpoint update frequencies to cope with the dynamic changes in the training environment.
[0013] (3)Model training iteration time and trend-aware checkpoint frequency decision method: Based on loss analysis, dynamically adjust the checkpoint update frequency to ensure that under the premise of relatively small failure losses, the overhead of checkpoint operations is reduced. By calibrating the tolerable loss and adjusting the frequency, optimize the checkpoint update strategy to improve the efficiency and reliability of the training process.
[0014] Beneficial effects In the existing large model distributed training technology, the traditional checkpoint recovery scheme faces problems such as slow recovery speed and large failure losses. Especially when dealing with large-scale data and long-time iterations, these problems become more prominent. In response to this situation, the heterogeneous priority-driven four-layer collaborative checkpoint placement architecture proposed by the present invention can dynamically determine the storage location of checkpoints through a deep reinforcement learning model agent, fully considering multi-dimensional factors such as resource load, failure mode, and storage reliability. This method not only improves the efficiency and reliability of checkpoint transmission but also effectively reduces the training time loss caused by the recovery process.
[0015] The existing training time prediction methods mainly focus on the performance analysis of a single training batch and cannot adapt to the dynamic changes in the large model distributed training scenario. Therefore, the present invention proposes a multi-step deep learning training time prediction method based on spatio-temporal joint modeling, which uses graph neural networks and long short-term memory networks to capture the topological space features and temporal dynamics during the training process. This innovation enables the early identification of future load change trends, thereby driving the adaptive optimization of the checkpoint update frequency to ensure efficient training while reducing the additional overhead caused by frequent checkpoint operations.
[0016] Most of the existing methods adopt a static checkpoint update frequency, which may cause a large amount of training state loss in case of a failure, while too high a frequency will increase operation redundancy. The checkpoint dual-frequency update method of the present invention realizes the online dynamic adjustment of the checkpoint update frequency by establishing a continuity representation model, quantifying the cost function under different update frequencies, and combining the identification of future training time change trends. This method not only enhances the fault tolerance ability but also significantly improves the training efficiency, solving the problem of low efficiency caused by improper frequency selection in the prior art. Through these innovations, the present invention effectively improves the reliability and efficiency of large model training and has important practical application value.
[0017] In summary, the present invention provides a more efficient and reliable solution to the deficiencies in the distributed training of existing large models by introducing a heterogeneous priority-driven four-layer collaborative checkpoint placement architecture, a multi-step training time prediction method based on spatio-temporal joint modeling, and a checkpoint dual-frequency update strategy. These innovations not only address the speed and loss issues in the traditional checkpoint recovery process but also optimize resource utilization and reduce redundant overhead during training by dynamically adjusting the checkpoint update frequency. These improvements enable large-scale deep learning models to more flexibly and efficiently handle failures in the face of complex training environments, ensuring the continuity and stability of training, thus providing new ideas and directions for the development of large model training technologies. BRIEF DESCRIPTION OF THE DRAWINGS
[0018] Figure 1 is a schematic diagram of the algorithm flow of the method for handling faults in the distributed training of large models based on the dynamic checkpoint strategy of the present invention; Figure 2 is a schematic diagram of the hierarchical storage architecture proposed by the present invention; Figure 3 is the neural network architecture design of the iterative time series prediction method proposed by the present invention; Figure 4 is the algorithm flow chart of the iterative time variation trend analysis method proposed by the present invention; Figure 5 is a schematic diagram of the fault loss modeling method proposed by the present invention; Figure 6 is the algorithm flow chart of the dynamic checkpoint frequency decision method proposed by the present invention; Figure 7 is a comparison chart of the total model training time between the present invention and other existing methods; Figure 8 is a comparison chart of the average loss and average recovery time for a single fault between the present invention and other existing methods; Figure 9 is a comparison chart of the additional overhead of checkpoint operations between the present invention and other existing methods. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0019] The technical solutions provided by the present application will be further described below in conjunction with specific embodiments and their accompanying drawings. The advantages and features of the present application will become clearer in light of the following description.
[0020] A method for handling faults in the distributed training of large models based on the dynamic checkpoint strategy, as Figure 1 , includes: (1)Checkpoint Distributed Access Strategy with Cluster Topology and Environment Dynamic Awareness: Design a four-layer access topology checkpoint distributed access method to dynamically sense the topology structure and resource status of the GPU cluster, analyze the access locations of multiple copies of checkpoints during the training process, and determine the optimal access strategy.
[0021] (2)Large Model Iteration Time Prediction Method with Change Trend Awareness: Implement time series prediction of large model training iterations, and analyze the future iteration time trend through historical data. Combine change trend analysis to improve the timeliness of checkpoint update frequency to cope with the dynamic changes of the training environment.
[0022] (3)Checkpoint Frequency Decision Method with Model Training Iteration Time and Trend Awareness: Dynamically adjust the checkpoint update frequency based on loss analysis to ensure that the overhead of checkpoint operations is reduced on the premise of a small failure loss. Optimize the checkpoint update strategy by calibrating the tolerable loss and adjusting the frequency to improve the efficiency and reliability of the training process.
[0023] Furthermore, the checkpoint distributed access strategy with cluster topology and environment dynamic awareness is based on the four-layer access topology checkpoint distributed access method. Through the dynamic awareness of the cluster topology and resources, a reinforcement learning decision model based on the Deep Q-Network (DQN) is constructed to make decisions on the access locations of multiple copies. It includes: (1)Cluster Topology Information Collection: Collect the topology information of the four-layer storage media in the GPU cluster, including local memory, physical machine memory, memory of other machines in the cluster, and external storage. As Figure 2 is the schematic diagram of the hierarchical storage architecture of the present invention.
[0024] (2)Environment Dynamic Awareness: Analyze the access conditions in the current environment by dynamically sensing the status of cluster resources.
[0025] (3)Access Location Decision: Based on the collected topology information and resource status, decide the access locations of multiple copies to achieve fast fault recovery.
[0026] Furthermore, the large model iteration time prediction method with change trend awareness is as Figure 3 、 4 , and by predicting the future iteration time series and analyzing the change trend, the timeliness of the checkpoint frequency update method is improved, including: (1)Historical Data Analysis: Collect and analyze historical training iteration time series data to identify the change patterns of the time series.
[0027] (2)Trend Prediction: Use the GNN-LSTM network to predict the future iteration time and output the future change trend.
[0028] (3)Classification and Response: Classify the predicted change trends (such as fluctuations and regularities), promptly identify the instability of the training environment, and provide support for the checkpoint strategy.
[0029] Furthermore, the model training iteration time and the checkpoint frequency decision method for trend awareness are as Figure 5 , 6 , and the dynamic hierarchical decision method for checkpoint update frequency based on loss analysis reduces the overhead of checkpoint operations on the premise of less fault loss, including: (1)Loss Tolerance Setting: Determine the tolerable training loss L base , and combine it with the change coefficient to calculate the future loss tolerance L1.
[0030] (2)Frequency Adjustment Strategy: On the basis of ensuring that the loss does not exceed L1, gradually increase the checkpoint acquisition interval Δ to find the maximum interval that meets the conditions.
[0031] (3)Overhead Optimization: Combine the overhead of the access strategy and dynamically adjust the checkpoint update frequency to reduce unnecessary checkpoint operations and ensure the efficiency of the training process.
[0032] Embodiment 1 This embodiment provides a method for handling faults in distributed training of large models based on a dynamic checkpoint strategy, as Figure 1 shown, including: Based on the four-layer access topology checkpoint distributed access method, make decisions on the access locations of multiple replicas through dynamic awareness of the cluster topology and resources.
[0033] The method for predicting the time series of large model training iterations and analyzing the change trends improves the timeliness of the checkpoint frequency update method through the prediction of the future iteration time series and the analysis of the change trends.
[0034] The dynamic hierarchical decision method for checkpoint update frequency based on loss analysis reduces the overhead of checkpoint operations on the premise of less fault loss.
[0035] The following will describe in detail the fault handling method provided in this embodiment with reference to the diagrams. As Figure 1 is the algorithm flow diagram of the method for handling faults in distributed training of large models based on a dynamic checkpoint strategy of the present invention, including the following steps: S1, Based on the four-layer access topology checkpoint distributed access method, make decisions on the access locations of multiple replicas through dynamic awareness of the cluster topology and environmental resources.
[0036] In this embodiment, for the distributed access method, by perceiving the information of the four-level storage medium topology and the environmental resource information in the GPU cluster, the access locations of multiple copies of the checkpoint during the training process are analyzed, and the access locations of each copy are output.
[0037] S11 is used for the infrastructure architecture for storing the checkpoint as Figure 2 shown.
[0038] The four-level access topology of the checkpoint is as follows: The storage media are, in order of transmission speed with the training node, local memory, physical machine memory, memory of other machines in the cluster, and external storage at four levels.
[0039] The first level is local memory. For the training process, the data in the local CPU memory can be obtained most quickly. Therefore, storing 1 copy of the checkpoint in local memory can be used to directly load the checkpoint from local when the training process itself fails, achieving the fastest failure recovery.
[0040] When the training virtual machine fails, the checkpoint in local memory will be damaged. For this reason, the memory level of other virtual machines within the same physical machine is introduced to solve the virtual machine failure.
[0041] The second level is the memory of other virtual machines located in the same physical machine as the training virtual machine. The advantage of this level is that when a single virtual machine fails or the checkpoint in local memory is damaged, the checkpoint can be transferred to the restarted faulty virtual machine relatively quickly through methods such as shared memory and the internal network of the physical machine. In terms of the number of copies, if two virtual machines in the same physical machine crash simultaneously, the failure is likely caused by the failure of the entire physical machine. This makes it meaningless to store 2 or more copies at this level, and instead will cause additional resource overhead. Therefore, it is only necessary to store one copy of the checkpoint on the virtual machine in the same physical machine with the lowest memory utilization.
[0042] When the physical machine fails, the copies of the first two levels will be lost. For this reason, the memory level of other virtual machines in the cluster is introduced to solve the physical machine failure.
[0043] The third level is the memory of other virtual machines in the cluster. This level can transfer the checkpoint copy to the training node using the internal network of the data center when the copies of the first two levels are lost. At the same time, since the external storage network overhead is very large, the existence of available copies at this level should be ensured as much as possible. By establishing a reinforcement learning model to perceive the changes in environmental resources, the number and access locations of the checkpoint copies at this level in the cluster are dynamically adjusted to ensure the effectiveness of this level under the premise of avoiding excessive impact on the training process.
[0044] The fourth level is external storage. Although there are significant delays and communication overheads in the process of using checkpoints in external storage such as cloud storage and external databases, the availability of checkpoint replicas can be guaranteed. When all replicas stored within the cluster are lost, the checkpoint can be transmitted to the training nodes via the external network of the data center.
[0045] S12. To address the issue of distributed storage of checkpoints in the memory of other virtual machines in the cluster in S11, a reinforcement learning decision-making model based on Deep Q-Network (DQN) was constructed to optimize the storage locations of multiple replicas.
[0046] The reinforcement learning decision-making model: There are N virtual machines in the cluster, and the checkpoint of the training process on a single training node is the smallest operating unit targeted by the DQN model.
[0047] State space: The relevant information of the checkpoint and the virtual machine resources in the cluster are respectively modeled to describe the state space on which the decision-making model depends. As follows: Regarding the resource information of N virtual machines in the cluster, using their bandwidth utilization , memory utilization , and number of failures as resource metrics to respectively describe the communication resource load, memory resource load, and the number of failures within the past time window of the th virtual machine.
[0048] Regarding the checkpoint information, using its replica size s , and basic number of replicas n as metrics to respectively describe the storage space occupied by its replicas and the number of replicas required. Each time the layer does not fail due to a cluster-wide failure, the value is incremented by 1 until the maximum threshold is reached.
[0049] Action space: Consists of an N-dimensional bool vector , where indicates storing a replica of the current decision-making object (checkpoint) in the memory of the th virtual machine VM i , and means not storing a replica of the current decision-making object (checkpoint) in the memory of the th virtual machine VM i .
[0050] When selecting an action, the following three conditions must be met first: First, at most one replica is stored on the same physical machine to avoid both replicas becoming invalid due to a physical machine failure; second, the remaining available memory in the selected virtual machine should be less than the replica size to ensure the operability of the storage solution; third, the total number of replicas . When no feasible solution can be found, the size of the basic number of replicas n is gradually reduced one by one.
[0051] When performing an action, the changes in the state space are described by storing the occupation of communication and memory resources during the operation and calculating the values after occupation, so as to support the calculation of the reward function.
[0052] Reward mechanism: The decision-making model scores the actions based on a reward function R to select a distributed storage scheme: R Among them, is a coefficient and satisfies . The reward function R consists of three parts: is the proportion of the historical failure information of the virtual machines with checkpoints in the entire cluster, used to describe the reliability of the storage scheme based on historical failure information, and the less the better; and are used to describe the impact of the checkpoint storage operation on the performance of the training process itself. is the average memory utilization rate of each storage virtual machine. is the average network bandwidth utilization rate of each storage virtual machine, respectively used to measure the impact degree of the storage checkpoint operation on the memory performance and communication ability of the virtual machine, and the less the better.
[0053] In terms of the network architecture, the model uses a feedforward neural network to evaluate the Q value of the state, and the dimension of the input parameters is , among which, are respectively N the historical failure, memory, and communication resource status of the virtual machines, and a dimension is used to describe the size of the checkpoint replica. The network has two hidden layers and uses the ReLU activation function. The output layer is used to output the Q value under the given current state, that is, the expectation of the future reward after the action is executed.
[0054] S2, a method for predicting the time series and analyzing the change trend of the large model training iteration, predicts the future iteration time series and analyzes the change trend, and the results are provided to S3 to improve the timeliness of the checkpoint frequency update method.
[0055] In this embodiment, for the iteration time analysis method, it consists of a time series prediction module and a change trend analysis module. The time series module will predict the future iteration time series through the GNN-LSTM network based on the historical iteration time series and the distributed architecture information of the model training. The change trend analysis module will output the prediction of the future change trend based on the iteration time series of a certain period of history.
[0056] S21, Time Series Prediction: To predict the future iteration time series of large model distributed training, feature modeling is performed on the training topology and historical iteration time series, and the prediction results of the network model are constructed. The design of the specific neural network prediction model is as follows Figure 3 shown.
[0057] For feature modeling, the feature modeling is divided into two parts, including the training topology and the historical iteration time series.
[0058] Modeling of the LLM Training Topology Structure: The nodes in the topology graph correspond to the workers in the large model distributed training. For the th node , the node features are described by the computational overhead of the node and the resource load of the node, that is, the number of floating-point operations performed by the node in one iteration and the CPU utilization rate of the node , memory utilization rate and GPU utilization rate . The edges in the topology graph represent the data transfer relationships between workers. The edge features include the parallel training relationships between nodes (including data parallelism, pipeline parallelism, and tensor parallelism), the size of the tensors transmitted between nodes, and the communication bandwidth between nodes.
[0059] Modeling of the Historical Iteration Time Series: Specify the number of iterations W as the window, and select the training iteration times of the W most recently executed ones to form the historical iteration time series.
[0060] In terms of the network architecture, the network in the prediction model includes a GCN network and an LSTM network. The GCN network takes the training topology as the input, aiming to analyze the data transfer relationships and worker features between workers during the LLM distributed training process to improve the accuracy of the LSTM model prediction results. The results of the GCN model and the historical event sequence are combined as the input of the LSTM model, and the prediction results of the future iteration time series of the specified window size are obtained through rolling prediction. Among them, the GCN network consists of three graph convolutional layers and a linear layer, while the LSTM network consists of a single-direction LSTM layer and a linear layer.
[0061] The prediction results obtained by the time series prediction module are provided to S22 for trend analysis and to S3 for analysis and decision-making on the checkpoint update frequency.
[0062] S22, Trend Analysis: For the iteration time series with a given window size of M Analyze the iterative time variation trend. This sequence is the output result of the time series prediction module. The specific analysis method is as Figure 4 shown.
[0063] First, divide the variation trend into two major categories: fluctuation and regularity according to the stability of the sequence. For the fluctuation trend, the sequence shows strong unpredictability and instability within the time window, indicating that a fault has occurred or is about to occur in the training environment. A certain checkpoint acquisition strategy needs to be adopted to deal with this situation. For the regularity trend, the change of the model training iteration time over time is relatively stable, and the change trend in the future for a long period of time, such as increasing, decreasing or remaining unchanged, has a certain credibility. By analyzing the degree of change of the regular sequence, quantitative data support can be provided for improving time effectiveness.
[0064] The coefficient of variation (CV) refers to the ratio of the standard deviation and the mean of a set of data, which is usually used in mathematics to reflect the degree of dispersion of a set of data and can measure the stability of the data. Based on the analysis of the coefficient of variation, the sequence T is divided into fluctuation and regularity. Take a threshold C, and let the coefficient of variation of T be , when , it is considered that the change trend of T is fluctuation, and the fluctuation should be directly returned as the analysis result. Otherwise, it is regularity. As an example, through experiments, it is shown that when the threshold C is taken as 0.2, different time series change patterns can be effectively distinguished, and a better checkpoint update frequency can be judged. When the change trend of the sequence T is regular, it is necessary to measure the change direction and degree of T to quantitatively describe the change of the training iteration time in the future for a period of time. Analyze by means of linear fitting. First, calculate the slope k and intercept b of the linear relationship between the iteration batch number and the iteration time based on the least squares method; then calculate the ratio of the iteration time of the future th batch and the first batch ( ), which is called the change coefficient , that is , which is used to measure the degree of change of the iteration time within the future window. The closer the value of is to
[0065] S3. A dynamic hierarchical decision-making method for checkpoint update frequency based on loss analysis, which reduces the overhead of checkpoint operations on the premise of less fault loss.
[0066] In this embodiment, for the checkpoint frequency decision method, first, the tolerable fault loss is calibrated based on the prediction of the future iteration time series and the analysis result of the change trend. Then, based on the calibration result and the overhead of the access policy, the frequency decision is made and the acquisition frequency of the checkpoint is output.
[0067] S31. A hierarchical checkpoint frequency decision method is proposed based on the hierarchical storage architecture. For storage locations with higher bandwidth (such as local memory, memory of other virtual machines within the same physical machine), the checkpoint is updated once per training batch to ensure effective recovery in case of small-scale faults; for storage locations with lower bandwidth (memory of other virtual machines in the data center, external storage), a dynamically adjusted and lower update frequency is adopted to reduce the time and communication overhead of the checkpoint process while minimizing the impact on the training progress as much as possible.
[0068] S32. The basic idea of the dynamic checkpoint frequency decision method based on loss amount analysis is to select the maximum checkpoint interval that can be adopted under the condition that the training loss is tolerable by modeling and analyzing the loss amount of the training progress when a possible fault occurs. The purpose is to minimize the overhead of checkpoint operations while ensuring that the training loss is acceptable.
[0069] Figure 5 It is a schematic diagram of loss amount analysis under a specific update frequency and iteration time series. The farther away from the checkpoint, the greater the loss amount when a training loss occurs.
[0070] First, the loss amount of the training progress when a fault occurs under a given checkpoint acquisition scheme is modeled. For the model iteration time series T, the checkpoint frequency is to perform a checkpoint every p iterations. At this time, a total of checkpoint operations are performed. Suppose the cost of one checkpoint operation is , first calculate the start time and end time of each checkpoint operation: Among them, indicates that in case of a fault before the first checkpoint operation, it is necessary to retrain from the initial model, while indicates the end time of training.
[0071] Next, define the training progress loss and the average training progress loss at any time t as follows: Among them, represents the average value of the fault loss caused when a fault may occur at any time within the future predicted time window under the given checkpoint operation interval p.
[0072] Figure 5 Fig. shows a schematic diagram for calculating the average fault loss with a training batch number of 7 and a checkpoint frequency of performing operations every 2 batches. The average loss is the quotient of the sum of the areas of all triangles divided by the total training time.
[0073] S33. Based on the analysis of the average loss of the training progress at different frequencies, a more reasonable evaluation of the checkpoint frequency decision result can be improved. The specific algorithm is as Figure 6 shown.
[0074] Define the baseline checkpoint acquisition interval Δ as 1, and adopt different decision-making strategies for different change trends.
[0075] If the change trend is fluctuating, it indicates that the state of the training cluster is very unstable and the training process may fail at any time. Therefore, the algorithm returns the fastest checkpoint frequency Δ = 1.
[0076] If the change trend is regular, it indicates that the change of the iteration time is stable and has a certain pattern. At this time, on the basis of the tolerable training loss multiply it by the coefficient of variation so that the decision-making basis can match the training situation in a longer time, that is: Next, unless the fastest checkpoint frequency still cannot meet the requirements, gradually increase the checkpoint acquisition interval Δ one by one until the maximum interval that satisfies the average loss of the training progress not being greater than is found as the return result of the algorithm. In this way, the maximum checkpoint interval that minimizes the total overhead of checkpoint operations is found.
[0077] Embodiment 2 This embodiment provides a verification test of a large model distributed training fault handling method (DesCheck, Des represents Dynamic methods, Check represents checkpoint) based on a dynamic checkpoint strategy of the present invention, and verifies and explains the technical effects adopted in this method.
[0078] Traditional technical methods: GEMINI method, Unicron method, and CheckFreq method. Most traditional technical methods only consider static checkpoint strategies, resulting in their usually lacking the fault handling ability in large model training scenarios.
[0079] To verify that the DesCheck method of the present invention can obtain the fault handling ability during the distributed training of large models compared with traditional methods, in this embodiment, the traditional GEMINI method, Unicron method, CheckFreq method, and this method will be used to perform checkpoint operations during the distributed training of large models respectively, and handle the faults during the training process.
[0080] Test environment: The open-source project DeepSpeedTutorial and related public fault trace data are used as test samples. Different methods are used to perform checkpoint operations during the training process. The solution is run on a real distributed training cluster for 50 batches of training, and test result data is obtained.
[0081] Using the above different methods, data is obtained according to the experimental results. The total training time, resource cost, and fault cost of each method are analyzed.
[0082] The results of the model training time are as Figure 7 shown. The shorter the total time cost, the faster the overall model training process and the higher the effect. It can be seen that among the four methods, the performance of this method is the highest, which is 25.79%, 11.86%, and 44.41% faster than the GEMINI method, Unicron method, and CheckFreq method respectively.
[0083] The DesCheck method of the present invention is compared with the traditional GEMINI method, Unicron method, and CheckFreq method respectively to measure and compare the resource metrics and fault costs per unit time during the training process. Figure 8 represents the resource metrics, including the usage of memory and CPU resources per unit time; Figure 9 represents the fault cost metrics, including the average training progress loss when a fault occurs and the average recovery time after a fault occurs; both factors need to be considered comprehensively. For example, in the CheckFreq method, although the resource cost is very small, the fault recovery requires a huge cost, resulting in the worst overall training efficiency. Among the five methods, the comprehensive performance of this method is the highest, and it can be found that the results of this method have better fault handling ability performance.
[0084] The above description is only a description of the preferred embodiments of the present application, and is not any limitation on the scope of the present application. Any change or modification made by any ordinary person skilled in the art according to the technical content disclosed above should be regarded as an equivalent effective embodiment, and all belong to the scope of protection of the technical solution of the present application.
Claims
1. A method for handling faults in the distributed training of large models based on a dynamic checkpointing strategy, characterized in that, Including: (1) Checkpoint distributed access strategy with cluster topology and environment dynamic awareness: Design a four-layer access topology checkpoint distributed access method to dynamically sense the topology structure and resource status of the GPU cluster, analyze the access locations of multiple copies of checkpoints during the training process, and determine the optimal access strategy; (2) Large model iteration time prediction method with change trend awareness: Implement time series prediction of large model training iterations, and analyze the future iteration time trend through historical data; Combine change trend analysis to improve the timeliness of checkpoint update frequency to cope with the dynamic changes of the training environment; (3) Checkpoint frequency decision method with model training iteration time and trend awareness: Dynamically adjust the checkpoint update frequency based on loss analysis to ensure that the overhead of checkpoint operations is reduced on the premise of a small failure loss; Optimize the checkpoint update strategy by calibrating the tolerable loss and adjusting the frequency to improve the efficiency and reliability of the training process.
2. The method for handling faults in the distributed training of large models based on the dynamic checkpoint strategy according to claim 1, wherein, The checkpoint distributed access strategy with cluster topology and environment dynamic awareness is based on the four-layer access topology checkpoint distributed access method. Through the dynamic awareness of the cluster topology and resources, a reinforcement learning decision model based on the deep Q network is constructed to make decisions on the access locations of multiple copies, including: (1) Cluster topology information collection: Collect the topology information of the four-layer storage media in the GPU cluster, including local memory, physical machine memory, memory of other machines in the cluster, and external storage; (2) Environment dynamic awareness: Analyze the access conditions in the current environment by dynamically sensing the status of cluster resources; (3) Access location decision: Based on the collected topology information and resource status, decide the access locations of multiple copies to achieve fast fault recovery.
3. The method for handling faults in the distributed training of large models based on the dynamic checkpoint strategy according to claim 2, wherein, The specific four-layer access topology is as follows: The storage media are, in order of transmission speed with the training node, local memory, physical machine memory, memory of other machines in the cluster, and external storage, which are four levels; The first level is local memory; The second level is the memory of other virtual machines located on the same physical machine as the training virtual machine; When a single virtual machine fails or the checkpoint in the local memory is damaged, the checkpoint is transmitted to the restarted faulty virtual machine through shared memory and the internal network of the physical machine; In terms of the number of copies, select the virtual machine with the lowest memory utilization on the same physical machine to store a copy of the checkpoint; The third level is the memory of other virtual machines in the cluster; When the copies in the first two levels are lost, the checkpoint copy is transmitted to the training node using the internal network of the data center; By establishing a reinforcement learning model to sense the changes in environmental resources, dynamically adjust the number of checkpoint copies and access locations at this level in the cluster to ensure the effectiveness of this level on the premise of avoiding excessive impact on the training process; The fourth level is external storage.
4. The method for handling faults in the distributed training of large models based on the dynamic checkpoint strategy according to claim 2, wherein The reinforcement learning decision model based on the deep Q network is used to optimize the storage locations of multiple copies; Specifically, there are N virtual machines in the cluster, and the checkpoint of the training process on a single training node is the smallest operation unit targeted by the DQN model; State space: Model the relevant information of the checkpoint and virtual machine resources in the cluster respectively to describe the state space on which the decision model depends; As follows: For the resource information of N virtual machines in the cluster, taking their bandwidth utilization rate , memory utilization rate , number of failures as resource metrics, respectively describe the communication resource load, memory resource load, and the number of failures in the past time window; For checkpoint information, using its replica size s and the basic number of replicas n as indicators, respectively describe the storage space occupied by its replicas and the number of replicas required. Each time the value is incremented by 1 after the layer fails due to a cluster-wide failure until the maximum threshold is reached; Action space: Consists of an N-dimensional boolean vector where represents storing a copy of the current decision object, i.e., the checkpoint, in the memory of the th virtual machine When selecting an action, the following three conditions are met: First, at most one copy is stored on the same physical machine to avoid both copies becoming invalid due to physical machine failures; second, the remaining available memory in the selected virtual machine should be less than the size of the copy to ensure the operability of the storage solution; third, for the total number of copies , when no feasible solution can be found, the size of the basic number of copies n is reduced one by one; When performing an action, describe the change in the state space by the occupation of communication and memory resources through storage operations and calculate the value after occupation to support the calculation of the reward function; Reward mechanism: The decision-making model scores the action based on a reward function to select a distributed storage scheme: R Among them, is a coefficient and satisfies ; The reward function includes three parts: is the proportion of the historical failure information of the virtual machines with checkpoints in the entire cluster, used to describe the reliability of the storage scheme based on historical failure information; and are used to describe the impact of the checkpoint storage operation on the performance of the training process itself, is the average memory utilization rate of each storage virtual machine, is the average network bandwidth utilization rate of each storage virtual machine, respectively used to measure the impact degree of the checkpoint storage operation on the memory performance and communication ability of the virtual machine; In terms of the network architecture, the model uses a feedforward neural network to evaluate the Q-value of the state, and the dimension of the input parameters is , where are the historical failure, memory, and communication resource statuses of N virtual machines respectively, and a dimension is used to describe the size of the checkpoint replica; the network has two hidden layers and uses the ReLU activation function; the output layer is used to output the Q-value under the given current state, that is, the expected future reward after the action is executed.
5. The method for handling faults in the distributed training of large models based on the dynamic checkpoint strategy according to claim 1, wherein, The method for predicting the iteration time of the large model with change trend awareness improves the timeliness of the checkpoint frequency update method through the prediction of the future iteration time series and the analysis of the change trend, including: (1) Historical data analysis: Collect and analyze the historical training iteration time series data to identify the change patterns of the time series; (2) Trend prediction: Use the GNN-LSTM network to predict the future iteration time and output the future change trend; (3) Classification and response: Classify the predicted change trend, timely identify the instability of the training environment, and provide support for the checkpoint strategy.
6. The method for handling faults in the distributed training of large models based on the dynamic checkpoint strategy according to claim 5, wherein The time series prediction: To predict the future iteration time series of the large model distributed training, feature modeling is performed on the training topology and the historical iteration time series, and the prediction result of the network model is constructed; For feature modeling, feature modeling is divided into two parts, including the training topology and the historical iteration time series; Modeling of the LLM Training Topology: Nodes in the Topology Graph correspond to workers in the distributed training of large models. For the th node , the node features are described by the computational overhead of the node and the resource load of the node, that is, the number of floating-point operations performed by the node in one iteration and the CPU utilization rate of the node , memory utilization rate and GPU utilization rate ; The edges in the topology graph represent the data transfer relationships between workers, and the edge features include the parallel training relationships between nodes, the size of the transmitted tensors between nodes, and the communication bandwidth between nodes; For the modeling of the historical iteration time series: Intercept the specified number of iterations W as the window, and select the W most recently executed training iteration times to form the historical iteration time series; In terms of the network architecture, the network in the prediction model includes a GCN network and an LSTM network; The GCN network takes the training topology as the input, aiming to analyze the data transmission relationship and worker features among various workers during the LLM distributed training process to improve the accuracy of the LSTM model prediction result; Combine the result of the GCN model and the historical event sequence as the input of the LSTM model, and obtain the prediction result of the future iteration time series of the specified window size through the rolling prediction method; Among them, the GCN network consists of three graph convolutional layers and a linear layer, while the LSTM network consists of a single-direction LSTM layer and a linear layer.
7. The method for handling faults in the distributed training of large models based on the dynamic checkpoint strategy according to claim 5, wherein The above-mentioned trend analysis: Analyze the trend of iterative time variation for a given iterative time series with a window size of M, which is the output result of the time series prediction module; First, divide the change trend into two major categories: fluctuating and regular according to the stability of the sequence; Divide the sequence T into fluctuating and regular based on the analysis of the coefficient of variation CV; take a threshold C, and let the coefficient of variation of T be , when , it is considered that the change trend of T is fluctuating, and the fluctuation should be directly returned as the analysis result; otherwise, it is regular; When the change trend of sequence T is regular, it is necessary to measure the change direction and degree of T to quantitatively describe the change of training iteration time in a future period of time. Through linear fitting analysis, first, calculate the slope k and intercept b of the linear relationship between the iteration batch number and iteration time based on the least squares method. Then calculate the ratio of the iteration time of the -th batch in the future to that of the first batch ( ), which is called the change coefficient , that is , used to measure the change degree of iteration time within the future window; The closer the value of is to , the more effective the analysis result of the checkpoint acquisition frequency based on the change coefficient will be in a relatively long future time, but it will also sacrifice the reliability of recent checkpoints.
8. The fault handling method for large model distributed training based on the dynamic checkpoint strategy according to claim 1, wherein The method for decision-making on the model training iteration time and checkpoint frequency with trend awareness is a dynamic hierarchical decision-making method for checkpoint update frequency based on loss analysis. On the premise of less failure loss, reduce the overhead of checkpoint operations, including: (1)Loss tolerance setting: Determine the tolerable training loss L base , and combine it with the coefficient of variation to calculate the future loss tolerance L1; (2) Frequency adjustment strategy: On the basis of ensuring that the loss does not exceed L1, gradually increase the checkpoint acquisition interval Δ to find the maximum interval that meets the conditions; (3) Overhead optimization: Combine the overhead of the access strategy to dynamically adjust the checkpoint update frequency to reduce unnecessary checkpoint operations and ensure the efficiency of the training process.
9. The method for handling faults in the distributed training of large models based on the dynamic checkpoint strategy according to claim 8, wherein, The model training iteration time and trend-aware checkpoint frequency decision method first calibrates the tolerable fault loss based on the prediction of the future iteration time series and the analysis result of the change trend; Then, based on the calibration result and the overhead of the access strategy, frequency decision is made and the acquisition frequency of the checkpoint is output; specifically including: S31, A hierarchical checkpoint frequency decision method is proposed based on the hierarchical storage architecture. For storage locations with higher bandwidth, the checkpoint is updated once per training batch to ensure effective recovery in case of small-scale faults; for storage locations with lower bandwidth, a dynamically adjusted and lower update frequency is adopted to reduce the time and communication overhead of the checkpoint process while minimizing the impact on the training progress as much as possible. S32, By modeling and analyzing the loss of the training progress when possible faults occur, the maximum checkpoint interval that can be adopted under the condition that the training loss is tolerable is selected; the purpose is to minimize the overhead of checkpoint operations while ensuring that the training loss is acceptable. First, the loss of training progress at the time of failure was modeled under a given checkpoint acquisition scheme. For the model iteration time series T, the checkpoint frequency is to execute a checkpoint every p iterations. At this time, a total of checkpoint operations are performed; assuming that the cost of one checkpoint operation is , first calculate the start time and end time of each checkpoint operation: Among them, It means that if a failure occurs before the first checkpoint operation, it is necessary to retrain from the initial model, while It represents the time when the training ends. Next, define the training progress loss at any time t and the average training progress loss as follows: Among them, represents the average value of the failure loss caused when a failure may occur at any time within the future predicted time window under the given checkpoint operation interval p. S33, Based on the analysis of the average loss of the training progress at different frequencies, a more reasonable evaluation of the checkpoint frequency decision result is improved; the specific algorithm is as follows: Define the baseline checkpoint acquisition interval Δ as 1, and adopt different decision strategies for different change trends; If the change trend is fluctuating, it indicates that the state of the training cluster is very unstable and the training process may fail at any time. Therefore, the algorithm returns the fastest checkpoint frequency Δ = 1; If the change trend is regular, it indicates that the change in the iteration time is stable and has a certain pattern. At this time, on the basis of the tolerable training loss , multiply it by the change coefficient to make the decision basis match the training situation at a farther time, that is: Next, unless the fastest checkpoint frequency still cannot meet the requirements, increase the checkpoint acquisition interval Δ one by one until the maximum interval that satisfies the condition that the average loss of the training progress is not greater than is found as the return result of the algorithm, thus finding the maximum checkpoint interval that minimizes the total overhead of the checkpoint operation.
Citation Information
Patent Citations
Design method of image classification loss function based on cosine space optimization
CN113052261A
Electric drive assembly fault diagnosis method and device based on large language model
CN118035757A
Load prediction method for computing power network
CN118277093A
Self-adaptive model partitioning method and system applied to distributed training
CN119004112A
Optimization of checkpoint operations for deep learning computing
US20190324856A1
Cited By
Training fault tolerance method and device applied to distributed training system and chip product
CN121257770A
Multi-machine multi-card distributed training optimization system and method oriented to credential heterogeneous environment
CN121603382A
A multi-machine multi-card distributed training optimization system and method for a signal creation heterogeneous environment
CN121603382B
Self-adaptive distributed training optimization method
CN121615720A
Distributed fault tolerance and stability guarantee method for large model training
CN122065986A