Large model training fault tolerance method and system, medium and program product

By using pipeline cavitation time and double buffer technology in a three-dimensional distributed parallel system, and blocking and collecting checkpoint partitions in a two-dimensional communication topology diagram, the problem that checkpoint backup operations affects the training performance of large-scale models is solved, and efficient and reliable large-scale model training fault tolerance is achieved.

CN119938407AActive Publication Date: 2025-05-06BEIJING UNIV OF POSTS & TELECOMM
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
CN202510024249.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-07
Publication Date
2025-05-06
Estimated Expiration
2045-01-07

AI Technical Summary

Technical Problem

In the prior art, checkpoint backup operations affect the training performance of large models, and the big model fault tolerance method does not fully consider the three-dimensional distributed parallel training scenario of large models, resulting in the communication operations of checkpoint backup conflict with the communication operations in the training task, resulting in the deterioration of the training performance of large models.

Method used

On a three-dimensional distributed parallel system, by obtaining pipeline vacuum time, the checkpoint partition is transmitted from the GPU to the CPU for storage, and written to the remote persistent storage, and the checkpoint partition is blocked and collected by using the two-dimensional communication topology diagram to avoid conflicts with communication operations during training.

Benefits of technology

It effectively improves the efficiency of checkpoint backup and model training, ensures the efficiency and reliability of the training process, avoids the situation where the training speed is affected, and achieves the effect of breakpoint training.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119938407A_ABST
    Figure CN119938407A_ABST
Patent Text Reader

Abstract

The invention provides a large model training fault tolerance method and system, a medium and a program product. The method is executed on a three-dimensional distributed parallel system. The system divides parameters of a target large model training process to a plurality of GPUs according to data parallelism, tensor parallelism and pipeline parallelism, and obtains a plurality of check point partitions containing model parameters responsible for the GPUs and optimizer state parameters; obtaining a plurality of assembly line cavitation time in the current training batch of the target large model, taking the check point partition of each GPU in the previous training batch as a self check point partition, and transmitting the check point partition to one buffer area of the corresponding CPU double buffer areas from the GPU, writing a self check point partition in the other buffer area into remote persistent storage, and exchanging the purposes of the two buffer areas; a plurality of processes are constructed into a two-dimensional communication topological graph, check points in a CPU are partitioned and blocked, neighbor process collection operation on a plurality of check point blocks is inserted into a plurality of assembly line cavitation bubble time, and check point backup is carried out by using system idle time.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of large model training technology, and in particular to a large model training fault-tolerant method, system, medium and program product. Background Art

[0002] In recent years, with the rapid growth of large model parameter scale, the demand for computing resources has increased. Training large models requires a lot of computing resources, storage capacity and computing time. The memory requirements of large models make it impossible for a single GPU accelerator card to accommodate the entire model. Large models need to be divided into multiple GPU cards for distributed parallel training. Pipeline parallelism divides large models into multiple pipeline levels by layer and fills the pipeline. The pipeline cavitation problem leads to low computing efficiency. In the actual scenario of large model distributed training, when there are many GPU cards, the frequency of hardware failures in the system is very high. Hardware failure of a single GPU card will cause the entire training system to pause or fail, resulting in serious waste of computing resources. In addition, algorithm or software problems can also cause training system failures. How to achieve efficient fault tolerance for large models and ensure the efficiency and reliability of the training process under limited computing resources has become an urgent problem to be solved.

[0003] In the prior art, checkpoint technology is a widely used fault-tolerant technology. Checkpoint technology implements the backup of model parameters and optimizer state parameters. When an exception occurs, loading the backup checkpoint can restore the training state of the model at the time of backup, achieving the effect of breakpoint resumption. However, the amount of data that needs to be transmitted for checkpoint backup is huge, the input / output network bandwidth connected to remote persistent storage is low, and the latency is high. Storing checkpoints in remote persistent storage cannot achieve high-frequency checkpoint backup, resulting in serious loss of computing power when the system fails; further, the large model fault-tolerant method does not fully consider the complex scenario of large model three-dimensional distributed parallel training. Its checkpoint backup communication operation can easily conflict with the communication operation in the training task, resulting in degradation of large model training performance. Summary of the invention

[0004] In view of this, the embodiments of the present invention provide a large model training fault-tolerant method, system, medium and program product to eliminate or improve one or more defects existing in the prior art, and solve the problem in the prior art that checkpoint backup operations affect the performance of large model training.

[0005] One aspect of the present invention provides a large model training fault tolerance method, which is executed on a three-dimensional distributed parallel system. The three-dimensional distributed parallel system divides the parameters in the target large model training process into multiple GPUs according to three parallel dimensions of data parallelism, tensor parallelism and pipeline parallelism, and obtains multiple checkpoint partitions. Each of the multiple checkpoint partitions contains model parameters and optimizer state parameters arranged according to training time. The method includes the following steps:

[0006] Acquire multiple pipeline cavitation times in the current training batch of the target large model; the pipeline cavitation time is the resource idle time caused by the parallel operation of the pipeline;

[0007] The checkpoint partition that each GPU is responsible for in the previous round of training batch of the target large model is used as its own checkpoint partition and the processes run by the multiple GPUs are obtained, and the self-checkpoint partitions in the multiple processes are transferred from the GPU to one of the double buffers deployed in the corresponding CPU through the PCIE bus for storage, and the self-checkpoint partition stored in the other buffer is written into the remote persistent storage, and after the self-checkpoint partition is written into the remote persistent storage, the uses of the two buffers are swapped;

[0008] The multiple processes are constructed into a two-dimensional communication topology diagram and each self-checkpoint partition is further divided into blocks to obtain checkpoint blocks. The checkpoint blocks perform neighbor process collection operations according to the positional relationship of each checkpoint partition in the two-dimensional communication to collect adjacent process checkpoint partitions into the CPU of the self-checkpoint partition. The neighbor process collection operations are sequentially inserted into the multiple pipeline bubble times.

[0009] In some embodiments, the method further comprises:

[0010] When a failure occurs in the three-dimensional distributed parallel system, a spare GPU is used to replace the failed GPU and a corresponding checkpoint partition is loaded into the spare GPU, including:

[0011] When the CPUs of the standby GPU and the faulty GPU are in the same computing node, the corresponding checkpoint partition is directly loaded from the CPU of its own checkpoint partition;

[0012] When the CPUs of the standby GPU and the faulty GPU are in different computing nodes, loading the corresponding checkpoint partition from the CPU of the adjacent process checkpoint partition;

[0013] When the CPU of the checkpoint partition of the faulty GPU itself and the CPU of the adjacent checkpoint partition are both in a faulty state, the corresponding checkpoint partition is loaded from the remote persistent storage.

[0014] In some embodiments, obtaining multiple pipeline cavitation times in a current training batch of the target large model includes:

[0015] Use execution time test functions or performance profiling tools to measure overall execution time, forward propagation time, and backward propagation time;

[0016] The plurality of pipeline cavitation times are obtained by calculating the total execution time minus the forward propagation time and the reverse propagation time; and the performance evaluation tool visualizes the plurality of pipeline cavitation times.

[0017] In some embodiments, the step of constructing the plurality of processes into a two-dimensional communication topology graph comprises:

[0018] Obtaining the number of the processes executed by the GPU, and setting a two-dimensional Cartesian communication topology structure according to the number of processes; each node in the two-dimensional Cartesian communication topology structure has four neighboring nodes: upper, lower, left, and right;

[0019] According to the two-dimensional Cartesian communication topology structure, the plurality of processes are connected according to a preset horizontal dimension connection requirement and a preset vertical dimension connection requirement to obtain the two-dimensional communication topology graph;

[0020] Check the integrity of the plurality of processes in the two-dimensional communication topology diagram, the correctness of the checkpoint network configuration and the security of the inter-process communication protocol.

[0021] In some embodiments, the self-checkpoint partitions in the plurality of processes are transferred from the GPU to one of the double buffers deployed in the corresponding CPU via the PCIE bus for storage, the self-checkpoint partitions stored in the other buffer are written to the remote persistent storage, and the uses of the two buffers are swapped after the self-checkpoint partitions are written to the remote persistent storage, the process comprising:

[0022] Opening the double buffer in the CPU memory for storing the self-checkpoint partition;

[0023] The double buffer is divided into two types: one buffer is used to temporarily store the self-checkpoint partition; the other buffer is used to transmit data of the stored self-checkpoint partition to the remote persistent storage through an input / output network;

[0024] The dual buffers adopt an alternating use mechanism, and the uses of the two buffers are exchanged after the buffer used to transmit data stored in the self-checkpoint partition to the remote persistent storage through the input / output network completes the task.

[0025] In some embodiments, further partitioning each self-checkpoint partition further includes:

[0026] The self-checkpoint partition is further divided into blocks according to the pipeline vacancy time, network bandwidth and delay performance parameters so that the time of the neighbor process collection operation of each checkpoint block is equal to one pipeline vacancy time.

[0027] In some embodiments, the method further comprises:

[0028] The failure conditions and solutions of the three-dimensional distributed parallel system are stored as logs, and the logs are fed back to the client for analysis to generate a failure analysis report.

[0029] On the other hand, the present invention also provides a large model training fault-tolerant system, comprising a processor, a memory, and a computer program / instructions stored in the memory, wherein the processor is used to execute the computer program / instructions, and when the computer program / instructions are executed, the system implements the steps of any one of the above methods.

[0030] On the other hand, the present invention further provides a computer-readable storage medium having a computer program / instruction stored thereon, which implements the steps of any of the above methods when the program / instruction is executed by a processor.

[0031] On the other hand, the present invention further provides a computer program product, comprising a computer program / instruction, which implements the steps of any of the above methods when executed by a processor.

[0032] The beneficial effects of the present invention are at least:

[0033] In the fault-tolerant method, system, medium and program product for large model training described in the present invention, the neighbor process collection operation is inserted into the pipeline bubble time to perform checkpoint backup, so as to avoid conflicts with communication operations in the training process; through the double buffering technology in the CPU, the remote persistence operation and the CPU checkpoint backup operation are separated to achieve asynchronous operation, and when data is written to the remote persistent storage, the target large model can continue to be carried out at a normal rhythm to avoid the situation where the training speed is affected, effectively improve the efficiency of checkpoint backup and model training, and ensure that both aspects of the work can be carried out relatively efficiently and without interfering with each other; when an exception occurs in the three-dimensional distributed parallel system, the loaded backup checkpoint can be restored to the training state of the model at the time of backup, so as to achieve the effect of breakpoint resumption of training.

[0034] Additional advantages, purposes, and features of the present invention will be described in part in the following description, and will become apparent to those skilled in the art after studying the following, or may be learned from the practice of the present invention. The purposes and other advantages of the present invention may be achieved and obtained by the structures specifically indicated in the specification and the accompanying drawings.

[0035] Those skilled in the art will appreciate that the objectives and advantages that can be achieved with the present invention are not limited to the above specific description, and the above and other objectives that can be achieved by the present invention will be more clearly understood from the following detailed description. BRIEF DESCRIPTION OF THE DRAWINGS

[0036] The drawings described herein are used to provide a further understanding of the present invention, constitute a part of the present application, and do not constitute a limitation of the present invention. In the drawings:

[0037] Figure 1 The figure is a flow chart of a large model training fault-tolerant method according to an embodiment of the present invention.

[0038] Figure 2 A schematic diagram of pipeline parallelism in a three-dimensional distributed parallel system according to an embodiment of the present invention.

[0039] Figure 3 A two-dimensional communication topology diagram during the large model training process described in one embodiment of the present invention.

[0040] Figure 4 A schematic diagram of the pipeline cavitation time during which a neighbor collection operation is inserted in the large model training fault-tolerant method according to an embodiment of the present invention.

[0041] Figure 5 This is a structural diagram of the large model training fault-tolerant method described in one embodiment of the present invention. DETAILED DESCRIPTION

[0042] In order to make the purpose, technical solution and advantages of the present invention more clearly understood, the present invention is further described in detail below in conjunction with the embodiments and the accompanying drawings. Here, the illustrative embodiments of the present invention and their descriptions are used to explain the present invention, but are not intended to limit the present invention.

[0043] It should also be noted that, in order to avoid obscuring the present invention due to unnecessary details, only structures and / or processing steps closely related to the solutions according to the present invention are shown in the accompanying drawings, while other details that are not closely related to the present invention are omitted.

[0044] It should be emphasized that the term “include / comprises” when used herein refers to the presence of features, elements, steps or components, but does not exclude the presence or addition of one or more other features, elements, steps or components.

[0045] It should also be noted that, unless otherwise specified, the term “connection” herein may refer not only to a direct connection but also to an indirect connection involving an intermediate.

[0046] Hereinafter, embodiments of the present invention will be described with reference to the accompanying drawings. In the accompanying drawings, the same reference numerals represent the same or similar components, or the same or similar steps.

[0047] In the prior art, the amount of data that needs to be transmitted for checkpoint backup is huge, the input / output network bandwidth connected to the remote persistent storage is low, and the latency is high. Storing checkpoints in the remote persistent storage cannot achieve high-frequency checkpoint backup, resulting in serious loss of computing power when the system fails; further, the large model fault-tolerant method does not fully consider the complex scenario of three-dimensional distributed parallel training of large models. The communication operation of its checkpoint backup can easily conflict with the communication operation in the training task, resulting in degradation of the large model training performance; the present invention proposes a large model training fault-tolerant method, system, medium and program product, the method is executed on a three-dimensional distributed parallel system, and the three-dimensional distributed parallel system divides the parameters in the training process of the target large model into multiple GPUs according to three parallel dimensions of data parallelism, tensor parallelism and pipeline parallelism, thereby obtaining multiple checkpoint partitions, each of the multiple checkpoint partitions contains model parameters and optimizer state parameters arranged according to training time; obtain multiple checkpoint partitions in the current training batch of the target large model The pipeline vacancy time is the resource idle time caused by the parallelization of the pipeline; the checkpoint partition responsible for each GPU in the previous round of training batch of the target large model is used as the self-checkpoint partition and the processes run by the multiple GPUs are obtained, and the self-checkpoint partitions in the multiple processes are transferred from the GPU to one of the double buffers deployed in the corresponding CPU through the PCIE bus for storage, and the self-checkpoint partition stored in the other buffer is written to the remote persistent storage, and the uses of the two buffers are exchanged after the self-checkpoint partition is written to the remote persistent storage; the multiple processes are constructed into a two-dimensional communication topology map and the self-checkpoint partition of each process is further divided into blocks, and the checkpoint blocks perform neighbor process collection operations according to the positional relationship of each checkpoint partition in the two-dimensional communication topology map to collect the adjacent process checkpoint partitions into the CPU of the self-checkpoint partition, and the neighbor process collection operations are sequentially inserted into the multiple pipeline vacancy times.

[0048] Figure 1The present invention is a flowchart of a large model training fault-tolerant method according to an embodiment of the present invention. Specifically, the present application provides a large model training fault-tolerant method, which is executed on a three-dimensional distributed parallel system. The three-dimensional distributed parallel system divides the parameters in the target large model training process into multiple GPUs according to three parallel dimensions of data parallelism, tensor parallelism, and pipeline parallelism, and obtains multiple checkpoint partitions. Each of the multiple checkpoint partitions contains model parameters and optimizer state parameters arranged according to training time. The method includes the following steps S101 to S103:

[0049] Step S101: Obtain multiple pipeline vacancy times in the current training batch of the target large model; the pipeline vacancy time is the resource idle time caused by pipeline parallelism.

[0050] Step S102: Use the checkpoint partition that each GPU in the previous round of training batch of the target large model is responsible for as its own checkpoint partition and obtain the processes run by multiple GPUs, transfer the self-checkpoint partitions in the multiple processes from the GPU to one of the double buffers deployed in the corresponding CPU through the PCIE bus for storage, write the self-checkpoint partition stored in the other buffer into the remote persistent storage, and after the operation of writing the self-checkpoint partition into the remote persistent storage is completed, swap the uses of the two buffers.

[0051] Step S103: construct multiple processes into a two-dimensional communication topology graph and further divide each checkpoint partition into blocks to obtain checkpoint blocks. The checkpoint blocks perform neighbor process collection operations according to the positional relationship of each checkpoint partition in the two-dimensional communication topology graph to collect adjacent process checkpoint partitions into the CPU of their own checkpoint partitions. The neighbor process collection operations are inserted into multiple pipeline bubble times in sequence.

[0052] In step S101, the three-dimensional distributed parallel system divides the training process of the target large model into multiple GPUs for distributed training, and the three parallel dimensions are data parallelism, tensor parallelism and pipeline parallelism; pipeline parallelism divides the target large model into multiple pipeline stages by layer and each pipeline stage runs on a GPU, and the pipeline stages transmit data through point-to-point communication. In the forward calculation of the target large model training, the intermediate result data of the pipeline stage is transmitted from top to bottom, and the intermediate result data of the reverse stage is transmitted from bottom to top. In order to meet the data dependency between the pipeline stages, pipeline cavitation is generated. The pipeline cavitation is the time period when the GPU is idle; data parallelism copies the pipeline stage to multiple GPUs; tensor parallelism further divides the multi-layer neural network model parameters in each pipeline stage and distributes them to multiple GPUs for parallel calculation. Further, the stability of the measurement results can be improved by measuring the pipeline cavitation time in the previous training batches and taking the average value; the execution time test function uses the timer function, and the performance evaluation tool uses the Profiling tool.

[0053] In some embodiments, obtaining multiple pipeline cavitation times in the current training batch of the target large model includes steps S1011 to S1012:

[0054] Step S1011: Use an execution time test function or a performance evaluation tool to measure the overall execution time, forward propagation time, and backward propagation time.

[0055] Step S1012: multiple pipeline cavitation times are obtained by subtracting the forward propagation time and the reverse propagation time from the total execution time; the performance evaluation tool visualizes the multiple pipeline cavitation times.

[0056] In step S102, the checkpoint includes model parameters and optimizer state parameters. The checkpoint is divided into multiple checkpoint partitions. Each checkpoint partition runs on a GPU and includes model parameters and optimizer state parameters arranged according to training time. After the checkpoint partition is divided into the corresponding GPU, it becomes the GPU's own checkpoint partition and obtains the process of the GPU processing its own checkpoint partition. The process is the process of executing operations on the GPU's own checkpoint partition. Each process includes a GPU and a CPU. The self-checkpoint partition is transferred from the GPU to the CPU for backup. The CPU on each process has a checkpoint partition backup. All programs in the CPU run in the CPU memory. All programs in the GPU run on the GPU video memory.

[0057] Further, in some embodiments, the self-checkpoint partitions in the multiple processes are transferred from the GPU to one of the double buffers deployed in the corresponding CPU via the PCIE bus for storage, and the self-checkpoint partitions stored in the other buffer are written to the remote persistent storage. When the self-checkpoint partition is written to the remote persistent storage, the uses of the two buffers are swapped. The process includes steps S1021 to S1023:

[0058] Step S1021: A double buffer is allocated in the CPU memory for storing its own checkpoint partition.

[0059] Step S1022: The dual buffers are divided into two types: one buffer is used to temporarily store the own checkpoint partitions; and the other buffer is used to transmit data of the stored own checkpoint partitions to the remote persistent storage through the input / output network.

[0060] Step S1023: The dual buffers adopt an alternating use mechanism, and the uses of the two buffers are swapped after the buffer used to transmit data to the remote persistent storage via the input / output network for storing its own checkpoint partitions completes the task.

[0061] Specifically, the CPU memory contains a double buffer, both of which store the self-checkpoint partitions transmitted from the GPU. One buffer is used to store the self-checkpoint partitions, and the other is used to asynchronously store the self-checkpoint partitions stored by the CPU to the remote persistent storage through the input / output network. The remote persistent storage includes but is not limited to disks, tapes, optical disks, and solid-state drives. When the remote persistence operation is completed, the uses of the two buffers are swapped. Through the double buffering technology in the CPU, the remote persistent storage operation and the checkpoint partition backup operation in the CPU are decoupled, and the checkpoint storage to the remote persistent storage is completed asynchronously.

[0062] In step S103, the present application constructs multiple processes into a two-dimensional Cartesian communication topology, wherein adjacent processes in the horizontal dimension of the two-dimensional Cartesian communication topology are interconnected and the first and last processes in each row are interconnected, and adjacent processes in the vertical dimension are interconnected and the first and last processes in each column are interconnected, and each process is a neighbor relationship with the four processes directly connected to it, above, below, left and right. In some embodiments, constructing multiple processes into a two-dimensional communication topology diagram includes steps S1031 to S1033:

[0063] Step S1031: Obtain the number and process number of the GPU, and set a two-dimensional Cartesian communication topology structure according to the number of processes; each node in the two-dimensional Cartesian communication topology structure has four neighboring nodes: upper, lower, left and right.

[0064] Step S1032: According to the two-dimensional Cartesian communication topology structure, multiple processes are connected according to preset horizontal dimension connection requirements and preset vertical dimension connection requirements to obtain a two-dimensional communication topology diagram.

[0065] Step S1033: Check the integrity of multiple processes in the two-dimensional communication topology diagram, check the correctness of the point network configuration and the security of the inter-process communication protocol.

[0066] Furthermore, in some embodiments, further dividing each self-checkpoint partition into blocks also includes: further dividing the self-checkpoint partition into blocks according to the pipeline vacuole time, network bandwidth and delay performance parameters so that the time of the neighbor process collection operation of each checkpoint block is equal to one pipeline vacuole time. After the self-checkpoint partition is divided into blocks, the checkpoint block is obtained, and it is ensured that the time of the checkpoint block executing the neighbor process collection operation is consistent with one pipeline vacuole time, so that the training process and the neighbor collection operation do not interfere with each other. When the pipeline vacuole time is full but there are still remaining neighbor collection operations that have not been inserted and backed up, the neighbor collection operation is continued to be executed until it is completed, and then the gradient synchronization and model parameter update are performed; the checkpoint block executes the neighbor process collection operation to store the checkpoint partitions of the four neighbor processes of the upper, lower, left and right into the CPU of the self-checkpoint partition, and the neighbor process collection operation is performed during the pipeline vacuole time to complete the backup of the training parameters of the previous round of training batches, and one CPU contains the backup data of five checkpoint partitions.

[0067] In some embodiments, the large model training fault tolerance method further includes:

[0068] When a failure occurs in the 3D distributed parallel system, a spare GPU is used to replace the failed GPU and the corresponding checkpoint partition is loaded into the spare GPU, including:

[0069] When the CPUs of the standby GPU and the failed GPU are in the same computing node, the corresponding checkpoint partition is directly loaded from the CPU of its own checkpoint partition.

[0070] When the CPUs of the standby GPU and the failed GPU are in different computing nodes, the corresponding checkpoint partition is loaded from the CPU of the adjacent process checkpoint partition.

[0071] When the CPU of the faulty GPU's own checkpoint partition and the CPU of the adjacent checkpoint partition are both in a faulty state, the corresponding checkpoint partition is loaded from the remote persistent storage.

[0072] Specifically, when a target large model fails, the backup checkpoint partitions can be retrieved from three locations: its own CPU, neighbor CPU, and remote persistent storage, and restored to the training state of the target large model at the time of backup, ensuring the continuity, efficiency, and reliability of the training process.

[0073] In some embodiments, the method further comprises:

[0074] The failure situations and solutions of the three-dimensional distributed parallel system are stored as logs and the logs are fed back to the client for analysis to generate a failure analysis report.

[0075] On the other hand, the present invention also provides a large model training fault-tolerant system, comprising a processor, a memory, and a computer program / instructions stored in the memory, wherein the processor is used to execute the computer program / instructions, and when the computer program / instructions are executed, the system implements the steps of any one of the above methods.

[0076] On the other hand, the present invention further provides a computer-readable storage medium having a computer program / instruction stored thereon, which implements the steps of any of the above methods when the program / instruction is executed by a processor.

[0077] On the other hand, the present invention further provides a computer program product, comprising a computer program / instruction, which implements the steps of any of the above methods when executed by a processor.

[0078] The present invention is described below in conjunction with a specific embodiment:

[0079] The present invention designs a fault-tolerant method, system, medium and program product for large model training, stores multiple checkpoint partition backups on the CPU memory, and fully considers the training scenario of the three-dimensional distributed parallel system, inserts the neighbor process collection operation of the checkpoint block into the pipeline bubble, thereby avoiding the impact on the model training and realizing efficient hiding of the checkpoint backup overhead; further, the present invention adopts a checkpoint asynchronous remote persistence technology based on a double buffer, and asynchronously stores the checkpoint partition from the CPU memory to the remote persistent storage. When a system failure occurs, the checkpoint partition in the CPU memory is preferentially used to restore the model training. If all the checkpoint partition backups in the CPU memory are unavailable, the checkpoint partition is loaded from the remote persistent storage to restore the model training.

[0080] Figure 2The figure is a schematic diagram of pipeline parallelism in a three-dimensional distributed parallel system described in an embodiment of the present invention. Pipeline parallelism divides a large model into multiple pipeline stages by layer, and batch training data is further divided into micro-batches to fill the pipeline. Data is transmitted between pipeline stages through point-to-point communication. The model parameters in the model training process are divided into pipeline stage 1, pipeline stage 2, pipeline stage 3 and pipeline stage 4 and placed on four GPUs in sequence. The four GPUs are device 1, device 2, device 3 and device 4 respectively. The batch training data of a single training iteration is further divided into 8 micro-batches. In the forward calculation of the model training, the intermediate result data of the pipeline stage is transmitted from top to bottom in sequence, and in the reverse calculation, the intermediate result data of the pipeline stage is transmitted from bottom to top in sequence, so as to meet the input and output data dependency. In order to meet the data dependency between pipeline stages, pipeline cavitation will inevitably occur, that is, the time period when the device is in an idle state. In actual execution, the total time of pipeline cavitation on each device is roughly equal, but may be distributed to different time periods. Tensor parallelism further divides the model parameters of each layer of the neural network in the pipeline level, thereby distributing them to more GPU devices for parallel computing. However, under tensor parallelism, each layer of the large model requires multiple collective communication operations on the input data before subsequent forward or reverse calculations can be performed. Collective communication operations include but are not limited to the global reduction operation Allreduce. Multiple collective communication operations are interspersed in the forward or reverse calculation of each micro-batch, and the number of collective communications is proportional to the number of neural network layers in the pipeline level. Data parallelism copies the pipeline level and pipeline scheduling scheme to more GPUs, but different micro-batch training data needs to be input. When each GPU device completes the calculation task of this training iteration batch n, it will obtain the local gradient, and then execute the global reduction operation Allreduce in the data parallel dimension to synchronize the local gradient and update the model parameters.

[0081] The checkpoint includes the model parameter matrix and the optimizer state matrix in the training process. The optimizer state matrix includes the first-order moment of the gradient and the second-order moment of the gradient. When the target large model is trained in a three-dimensional distributed parallel system, the checkpoints of the entire target large model training process are divided into the video memory of each GPU, and each GPU is responsible for one checkpoint partition. The present invention stores the checkpoint partitions on the CPU memory, and forms multiple checkpoint partition backups on multiple CPU memories. The checkpoint partition backup on the CPU memory will be efficiently implemented using the pipeline cavitation time. In addition, the checkpoint partitions are further asynchronously stored from the CPU memory to the remote persistent storage. When all the checkpoint partition backups in the CPU memory are unavailable, the training parameters of the target large model can be restored from the checkpoint partitions in the remote persistent storage.

[0082] 1. Measure the pipeline cavitation time in the three-dimensional distributed parallel system of large models. In the first iteration of the target large model training, each process measures the pipeline cavitation time. Furthermore, in order to improve the stability of the measurement results, the first few training iterations are selected to test the pipeline cavitation time and then take the measurement average. Use a simple execution time test function or a performance evaluation (Profiling) tool provided by the hardware manufacturer. The execution time test function uses the timer function. Specifically, measure the overall execution time, and then subtract the calculation and communication time of forward and reverse propagation to obtain the pipeline cavitation time. The performance evaluation tool can directly visualize the pipeline cavitation time and perform statistical analysis.

[0083] 2. Construct a two-dimensional communication topology. In a three-dimensional distributed parallel system, the checkpoints in the entire target large model training process are divided into each process; each process is constructed as a two-dimensional Cartesian communication topology, and the two dimensions are made as close as possible; Figure 3 A 4*4 two-dimensional Cartesian communication topology diagram of 16 processes P is shown in FIG. Figure 3 As shown, adjacent processes in the horizontal dimension are interconnected, and the first and last processes in each row are interconnected, and adjacent processes in the vertical dimension are interconnected, and the first and last processes in each column are interconnected. In the two-dimensional Cartesian communication topology, each process is a neighbor to the four processes directly connected to it, above, below, left, and right. The GPU memory of each process contains a self-checkpoint partition. Each process first writes its own checkpoint partition from the GPU memory to its own CPU memory through the PCIE bus, and then performs a neighbor process collection operation based on the constructed two-dimensional Cartesian communication topology graph, that is, each process collects the checkpoint partition in the CPU memory of the neighbor process to its own CPU memory through the Internet. Checkpoint backups of five checkpoint partitions will be stored in the CPU memory of all processes; when the system fails, the training is restored from the checkpoint backup in the available CPU memory.

[0084] 3. Figure 4The present invention is a schematic diagram of the insertion of checkpoint backup into pipeline bubble time in the large model training fault-tolerant method described in one embodiment of the present invention. The neighbor process collection operation executed after the checkpoint partition of the CPU memory is divided into blocks is inserted into the pipeline bubble. According to the measured pipeline bubble time, Internet bandwidth and delay performance parameters, the neighbor process collection operation time of each checkpoint block is equal to one pipeline bubble time, and then the neighbor process collection operation on each checkpoint block is inserted into the pipeline bubble in sequence; during the execution of the current training iteration n, the present invention backs up the checkpoint of the previous training iteration n-1, and inserts the neighbor process collection operation on the checkpoint block into the pipeline bubble of the current training iteration n. In most cases, the pipeline bubble can fully accommodate the checkpoint backup. When the pipeline bubble is full but there are still remaining checkpoint blocks that have not completed the checkpoint backup, the backup operation of the remaining checkpoint blocks needs to be completed before the gradient synchronization of this training iteration n.

[0085] 4. Further, the checkpoint partitions are asynchronously stored from the CPU memory to the remote persistent storage. When all the checkpoint backups in the CPU memory are unavailable, the checkpoint partitions stored in the remote persistent storage are restored. Each process writes its own checkpoint partitions in the CPU memory to the remote persistent storage in parallel, so that the remote persistent storage saves a complete model checkpoint. Furthermore, a double buffer is opened in the CPU memory to store the process's own checkpoint partitions, one of which is used to store the process's own checkpoint partitions, and the checkpoint partitions in the other buffer are asynchronously stored in the remote persistent storage. The two buffers are used alternately. When the remote persistence operation is completed, the uses of the two buffers are swapped. Through the double buffering technology in the CPU memory, the remote persistence operation is decoupled from the checkpoint backup operation of the CPU memory, and the remote persistence of the checkpoint is completed asynchronously.

[0086] 5. When a system failure occurs, a backup GPU is used to replace the faulty GPU. At this time, the checkpoint partition needs to be loaded into the video memory of the backup GPU. The backup GPU preferentially loads the checkpoint partition from the CPU memory. When the CPU memories of the backup GPU and the faulty GPU checkpoint partition are in the same computing node, the corresponding checkpoint partition is directly loaded from the CPU memory of the current node; when the CPU memories of the backup GPU and the faulty GPU checkpoint partition are in different computing nodes, the corresponding checkpoint partition is loaded through the Internet network between the nodes; when the CPUs of the checkpoint partitions of the faulty GPU are all in a faulty state, the checkpoint partition cannot be loaded from the CPU memory. At this time, the checkpoint partition needs to be loaded from the remote persistent storage.

[0087] Figure 5This is a structural diagram of the fault-tolerant method for large model training described in an embodiment of the present invention. Node 1 has four neighbors, namely, node 2, node 3, node 4 and node 5, and through the computing node interconnection network, node 1 collects the checkpoint partitions of the neighboring nodes, and stores the checkpoint partitions 2, checkpoint partition 3 and checkpoint partition 4 in the CPU of checkpoint partition 1; each node includes a CPU memory and a GPU video memory, and the buffer 1 in the CPU memory is used to store the checkpoint partitions transmitted from the GPU video memory through the PCIE bus, and the buffer 2 is used to store the checkpoint partitions transmitted from the GPU video memory and transmitted to the remote persistent storage, and transmitted to the remote persistent storage through the input / output network (I / O network).

[0088] In summary, the present invention provides a large model training fault-tolerant method, system, medium and program product, wherein the method is executed on a three-dimensional distributed parallel system, wherein the three-dimensional distributed parallel system divides the parameters in the target large model training process according to three parallel dimensions of data parallelism, tensor parallelism and pipeline parallelism to obtain multiple checkpoint partitions including model parameters and optimizer state parameters arranged according to training time, wherein each GPU is responsible for one checkpoint partition; multiple pipeline cavitation times in the current training batch of the target large model are obtained; the pipeline cavitation time is the resource idle time generated by pipeline parallelism; the checkpoint partition responsible for each GPU in the previous round of training batch of the target large model is used as its own checkpoint partition and multiple processes running on the GPU are obtained, and the checkpoint partition is connected to the GPU through the PCIE bus. The self-checkpoint partitions in the multiple processes are transmitted from the GPU to one of the double buffers deployed in the corresponding CPU for storage, and the self-checkpoint partitions stored in the other buffer are written to the remote persistent storage. After the self-checkpoint partition is written to the remote persistent storage, the uses of the two buffers are exchanged; the multiple processes are constructed into a two-dimensional communication topology map and the self-checkpoint partitions of each process are further divided into blocks. The checkpoint blocks perform neighbor process collection operations according to the positional relationship of each checkpoint partition in the topology map to collect adjacent process checkpoint partitions into the CPU of the self-checkpoint partition, and the neighbor process collection operations are sequentially inserted into the multiple pipeline bubble times, so as to use the system idle time for checkpoint backup.

[0089] Corresponding to the above method, the present invention also provides a large model training fault-tolerant system, which includes a computer device, the computer device includes a processor and a memory, the memory stores computer instructions, the processor is used to execute the computer instructions stored in the memory, and when the computer instructions are executed by the processor, the system implements the steps of the method described above.

[0090] The embodiment of the present invention also provides a computer-readable storage medium on which a computer program is stored, and when the computer program is executed by a processor, the steps of the aforementioned edge computing server deployment method are implemented. The computer-readable storage medium can be a tangible storage medium, such as a random access memory (RAM), a memory, a read-only memory (ROM), an electrically programmable ROM, an electrically erasable programmable ROM, a register, a floppy disk, a hard disk, a removable storage disk, a CD-ROM, or any other form of storage medium known in the technical field.

[0091] It should be understood by those skilled in the art that the exemplary components, systems and methods described in conjunction with the embodiments disclosed herein can be implemented in hardware, software or a combination of the two. Whether it is performed in hardware or software depends on the specific application and design constraints of the technical solution. Professional and technical personnel can use different methods to implement the described functions for each specific application, but such implementation should not be considered to be beyond the scope of the present invention. When implemented in hardware, it can be, for example, an electronic circuit, an application specific integrated circuit (ASIC), appropriate firmware, a plug-in, a function card, etc. When implemented in software, the elements of the present invention are programs or code segments used to perform the required tasks. The program or code segment can be stored in a machine-readable medium, or transmitted on a transmission medium or a communication link via a data signal carried in a carrier.

[0092] It should be clear that the present invention is not limited to the specific configuration and processing described above and shown in the figures. For the sake of simplicity, a detailed description of the known method is omitted here. In the above embodiments, several specific steps are described and shown as examples. However, the method process of the present invention is not limited to the specific steps described and shown, and those skilled in the art can make various changes, modifications and additions, or change the order between the steps after understanding the spirit of the present invention.

[0093] In the present invention, features described and / or illustrated for one embodiment may be used in the same or similar manner in one or more other embodiments, and / or combined with or replace features of other embodiments.

[0094] The above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. For those skilled in the art, the embodiments of the present invention may have various modifications and variations. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present invention shall be included in the protection scope of the present invention.

Claims

1. A large model training fault tolerance method, characterized in that: The method is executed on a three-dimensional distributed parallel system, which divides the parameters in the training process of the target large model into multiple GPUs according to three parallel dimensions of data parallelism, tensor parallelism and pipeline parallelism, and obtains multiple checkpoint partitions, each of which contains model parameters and optimizer state parameters arranged according to training time. The method includes the following steps: Acquire multiple pipeline cavitation times in the current training batch of the target large model; the pipeline cavitation time is the resource idle time caused by the parallel operation of the pipeline; The checkpoint partition that each GPU is responsible for in the previous round of training batch of the target large model is used as its own checkpoint partition and the processes run by the multiple GPUs are obtained, and the self-checkpoint partitions in the multiple processes are transferred from the GPU to one of the double buffers deployed in the corresponding CPU through the PCIE bus for storage, and the self-checkpoint partition stored in the other buffer is written into the remote persistent storage, and after the self-checkpoint partition is written into the remote persistent storage, the uses of the two buffers are swapped; The multiple processes are constructed into a two-dimensional communication topology diagram and each self-checkpoint partition is further divided into blocks to obtain checkpoint blocks. The checkpoint blocks perform neighbor process collection operations according to the positional relationship of each checkpoint partition in the two-dimensional communication to collect adjacent process checkpoint partitions into the CPU of the self-checkpoint partition. The neighbor process collection operations are sequentially inserted into the multiple pipeline bubble times.

2. The large model training fault-tolerant method according to claim 1, characterized in that: The method further comprises: When a failure occurs in the three-dimensional distributed parallel system, a spare GPU is used to replace the failed GPU and a corresponding checkpoint partition is loaded into the spare GPU, including: When the CPUs of the standby GPU and the faulty GPU are in the same computing node, the corresponding checkpoint partition is directly loaded from the CPU of its own checkpoint partition; When the CPUs of the standby GPU and the faulty GPU are in different computing nodes, loading the corresponding checkpoint partition from the CPU of the adjacent process checkpoint partition; When the CPU of the checkpoint partition of the faulty GPU itself and the CPU of the adjacent checkpoint partition are both in a faulty state, the corresponding checkpoint partition is loaded from the remote persistent storage.

3. The large model training fault-tolerant method according to claim 1, characterized in that: Obtaining multiple pipeline cavitation times in the current training batch of the target large model includes: Use execution time test functions or performance profiling tools to measure overall execution time, forward propagation time, and backward propagation time; The plurality of pipeline cavitation times are obtained by calculating the total execution time minus the forward propagation time and the reverse propagation time; and the performance evaluation tool visualizes the plurality of pipeline cavitation times.

4. The large model training fault-tolerant method according to claim 1, characterized in that: The step of constructing a plurality of the processes into a two-dimensional communication topology graph comprises: Obtaining the number of the processes executed by the GPU, and setting a two-dimensional Cartesian communication topology structure according to the number of processes; each node in the two-dimensional Cartesian communication topology structure has four neighboring nodes: upper, lower, left, and right; According to the two-dimensional Cartesian communication topology structure, the plurality of processes are connected according to a preset horizontal dimension connection requirement and a preset vertical dimension connection requirement to obtain the two-dimensional communication topology graph; Check the integrity of the plurality of processes in the two-dimensional communication topology diagram, the correctness of the checkpoint network configuration and the security of the inter-process communication protocol.

5. The large model training fault tolerance method according to claim 1, characterized in that: The self-checkpoint partitions in the plurality of processes are transmitted from the GPU to one of the double buffers deployed in the corresponding CPU via the PCIE bus for storage, the self-checkpoint partitions stored in the other buffer are written into the remote persistent storage, and the uses of the two buffers are swapped after the self-checkpoint partitions are written into the remote persistent storage, the process comprising: Opening the double buffer in the CPU memory for storing the self-checkpoint partition; The double buffer is divided into two types: one buffer is used to temporarily store the self-checkpoint partition; the other buffer is used to transmit data of the stored self-checkpoint partition to the remote persistent storage through an input / output network; The dual buffers adopt an alternating use mechanism, and the uses of the two buffers are exchanged after the buffer used to transmit data stored in the self-checkpoint partition to the remote persistent storage through the input / output network completes the task.

6. The large model training fault tolerance method according to claim 1, characterized in that: Further partitioning of each self-checkpoint partition also includes: The self-checkpoint partition is further divided into blocks according to the pipeline vacancy time, network bandwidth and delay performance parameters so that the time of the neighbor process collection operation of each checkpoint block is equal to one pipeline vacancy time.

7. The large model training fault-tolerant method according to claim 1, characterized in that: The method further comprises: The failure conditions and solutions of the three-dimensional distributed parallel system are stored as logs, and the logs are fed back to the client for analysis to generate a failure analysis report.

8. A large model training fault-tolerant system, comprising a processor, a memory, and a computer program / instruction stored in the memory, characterized in that: The processor is used to execute the computer program / instructions, and when the computer program / instructions are executed, the system implements the steps of the method according to any one of claims 1 to 7.

9. A computer-readable storage medium having a computer program / instruction stored thereon, characterized in that: When the computer program / instructions are executed by a processor, the steps of the method as claimed in any one of claims 1 to 7 are implemented.

10. A computer program product comprising a computer program / instructions, characterized in that When the computer program / instructions are executed by a processor, the steps of the method according to any one of claims 1 to 7 are implemented.

Citation Information

Patent Citations

  • Memory optimization method and system for distributed training of deep learning model

    CN116452404A

  • Single-GPU large model training method and system based on multiple SSDs

    CN118939434A

  • Distributed training fault recovery method and device, medium and computer program product

    CN119201553A

  • Optimization of checkpoint operations for deep learning computing

    US20190324856A1