A Fault Tolerance Method, System, Medium and Program Product for Large Model Training
By using pipeline vacuum time in a three-dimensional distributed parallel system, checkpoint backup is performed using pipeline vacuum time, and combining the double buffers of CPU memory and two-dimensional communication topology diagram, the impact of checkpoint backup on training performance is solved, and efficient breakpoint training and system reliability are achieved.
Patent Information
- Application Number
- CN202510024249.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-07
- Publication Date
- 2025-07-25
- Estimated Expiration
- 2045-01-07
AI Technical Summary
In the prior art, large model training affects the training performance in a three-dimensional distributed parallel system, and the checkpoint backup and the communication operations of the training task are prone to conflict, resulting in serious computing power loss in the event of system failure.
On a three-dimensional distributed parallel system, the parameters of the target large model are divided into multiple GPUs according to the data parallelism, tensor parallelism and pipeline parallelism dimensions, and checkpoint backup is performed using pipeline vacuum time, and the double buffers in the CPU memory are asynchronously stored to remote persistent storage. A two-dimensional communication topology diagram is built to perform neighbor process collection operations to avoid communication conflicts with the training process.
The checkpoint backup efficiency is improved, the efficiency and reliability of the training process is ensured, the training speed is avoided, and the effect of breakpoint training is achieved.
Smart Images

Figure CN119938407B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of large model training, and in particular, to a fault tolerance method, system, medium and program product for large model training. Background Art
[0002] In recent years, with the rapid growth of the parameter scale of large models, the demand for computing resources has increased. Training a large model requires a large amount of computing resources, storage capacity, and computing time. The memory requirements of large models make it impossible for a single GPU acceleration card to accommodate the entire model. Large models need to be divided into multiple GPU cards for distributed parallel training. Pipeline parallelism divides the large model into multiple pipeline stages by layer and fills the pipeline. The pipeline bubble problem results in low computing efficiency. In the actual scenario of large model distributed training, when there are many GPU cards in the system, the frequency of hardware failures is very high. A hardware failure of a single GPU card will cause the entire training system to pause or fail, resulting in serious waste of computing resources. In addition, algorithm or software problems can also cause the training system to fail. How to achieve efficient fault tolerance for large models and ensure the efficiency and reliability of the training process under limited computing resources has become an urgent problem to be solved.
[0003] In the prior art, the checkpoint technology is a widely adopted fault tolerance technology. The checkpoint technology realizes the backup of model parameters and optimizer state parameters. When an exception occurs, loading the backup checkpoint can restore the training state of the model at the time of backup, achieving the effect of resuming training from the breakpoint. However, the amount of data that needs to be transmitted for checkpoint backup is huge, the input / output network bandwidth connecting to remote persistent storage is low, and the latency is high. Storing the checkpoint in remote persistent storage cannot achieve high-frequency checkpoint backup, resulting in serious loss of computing power when the system fails. Further, the large model fault tolerance method does not fully consider the complex scenario of three-dimensional distributed parallel training of large models, and its communication operation for checkpoint backup is very likely to conflict with the communication operation in the training task, resulting in deterioration of the training performance of large models. Summary of the Invention
[0004] In view of this, embodiments of the present invention provide a fault tolerance method, system, medium and program product for large model training to eliminate or improve one or more defects existing in the prior art and solve the problem that the checkpoint backup operation in the prior art affects the training performance of large models.
[0005] One aspect of the present invention provides a fault tolerance method for large model training. The method is executed on a three-dimensional distributed parallel system, which divides the parameters in the target large model training process into multiple GPUs according to three parallel dimensions: data parallelism, tensor parallelism, and pipeline parallelism, and obtains multiple checkpoint partitions. Each checkpoint partition in the multiple checkpoint partitions contains model parameters and optimizer state parameters arranged according to the training time. The method includes the following steps:
[0006] Obtain multiple pipeline bubble times in the current training batch of the target large model; the pipeline bubble time is the resource idle time generated due to pipeline parallelism;
[0007] Take the checkpoint partition responsible for each GPU in the previous training batch of the target large model as its own checkpoint partition and obtain the processes running on the multiple GPUs. Transmit the own checkpoint partition in the multiple processes from the GPU to one buffer in the double buffer deployed in the corresponding CPU through the PCIE bus for storage, write the own checkpoint partition stored in the other buffer into the remote persistent storage, and exchange the uses of the two buffers after the own checkpoint partition is written into the remote persistent storage;
[0008] Construct a two-dimensional communication topology graph for the multiple processes and further divide each own checkpoint partition to obtain checkpoint blocks. The checkpoint blocks perform neighbor process collection operations according to the positional relationship of each checkpoint partition in the two-dimensional communication to collect adjacent process checkpoint partitions into the CPU of its own checkpoint partition. The neighbor process collection operations are sequentially inserted into the multiple pipeline bubble times.
[0009] In some embodiments, the method further includes:
[0010] When the three-dimensional distributed parallel system fails, use a spare GPU to replace the faulty GPU and load the corresponding checkpoint partition into the spare GPU, including:
[0011] When the CPU of the spare GPU and the faulty GPU are in the same computing node, directly load the corresponding checkpoint partition from the CPU of its own checkpoint partition;
[0012] When the CPU of the spare GPU and the faulty GPU are in different computing nodes, load the corresponding checkpoint partition from the CPU of the adjacent process checkpoint partition;
[0013] When the CPUs of the own checkpoint partition and the adjacent checkpoint partition of the faulty GPU are both in a faulty state, load the corresponding checkpoint partition from the remote persistent storage.
[0014] In some embodiments, obtaining multiple pipeline bubble times in the current training batch of the target large model includes:
[0015] Measuring the overall execution time, forward propagation time, and backward propagation time using an execution time test function or a performance evaluation tool;
[0016] Obtaining multiple pipeline bubble times by subtracting the forward propagation time and the backward propagation time from the overall execution time; the performance evaluation tool visually displays multiple pipeline bubble times.
[0017] In some embodiments, the steps of constructing multiple processes into a two-dimensional communication topology graph include:
[0018] Obtaining the number of processes executed by the GPU, and setting a two-dimensional Cartesian communication topology structure according to the number of processes; each node in the two-dimensional Cartesian communication topology structure has four neighbor nodes: up, down, left, and right;
[0019] Connecting multiple processes according to the preset horizontal dimension connection requirements and preset vertical dimension connection requirements according to the two-dimensional Cartesian communication topology structure to obtain the two-dimensional communication topology graph;
[0020] Checking the integrity of multiple processes, the correctness of the checkpoint network configuration, and the security of the inter-process communication protocol in the two-dimensional communication topology graph.
[0021] In some embodiments, transferring the self-checkpoint partition of multiple processes from the GPU to one buffer of the double buffer deployed in the corresponding CPU through the PCIE bus for storage, writing the self-checkpoint partition stored in the other buffer into the remote persistent storage, and swapping the uses of the two buffers after the self-checkpoint partition is written into the remote persistent storage. The process includes:
[0022] Allocating the double buffer for storing the self-checkpoint partition in the CPU memory;
[0023] Dividing the uses of the double buffer, one buffer is used to temporarily store the self-checkpoint partition; the other buffer is used to transmit data of the stored self-checkpoint partition to the remote persistent storage through the input / output network;
[0024] The double buffer adopts an alternating use mechanism, and when the buffer used to transmit data of the stored self-checkpoint partition to the remote persistent storage through the input / output network completes the task, the uses of the two buffers are swapped.
[0025] In some embodiments, further partitioning each self-checkpoint partition further includes:
[0026] Further partitioning the self-checkpoint partition according to the pipeline bubble time, network bandwidth, and latency performance parameters such that the time for the neighbor process to collect operations for each checkpoint block is equal to one pipeline bubble time.
[0027] In some embodiments, the method further includes:
[0028] Storing the situation and solutions of the three-dimensional distributed parallel system failure as a log and feeding back the log to the client for analysis to generate a failure analysis report.
[0029] On the other hand, the present invention also provides a large model training fault tolerance system, including a processor, a memory, and a computer program / instruction stored on the memory. The processor is configured to execute the computer program / instruction, and when the computer program / instruction is executed, the system implements the steps of the method described in any one of the above.
[0030] On the other hand, the present invention also provides a computer-readable storage medium, on which a computer program / instruction is stored. When the program / instruction is executed by a processor, the steps of the method described in any one of the above are implemented.
[0031] On the other hand, the present invention also provides a computer program product, including a computer program / instruction. When the computer program / instruction is executed by a processor, the steps of the method described in any one of the above are implemented.
[0032] The beneficial effects of the present invention are at least:
[0033] In the large model training fault tolerance method, system, medium, and program product of the present invention, inserting the neighbor process collection operation into the pipeline bubble time for checkpoint backup to avoid conflicts with communication operations during the training process; through the double buffering technology in the CPU, separating the remote persistence operation and the CPU checkpoint backup operation to achieve asynchronous execution. When writing data to the remote persistent storage, the target large model can continue to operate at a normal pace, avoiding the situation where the training speed is affected, effectively improving the efficiency of checkpoint backup and model training, and ensuring that both aspects of work can be promoted relatively efficiently and without interference; when the three-dimensional distributed parallel system has an exception, loading the backup checkpoint can restore the training state of the model at the time of backup, achieving the effect of resuming training from a breakpoint.
[0034] Additional advantages, objects, and features of the present invention will be partly set forth in the description which follows, and will partly become obvious to those of ordinary skill in the art upon examination of the following, or may be learned by practice of the present invention. The objects and other advantages of the present invention may be realized and obtained by means of the structures particularly pointed out in the specification and the drawings.
[0035] Those skilled in the art will understand that the objects and advantages that can be achieved by the present invention are not limited to those specifically described above, and the above and other objects that the present invention can achieve will be more clearly understood according to the following detailed description. BRIEF DESCRIPTION OF THE DRAWINGS
[0036] The drawings described herein are for further understanding of the present invention, form a part of this application, and do not limit the present invention. In the drawings:
[0037] Figure 1 It is a schematic flowchart of the large model training fault tolerance method according to an embodiment of the present invention.
[0038] Figure 2 It is a schematic diagram of pipeline parallelism in a three-dimensional distributed parallel system according to an embodiment of the present invention.
[0039] Figure 3 It is a two-dimensional communication topology diagram during the large model training process according to an embodiment of the present invention.
[0040] Figure 4 It is a schematic diagram of inserting pipeline bubble time for the collection operation of neighbors in the large model training fault tolerance method according to an embodiment of the present invention.
[0041] Figure 5 It is a schematic diagram of the structure in the large model training fault tolerance method according to an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0042] To make the objects, technical solutions, and advantages of the present invention clearer and more understandable, the present invention will be further described in detail below in conjunction with the embodiments and the drawings. Here, the illustrative embodiments of the present invention and their descriptions are used to explain the present invention, but do not limit the present invention.
[0043] Here, it should also be noted that in order to avoid obscuring the present invention due to unnecessary details, only the structures and / or processing steps closely related to the solution of the present invention are shown in the drawings, while other details less related to the present invention are omitted.
[0044] It should be emphasized that the term "comprising / including" when used herein refers to the presence of features, elements, steps, or components, but does not exclude the presence or addition of one or more other features, elements, steps, or components.
[0045] Here, it should also be noted that, unless otherwise specified, the term "connection" in this text can not only refer to direct connection, but also indirect connection with intermediate substances.
[0046] In the following, embodiments of the present invention will be described with reference to the accompanying drawings. In the drawings, the same reference numerals represent the same or similar components, or the same or similar steps.
[0047] In the prior art, the amount of data that needs to be transmitted for checkpoint backup is huge, the input / output network bandwidth connecting to remote persistent storage is low, and the latency is high. Storing checkpoints in remote persistent storage cannot achieve high-frequency checkpoint backup, resulting in serious loss of computing power when the system fails; further, the large model fault tolerance method does not fully consider the complex scenario of three-dimensional distributed parallel training of large models, and its checkpoint backup communication operations are prone to conflict with the communication operations in the training task, resulting in deterioration of the large model training performance; the present invention proposes a large model training fault tolerance method, system, medium and program product. The method is executed on a three-dimensional distributed parallel system. The three-dimensional distributed parallel system divides the parameters in the target large model training process into multiple GPUs according to three parallel dimensions of data parallelism, tensor parallelism and pipeline parallelism, so as to obtain multiple checkpoint partitions. Each checkpoint partition in the multiple checkpoint partitions contains model parameters and optimizer state parameters arranged according to the training time; obtaining multiple pipeline bubble times in the current training batch of the target large model; the pipeline bubble time is the resource idle time generated by pipeline parallelism; using the checkpoint partition responsible for each GPU in the previous training batch of the target large model as its own checkpoint partition and obtaining the processes running on the multiple GPUs, and transferring the own checkpoint partition in the multiple processes from the GPU to one buffer in the double buffer deployed in the corresponding CPU through the PCIE bus for storage, writing the own checkpoint partition stored in the other buffer into the remote persistent storage, and swapping the uses of the two buffers after the operation of writing the own checkpoint partition into the remote persistent storage is completed; constructing the multiple processes into a two-dimensional communication topology graph and further dividing the own checkpoint partition of each process, and performing neighbor process collection operations on the own checkpoint partition according to the position relationship of each checkpoint partition in the two-dimensional communication topology graph to collect the adjacent process checkpoint partitions into the CPU of the own checkpoint partition, and inserting the neighbor process collection operations into the multiple pipeline bubble times in turn.
[0048] Figure 1Schematic flowchart of the large model training fault tolerance method according to an embodiment of the present invention. Specifically, the present application provides a large model training fault tolerance method, which is executed on a three-dimensional distributed parallel system. The three-dimensional distributed parallel system divides the parameters in the target large model training process into multiple GPUs according to three parallel dimensions of data parallelism, tensor parallelism, and pipeline parallelism, and obtains multiple checkpoint partitions. Each checkpoint partition in the multiple checkpoint partitions contains model parameters and optimizer state parameters arranged according to the training time. The method includes the following steps S101 to S103:
[0049] Step S101: Obtain multiple pipeline bubble times in the current training batch of the target large model; the pipeline bubble time is the resource idle time generated by pipeline parallelism.
[0050] Step S102: Take the checkpoint partition responsible for each GPU in the previous training batch of the target large model as its own checkpoint partition, and obtain multiple processes running on the GPUs. Transfer the own checkpoint partition in multiple processes from the GPU to one buffer of the double buffer deployed in the corresponding CPU through the PCIE bus for storage, write the own checkpoint partition stored in the other buffer into the remote persistent storage, and swap the uses of the two buffers after the operation of writing the own checkpoint partition into the remote persistent storage is completed.
[0051] Step S103: Construct a two-dimensional communication topology graph for multiple processes, and further divide each own checkpoint partition to obtain checkpoint blocks. The checkpoint blocks perform neighbor process collection operations according to the positional relationship of each checkpoint partition in the two-dimensional communication topology graph to collect adjacent process checkpoint partitions into the CPU of their own checkpoint partition. The neighbor process collection operations are sequentially inserted into multiple pipeline bubble times.
[0052] In step S101, the three-dimensional distributed parallel system divides the training process of the target large model to be performed on multiple GPUs for distributed training. The three parallel dimensions are data parallelism, tensor parallelism, and pipeline parallelism. Pipeline parallelism divides the target large model into multiple pipeline stages layer by layer, and each pipeline stage runs on one GPU. Data is transmitted between pipeline stages through point-to-point communication. During the forward calculation of the target large model training, the intermediate result data of the pipeline stages is transmitted from top to bottom, and the intermediate result data of the backward stage is transmitted from bottom to top. To satisfy the data dependency relationship between pipeline stages, pipeline bubbles are generated. A pipeline bubble is a time period when the GPU is idle. Data parallelism replicates the pipeline stages onto multiple GPUs. Tensor parallelism further divides the multi-layer neural network model parameters in each pipeline stage and distributes them onto multiple GPUs for parallel computing. Further, the stability of the measurement results can be improved by measuring the pipeline bubble times in the previous training batches and taking the average. The execution time test function uses the timer function, and the performance evaluation tool uses the Profiling tool.
[0053] In some embodiments, obtaining the multiple pipeline bubble times in the current training batch of the target large model includes steps S1011 to S1012:
[0054] Step S1011: Measure the total execution time, forward propagation time, and backward propagation time using the execution time test function or the performance evaluation tool.
[0055] Step S1012: Obtain the multiple pipeline bubble times by subtracting the forward propagation time and the backward propagation time from the total execution time; the performance evaluation tool visually displays the multiple pipeline bubble times.
[0056] In step S102, the checkpoint includes model parameters and optimizer state parameters. The checkpoint is divided into multiple checkpoint partitions, and each checkpoint partition runs on one GPU and includes model parameters and optimizer state parameters arranged according to the training time. After the checkpoint partition is divided to the corresponding GPU, it becomes the self-checkpoint partition of that GPU and obtains the process for the GPU to process its self-checkpoint partition. A process is the process for the self-checkpoint partition in the GPU to execute operations. Each process includes a GPU and a CPU. The self-checkpoint partition is transmitted from the GPU to the CPU for backup, and each CPU on each process has a backup of the checkpoint partition. All programs in the CPU run in the CPU memory. All programs in the GPU run on the GPU video memory.
[0057] Further, in some embodiments, the self-checkpoint partitions in multiple processes are transmitted from the GPU to one buffer in the double buffer deployed in the corresponding CPU through the PCIE bus for storage, and the self-checkpoint partitions stored in the other buffer are written into the remote persistent storage. After the operation of writing the self-checkpoint partitions into the remote persistent storage is completed, the uses of the two buffers are swapped. The process includes steps S1021 to S1023:
[0058] Step S1021: Create a double buffer in the CPU memory for storing the self-checkpoint partitions.
[0059] Step S1022: Divide the uses of the double buffer. One buffer is used to temporarily store the self-checkpoint partitions; the other buffer is used to transmit the stored self-checkpoint partitions to the remote persistent storage through the input / output network.
[0060] Step S1023: The double buffer adopts an alternating use mechanism. When the buffer used to transmit the stored self-checkpoint partitions to the remote persistent storage through the input / output network completes the task, the uses of the two buffers are swapped.
[0061] Specifically, the CPU memory contains a double buffer. Both buffers store the self-checkpoint partitions transmitted from the GPU. One buffer is used to store the self-checkpoint partitions, and the other is used to asynchronously store the self-checkpoint partitions stored in the CPU into the remote persistent storage through the input / output network. The remote persistent storage includes but is not limited to disks, tapes, optical discs, and solid-state drives; when the remote persistent operation is completed, the uses of the two buffers are swapped. Through the double-buffer technology in the CPU, the remote persistent storage operation and the checkpoint partition backup operation in the CPU are decoupled, and the checkpoint is asynchronously stored in the remote persistent storage.
[0062] In step S103, the present application constructs multiple processes into a two-dimensional Cartesian communication topology. For the two-dimensional Cartesian communication topology, adjacent processes in the horizontal dimension are connected to each other and the head and tail processes in each row are connected to each other. Adjacent processes in the vertical dimension are connected to each other and the head and tail processes in each column are connected to each other. Each process has a neighbor relationship with the four processes directly connected to its upper, lower, left, and right. In some embodiments, constructing multiple processes into a two-dimensional communication topology diagram includes steps S1031 to S1033:
[0063] Step S1031: Obtain the number of processes executed by the GPU, and set the two-dimensional Cartesian communication topology structure according to the number of processes; each node in the two-dimensional Cartesian communication topology structure has four neighbor nodes, namely, upper, lower, left, and right.
[0064] Step S1032: Connect multiple processes according to the two-dimensional Cartesian communication topology structure according to the preset horizontal dimension connection requirements and preset vertical dimension connection requirements to obtain a two-dimensional communication topology diagram.
[0065] Step S1033: Check the integrity of multiple processes in the two-dimensional communication topology diagram, the correctness of the checkpoint network configuration, and the security of the inter-process communication protocol.
[0066] Further, in some embodiments, further partitioning each own checkpoint partition further includes: further partitioning the own checkpoint partition according to the pipeline bubble time, network bandwidth, and latency performance parameters so that the time for the neighbor process collection operation of each checkpoint partition is equal to one pipeline bubble time. After the own checkpoint partition is partitioned, checkpoint partitions are obtained, ensuring that the time for the checkpoint partition to perform the neighbor process collection operation is consistent with one pipeline bubble time, so that the training process and the neighbor collection operation do not interfere with each other. When the pipeline bubble time is full but there are still remaining neighbor collection operations not inserted and backed up, continue to execute the neighbor collection operation until it is completed and then perform gradient synchronization and model parameter update; the checkpoint partition performs the neighbor process collection operation to store the checkpoint partitions of the four neighbor processes above, below, left, and right in the CPU of its own checkpoint partition. The neighbor process collection operation is performed during the pipeline bubble time to complete the backup of the training parameters of the previous training batch. One CPU contains the backup data of five checkpoint partitions.
[0067] In some embodiments, the large model training fault tolerance method further includes:
[0068] When a three-dimensional distributed parallel system fails, use a spare GPU to replace the faulty GPU and load the corresponding checkpoint partition into the spare GPU, including:
[0069] When the CPUs of the spare GPU and the faulty GPU are in the same computing node, directly load the corresponding checkpoint partition from the CPU of its own checkpoint partition.
[0070] When the CPUs of the spare GPU and the faulty GPU are in different computing nodes, load the corresponding checkpoint partition from the CPU of the adjacent process checkpoint partition.
[0071] When the CPUs of the own checkpoint partition of the faulty GPU and the adjacent checkpoint partition are both in a faulty state, load the corresponding checkpoint partition from the remote persistent storage.
[0072] Specifically, when the target large model fails, the backup checkpoint partition can be retrieved from three locations: its own CPU, the neighbor CPU, and the remote persistent storage, and restored to the training state of the target large model at the time of backup, ensuring the continuity, efficiency, and reliability of the training process.
[0073] In some embodiments, the method further includes:
[0074] Storing the situation and solutions of failures occurring in the three-dimensional distributed parallel system as a log and feeding back the log to the client for analysis to generate a failure analysis report.
[0075] On the other hand, the present invention also provides a large model training fault tolerance system, including a processor, a memory, and computer programs / instructions stored on the memory. The processor is configured to execute the computer programs / instructions, and when the computer programs / instructions are executed, the system implements the steps of any one of the above methods.
[0076] On the other hand, the present invention also provides a computer-readable storage medium, on which computer programs / instructions are stored. When the programs / instructions are executed by a processor, the steps of any one of the above methods are implemented.
[0077] On the other hand, the present invention also provides a computer program product, including computer programs / instructions. When the computer programs / instructions are executed by a processor, the steps of any one of the above methods are implemented.
[0078] The present invention will be described below in conjunction with a specific embodiment:
[0079] The present invention designs a large model training fault tolerance method, system, medium, and program product. Multiple checkpoint partition backups are stored in the CPU memory. At the same time, fully considering the training scenario of the three-dimensional distributed parallel system, the operation of collecting neighbor processes of checkpoint blocks is inserted into the pipeline bubbles, thereby avoiding the impact on model training and achieving efficient hiding of checkpoint backup overhead. Further, the present invention adopts a checkpoint asynchronous remote persistence technology based on a double buffer to asynchronously store checkpoint partitions from the CPU memory to remote persistent storage. When a system failure occurs, the checkpoint partitions in the CPU memory are preferentially used to resume model training. If all the checkpoint partition backups in the CPU memory are unavailable, the checkpoint partitions are loaded from the remote persistent storage to resume model training.
[0080] Figure 2It is a schematic diagram of pipeline parallelism in the three-dimensional distributed parallel system according to an embodiment of the present invention. In pipeline parallelism, the large model is divided into multiple pipeline stages by layer, and the batch training data is further divided into micro-batches to fill the pipeline. Data is transmitted between pipeline stages through point-to-point communication; the model parameters during model training are sliced into pipeline stage 1, pipeline stage 2, pipeline stage 3, and pipeline stage 4 and placed on four GPUs in sequence. The four GPUs are device 1, device 2, device 3, and device 4 respectively; the batch training data for a single training iteration is further divided into 8 micro-batches. In the forward calculation of model training, the intermediate result data of the pipeline stages is transmitted sequentially from top to bottom, and in the reverse calculation, the intermediate result data of the pipeline stages is transmitted sequentially from bottom to top to satisfy the input-output data dependency relationship; in order to satisfy the data dependency relationship between pipeline stages, pipeline bubbles will inevitably occur, that is, the time period when the device is in an idle state. In actual execution, the total time of pipeline bubbles on each device is approximately equal, but it may be distributed in different time periods. Tensor parallelism further slices the model parameters of each layer of the neural network in the pipeline stage, so as to be distributed to more GPU devices for parallel computing. However, under tensor parallelism, each layer of the large model requires multiple collective communication operations on the input data to perform subsequent forward or reverse calculations. The collective communication operations include but are not limited to the global reduction operation Allreduce. Multiple collective communication operations are interspersed in the forward or reverse calculation of each micro-batch, and the number of collective communications is proportional to the number of neural network layers in the pipeline stage. Data parallelism copies the pipeline stage and the pipeline scheduling scheme to more GPUs, but different micro-batch training data needs to be input. When each GPU device finishes the calculation task of batch n of this training iteration, local gradients will be obtained, and then a global reduction operation Allreduce will be performed in the data parallel dimension to synchronize the local gradients and update the model parameters.
[0081] The checkpoint includes the model parameter matrix and the optimizer state matrix during the training process. The optimizer state matrix includes the first moment of the gradient and the second moment of the gradient. When the target large model is trained in the three-dimensional distributed parallel system, the checkpoint of the entire target large model training process is divided and stored on the video memory of each GPU, and each GPU is responsible for 1 checkpoint partition; the present invention stores the checkpoint partition on the CPU memory and forms multiple checkpoint partition backups on multiple CPU memories. The checkpoint partition backups on the CPU memory will be efficiently implemented by using the pipeline bubble time; in addition, the checkpoint partition is further asynchronously stored from the CPU memory to the remote persistent storage. When all the checkpoint partition backups in the CPU memory are unavailable, the training parameters of the target large model can be restored from the checkpoint partition in the remote persistent storage.
[0082] 1. Measure the pipeline bubble time in the three-dimensional distributed parallel system of the large model. In the first iteration of the target large model training, each process measures the pipeline bubble time. Further, in order to improve the stability of the measurement results, the previous several training iterations are selected for the pipeline bubble time test and then the average measurement value is taken. Use a simple execution time test function or the performance profiling tool provided by the hardware manufacturer. The execution time test function uses the timer function. Specifically, by measuring the total execution time and then subtracting the calculation and communication time of the forward and backward propagation, the pipeline bubble time can be obtained. Using the performance profiling tool can directly visualize and statistically analyze the pipeline bubble time.
[0083] 2. Construct a two-dimensional communication topology graph. In the three-dimensional distributed parallel system, the checkpoints during the entire target large model training process are divided among each process; each process is constructed into a two-dimensional Cartesian communication topology graph, and the two dimensions are made as close as possible; Figure 3 This is the two-dimensional communication topology graph during the large model training process described in an embodiment of the present invention. A 4×4 two-dimensional Cartesian communication topology graph of 16 processes P is as Figure 3 shown. Among them, adjacent processes in the horizontal dimension are connected to each other, and the first and last processes in each row are connected to each other. Adjacent processes in the vertical dimension are connected to each other, and the first and last processes in each column are connected to each other. In the two-dimensional Cartesian communication topology, each process has a neighbor relationship with the 4 processes directly connected to its top, bottom, left, and right. Each process's GPU video memory contains a self-checkpoint partition. Each process first writes its self-checkpoint partition from the GPU video memory to its own CPU memory through the PCIE bus, and then based on the constructed two-dimensional Cartesian communication topology graph, performs a neighbor process collection operation, that is, each process collects the checkpoint partitions in the neighbor process's CPU memory to its own CPU memory through the interconnection network. There will be checkpoint backups of five checkpoint partitions stored in the CPU memory of all processes; when the system fails, the training is restored from the available checkpoint backups in the CPU memory.
[0084] 3. Figure 4Schematic diagram of the bubble time of the checkpoint backup insertion pipeline in the large model training fault tolerance method according to an embodiment of the present invention. The neighbor process collection operation performed after partitioning the checkpoint partition in the CPU memory is inserted into the pipeline bubble. According to the measured pipeline bubble time, the interconnection network bandwidth, and the delay performance parameters, the neighbor process collection operation time for each checkpoint block is made equal to a pipeline bubble time, and then the neighbor process collection operations on each checkpoint block are inserted into the pipeline bubble in sequence; in the present invention, during the execution of the current training iteration n, the checkpoint of the previous training iteration n - 1 is backed up, and the neighbor process collection operations on the checkpoint blocks are inserted into the pipeline bubble of the current training iteration n. In most cases, the pipeline bubble can completely accommodate the checkpoint backup. When the pipeline bubble is already full but there are still remaining checkpoint blocks for which the checkpoint backup has not been completed, the backup operations for the remaining checkpoint blocks need to be completed before the gradient synchronization of the current training iteration n.
[0085] 4. Further, asynchronously store the checkpoint partition from the CPU memory to the remote persistent storage. When all the checkpoint backups in the CPU memory are unavailable, recover from the checkpoint partition stored in the remote persistent storage; each process writes its own checkpoint partition in the CPU memory to the remote persistent storage in parallel, so that the remote persistent storage stores a complete model checkpoint; further, a double buffer is opened in the CPU memory to store the process's own checkpoint partition. One buffer is used to store the process's own checkpoint partition, and the checkpoint partition in the other buffer is asynchronously stored to the remote persistent storage. The two buffers are used alternately. When the remote persistent operation is completed, the uses of the two buffers are swapped. Through the double buffer technology in the CPU memory, the remote persistent operation is decoupled from the checkpoint backup operation in the CPU memory, and the remote persistence of the checkpoint is completed asynchronously.
[0086] 5. When a system failure occurs, use a spare GPU to replace the faulty GPU. At this time, it is necessary to load the checkpoint partition into the video memory of the spare GPU. The spare GPU preferentially loads the checkpoint partition from the CPU memory. When the CPU memory of the checkpoint partitions of the spare GPU and the faulty GPU is on the same computing node, directly load the corresponding checkpoint partition from the CPU memory of the current node; when the CPU memory of the checkpoint partitions of the spare GPU and the faulty GPU is on different computing nodes, load the corresponding checkpoint partition through the interconnection network between the nodes; when the CPUs of the checkpoint partitions of the faulty GPU are all in a faulty state, it is impossible to load the checkpoint partition from the CPU memory. At this time, it is necessary to load the checkpoint partition from the remote persistent storage.
[0087] Figure 5It is a schematic structural diagram in the large model training fault tolerance method according to an embodiment of the present invention. Node 1 has four neighbors, namely Node 2, Node 3, Node 4, and Node 5, and is interconnected through a computing node network. Node 1 collects the checkpoint partitions of the neighbor nodes and stores the checkpoint partitions 2, 3, and 4 in the CPU of the checkpoint partition 1. Each node contains a CPU memory and a GPU video memory. Buffer 1 in the CPU memory is used to store the checkpoint partitions transferred from the GPU video memory through the PCIE bus, and buffer 2 is used to store the checkpoint partitions transferred from the GPU video memory and then transferred to the remote persistent storage, and is transmitted into the remote persistent storage through the input / output network (I / O network).
[0088] In summary, the present invention provides a large model training fault tolerance method, system, medium, and program product. The method is executed on a three-dimensional distributed parallel system. The three-dimensional distributed parallel system divides the parameters in the target large model training process into multiple checkpoint partitions containing model parameters and optimizer state parameters arranged according to training time in three parallel dimensions: data parallelism, tensor parallelism, and pipeline parallelism. Each GPU is responsible for 1 checkpoint partition; obtaining multiple pipeline bubble times in the current training batch of the target large model; the pipeline bubble time is the resource idle time generated by pipeline parallelism; using the checkpoint partitions responsible for each GPU in the previous training batch of the target large model as its own checkpoint partitions and obtaining multiple processes running on the GPUs, and transferring the own checkpoint partitions in the multiple processes from the GPUs to one of the double buffers deployed in the corresponding CPUs through the PCIE bus for storage, writing the own checkpoint partitions stored in the other buffer into the remote persistent storage, and swapping the uses of the two buffers after the operation of writing the own checkpoint partitions into the remote persistent storage is completed; constructing the multiple processes into a two-dimensional communication topology graph and further dividing the own checkpoint partitions of each process, and performing neighbor process collection operations on the own checkpoint partitions according to the positional relationship of the checkpoint partitions in the topology graph to collect the adjacent process checkpoint partitions into the CPUs of the own checkpoint partitions, and inserting the neighbor process collection operations into the multiple pipeline bubble times in sequence, so as to perform checkpoint backup using the system idle time.
[0089] Corresponding to the above method, the present invention also provides a large model training fault tolerance system. The system includes a computer device, the computer device includes a processor and a memory, the memory stores computer instructions, and the processor is used to execute the computer instructions stored in the memory. When the computer instructions are executed by the processor, the system implements the steps of the method described above.
[0090] An embodiment of the present invention also provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the steps of the foregoing edge computing server deployment method are implemented. The computer-readable storage medium may be a tangible storage medium, such as a random access memory (RAM), internal memory, read-only memory (ROM), electrically programmable ROM, electrically erasable programmable ROM, register, floppy disk, hard disk, removable storage disk, CD-ROM, or any other form of storage medium well-known in the technical field.
[0091] Those of ordinary skill in the art should understand that the various exemplary components, systems, and methods described in connection with the embodiments disclosed herein can be implemented in hardware, software, or a combination of both. Specifically, whether to implement in hardware or software depends on the specific application and design constraints of the technical solution. A professional technician can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of the present invention. When implemented in hardware, it can be, for example, an electronic circuit, an application-specific integrated circuit (ASIC), appropriate firmware, a plug-in, a functional card, and so on. When implemented in software, the elements of the present invention are programs or code segments used to perform the required tasks. The program or code segment can be stored in a machine-readable medium or transmitted through a data signal carried in a carrier wave on a transmission medium or a communication link.
[0092] It should be clear that the present invention is not limited to the specific configurations and processes described above and shown in the figures. For the sake of brevity, the detailed description of known methods is omitted here. In the above embodiments, several specific steps are described and shown as examples. However, the method process of the present invention is not limited to the specific steps described and shown. Those skilled in the art can make various changes, modifications, and additions, or change the order between steps after understanding the spirit of the present invention.
[0093] In the present invention, the features described and / or illustrated for one embodiment can be used in the same way or in a similar way in one or more other embodiments, and / or combined with the features of other embodiments or replace the features of other embodiments.
[0094] The foregoing is only the preferred embodiment of the present invention and is not used to limit the present invention. For those skilled in the art, various changes and modifications can be made to the embodiments of the present invention. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present invention shall be included within the protection scope of the present invention.
Claims
1. A large model training fault tolerance method, characterized in that The method is executed on a three-dimensional distributed parallel system, which divides the parameters in the training process of the target large model into multiple GPUs according to three parallel dimensions: data parallelism, tensor parallelism, and pipeline parallelism, and obtains multiple checkpoint partitions. Each checkpoint partition in the multiple checkpoint partitions contains model parameters and optimizer state parameters arranged according to the training time. The method includes the following steps: Obtain multiple pipeline bubble times in the current training batch of the target large model; the pipeline bubble time is the resource idle time generated by pipeline parallelism; Use the checkpoint partition responsible for each GPU in the previous training batch of the target large model as its own checkpoint partition and obtain the processes running on the multiple GPUs. Transfer the own checkpoint partition in the multiple processes from the GPU to one buffer in the double buffer deployed in the corresponding CPU through the PCIE bus for storage, write the own checkpoint partition stored in the other buffer to the remote persistent storage, and swap the uses of the two buffers after the own checkpoint partition is written to the remote persistent storage; Construct the multiple processes into a two-dimensional communication topology graph and further divide each own checkpoint partition to obtain checkpoint blocks. The checkpoint blocks perform neighbor process collection operations according to the position relationship of each checkpoint partition in the two-dimensional communication to collect the adjacent process checkpoint partitions into the CPU of its own checkpoint partition. The neighbor process collection operations are sequentially inserted into the multiple pipeline bubble times.
2. The large model training fault tolerance method according to claim 1, wherein The method further includes: When a failure occurs in the three-dimensional distributed parallel system, use a spare GPU to replace the failed GPU and load the corresponding checkpoint partition into the spare GPU, including: When the CPUs of the spare GPU and the failed GPU are in the same computing node, directly load the corresponding checkpoint partition from the CPU of its own checkpoint partition; When the CPUs of the spare GPU and the failed GPU are in different computing nodes, load the corresponding checkpoint partition from the CPU of the adjacent process checkpoint partition; When the CPUs of the own checkpoint partition and the adjacent checkpoint partition of the failed GPU are both in a failed state, load the corresponding checkpoint partition from the remote persistent storage.
3. The large model training fault tolerance method according to claim 1, characterized in that, Obtaining multiple pipeline bubble times in the current training batch of the target large model includes: Use an execution time test function or a performance evaluation tool to measure the overall execution time, forward propagation time, and backward propagation time; Calculate the multiple pipeline bubble times by subtracting the forward propagation time and the backward propagation time from the overall execution time; the performance evaluation tool visually displays the multiple pipeline bubble times.
4. The large model training fault tolerance method according to claim 1, characterized in that The step of constructing the multiple processes into a two-dimensional communication topology graph includes: Obtain the number of processes executed by the GPU, and set a two-dimensional Cartesian communication topology structure according to the number of processes; each node in the two-dimensional Cartesian communication topology structure has four neighbor nodes: up, down, left, and right. Connect multiple of the processes according to the two-dimensional Cartesian communication topology structure and the preset horizontal dimension connection requirements and preset vertical dimension connection requirements to obtain the two-dimensional communication topology diagram; Check the integrity of multiple of the processes in the two-dimensional communication topology diagram, the correctness of the checkpoint network configuration, and the security of the inter-process communication protocol.
5. The large model training fault tolerance method according to claim 1, characterized in that Transfer the self-checkpoint partition in multiple of the processes from the GPU to one of the double buffers deployed in the corresponding CPU through the PCIE bus for storage, write the self-checkpoint partition stored in the other buffer into the remote persistent storage, and swap the uses of the two buffers after the self-checkpoint partition is written into the remote persistent storage. The process includes: Allocate the double buffer in the CPU memory for storing the self-checkpoint partition; Divide the uses of the double buffer. One buffer is used to temporarily store the self-checkpoint partition; the other buffer is used to transfer the stored self-checkpoint partition to the remote persistent storage through the input / output network; The double buffer adopts an alternating use mechanism. When the buffer used to transfer the stored self-checkpoint partition to the remote persistent storage through the input / output network completes the task, swap the uses of the two buffers.
6. The large model training fault tolerance method according to claim 1, characterized in that, Further dividing each self-checkpoint partition further includes: Further divide the self-checkpoint partition according to the pipeline bubble time, network bandwidth, and latency performance parameters so that the time for the neighbor process collection operation of each checkpoint block is equal to one pipeline bubble time.
7. The large model training fault tolerance method according to claim 1, wherein The method further includes: Store the failure situations and solutions of the three-dimensional distributed parallel system as a log and feedback the log to the client for analysis to generate a failure analysis report.
8. A large model training fault tolerance system, comprising a processor, a memory, and computer programs / instructions stored on the memory, characterized in that, The processor is used to execute the computer program / instructions. When the computer program / instructions are executed, the system implements the steps of the method according to any one of claims 1 to 7.
9. A computer-readable storage medium having computer programs / instructions stored thereon, characterized in that, When the computer program / instructions are executed by the processor, the steps of the method according to any one of claims 1 to 7 are implemented.
10. A computer program product, comprising a computer program / instructions, characterized in that, When the computer program / instructions are executed by the processor, the steps of the method according to any one of claims 1 to 7 are implemented.
Citation Information
Patent Citations
Memory optimization method and system for distributed training of deep learning model
CN116452404A
Single-GPU large model training method and system based on multiple SSDs
CN118939434A