Model training method and system
By decomposing the MoE large model training into multiple layers for parallel execution and scheduling computation and communication in parallel, the problem of long GPU computation waiting time in high-latency, low-bandwidth network environments is solved, thereby improving GPU utilization and model training efficiency.
Patent Information
- Application Number
- CN202511159577.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-08-18
- Publication Date
- 2025-12-05
AI Technical Summary
Existing parallel acceleration methods for large MoE models with overlapping computation and communication pipelines result in increased GPU computation latency in high-latency, low-bandwidth network environments, leading to low GPU computing power utilization and low model training efficiency.
By decomposing the model training into multiple layers and executing them in parallel, and scheduling computation and communication operations in parallel, especially by finely controlling the number of micro-batches during the warm-up, stabilization and cooling phases, a strategy of parallel execution of computation and communication is adopted to avoid communication operations that are delayed during serial execution, thereby achieving the overlap of GPU computation and communication.
In high-latency, low-bandwidth network environments, it significantly improves GPU utilization, reduces waiting time during training iterations, and enhances model training efficiency.
Smart Images

Figure CN121070547A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the field of model training, in particular to a model training method and system. BACKGROUND
[0002] The existing parallel acceleration method of the computing and communication overlapping pipeline of the MoE large model creates GPU computing flow and GPU communication flow on the same GPU card to realize the overlapping execution of GPU computing and GPU communication. In the GPU communication flow, both all-to-all communication and send / recv communication between GPU servers need to be performed, and the all-to-all communication and the send / recv communication between the GPU servers are in a serial execution relationship.
[0003] This serial execution mode has a strong dependence on a low-latency and high-bandwidth network environment. Only in the low-latency and high-bandwidth network environment can the complete overlap of GPU computing and GPU communication be realized. When the existing technology is used in a high-latency and low-bandwidth network, the send / recv communication will block the subsequent all-to-all communication and GPU computing, so that the waiting time of the GPU computing is increased, thereby slowing down the entire large model training and inference process and significantly reducing the GPU computing power utilization rate.
[0004] In view of the above problems, no effective solution has been proposed so far. SUMMARY
[0005] The embodiments of the present application provide a model training method and system to at least solve the technical problem of low model training efficiency in the related art.
[0006] According to an aspect of some embodiments of the present application, there is provided a model training method applied to a target GPU, wherein the target GPU is configured to drive a corresponding model layer segment included in an initial model, the method comprising: receiving a model training request, wherein the model training request carries a target strategy, the target strategy being a training strategy for training the initial model; in response to the model training request, determining, according to the target strategy, a number of micro-batches processed by each of a plurality of pipeline stages, wherein the plurality of pipeline stages comprises a warm-up stage, a stabilization stage, and a cooling-down stage; performing, in a manner of computing and communicating in parallel, a corresponding computing and communicating operation of each of the plurality of pipeline stages in a sequence of the pipeline stages, to obtain a current back-propagation result corresponding to each of a plurality of micro-batches, wherein the corresponding computing and communicating operation comprises a computing operation of computing a propagation result of the corresponding micro-batch according to the number of micro-batches of the corresponding stage, and a communicating operation of transmitting the propagation result of the corresponding micro-batch; and training the model layer segment according to the current back-propagation result corresponding to each of the plurality of micro-batches, to update a model parameter corresponding to the model layer segment, to obtain a target model.
[0007] According to an aspect of some embodiments of the present application, there is provided a hybrid model of experts (MOE) target model, the target model comprising a plurality of model layer segments, wherein each of the plurality of model layer segments is driven by a corresponding GPU, and the corresponding GPU is configured to train the corresponding model layer segment using the method described above to update a model parameter of the corresponding model layer segment, to obtain the target model.
[0008] According to an aspect of some embodiments of the present application, there is provided a model training system comprising a target model and a plurality of GPUs, wherein the target model comprises a plurality of model layer segments, each of the plurality of model layer segments is driven by a corresponding GPU, and the corresponding GPU is configured to train the corresponding model layer segment using the method described above to update a model parameter of the corresponding model layer segment, to obtain the target model.
[0009] According to an aspect of some embodiments of the present application, there is provided an electronic device comprising a processor and a memory storing instructions executable by the processor, wherein the processor is configured to execute the instructions to implement any of the methods described above.
[0010] According to an aspect of some embodiments of the present application, there is provided a computer-readable storage medium storing instructions which, when executed by a processor of an electronic device, cause the electronic device to perform any of the methods described above.
[0011] In the embodiment of the present application, a model training request is received, wherein the target strategy is carried in the model training request, and the target strategy is a training strategy for training an initial model; in response to the model training request, the number of micro-batches processed by each pipeline stage is determined according to the target strategy, wherein the plurality of pipeline stages include a warm-up stage, a stable stage, and a cooling stage; in a manner of parallel execution of calculation and communication, the calculation and communication operations corresponding to each pipeline stage are sequentially executed in the order of execution of the pipeline stages to obtain the current backward propagation results corresponding to the plurality of micro-batches, wherein the corresponding calculation and communication operations include a calculation operation for calculating the propagation results of the corresponding micro-batches according to the number of micro-batches of the corresponding stage, and a communication operation for transmitting the propagation results of the corresponding micro-batches; the model layer segment is trained according to the current backward propagation results corresponding to the plurality of micro-batches to update the model parameters corresponding to the model layer segment, and the target model is obtained. The present application finely adjusts the number of micro-batches in the pipeline warm-up, stable, and cooling stages of model training, and adopts a strategy of parallel execution of calculation and communication, thereby effectively solving the problem of waste of computing resources caused by waiting for communication in the prior art. Especially in a high-latency and low-bandwidth network environment, the communication process and the calculation process are scheduled in parallel, thereby avoiding the blocking of subsequent calculation caused by the delay of the communication operation in serial execution, and at the same time, the parallel scheduling also significantly improves the GPU utilization, reduces the waiting time in the training iteration, and further speeds up the model training process, thereby improving the overall training efficiency, and further solving the technical problem of low model training efficiency in the related art. BRIEF DESCRIPTION OF DRAWINGS
[0012] The accompanying drawings, which are included to provide a further understanding of the present application and are incorporated in and constitute a part of this application, illustrate embodiments of the present application and together with the description serve to explain the present application. In the drawings:
[0013] Figure 1 is a flowchart of a model training method according to an embodiment of the present application;
[0014] Figure 2 is a schematic diagram of a pipeline parallel acceleration method using an existing calculation and communication overlap in a low-latency and high-bandwidth network environment;
[0015] Figure 3 is a schematic diagram of a pipeline parallel acceleration method using an existing calculation and communication overlap in a high-latency and low-bandwidth network environment;
[0016] Figure 4 is a schematic diagram of a pipeline parallel acceleration method using a new calculation and communication overlap in a high-latency and low-bandwidth network environment according to an optional embodiment of the present application;
[0017] Figure 5is a timing diagram of a common pipelined parallel acceleration method in the prior art;
[0018] Figure 6 is a diagram of intra-group computation and communication overlap in the prior art;
[0019] Figure 7 is another diagram of intra-group computation and communication overlap in the prior art;
[0020] Figure 8 is a timing diagram of a new computation and communication overlap pipeline for MoE large models provided by an optional embodiment of the present application;
[0021] Figure 9 is a flowchart of computation and communication operations in a warm-up stage of the pipeline provided by an optional embodiment of the present application;
[0022] Figure 10 is a flowchart of computation and communication operations in a stable stage of the pipeline provided by an optional embodiment of the present application;
[0023] Figure 11 is a flowchart of computation and communication operations in a cooling stage of the pipeline provided by an optional embodiment of the present application. DETAILED DESCRIPTION
[0024] In order to enable persons skilled in the art to better understand the present application, the technical solutions in the embodiments of the present application will be described clearly and completely below with reference to the drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by persons skilled in the art without creative labor should fall within the scope of protection of the present application.
[0025] It should be noted that the terms "first", "second", and the like in the specification and claims of the present application and the above-described drawings are used to distinguish similar objects, and do not necessarily indicate a specific order or a chronological sequence. It should be understood that the data thus used can be interchanged under appropriate circumstances, so that the embodiments of the present application described herein can be implemented in an order other than those illustrated or described herein. In addition, the terms "include" and "have" and any variations thereof are intended to cover non-exclusive inclusion, for example, a process, method, system, product or device that includes a series of steps or units does not necessarily have to be limited to those steps or units clearly listed, but can include other steps or units that are not clearly listed or inherent to these processes, methods, products or devices.
[0026] First, some nouns or terms appearing in the description of the embodiments of the present application are applicable to the following explanations:
[0027] All-to-All communication: All-to-All communication is also known as All-to-All communication. All-to-All communication is a communication mode in distributed computing, in which every node in the network (in this scenario, every GPU server) needs to exchange data directly with all other nodes. This means that every node in the network will send data packets to all other nodes and receive data packets from all nodes.
[0028] Send / Recv communication: Send / Recv communication is a more traditional point-to-point communication mode, in which one node in the network (usually a GPU server) sends data to another specific node, while also receiving data from another node. In this mode, communication is directional, and data is only sent to the intended target node, similarly, the node only receives data from its intended source.
[0029] MoE large model: Mixture of Experts (MoE) large model, MoE large model is a special architecture of neural network model, which contains multiple "expert" modules, each of which is responsible for processing a part of input data. In the training and inference process, input data is dynamically routed to the expert module that is most suitable for processing the data, which can significantly improve the processing capacity and efficiency of the model. MoE large model is widely used in natural language processing (NLP), computer vision (CV), speech recognition and other fields, especially in deep learning tasks that process massive data and long text, images, audio sequences.
[0030] GPU: GPU server is a high-performance computing server equipped with one or more graphics processing units (GPU). GPU was originally designed to handle graphics rendering and video output, but due to its powerful parallel computing capabilities, it has been widely used in scientific computing, deep learning, big data processing, virtual reality, game development and other fields.
[0031] Embodiment 1
[0032] According to an embodiment of the present application, an embodiment of a model training method is provided. It should be noted that the steps shown in the flowchart of the accompanying drawings can be executed in a computer system such as a set of computer executable instructions, and although the logical order is shown in the flowchart, in some cases, the steps shown or described herein can be executed in a different order than that shown here.
[0033] Firstly, the application scenarios provided by the present application are introduced. The present application provides a model training method, which is applied to a target GPU. The target GPU is used to drive a corresponding model layer segment. The model layer segment is included in an initial model. The core of this method is to decompose a complex model into multiple layer segments. The corresponding layer segments are executed on the corresponding GPUs. The execution can be performed in parallel. Meanwhile, a carefully designed communication mechanism is used to ensure the collaborative work between the layer segments and improve the overall training efficiency.
[0034] Among them, the target GPU is involved. In a parallel computing architecture, the target GPU refers to a GPU that undertakes a specific model layer segment computing task. In a multi-GPU parallel training scenario, different GPUs are responsible for different parts of the model, i.e., different layer segments of the model, so as to drive the execution of the corresponding layer segments using and training.
[0035] Among them, the model layer segment (Model Segments) is involved. The model layer segment refers to the segmentation of the entire model into several independent or relatively independent layer segments. Each layer segment performs its own corresponding task. For example, when the model is an image recognition model, it can be divided into an image preprocessing layer segment, a feature extraction layer segment, an object recognition layer segment, etc. This segmentation helps distributed training, which can enable the parallel execution of each part of the model on a respective GPU, thereby improving the efficiency of model training and also improving the processing efficiency when using the model.
[0036] Figure 1 The flowchart of the model training method according to the embodiment of the present application is shown in FIG. 1. Figure 1 As shown in FIG. 1, the method comprises the following steps:
[0037] Step S102, receiving a model training request, wherein the model training request carries a target strategy, and the target strategy is a training strategy for training an initial model;
[0038] In step S102 provided in the present application, the model training request is received.
[0039] Among them, the model training request is involved. The model training request is an instruction or request received by the system to start the model training process, which can contain information such as model type, training data set, parameter configuration, etc.
[0040] Among them, the target strategy is involved. The target strategy includes a specific execution strategy for training the model or a training target. For example, the specific execution strategy can be a parallel strategy, and in the case of an image recognition model, the highest recognition accuracy can be used as the training target, which can be used to guide how to handle the parallel training strategy of the model.
[0041] In this step, the system first receives a model training request from the user, which contains the target strategy for guiding the model training. Through this request, the entire model training process is started, and the main parameters and strategies in the training process are determined.
[0042] Step S104, in response to the model training request, determines the number of micro-batches processed by each pipeline stage according to the target strategy, wherein the multiple pipeline stages include: warm-up stage, stable stage, cooling stage;
[0043] In step S104 provided in the present application, in response to the model training request, the number of micro-batches processed by each pipeline stage is determined according to the target strategy.
[0044] Wherein, the micro-batch is involved, in parallel training, a large training batch is subdivided into multiple smaller micro-batches to adapt to the parallel execution of computation and communication in the pipeline parallel acceleration method.
[0045] It should be noted that several concepts related to batches in the present application are explained here:
[0046] 1) Batch (Mini-batch):
[0047] In traditional model training, each training iteration processes a batch (mini-batch) of training samples. This is the basic unit of model training, usually containing multiple samples.
[0048] 2) Micro-batch (Micro-batch):
[0049] After introducing the pipeline parallel training technology, in order to better utilize parallel computing resources, each mini-batch is further divided into multiple micro-batches. Each micro-batch contains fewer samples, which allows for more fine-grained allocation of computing tasks. By dividing the mini-batch into multiple micro-batches, each micro-batch can be independently processed in parallel on multiple GPU servers, greatly improving the parallel efficiency of model training. For example, a mini-batch can be divided into 4 micro-batches, and in the pipeline, the forward propagation and backward propagation calculations of each micro-batch can be interleaved, reducing the waiting time of the GPU and speeding up the overall training speed.
[0050] 3) Number of Micro-batches:
[0051] Total number of micro-batches that each mini-batch is split into. For example, if a mini-batch is split into 4 micro-batches, then the micro-batch number is 4. The micro-batch number determines the granularity and complexity of parallel computation. Reasonable micro-batch number determination is crucial for balancing the efficiency of parallel computation and communication cost.
[0052] 4) Micro-batch Index:
[0053] Used to identify the order of each micro-batch in the entire mini-batch. If the index starts from 0, then the first micro-batch has an index of 0, the second micro-batch has an index of 1, and so on. The micro-batch index helps manage the flow and scheduling of parallel computation. For example, in pipeline parallelism, each GPU or server can determine when to start computing the next micro-batch and when to send and receive data based on the micro-batch index. In addition, the micro-batch index is also used to track and confirm the order of completion of communication operations, ensuring data accuracy and training consistency.
[0054] 5) Batch Group:
[0055] Batch group is a concept introduced in parallel model training to achieve more efficient data processing and resource allocation. A batch group usually contains one or more micro-batches designed as computational units that can be executed simultaneously or staggered to take advantage of parallel computing resources and reduce the impact of communication delay on computing efficiency. For example, performing M+N batch group operations means performing M micro-batch and N micro-batch operations. A batch group can contain forward propagation calculations of one micro-batch and backward propagation calculations of another micro-batch, through such combination, efficient operation of the computing process and balanced use of resources are achieved.
[0056] In this step, the pipeline stage is also involved, which refers to different stages in the model training process, including warm-up stage, stable stage, and cooling-down stage, each stage has different micro-batch processing strategies according to network status and computing requirements.
[0057] In the warm-up stage, the system can perform as many forward propagation calculations as possible, that is, increase the number of pipeline warm-up stage micro-batches as much as possible. By performing more forward propagation calculations in this stage, the system can start and warm up the communication links between the GPU servers, laying the foundation for efficient communication in the subsequent stable stage and cooling stage. The GPU servers can start to continuously receive data from the previous micro-batch without stalling due to waiting for reception and calculation. In this way, even while processing the current micro-batch calculation, data can be prepared for subsequent micro-batch calculations, improving the parallelism of communication. Moreover, more forward propagation calculations mean that communication operations are performed earlier, especially in high-latency network environments, and pre-sending and pre-receiving of data can effectively alleviate the impact of network latency. By occupying network bandwidth in advance, the waiting time for subsequent communication operations is reduced, improving overall communication efficiency.
[0058] It should be noted that the number of micro-batches processed in the warm-up stage can be a predetermined number, which is determined according to the number of GPUs in the system, the arrangement order of the target GPU (current GPU) in the multiple GPUs, and the total number of micro-batches of iterative training. Among them, the total number of micro-batches of iterative training refers to the number of all micro-batches that need to be processed in a complete training iteration. Among them, the arrangement order is determined according to the arrangement order of the model layer segments respectively driven by the multiple GPUs in the model. For example, the target model is divided into four layer segments, which are referred to as stage-0, stage-1, stage-2, and stage-3. Different layer segments are driven by different GPUs, such as GPU-0 driving stage-0, GPU-1 driving stage-1, GPU-2 driving stage-2, and GPU-3 driving stage-3. The sorting of the GPU responsible for the frontmost layer segment of the model is the first, and the sorting of the GPU responsible for the last layer segment of the model is the last.
[0059] Assuming there are P GPUs, the current GPU in the pipeline is arranged in rank, and the total number of micro batches in the training iteration is Q, the predetermined number can be (P-rank-1)*4+1. That is, the GPU arranged at the end can also receive the upstream forward propagation result corresponding to the first micro batch it needs in the warm-up stage, and each GPU in front can send several upstream forward propagation results corresponding to the micro batch in advance. It should be noted that if the total number of micro batches is less than P-rank-1, it cannot run more, that is, Q is the upper limit, and as much as possible, but not more than the limit of the number of batches, where 0≤rank≤P-1. Suppose: P=4, Q=10: GPU 0: min((4-0-1)*4+1,10)=10, GPU 1: min((4-1-1)*4+1,10)=9, etc. Therefore, each GPU runs 10, 9, etc. in the warm-up stage.
[0060] Among them, the stable stage is the most core part of the model training process, and the calculation and communication are executed in parallel to maximize the GPU utilization and training speed.
[0061] Among them, the cooling stage is the last stage of the model training, which processes the remaining micro batches and completes the communication operation to update the final model parameters.
[0062] In this step, after receiving the model training request, the system determines the number of micro batches that each of the three pipeline stages of warm-up, stabilization, and cooling should handle according to the target strategy. This strategy takes into account network environment factors and parallel computing needs, aiming to optimize resource allocation, reduce the negative impact of communication delay on calculation, and improve GPU computing power utilization. By accurately planning the number of micro batches in each stage, the system can more efficiently manage the parallel execution of calculation and communication, avoiding waste in resource allocation, especially in cases of high network latency or limited bandwidth. This method can significantly reduce the total training time.
[0063] Step S106, in a manner of calculation and communication parallel execution, in the order of execution of the pipeline stages, sequentially executes the calculation and communication operations corresponding to the multiple pipeline stages respectively, to obtain the current backward propagation results corresponding to the multiple micro batches respectively, wherein the corresponding calculation and communication operations include calculation operations for calculating the propagation results of the corresponding micro batches according to the number of micro batches of the corresponding stage, and communication operations for transmitting the propagation results of the corresponding micro batches.
[0064] In step S106 provided in the present application, in a manner of calculation and communication parallel execution, in the order of execution of the pipeline stages, sequentially executes the calculation and communication operations corresponding to the multiple pipeline stages respectively, to obtain the current backward propagation results corresponding to the multiple micro batches respectively.
[0065] wherein, the computing operations involve forward propagation and backward propagation computations for generating and updating the outputs and gradients of the model layer segment.
[0066] wherein, the communication operations involve sending and receiving of the micro-batch propagation results, including sending and receiving of the forward propagation results and the backward propagation results, to support the computing requirements of the stages in the pipeline.
[0067] wherein, the way of performing the computing and the communication in parallel is involved, wherein, the parallel performance can also be called overlapping performance, which can be understood as that the computing operations and the communication operations can be performed in parallel, such as can be performed in parallel by the following operations: the processor of the GPU controls a computing stream to start a computing process, and then returns, so that the computing stream is started and performs the computing operations, the processor controls a communication stream to start a communication process (such as a sending process or a receiving process), and then returns, so that the communication stream is started and performs the communication operations, at this time, in the GPU, the computing stream performs the corresponding computing operations, and the communication stream performs the corresponding communication operations, and the two are performed in parallel. Alternatively, different communication processes can also control different communication streams to do, so as to further improve the efficiency.
[0068] wherein, the current backward propagation results corresponding to the plurality of micro-batches are involved, which can also be called the gradients of the model layer segment, that is, after the backward propagation computation of each micro-batch is completed, a gradient is obtained. The gradient is obtained according to the gradients of all micro-batches. The gradient is obtained by obtaining the gradient, so as to realize the update of the model layer segment by gradient accumulation. The gradient accumulation can be understood as that the gradient values obtained by multiple times of computation are accumulated, and then the parameter update is performed at one time, so as to achieve the purpose of model training.
[0069] In this step, the system performs the corresponding computing and communication operations in the order of the warm-up stage, the stable stage, and the cooling stage in the way of performing the computing and the communication in parallel. In any stage, while performing the propagation result computation of the next micro-batch, the transmission of the propagation result of the previous micro-batch can be performed in parallel, so as to overlap the computing and the communication, to fully utilize the computing resources and the network bandwidth, to avoid too much invalid waiting, to reduce the time of the GPU waiting for the completion of the communication, and to significantly improve the efficiency and the speed of the model training.
[0070] Step S108: training the model layer segment according to the current backward propagation results corresponding to the plurality of micro-batches, to update the model parameters corresponding to the model layer segment, and obtaining a target model.
[0071] In step S108 provided in the present application, the model layer segment is trained according to the current backward propagation results corresponding to the plurality of micro-batches, to update the model parameters corresponding to the model layer segment, and a target model is obtained.
[0072] The model parameters are the basic values of the model layers, such as the weights and biases of the neural network, which are gradually optimized through the training process.
[0073] In this step, after the parallel execution of the parallel computing and communication, the current back propagation results corresponding to each micro batch are obtained. The system uses these back propagation results to train the model layer, updates the model parameters, until the training target is reached, and generates the target model. Through the above parallel computing and communication parallel execution method, the back propagation results obtained are more timely and accurate, which helps to quickly and effectively update the model parameters, accelerate the model convergence, and improve the training quality. In addition, this method can also effectively deal with the common data imbalance problem in large model training, and through the micro batch and group strategy optimization, the resource allocation is more reasonable, and the waste of computing resources is reduced.
[0074] Through the above steps S102-S108, a model training request is received, wherein the model training request carries a target strategy, and the target strategy is a training strategy for training an initial model. In response to the model training request, the number of micro batches processed by each pipeline stage is determined according to the target strategy, wherein the plurality of pipeline stages include: a warm-up stage, a stable stage, and a cooling stage. In a manner of parallel execution of computing and communication, the corresponding computing and communication operations of the plurality of pipeline stages are executed in sequence according to the execution order of the pipeline stages, to obtain current back propagation results corresponding to the plurality of micro batches. The corresponding computing and communication operations include a computing operation for calculating the propagation results of the corresponding micro batch according to the number of micro batches of the corresponding stage, and a communication operation for transmitting the propagation results of the corresponding micro batch. The model layer is trained according to the current back propagation results corresponding to the plurality of micro batches, to update the model parameters corresponding to the model layer, and to obtain the target model. The present application finely adjusts the number of micro batches in the pipeline warm-up, stable, and cooling stages of model training, and adopts the strategy of parallel execution of computing and communication, which effectively solves the problem of waste of computing resources caused by GPU waiting for communication in the traditional technology. Especially in a high-latency and low-bandwidth network environment, the communication process and the computing process are scheduled in parallel, which avoids the blocking of subsequent calculations caused by the delayed communication operation in serial execution, and at the same time, the parallel scheduling also significantly improves the GPU utilization rate, reduces the waiting time in the training iteration, and further speeds up the model training process, improves the overall training efficiency, and solves the technical problem of low model training efficiency in related technologies.
[0075] Before introducing the following embodiments, the forward propagation result, the back propagation result, the current forward propagation result, the upstream forward propagation result, the current back propagation result, and the upstream back propagation result are introduced.
[0076] Wherein, the forward propagation result is involved, in model training, the forward propagation refers to the transmission process of data from the input layer to the output layer. The output of each layer (i.e. the input of the next layer) can be referred to as the forward propagation result. The forward propagation result is used to calculate the predicted output of the model under the current parameter setting, and these predicted outputs are compared with the true labels to generate the value of the loss function, which in turn guides the adjustment of the model parameters.
[0077] Wherein, the current forward propagation result is involved, the current forward propagation result refers to the current forward propagation output result obtained by the target GPU based on the upstream forward propagation result when processing a specific micro-batch of data.
[0078] Wherein, the upstream forward propagation result is involved, the upstream forward propagation result refers to the input of each layer, i.e. the output of the previous layer, in the pipeline parallel architecture, which can also be understood as the forward propagation result of the corresponding output of the upstream GPU. That is, the upstream forward propagation result refers to the forward propagation output result received by the target GPU from the previous GPU, which is used for forward propagation calculation of the current layer segment.
[0079] Wherein, the backward propagation result is involved, backward propagation is another stage in model training, which calculates the gradient of the loss function with respect to the parameters of each layer in reverse from the output layer to the input layer. The backward propagation result contains gradient information about the model parameters, which is used to update the model parameters to reduce prediction errors and optimize model performance.
[0080] Wherein, the current backward propagation result is involved, the current backward propagation result specifically refers to the current backward propagation result obtained by the target GPU based on the upstream backward propagation result when processing a specific micro-batch of training data.
[0081] Wherein, the upstream backward propagation result is involved, the upstream backward propagation result refers to the backward propagation result of the corresponding output of the upstream GPU in the pipeline parallel training. That is, the upstream backward propagation result refers to the gradient information received by the target GPU from the next stage of the pipeline, which is used for backward propagation calculation of the current layer segment.
[0082] The optional embodiments of the present application are introduced as follows:
[0083] As an optional embodiment, the execution of the calculation and communication operation corresponding to the warm-up stage comprises: starting and executing a first receiving operation, wherein the first receiving operation is used to receive an upstream forward propagation result corresponding to an Nth micro-batch, and N is set to an initial value; executing a warm-up stage loop operation: determining a first micro-batch order number currently to be processed by the target GPU; starting and executing a corresponding first warm-up stage operation according to a sequence number relationship between the first micro-batch order number and a micro-batch order number to be processed by the iterative training, wherein the first warm-up stage operation comprises at least one of the following: a second receiving operation, a first calculation operation, the second receiving operation is used to receive an upstream forward propagation result corresponding to an (N+1)th micro-batch, and the first calculation operation is used to determine a current forward propagation result corresponding to the Nth micro-batch according to the upstream forward propagation result corresponding to the Nth micro-batch; determining a second micro-batch order number currently to be processed by the target GPU; starting and executing a corresponding second warm-up stage operation according to a sequence number relationship between the first micro-batch order number and a micro-batch order number to be processed by the warm-up stage, wherein the second warm-up stage operation comprises at least one of the following: a first sending operation, a third receiving operation, a counting operation, an entering stable stage operation, the first sending operation is used to send the current forward propagation result corresponding to the Nth micro-batch, the third receiving operation is used to receive an upstream backward propagation result corresponding to an Mth micro-batch, M is set to an initial value, the counting operation is used to control N to be equal to N+1 in the case that the second receiving operation is executed, and the entering stable stage operation is used to control N to be equal to N+1 and control the target GPU to enter a stable stage and execute a calculation and communication operation corresponding to the stable stage, the micro-batch order number to be processed by the warm-up stage is determined according to a number of micro-batches processed by a plurality of pipeline stages respectively; and the warm-up stage loop operation is executed circularly until the target GPU enters the stable stage and executes the calculation and communication operation corresponding to the stable stage.
[0084] In this embodiment, how to execute the calculation and communication operation corresponding to the warm-up stage is illustrated.
[0085] In this embodiment, how to execute the calculation and communication operation corresponding to the warm-up stage is illustrated.
[0086] In this embodiment, how to execute the calculation and communication operation corresponding to the warm-up stage is illustrated.
[0087] wherein the micro batch order number to be processed in the iteration training is referred to, and the micro batch order number to be processed in the iteration training refers to the serial number of all micro batches to be processed in turn in one complete iteration training. For example, there are 12 micro batches in one iteration, and the serial numbers are 0, 1, 2, …, 11.
[0088] wherein the micro batch order number to be processed in the warm-up stage is referred to, and the micro batch order number to be processed in the warm-up stage refers to the serial number of the micro batch to be processed in the warm-up stage. It is used to define which micro batches need to be processed in the warm-up stage. Following the above example, if the warm-up stage only needs to process the first 3 micro batches, the serial numbers are 0, 1, 2.
[0089] wherein the micro batch number is referred to, and the micro batch number refers to the number of micro batches to be processed in each pipeline stage. The number of micro batches to be processed in each pipeline stage determines the number of plans to be processed in different stages. For example, for the warm-up stage, following the above example, assuming that the number of micro batches to be processed in the warm-up stage is determined to be 3, and the serial number starts from 0, the last micro batch serial number of the warm-up stage is 2.
[0090] wherein the second receiving operation is referred to, and the second receiving operation is used to receive the upstream forward propagation result of the next (N+1 micro batch) in the warm-up stage, for subsequent calculation operation.
[0091] wherein the first calculation operation is referred to, and the first calculation operation is used to calculate the forward propagation result of the current micro batch according to the received forward propagation result.
[0092] wherein the first sending operation is referred to, and the first sending operation sends the current forward propagation result corresponding to the N micro batch, and transmits it to the downstream GPU.
[0093] wherein the third receiving operation is referred to, and the third receiving operation is used to receive the upstream backward propagation result of the M micro batch in the later stage of the warm-up stage, for the calculation preparation of the backward propagation. M is also provided with an initial value, which can be 1, and is not limited here, and can be adaptively set according to the actual application and scene.
[0094] wherein the counting operation is referred to, and the counting operation is used to manage and control the micro batch processing flow, to ensure that N is incremented after the completion of the second receiving operation, so as to continue the processing of the next micro batch, and to execute the loop operation to achieve the corresponding condition.
[0095] In this step, the first receiving operation is started and executed, indicating the start of the warm-up phase, and the target GPU first receives the forward propagation result of the first micro-batch, starting the calculation process of model training. Receiving the forward propagation result in time avoids waiting time and improves the processing efficiency of the GPU and the overall speed of model training. Then the warm-up phase loop operation is executed: determining the first micro-batch order number currently to be processed by the target GPU. According to the order number relationship between the first micro-batch order number and the micro-batch order number to be processed in the iterative training, the corresponding first warm-up phase operation is started and executed. And according to the order number relationship between the first micro-batch order number and the micro-batch order number to be processed in the warm-up phase, the corresponding second warm-up phase operation is started and executed. Through the above steps, the micro-batch order number currently processed is dynamically determined, and the corresponding calculation or communication operation is executed accordingly. The loop operation ensures the seamless connection of calculation and communication, and reduces waiting time and communication congestion through continuous reception and transmission of forward and backward propagation results. By determining the subsequent operation according to the order number relationship, the boundary of processing in different phases can be clearly defined, ensuring orderly processing while improving the continuity and efficiency of calculation through dynamic scheduling, ensuring that the model layer calculation can quickly advance and not be stagnant due to communication delay. The warm-up phase loop operation is executed in a loop until the target GPU enters the stable phase and executes the calculation and communication operations corresponding to the stable phase. Through the loop operation in the warm-up phase, the calculation process is quickly started and initialized, and the execution time of calculation and communication is optimally allocated by the GPU to minimize waiting time. During the loop operation, the GPU can pre-receive the forward propagation result of the next micro-batch while performing the calculation of the current micro-batch, thereby realizing efficient overlap of calculation and communication and reducing the overall training time.
[0096] Through the above steps, especially the dynamic determination of the first micro-batch order number in the loop operation of the warm-up phase, and the timely execution of receiving, sending operations and intermediate calculation tasks, the problem of low efficiency caused by waiting for communication in traditional model training can be effectively solved. The dynamic scheduling strategy of the warm-up phase helps to quickly start the calculation process of the GPU and reduce the initial waiting time. The overlapping execution of calculation and communication continues to play a role in the subsequent stable phase, ensuring that the model training can run efficiently even in poor network conditions, and finally realizing the rapid update of model parameters and the significant improvement of model training efficiency.
[0097] As an optional embodiment, according to the sequence number relationship between the first micro-batch sequence number and the micro-batch sequence number to be processed in the iteration training, the corresponding first warm-up stage operation is started and executed, including: in the case that the first micro-batch sequence number is not the last micro-batch sequence number of the iteration training, starting and executing the second receiving operation and the first calculation operation, wherein the second receiving operation and the first calculation operation are executed in parallel; in the case that the first micro-batch sequence number is the last micro-batch sequence number of the iteration training, starting and executing the first calculation operation.
[0098] In this embodiment, how to start and execute the corresponding first warm-up stage operation according to the sequence number relationship between the first micro-batch sequence number and the micro-batch sequence number to be processed in the iteration training is illustrated.
[0099] Among them, the last micro-batch sequence number of the warm-up stage is involved, which is the micro-batch sequence number of the end of the warm-up stage. When the currently processed micro-batch reaches this number, the warm-up stage operation is completed.
[0100] Among them, the second receiving operation and the first calculation operation are executed in parallel, that is, at this time, two processes can be started, one for controlling the communication flow to start the corresponding process, thereby starting and executing the second receiving operation, and one for controlling the calculation flow to start the corresponding process, thereby starting and executing the second calculation operation. They can be executed in parallel in the case of starting time alignment or misalignment in order to improve the effect of communication and calculation.
[0101] In this step, two cases are illustrated, one is in the case that the first micro-batch sequence number is not the last micro-batch sequence number of the iteration training, starting and executing the second receiving operation and the first calculation operation. That is, as long as the current micro-batch sequence number has not reached the last sequence number of the warm-up stage, the GPU will start the second receiving operation (receive the forward propagation result of the next micro-batch) and execute the first calculation operation (calculate the forward propagation result) of the current micro-batch at the same time. This way of parallel execution of receiving and calculation operations greatly improves the utilization rate of GPU and the initialization speed of model training. By receiving the data of the next micro-batch in advance, the GPU can start the calculation of the next micro-batch immediately after completing the current calculation, reducing the waiting time.
[0102] The other is in the case that the first micro-batch sequence number is the last micro-batch sequence number of the iteration training, starting and executing the first calculation operation. That is, when the GPU recognizes that the current processing is the last micro-batch of the warm-up stage, it only executes the first calculation operation to complete the forward propagation calculation of the current micro-batch, and no longer starts additional receiving operations, preparing for the transition to the stable stage. This strategy prevents unnecessary data transmission and resource consumption, ensures a perfect conclusion of the warm-up stage, and reserves sufficient resources and time for the upcoming stable stage operation.
[0103] In this step, the operation design of the warm-up phase aims to quickly start the calculation process of model training by carefully arranging the processing of the first micro-batch order number, while performing the second receiving operation and the first calculation operation, significantly reducing the idle waiting time of the GPU and improving the calculation efficiency. In particular, the special processing of the last micro-batch of the warm-up phase ensures smooth switching of calculation and communication. This strategy not only speeds up the initialization process of model training, but also optimizes the utilization of computing resources, which is particularly beneficial for large-scale and computationally intensive model training. In this way, the utilization of GPU computing power can be effectively improved, the training time can be reduced, and the model can reach the expected performance level more quickly and efficiently.
[0104] As an optional embodiment, according to the order number relationship between the first micro-batch order number and the micro-batch order number to be processed in the warm-up phase, the corresponding second warm-up phase operation is started and executed, including: in the case that the second micro-batch order number is the last micro-batch order number of the warm-up phase, controlling N = N + 1, and controlling the target GPU to enter the stable phase to perform the corresponding calculation and communication operations of the stable phase; in the case that the first micro-batch order number is neither the last micro-batch order number nor the second last micro-batch order number of the warm-up phase, starting and executing the first sending operation and the counting operation; in the case that the first micro-batch order number is the second last micro-batch order number of the warm-up phase, starting and executing the first sending operation, the third receiving operation and the counting operation, wherein the first sending operation and the third receiving operation are executed in parallel.
[0105] In this embodiment, how to start and execute the corresponding second warm-up phase operation according to the order number relationship between the first micro-batch order number and the micro-batch order number to be processed in the warm-up phase is explained.
[0106] Among them, the first sending operation and the third receiving operation are executed in parallel, that is, two processes can be started at this time, one for controlling the communication stream to start the corresponding process to control the start and execution of the first sending operation, and the other for controlling the communication stream to start the corresponding process to control the start and execution of the third receiving operation, which can be executed in parallel in the case of starting time alignment or not, in order to improve the effect of communication and calculation. It should be noted that when controlling the communication stream to start and execute the corresponding operation, different communication streams can be used to start different processes to further improve the efficiency.
[0107] In this step, three cases are explained. One is that when the first micro batch order number is the last micro batch order number of the warm-up stage, the control N = N + 1 (indicating that the value of N + 1 is assigned to N) and the target GPU enters the stable stage. That is, when the target GPU processes the last micro batch of the warm-up stage, it updates the value of N (N = N + 1) and immediately enters the stable stage to start performing the calculation and communication operations corresponding to the stable stage. This is the transition step between the warm-up stage and the stable stage. By controlling the update of N, the GPU can mark the currently processed micro batch and prepare for entering the stable stage.
[0108] Another is that when the first micro batch order number is not the last micro batch order number of the warm-up stage, nor the second last micro batch order number, the first sending operation and the counting operation are started and performed. For all micro batches in the warm-up stage except the last and the second last batches, the GPU performs the first sending operation (sending the forward propagation result of the current micro batch) and the counting operation (updating the processed micro batch order number under the condition that the second receiving operation is performed), to maintain the continuous transmission of data and the orderly performance of operations. By performing the sending operation and the counting operation in parallel, the GPU can send data in advance, reduce the waiting time, and update the processing state at the same time, keeping the calculation process efficient.
[0109] The third is that when the first micro batch order number is the second last micro batch order number of the warm-up stage, the first sending operation, the third receiving operation and the counting operation are started and performed. When processing the second last micro batch, in addition to performing the first sending operation and the counting operation, the GPU also adds the third receiving operation for receiving the backward propagation result of the next micro batch, to prepare for the overlapping execution of the last calculation and communication of the warm-up stage. That is, by performing the additional receiving operation in the second last micro batch of the warm-up stage, the GPU can realize the overlap of calculation and communication in the last micro batch, ensuring the seamless connection of calculation and communication operations and reducing the idle waiting time of the GPU. This operation plays a role in connecting the past and the future at the end of the warm-up stage, not only completing the timely sending of forward propagation data, but also providing the necessary data input for the overlapping execution of the last micro batch calculation and communication, thereby realizing the efficient transition from the warm-up stage to the stable stage.
[0110] It should be noted that in the above judgment process, it can be first judged whether it is the last micro batch order number of the warm-up stage, and if not, it can be further judged whether it is the second last micro batch order number of the warm-up stage, to ensure the orderliness of the judgment.
[0111] In this step, through the operation of the second warm-up stage, including precise control of micro-batch order number, timely data sending and receiving, and timely stage transition, model training can smoothly transition from the initialization stage of the GPU to efficient calculation and communication in the stable stage. This mechanism not only improves the utilization of GPU computing power and reduces waiting time, but also ensures the continuity of the calculation process and the accuracy of data transmission, providing a solid foundation for the subsequent stages of model training. In addition, by performing an additional receiving operation on the second-to-last micro-batch in the warm-up stage, the overlap execution of calculation and communication is further optimized, realizing seamless connection from the initial warm-up to the main calculation stage, significantly improving the overall efficiency and performance of model training.
[0112] As an optional embodiment, the calculation and communication operations corresponding to the stable stage are executed, including: starting and executing a second sending operation and a fourth receiving operation, wherein the second sending operation is used to send the current forward propagation result corresponding to the N-1th micro-batch, and the fourth receiving operation is used to receive the upstream backward propagation result corresponding to the M+1th micro-batch; executing a stable stage loop operation: in the case that the fifth receiving operation and the sixth receiving operation are both completed, determining the batch group currently to be processed by the target GPU, wherein the fifth receiving operation is used to receive the upstream forward propagation result corresponding to the Nth micro-batch, and the sixth receiving operation is used to receive the upstream backward propagation result corresponding to the Mth micro-batch; according to the group order relationship between the batch group and the batch group to be processed in the stable stage, starting and executing corresponding stable transceiving operations, wherein the stable transceiving operations include at least one of the following: a third sending operation, a seventh receiving operation, the third sending operation is used to send the current backward propagation result corresponding to the M-1th micro-batch, and the seventh receiving operation is used to receive the upstream forward propagation result corresponding to the N+1th micro-batch, and the batch group to be processed in the stable stage is determined according to the number of micro-batches processed by each pipeline stage; starting and executing a second calculation operation, wherein the second calculation operation is used to perform propagation result calculation of the N+M batch group, including determining the current forward propagation result corresponding to the Nth micro-batch according to the upstream forward propagation result corresponding to the Nth micro-batch, and determining the current backward propagation result corresponding to the Mth micro-batch according to the upstream backward propagation result corresponding to the Mth micro-batch; determining whether the batch group currently to be processed by the target GPU is the last batch group of the stable stage; in the case that it is not the last batch group, controlling N=N+1, M=M+1, and executing the stable stage loop operation in a loop until it is the last batch group, controlling M=M+1, and controlling the target GPU to enter the cooling stage to execute the cooling stage operation corresponding to the cooling stage.
[0113] In this embodiment, how to execute the calculation and communication operations corresponding to the stable stage is explained.
[0114] The second sending operation refers to sending the current forward propagation result corresponding to the N-1th micro batch, to ensure that the downstream GPU receives and starts the calculation process of the corresponding batch.
[0115] The fourth receiving operation refers to receiving the upstream backward propagation result corresponding to the M+1th micro batch, to provide necessary data for backward propagation calculation.
[0116] The stable stage loop operation is used to loop the operations required in the stable stage until the cooling stage is entered, to realize continuous and efficient calculation and communication.
[0117] The fifth receiving operation refers to receiving the upstream forward propagation result of the Nth micro batch, to be used for calculating the forward propagation output of the current layer.
[0118] The sixth receiving operation refers to receiving the upstream backward propagation result of the Mth micro batch, to prepare for backward propagation calculation.
[0119] The batch group to be processed in the stable stage includes a subset of consecutive micro batches allocated to parallel processing in the stable stage. It defines which batch groups need to be processed in the stable stage, and different batch groups can contain forward propagation of one micro batch and backward propagation calculation of another micro batch, to realize overlapping of calculation and communication.
[0120] The third sending operation refers to sending the current backward propagation result of the M-1th micro batch to the upstream GPU after calculation is completed, to support its backward propagation calculation.
[0121] The seventh receiving operation refers to receiving the upstream forward propagation result of the N+1th micro batch, to prepare for calculation of the next batch group.
[0122] The second calculation operation is used to perform propagation result calculation of the N+M batch group, including generation of the current forward propagation result and the current backward propagation result. It should be noted that the second calculation operation can be executed in parallel with the stable sending and receiving operation described above. For example, when the stable sending and receiving operation includes the seventh receiving operation, the seventh receiving operation and the second calculation operation are executed in parallel, that is, two processes can be started at this time, one for controlling the communication flow to start the corresponding process, thereby starting and executing the seventh receiving operation, and one for controlling the calculation flow to start the corresponding process, thereby starting and executing the second calculation operation. They can be executed in parallel with or without time alignment to improve the effect of communication and calculation. The case of the stable sending and receiving operation including other operations is not described here.
[0123] wherein batch groups are involved, which contain forward propagation of one micro batch and backward propagation of another micro batch computation to enable overlap of computation and communication.
[0124] In this step, the second sending operation and the fourth receiving operation are started and executed, which ensures the data flow between the upstream and downstream, helps to maintain the continuity of the stable stage and the overlapping execution of calculation and communication. By sending the current forward propagation result in time and receiving the next backward propagation result, seamless data transfer is achieved, greatly reducing the waiting time of the GPU and improving the efficiency of calculation and communication. In the case that the fifth receiving operation and the sixth receiving operation are both completed, the batch group that the target GPU currently needs to process is determined, the prerequisite for calculation start is clear, and the ordered execution of calculation and communication operations is ensured. According to the group order relationship between the batch group and the batch group to be processed in the stable stage, the corresponding stable transceiving operation is started and executed, and through the group order relationship between the batch group and the stable stage micro-batch group, the system can accurately locate the global batch order number that should enter the stable stage at present, to start and execute the corresponding calculation and communication operation, including forward propagation, backward propagation calculation and related data sending and receiving. This dynamic scheduling strategy based on batch group can flexibly adapt to different micro-batch processing needs, ensure the efficient overlap of calculation and communication, especially in a multi-GPU environment, which helps to maintain the continuity and stability of the calculation process, further improving the speed of model training. The second calculation operation is started and executed to perform the propagation result calculation of the batch group, including the forward propagation calculation and the backward propagation calculation of the micro-batch. By integrating forward and backward propagation calculation, not only the calculation process of a single GPU is optimized, but also the efficient data exchange between the upper and lower GPUs is promoted, reducing the overall training time and improving the efficiency and quality of model training. Finally, it is determined whether the batch group that the target GPU currently needs to process is the last batch group of the stable stage, to judge whether the end of the stable stage is reached, to decide the next action. This judgment mechanism provides a clear signal for smooth transition to the cooling stage, avoiding unnecessary calculation or communication operations, ensuring the integrity and efficiency of model training. Control N=N+1, M=M+1, and execute the stable stage loop operation in a loop. By dynamically updating the micro-batch order number, the calculation and communication cycle of the stable stage is maintained. By updating the micro-batch order number, the continuity of the calculation process is ensured, and the GPU is provided with clear operation instructions, ensuring the smooth progress of model training. Until the last batch group, control M=M+1, and control the target GPU to enter the cooling stage, that is, when all batch groups of the stable stage are processed, the GPU enters the cooling stage to complete the remaining data transmission and processing. The entry of the cooling stage marks an important turning point in model training, from efficient parallelism to completion stage processing. This process helps to ensure that all micro-batch data is properly processed and model parameters are fully updated, providing a guarantee for the output of the final model.
[0125] In this step, through the calculation and communication operation design of the stable stage, precise control of data transmission and reception, and dynamic scheduling of micro-batches, efficient overlap of calculation and communication is achieved, overcoming the low efficiency problem caused by data transmission delay in traditional model training. The operation of this stage not only greatly improves the utilization efficiency of GPU and reduces the training time, but also helps the model to converge faster and improves the overall quality of model training. In addition, the cycle mechanism of the stable stage ensures the stability and continuity of model training, and can adapt well and efficiently in both high-performance and low-performance network environments.
[0126] As an optional embodiment, according to the group order relationship between the batch group and the batch group to be processed in the stable stage, the corresponding stable transceiving operation is started and executed, including: in the case that the batch group is the first batch group of the stable stage, the seventh receiving operation is started and executed; in the case that the batch group is not the first batch group of the stable stage and is not the last batch group of the stable stage, the seventh receiving operation and the third sending operation are started and executed, wherein the seventh receiving operation and the third sending operation are executed in parallel; in the case that the batch group is the last batch group of the stable stage, the third sending operation is started and executed.
[0127] In this embodiment, how to start and execute the corresponding stable transceiving operation according to the group order relationship between the batch group and the batch group to be processed in the stable stage is explained.
[0128] Among them, the first batch group of the stable stage is referred to, which refers to the batch group processed at the beginning of the stable stage, which needs special processing to start the calculation and communication overlap process.
[0129] Among them, the last batch group of the stable stage is referred to, which refers to the batch group processed at the end of the stable stage, which has different operation modes from the intermediate batch group and the first batch group, and is used to smoothly transition to the cooling stage.
[0130] Among them, the third sending operation and the seventh receiving operation are executed in parallel, that is, two processes can be started at this time, one is used to control the communication stream to start the corresponding process, thereby controlling the seventh receiving operation to be started and executed, and the other is used to control the communication stream to start the corresponding process, thereby controlling the third sending operation to be started and executed, which can be executed in parallel in the case of starting time alignment or not, in order to improve the effect of communication and calculation. It should be noted that when controlling the communication stream to start and execute the corresponding operation, different communication streams can be used to start different processes to further improve the efficiency.
[0131] In this step, three cases are described. One is when the batch group is the first batch group of the warm-up phase, the seventh receiving operation is started and executed. That is, when the GPU starts processing the first batch group of the warm-up phase, it first executes the seventh receiving operation to receive the forward propagation results of the next batch (N+1 micro-batch). This is because the calculation and communication overlap execution of the first batch group needs the forward propagation results upstream as a starting point. By receiving data in advance, the GPU can immediately start the overlap execution of calculation and communication, reducing the waiting time of the initialization phase, improving the continuity of subsequent calculation, and ensuring that the warm-up phase achieves high efficiency from the beginning.
[0132] Another is when the batch group is neither the first batch group of the warm-up phase nor the last batch group of the warm-up phase, the seventh receiving operation and the third sending operation are started and executed. That is, for any non-first and non-last batch group in the warm-up phase, the GPU simultaneously starts the seventh receiving operation (receiving the forward propagation results of N+1 micro-batch) and the third sending operation (sending the backward propagation results of M-1 micro-batch) to maintain the continuous overlap of calculation and communication operations. This operation ensures the efficient cooperation of calculation and communication, reduces the waiting time. By executing receiving and sending operations in parallel, the GPU can continuously process micro-batches, maintain the high-speed operation of the calculation process, and significantly improve the overall efficiency of model training.
[0133] The third is when the batch group is the last batch group of the warm-up phase, the third sending operation is started and executed. When the GPU processes the last batch group of the warm-up phase, it only executes the third sending operation to send the backward propagation results to the upstream GPU, preparing for the entry into the cooling phase. By only executing the sending operation, the system can ensure the complete execution of the calculation and communication of the last batch group, while smoothly transitioning to the cooling phase, avoiding unnecessary receiving operations, reducing resource waste, and ensuring the integrity of model training. The processing of the last batch group is a key turning point in the entire model training process. By only executing the third sending operation, the system not only completes the calculation and communication tasks of the batch group, but also provides a clear signal for the start of the cooling phase, ensuring a seamless transition from the warm-up phase to the cooling phase.
[0134] In this step, the operation design of the stable stage is the cornerstone of optimizing the entire model training process. By carefully designing the execution process, the GPU can achieve efficient overlap of computation and communication in mini-batch processing, significantly improving computational efficiency and resource utilization. This mechanism not only applies to high-performance network environments, but even in poor network conditions, it can maintain high-speed operation of model training through parallel execution of operations and precise timing processing, ensuring maximum utilization of computing resources. In addition, special processing of the first and last batch groups in the stable stage ensures smooth start and end of the computing process, avoiding confusion in computation and communication operations, providing strong protection for the efficiency and quality of model training. In this way, model training can achieve the expected training effect with minimal time delay and resource waste, promoting the development and application of large-scale parallel model training technology.
[0135] As an optional embodiment, the execution of the calculation and communication operation corresponding to the cooling stage includes: executing a cooling stage loop operation: determining the second micro-batch order number currently to be processed by the target GPU; according to the order number relationship between the second micro-batch order number and the micro-batch order number to be processed in the cooling stage, starting and executing the corresponding cooling transceiver operation, wherein the cooling transceiver operation includes at least one of the following: the fourth sending operation, the eighth receiving operation, the eighth receiving operation is used to receive the current back propagation result corresponding to the M+1 micro-batch, the fourth sending operation is used to send the upstream forward propagation result corresponding to the N micro-batch, and the micro-batch order number to be processed in the cooling stage is determined according to the number of micro-batches processed by each pipeline stage; under the condition that the ninth receiving operation is completed, starting and executing the fifth sending operation, wherein the ninth receiving operation is used to receive the upstream back propagation result corresponding to the M micro-batch, and the fifth sending operation is used to send the current back propagation result corresponding to the M-1 micro-batch; starting and executing the third calculation operation, wherein the third calculation operation is used to determine the current back propagation result corresponding to the M micro-batch according to the upstream back propagation result corresponding to the M micro-batch; determining whether the micro-batch order number currently to be processed by the target GPU is the last micro-batch order number of the cooling stage; in the case of not being the last micro-batch order number, controlling M=M+1, and executing the cooling stage loop operation in a loop until the last micro-batch order number is reached, and starting and executing the sixth sending operation, wherein the sixth sending operation is used to send the current back propagation result corresponding to the M micro-batch under the condition that the fifth sending operation is completed.
[0136] In this embodiment, how to execute the calculation and communication operation corresponding to the cooling stage is explained.
[0137] Among them, the cooling stage loop operation is involved, which represents the core process of the cooling stage and is used to dynamically schedule and execute calculation and communication tasks until all micro-batches are processed.
[0138] wherein the second micro batch order number is involved, the second micro batch order number is the micro batch order number being processed in the cooling phase, to perform subsequent operations according to the current order number.
[0139] wherein the micro batch order number to be processed in the cooling phase is involved, the micro batch order number to be processed in the cooling phase refers to the number of the micro batch to be processed in the cooling phase. It is defined which micro batches need to be processed in the cooling phase.
[0140] wherein the cooling transceiving operation is involved, the cooling transceiving operation refers to the receiving and sending operations performed for a specific micro batch order number in the cooling phase, including at least one of the fourth sending operation and the eighth receiving operation.
[0141] wherein the eighth receiving operation is involved, the eighth receiving operation represents receiving the current backward propagation result corresponding to the M+1th micro batch, for completing the last stage of model layer segment calculation.
[0142] wherein the fourth sending operation is involved, the fourth sending operation represents sending the upstream forward propagation result corresponding to the Nth micro batch, so that the downstream GPU receives and processes it.
[0143] wherein the ninth receiving operation is involved, the ninth receiving operation represents receiving the upstream backward propagation result corresponding to the Mth micro batch, to provide the necessary data input for the backward propagation calculation of the current layer segment.
[0144] wherein the fifth sending operation is involved, the fifth sending operation is used to send the current backward propagation result corresponding to the M-1th micro batch after the completion of the ninth receiving operation, so that the downstream GPU receives and processes it.
[0145] wherein the third calculation operation is involved, the third calculation operation is used to calculate the current backward propagation result corresponding to the Mth micro batch based on the received upstream backward propagation result. It should be noted that the third calculation operation and the fifth sending operation can be executed in parallel, that is, at this time, two processes can be started, one is used to control the communication flow to start the corresponding process, thereby starting and executing the fifth sending operation, and the other is used to control the calculation flow to start the corresponding process, thereby starting and executing the third calculation operation, which can be executed in parallel in the case of starting time alignment or misalignment, in order to improve the effect of communication and calculation.
[0146] wherein the sixth sending operation is involved, the sixth sending operation represents sending the current backward propagation result corresponding to the Mth micro batch in the last micro batch of the cooling phase, to ensure that all data is correctly processed, and the model training process is completed.
[0147] In this step, the cooling phase cycle operation is performed, and the second micro-batch order number currently to be processed by the target GPU is determined. The execution of the cooling phase cycle operation ensures that the system processes each batch in an orderly manner, avoiding confusion and omission of data processing. According to the order relationship between the second micro-batch order number and the micro-batch order number to be processed in the cooling phase, it can be determined when the GPU should receive data, when it should perform calculations, and what operation should be performed next, which helps to process in an orderly manner. By using the order number to determine subsequent operations, omission and errors can be avoided, which can cause the model training to fail. Through the cycle operation, the system can process data batch by batch and accurately, ensuring the integrity of the model training. After the ninth receiving operation is completed, the fifth sending operation is started and executed. By controlling the timing of sending the backpropagation result, it ensures that the downstream GPU can start its calculation at the correct time point, avoiding unnecessary waiting. At the same time, this also guarantees the correctness of the model parameter update, because the backpropagation result must be received before the calculation layer is completed. The third calculation operation is started and executed again, which is the core of the model parameter update in the cooling phase. Based on the received backpropagation result, it calculates to ensure that the final optimization of the model during training is achieved, to ensure the accuracy and training efficiency of the model. Finally, it is determined whether the micro-batch order number currently to be processed by the target GPU is the last micro-batch order number of the cooling phase. This judgment step is crucial for controlling the timing of the end of model training. Only when it is determined that all micro-batches have been processed, the system can safely enter the last step of the cooling phase, i.e., the execution of the sixth sending operation, to complete the model training. The orderly processing of micro-batches is the key to ensuring training efficiency and resource utilization. In the cooling phase, by judging whether the last micro-batch has been reached, the system can accurately control the end point of training, avoiding excessive calculation or resource waste.
[0148] In this step, through the design of the cooling phase, especially the fine control of the cooling phase cycle operation and the efficient execution of the cooling receiving and sending operations, it plays a crucial role in ensuring the integrity of the model training process, improving training efficiency, and reducing resource waste. By maintaining the orderly scheduling of calculations and communications in the cooling phase, even in the last stage of model training, the accuracy of data processing can be maintained, and the final optimization of model parameters can be ensured.
[0149] As an optional embodiment, according to the sequence number relationship between the second micro-batch sequence number and the micro-batch sequence number to be processed in the cooling stage, the corresponding cooling transceiving operation is started and executed, including: in the case that the second micro-batch sequence number is the first micro-batch sequence number of the cooling stage, the fourth sending operation and the eighth receiving operation are started and executed, wherein the fourth sending operation and the eighth receiving operation are executed in parallel; in the case that the second micro-batch sequence number is neither the first micro-batch sequence number nor the last micro-batch sequence number of the cooling stage, the eighth receiving operation is started and executed.
[0150] In this embodiment, how to start and execute the corresponding cooling transceiving operation according to the sequence number relationship between the second micro-batch sequence number and the micro-batch sequence number to be processed in the cooling stage is illustrated.
[0151] Among them, the first micro-batch sequence number of the cooling stage is involved, which represents the sequence number of the first micro-batch processed at the beginning of the cooling stage, marking the beginning of the transition from the stable stage to the cooling stage.
[0152] Among them, the last micro-batch sequence number of the cooling stage is involved, which represents the sequence number of the last micro-batch processed at the end of the cooling stage, indicating that the entire calculation and communication operation of model training is about to be completed.
[0153] Among them, the fourth sending operation and the eighth receiving operation are executed in parallel, that is, two processes can be started at this time, one is used to control the communication flow to start the corresponding process, thereby controlling the start and execution of the eighth receiving operation, and the other is used to control the communication flow to start the corresponding process, thereby controlling the start and execution of the fourth sending operation. They can be executed in parallel in the case of starting time alignment or misalignment in order to improve the effect of communication and calculation. It should be noted that when controlling the communication flow to start and execute the corresponding operation, different communication flows can be used to start different processes in order to further improve the efficiency.
[0154] In this step, two cases are described. One is when the second micro-batch order number is the first micro-batch order number of the cooling phase, the fourth sending operation and the eighth receiving operation are started and executed. That is, when the GPU starts processing the first micro-batch of the cooling phase, it simultaneously performs the fourth sending operation (sending the forward propagation result of the current micro-batch) and the eighth receiving operation (receiving the backward propagation result of the next micro-batch), which is the initial signal of the cooling phase, marking the beginning of the end of the model training. By executing the sending and receiving operations immediately, the GPU ensures the continuous transfer of data, even in the final stage of training, to maintain efficient computation and communication, providing the necessary data for the final adjustment of model parameters. Not only can the forward and backward propagation results be processed in time, but also the end of the entire model training process can be laid the foundation to ensure seamless connection in the last stage of computation and communication, improving the continuity of data processing and the integrity of the training process.
[0155] The other is when the second micro-batch order number is neither the first micro-batch order number of the cooling phase nor the last micro-batch order number, the eighth receiving operation is started and executed. In the middle of the cooling phase micro-batch processing, the GPU mainly focuses on the eighth receiving operation for receiving the backward propagation result. At this time, the sending operation has been completed, and the main task is to collect the necessary calculation feedback data for the update of model parameters. By focusing on receiving the backward propagation result, the GPU can ensure that all necessary data is correctly received and processed to help the model parameters to be calibrated, ensuring the accuracy of the training result and the integrity of the training process. In the middle of the cooling phase, the calculation focus of the GPU gradually shifts from forward propagation to backward propagation, and by executing the eighth receiving operation, the backward propagation result from the upstream GPU is collected, which helps the final adjustment of the model parameters and provides data support for the end of the model training. This mechanism ensures the continuity of the computation process, even in the final stage of training, to maintain efficient data processing and contribute to the optimization of model quality.
[0156] In this step, the operation design of the cooling phase aims to ensure that the final stage of model training can be completed efficiently and orderly. By precisely controlling the data sending and receiving operations under the second micro-batch order number, the optimization of data flow is achieved, reducing unnecessary waiting and communication delay. The concurrent execution of the fourth sending operation and the eighth receiving operation at the beginning of the cooling phase marks the start of the GPU to close the cycle of computation and communication. In the middle part of the cooling phase, the eighth receiving operation becomes the main task of the GPU, focusing on collecting the backpropagation results to collect the parameters for the final correction. Through this carefully designed operation process, model training can run efficiently while ensuring data processing integrity and accuracy until it is completely finished. This cooling phase transmission operation strategy ensures that even when the training process is close to the end, high levels of data processing efficiency can be maintained, providing strong support for the successful training of the model.
[0157] Based on the above embodiments and optional embodiments, an optional implementation is provided, which is described in detail as follows.
[0158] In related technologies, the existing MoE-oriented large model computing and communication overlapping pipeline parallel acceleration method creates GPU computing flow and GPU communication flow on the same GPU card to realize the overlapping execution of GPU computing and GPU communication. In the GPU communication flow, both all-to-all communication between GPU servers (for expert parallel acceleration) and send / recv communication between GPU servers (for pipeline parallel acceleration) are performed, and the all-to-all communication and the send / recv communication between the GPU servers are in a serial execution relationship, Figure 2 is a schematic diagram of applying the existing computing and communication overlapping pipeline parallel acceleration method in a low-latency, high-bandwidth network environment, as Figure 2 shown, this serial execution mode has a strong dependence on the low-latency, high-bandwidth network environment. Only in the low-latency, high-bandwidth network environment can the GPU computing and GPU communication be completely overlapped. When using the existing technology in a high-latency, low-bandwidth network, the send / recv communication will block the subsequent all-to-all communication and GPU computing, increasing the waiting time of the GPU computing, Figure 3 is a schematic diagram of applying the existing computing and communication overlapping pipeline parallel acceleration method in a high-latency, low-bandwidth network environment, as Figure 3 shown, slowing down the entire large model training and inference process, and significantly reducing the GPU computing power utilization rate.
[0159] In view of this, the model training method provided in the optional embodiment of the present application provides a new pipeline parallel acceleration method of overlapping calculation and communication. In the scenario provided in the optional embodiment of the present application, the target model is a MoE large model, the target model is divided into multiple model layer segments, and the system includes multiple GPUs (which can also be referred to as workers). The multiple GPUs drive the corresponding model layer segments to train the model layer segments and update the model parameters corresponding to the model layer segments, thereby obtaining the trained target model.
[0160] Figure 4 is a schematic diagram of the pipeline parallel acceleration method of overlapping calculation and communication provided in the optional embodiment of the present application in a high-latency and low-bandwidth network environment, as shown in Figure 4 , in which method, the all-to-all communication and the send / recv communication between the GPU servers are executed in an overlapping manner, that is, the send / recv communication does not block the subsequent all-to-all communication and GPU calculation. This overlapping execution can greatly reduce the requirements of the network bandwidth and latency between the GPU servers for the send / recv communication between the GPU servers. When this overlapping execution mode is used in a low-bandwidth and high-latency network environment, the waiting time of the GPU calculation can be reduced, thereby improving the large model training efficiency and the GPU computing power utilization rate.
[0161] First, the application background of the method provided in the optional embodiment of the present application is introduced:
[0162] (1) In order to improve the GPU utilization rate, a batch of training data is further divided into multiple micro-batches in each large model training iteration. For each micro-batch, forward propagation calculation is performed first and then backward propagation calculation is performed. The forward and backward propagation calculations of each micro-batch are performed in multiple steps, and each GPU server is only responsible for the calculation of one step. If 4 GPU servers are required to participate in the calculation in the pipeline, the calculation process of a micro-batch is as follows: GPU server 1 performs forward propagation calculation -> GPU server 2 performs forward propagation calculation -> GPU server 3 performs forward propagation calculation -> GPU server 4 performs forward propagation calculation -> GPU server 4 performs backward propagation calculation -> GPU server 3 performs backward propagation calculation -> GPU server 2 performs backward propagation calculation -> GPU server 1 performs backward propagation calculation. Figure 5 is a timing diagram of a common pipeline parallel acceleration method in the prior art, as shown in Figure 5 , it is assumed that the bandwidth between the servers is very large, and the communication time can be ignored. As shown in Figure 5This conclusion can be easily drawn from all the blue and green squares with the same number. It can also be seen that the execution order of the forward and backward propagation computations is reversed. When a large model training iteration contains multiple micro-batches, the forward and backward propagation computations of these multiple micro-batches overlap in time (see...). Figure 5 ). Figure 5 The pipeline has four stages, and each stage requires a GPU server. Figure 5 Each large model training iteration contains 12 micro-batches. Blue squares represent forward propagation computations, and green squares represent backward propagation computations. The numbers on the blue and green squares represent the micro-batch numbers (from 1 to 12). The horizontal axis is the time axis. The pipeline execution process can be divided into three stages: warm-up, stabilization, and cool-down. Figure 5 The production line shown in the video operates on a micro-batch basis.
[0163] (2) To avoid GPU communication blocking GPU computation, existing large-scale MoE models generally employ a pipelined parallel acceleration method that overlaps computation and communication during training. Regarding this point, Figure 2 This has already been demonstrated. In addition, to achieve computational and communication overlap, GPU computation in existing MoE large model training is performed on a group basis, rather than on a micro-batch basis. A group contains one micro-batch of forward propagation computation and another micro-batch of backward propagation computation. Figure 6 This is a schematic diagram of overlapping intra-group computing and communication in existing technology, such as... Figure 6 As shown, a characteristic of GPU computation performed on a group basis is that the forward propagation computation of one micro-batch and the backward propagation computation of another micro-batch belonging to the same group are performed alternately (see...). Figure 6 Alternating GPU computations is intended to achieve the aforementioned overlap in computation and communication. It's important to note that although GPU computations within different micro-batches of the same group are alternating, the computation process remains unchanged for each micro-batch.
[0164] (3) Strictly speaking, a group contains most of the forward propagation computation of a microbatch and the backward propagation computation of another microbatch, in addition to a small portion of the forward propagation computation of a third microbatch. Figure 7 This is another schematic diagram of overlapping intra-group computing and communication in the prior art, such as... Figure 7 As shown, but for the sake of explaining the technical solution, it will still be said that a group contains a forward propagation computation of one micro-batch and a backward propagation computation of another micro-batch. At the same time, a group is identified by the number of the micro-batch that performs most of the forward propagation computation and the number of the micro-batch that performs the backward propagation computation.
[0165] Figure 8is a timing diagram of a new computing and communication overlapping pipeline for MoE large model provided by the optional embodiment of the present application, as shown in Figure 8 In the stable stage, the pipeline performs GPU computing and related GPU communication in groups. For the second worker (or GPU server), the identification of the first micro-batch group in the stable stage is 10+1, and the premise of performing GPU computing and related GPU communication of this micro-batch group is to obtain the forward propagation calculation result of micro-batch 10 (sent by the first worker) and the backward propagation calculation result of micro-batch 1 (sent by the third worker).
[0166] The running process of each worker in the pipeline includes the following four steps:
[0167] 1. Determine the number of micro-batches in the warm-up, stable, and cooling stages of the pipeline according to the pre-agreed strategy.
[0168] Among them, this step is the same as receiving the model training request described above, wherein the model training request carries a target strategy, which is the training strategy for training the initial model; in response to the model training request, the number of micro-batches processed by each of the plurality of pipeline stages is determined according to the target strategy, wherein the plurality of pipeline stages include a warm-up stage, a stable stage, and a cooling stage.
[0169] 2. Perform the computing and communication operations of the warm-up stage of the pipeline in a computing and communication overlapping manner.
[0170] 3. If the number of micro-batch groups in the stable stage of the pipeline is not 0, perform the computing and communication operations of the stable stage of the pipeline in a computing and communication overlapping manner.
[0171] 4. Perform the computing and communication operations of the cooling stage of the pipeline in a computing and communication overlapping manner.
[0172] Among them, steps 2, 3, and 4 are the same as the above-mentioned computing and communication parallel execution manner, and the computing and communication operations corresponding to the plurality of pipeline stages are sequentially executed in the order of pipeline stage execution, to obtain the current backward propagation results corresponding to the plurality of micro-batches, wherein the corresponding computing and communication operations include the computing operation of calculating the propagation results of the corresponding micro-batches according to the number of micro-batches of the corresponding stage, and the communication operation of transmitting the propagation results of the corresponding micro-batches. In order to subsequently train the model layer segment according to the current backward propagation results corresponding to the plurality of micro-batches to update the model parameters corresponding to the model layer segment to obtain the target model.
[0173] Figure 9 is a flowchart of the computing and communication operations of the warm-up stage of the pipeline provided by the optional embodiment of the present application, as shown in Figure 9As shown, the pipeline warm-up phase contains the following computation and communication operations:
[0174] 1) Initialization, i.e., let N = 1, M = 1.
[0175] 2) (Blocking) receive the forward propagation computation result of the Nth micro-batch (same as the above-mentioned first receiving operation is started and executed).
[0176] 3) For each micro-batch of the warm-up phase, the following operations are performed (same as the above-mentioned warm-up phase loop operation is executed):
[0177] (1) If the current micro-batch is not the last micro-batch of the current training iteration, start the non-blocking receiving process of the forward propagation computation result of the N+1th micro-batch.
[0178] (2) Perform the forward propagation computation of the Nth micro-batch.
[0179] (Same as the above (1) (2) determines the first micro-batch order number currently to be processed by the target GPU; according to the order number relationship between the first micro-batch order number and the micro-batch order number to be processed by the iterative training, the corresponding first warm-up phase operation is started and executed, including: in the case that the first micro-batch order number is not the last micro-batch order number of the iterative training, the second receiving operation and the first computation operation are started and executed, wherein the micro-batch order number to be processed by the warm-up phase is determined according to the number of micro-batches processed by the plurality of pipeline phases respectively)
[0180] (3) Start the non-blocking sending process of the forward propagation computation result of the Nth micro-batch.
[0181] (Same as the above, according to the order number relationship between the first micro-batch order number and the micro-batch order number to be processed by the warm-up phase, the corresponding second warm-up phase operation is started and executed, including: in the case that the first micro-batch order number is not the last micro-batch order number of the warm-up phase, nor the second last micro-batch order number, the first sending operation is started and executed)
[0182] (3) If the current micro-batch is the last micro-batch of the pipeline warm-up phase, let N = N + 1, and then enter the pipeline stable phase.
[0183] (Same as the above, according to the order number relationship between the first micro-batch order number and the micro-batch order number to be processed by the warm-up phase, the corresponding second warm-up phase operation is started and executed, including: in the case that the second micro-batch order number is the last micro-batch order number of the warm-up phase, control N = N + 1, and control the target GPU to enter the stable phase and perform the computation and communication operations corresponding to the stable phase)
[0184] (4) If the current micro-batch is the second last micro-batch of the warm-up stage, start the non-blocking sending process of the forward propagation computation results of the Nth micro-batch and the non-blocking receiving process of the backward propagation computation results of the Mth micro-batch; otherwise, start the non-blocking sending process of the forward propagation computation results of the Nth micro-batch.
[0185] (Same as the above, according to the sequence number relationship between the first micro-batch sequence number and the micro-batch sequence number to be processed in the warm-up stage, start and execute the corresponding second warm-up stage operation, including: in the case that the first micro-batch sequence number is the second last micro-batch sequence number of the warm-up stage, start and execute the first sending operation, the third receiving operation and the counting operation; in the case that the first micro-batch sequence number is neither the last micro-batch sequence number nor the second last micro-batch sequence number of the warm-up stage, start and execute the first sending operation)
[0186] (5) Wait for the non-blocking receiving of the forward propagation computation results of the N+1th micro-batch to be completed.
[0187] (6) Let N = N + 1, and then go to step (1) to process the next micro-batch.
[0188] (Same as the above, the counting operation in the above (5) and (6) is the same as the above)
[0189] Figure 10 is the flowchart of the computing and communication operations of the pipeline stable stage provided by the optional embodiment of the present application, as shown in Figure 10 The pipeline stable stage contains the following computing and communication operations:
[0190] For each micro-batch group of the pipeline stable stage, the following operations are performed:
[0191] (1) Start the non-blocking sending process of the forward propagation computation results of the N-1th micro-batch and the non-blocking receiving process of the backward propagation computation results of the M+1th micro-batch.
[0192] (Same as the above, start and execute the second sending operation and the fourth receiving operation)
[0193] (2) Wait for the non-blocking receiving of the forward propagation computation results of the Nth micro-batch and the non-blocking receiving of the backward propagation computation results of the Mth micro-batch to be completed.
[0194] (3) If the current micro-batch group is the first micro-batch group of the warm-up stage of the pipeline, start the non-blocking receiving process of the forward propagation calculation result of the N+1th micro-batch; if the current micro-batch is the last micro-batch group of the warm-up stage of the pipeline, start the non-blocking sending process of the backward propagation calculation result of the M-1th micro-batch; if it is other cases, start the non-blocking receiving process of the forward propagation calculation result of the N+1th micro-batch and the non-blocking sending process of the backward propagation calculation result of the M-1th micro-batch.
[0195] (2) and (3) above, determine the batch group currently to be processed by the target GPU, and start and execute the corresponding stable receiving and sending operation according to the group order relationship between the batch group and the batch group to be processed in the stable stage: in the case that the batch group is the first batch group of the stable stage, start and execute the seventh receiving operation, wherein the batch group to be processed in the stable stage is determined according to the number of micro-batches processed by the plurality of pipeline stages respectively; in the case that the batch group is neither the first batch group of the stable stage nor the last batch group of the stable stage, start and execute the seventh receiving operation and the third sending operation; in the case that the batch group is the last batch group of the stable stage, start and execute the third sending operation)
[0196] (4) Execute the GPU calculation and related GPU communication of the group identifier “N+M”.
[0197] (2) and (3) above, determine the batch group currently to be processed by the target GPU, and start and execute the corresponding stable receiving and sending operation according to the group order relationship between the batch group and the batch group to be processed in the stable stage: in the case that the batch group is the first batch group of the stable stage, start and execute the seventh receiving operation, wherein the batch group to be processed in the stable stage is determined according to the number of micro-batches processed by the plurality of pipeline stages respectively; in the case that the batch group is neither the first batch group of the stable stage nor the last batch group of the stable stage, start and execute the seventh receiving operation and the third sending operation; in the case that the batch group is the last batch group of the stable stage, start and execute the third sending operation)
[0198] (5) If the current micro-batch group is the last micro-batch of the stable stage of the pipeline, let M=M+1, and then go to the cooling stage; otherwise, let N=N+1, M=M+1, and then go to the first step to process the next group of micro-batches.
[0199] (2) and (3) above, determine the batch group currently to be processed by the target GPU, and start and execute the corresponding stable receiving and sending operation according to the group order relationship between the batch group and the batch group to be processed in the stable stage: in the case that the batch group is the first batch group of the stable stage, start and execute the seventh receiving operation, wherein the batch group to be processed in the stable stage is determined according to the number of micro-batches processed by the plurality of pipeline stages respectively; in the case that the batch group is neither the first batch group of the stable stage nor the last batch group of the stable stage, start and execute the seventh receiving operation and the third sending operation; in the case that the batch group is the last batch group of the stable stage, start and execute the third sending operation)
[0200] Figure 11 is the flowchart of the calculation and communication operation of the pipeline cooling stage provided by the optional embodiment of the present application, as shown in Figure 11 The pipeline cooling stage contains the following calculation and communication operations:
[0201] For each micro-batch of the pipeline cooling stage, the following operations are performed:
[0202] (1) If the current micro-batch is the first micro-batch of the pipeline cooling stage, start the non-blocking receiving process of the forward propagation calculation result of the Nth micro-batch and the non-blocking sending process of the backward propagation calculation result of the M+1th micro-batch; if the current micro-batch is neither the first micro-batch of the pipeline cooling stage nor the last micro-batch of the pipeline cooling stage, start the non-blocking sending process of the backward propagation calculation result of the M+1th micro-batch.
[0203] (As described above, determine the second micro-batch order number currently to be processed by the target GPU; according to the order number relationship between the second micro-batch order number and the micro-batch order number to be processed in the cooling stage, start and execute the corresponding cooling transceiving operation: in the case that the second micro-batch order number is the first micro-batch order number of the cooling stage, start and execute the fourth sending operation and the eighth receiving operation, wherein the micro-batch order number to be processed in the cooling stage is determined according to the number of micro-batches processed by each of the plurality of pipeline stages; in the case that the second micro-batch order number is neither the first micro-batch order number of the cooling stage nor the last micro-batch order number, start and execute the eighth receiving operation)
[0204] (2) Wait for the non-blocking receiving of the backward propagation calculation result of the Mth micro-batch to be completed.
[0205] (3) Start the non-blocking sending process of the backward propagation calculation result of the M-1th micro-batch.
[0206] (As described above in (2) and (3), start and execute the fifth sending operation in the case that the ninth receiving operation is completed)
[0207] (4) Execute the backward propagation calculation of the Mth micro-batch.
[0208] (As described above, start and execute the third calculation operation)
[0209] (5) If the current micro-batch is the last micro-batch of the pipeline cooling stage, wait for the non-blocking sending of the backward propagation calculation result of the M-1th micro-batch to be completed, then send the backward propagation calculation result of the Mth micro-batch (in a blocking manner), and finally stop the pipeline; otherwise, let M=M+1, and then go to step (1) to process the next micro-batch.
[0210] (As described above, determine whether the micro-batch order number currently to be processed by the target GPU is the last micro-batch order number of the cooling stage; in the case that it is not the last micro-batch order number, control M=M+1, and execute the cooling stage loop operation cyclically until it is the last micro-batch order number, and start and execute the sixth sending operation, wherein the sixth sending operation is used to send the current backward propagation result corresponding to the Mth micro-batch in the case that the fifth sending operation is completed)
[0211] It should be noted that the receiving forward propagation calculation result in the optional embodiment of the present application refers to receiving the forward propagation calculation result sent by the upper worker. If the current training process is the first worker, the operations of receiving the forward propagation calculation result, starting the non-blocking reception of the forward propagation calculation result, waiting for the completion of the non-blocking reception of the forward propagation calculation result and the like are empty operations, which can be removed from the above process.
[0212] The sending forward propagation calculation result in the optional embodiment of the present application refers to sending the forward propagation calculation result to the lower worker. If the current training process is the last worker, the operations of sending the forward propagation calculation result, starting the non-blocking sending of the forward propagation calculation result, waiting for the completion of the non-blocking sending of the forward propagation calculation result and the like are empty operations, which can be removed from the above process.
[0213] The receiving backward propagation calculation result in the optional embodiment of the present application refers to receiving the backward propagation calculation result sent by the lower worker. If the current training process is the last worker, the operations of receiving the backward propagation calculation result, starting the non-blocking reception of the backward propagation calculation result, waiting for the completion of the non-blocking reception of the backward propagation calculation result and the like are empty operations, which can be removed from the above process.
[0214] The sending backward propagation calculation result in the optional embodiment of the present application refers to sending the backward propagation calculation result to the upper worker. If the current training process is the first worker, the operations of sending the backward propagation calculation result, starting the non-blocking sending of the backward propagation calculation result, waiting for the completion of the non-blocking sending of the backward propagation calculation result and the like are empty operations, which can be removed from the above process.
[0215] Through the above optional embodiment, at least the following beneficial effects can be achieved:
[0216] (1) Firstly, as much forward propagation calculation as possible is performed in the pipeline warm-up stage, that is, the number of micro-batches in the pipeline warm-up stage is increased as much as possible. Increasing the number of micro-batches in the warm-up stage means that the GPU can start the continuous calculation process earlier, avoiding the idle of the initial calculation resources and improving the continuity and efficiency of the calculation. Performing more calculations in the warm-up stage can reduce the waiting time of the GPU at the beginning of the stable stage. With the gradual overlap of GPU calculation and communication, the waiting time of the entire training process is reduced, and the utilization rate of the GPU is improved.
[0217] (2) The second is to overlap the execution of the front and back propagation GPU calculations, all-to-all communication between GPU servers, and GPU communication within the pipeline. Before each front and back propagation GPU calculation starts, one or more non-blocking GPU communication (i.e., receiving and sending front and back propagation calculation results) processes are started. Through the overlapping execution of calculation and communication, the GPU can receive or send data while performing calculations, reducing the time waiting for communication and improving the overall calculation and communication efficiency. Even in the presence of network delays, through overlapping execution, GPU calculations can continue to advance and will not be affected by long waiting for communication, improving the robustness of model training to network environment. And non-blocking communication allows the GPU to continue to perform calculation tasks while waiting for data transmission to complete, thereby maximizing the utilization of the GPU and reducing resource waste.
[0218] It should be noted that for the foregoing method embodiments, in order to simply describe, they are all expressed as a series of action combinations, but those skilled in the art should know that the present application is not limited by the order of the described actions, because according to the present application, certain steps can be performed in other orders or performed. Secondly, those skilled in the art should know that the embodiments described in the specification are all preferred embodiments, and the actions and modules involved are not necessarily necessary for the present application.
[0219] From the above description of the embodiments, those skilled in the art can clearly understand that the method according to the above embodiments can be realized by means of software and the necessary general hardware platform, of course, it can also be realized by hardware, but in many cases the former is a better embodiment. Based on such understanding, the technical solutions of the present application can be embodied in the form of a software product, which is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk) and includes a plurality of instructions for causing a terminal device (which can be a mobile phone, computer, server, or network device, etc.) to execute the method of each embodiment of the present application.
[0220] Embodiment 2
[0221] According to the embodiments of the present application, a mixed expert MoE target model is also provided, wherein the target model includes a plurality of model layer segments, and the plurality of model layer segments are driven by corresponding GPUs respectively, and the corresponding GPUs are used to train the corresponding model layer segments using the above method to update the model parameters of the corresponding model layer segments to obtain the target model.
[0222] It should be noted that the training method of the above model corresponds to steps S102 to S108 in the implementation of the model training method, and the instances and application scenarios realized by the corresponding steps are the same, but are not limited to the content disclosed in the above embodiment 1.
[0223] Embodiment 3
[0224] According to an embodiment of the present application, a system for implementing the above model training method is also provided, which comprises a target model and a plurality of GPUs, wherein the target model comprises a plurality of model layer segments, the plurality of model layer segments are respectively driven by corresponding GPUs, and the corresponding GPUs are used to train the corresponding model layer segments using the above method to update the model parameters of the corresponding model layer segments to obtain the target model.
[0225] It should be noted that the training method of the model in the above system corresponds to steps S102 to S108 in the implementation of the model training method, and the same instances and application scenarios are realized as the corresponding steps, but are not limited to the content disclosed in Embodiment 1.
[0226] Embodiment 4
[0227] According to another aspect of an embodiment of the present application, an electronic device is also provided, which comprises a processor and a memory for storing processor-executable instructions, wherein the processor is configured to execute the instructions to implement the model training method of any one of the above.
[0228] Embodiment 5
[0229] According to another aspect of an embodiment of the present application, a computer-readable storage medium is also provided, when the instructions in the computer-readable storage medium are executed by the processor of an electronic device, the electronic device can execute the model training method of any one of the above.
[0230] The above embodiment numbers of the present application are only for description, and do not represent the advantages and disadvantages of the embodiments.
[0231] In the above embodiments of the present application, the description of each embodiment has its own focus, and the parts not described in detail in a certain embodiment can be referred to the related description of other embodiments.
[0232] In several embodiments provided in the present application, it should be understood that the disclosed technical contents can be implemented by other ways. Among them, the above-described device embodiments are only schematic, for example, the division of the units can be a logical function division, and actual implementation can have another division way, for example, a plurality of units or components can be combined or integrated into another system, or some features can be ignored or not executed. In addition, the coupling or direct coupling or communication connection between the displayed or discussed each other can be through some interface, indirect coupling or communication connection between units or modules, which can be electrical or other forms.
[0233] The units described as separate components may or may not be physically separate, and the components displayed as units may or may not be physical units, that is, may be located in one place, or may be distributed to multiple units. Part or all of the units can be selected according to actual needs to achieve the purpose of the embodiment scheme.
[0234] In addition, each functional unit in each embodiment of the present application can be integrated in one processing unit, or each unit can exist physically, or two or more units can be integrated in one unit. The integrated unit can be realized in the form of hardware or in the form of a software functional unit.
[0235] The integrated unit, if realized in the form of a software functional unit and sold or used as an independent product, can be stored in a computer readable storage medium. Based on such understanding, the technical scheme of the present application or the part of the present application which contributes to the prior art or the whole or part of the technical scheme can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes a plurality of instructions for causing a computer device (which can be a personal computer, a server or a network device, etc.) to execute all or part of the steps of the method described in each embodiment of the present application. The foregoing storage medium includes various media that can store program codes, such as a U disk, a read-only memory (ROM), a random access memory (RAM), a mobile hard disk, a magnetic disk or an optical disk.
[0236] The above is only the preferred embodiment of the present application, and it should be pointed out that for ordinary skilled in the art, without departing from the principles of the present application, a number of improvements and refinements can be made, and these improvements and refinements should be considered as the protection scope of the present application.
Claims
1. A model training method, characterized in that, The application is applied to a target GPU, wherein the target GPU is used to drive a corresponding model layer segment included in an initial model, comprising: receiving a model training request, wherein the target strategy is carried in the model training request, and the target strategy is a training strategy for training the initial model; in response to the model training request, determining the number of micro-batches processed by each of the plurality of pipeline stages according to the target strategy, wherein the plurality of pipeline stages include a warm-up stage, a stabilization stage, and a cooling stage; in a manner of performing computation and communication in parallel, sequentially performing computation and communication operations corresponding to each of the plurality of pipeline stages according to the execution order of the pipeline stages to obtain current back propagation results corresponding to the plurality of micro-batches, wherein the corresponding computation and communication operations include a computation operation for calculating the propagation result of the corresponding micro-batch according to the number of micro-batches of the corresponding stage, and a communication operation for transmitting the propagation result of the corresponding micro-batch; training the model layer segment according to the current back propagation results corresponding to the plurality of micro-batches to update the model parameters corresponding to the model layer segment to obtain a target model.
2. The method of claim 1, wherein, performing the computation and communication operations corresponding to the warm-up stage, comprising: starting and executing a first receiving operation, wherein the first receiving operation is used to receive upstream forward propagation results corresponding to the Nth micro-batch, and the N is set with an initial value; performing a warm-up stage loop operation: determining a first micro-batch order number currently to be processed by the target GPU; starting and executing a corresponding first warm-up stage operation according to the order number relationship between the first micro-batch order number and the micro-batch order number to be processed by the iterative training, wherein the first warm-up stage operation includes at least one of a second receiving operation, a first calculation operation, the second receiving operation is used to receive upstream forward propagation results corresponding to the N+1th micro-batch, and the first calculation operation is used to determine the current forward propagation result corresponding to the Nth micro-batch according to the upstream forward propagation result corresponding to the Nth micro-batch, and the micro-batch order number to be processed by the iterative training is determined according to the total number of micro-batches; starting and executing a corresponding second warm-up stage operation according to the order number relationship between the first micro-batch order number and the micro-batch order number to be processed by the warm-up stage, wherein the second warm-up stage operation includes at least one of a first sending operation, a third receiving operation, a counting operation, an entering stabilization stage operation, the first sending operation is used to send the current forward propagation result corresponding to the Nth micro-batch, the third receiving operation is used to receive upstream back propagation results corresponding to the Mth micro-batch, the M is set with an initial value, the counting operation is used to control N to be equal to N+1 in the case that the second receiving operation is executed, and the entering stabilization stage operation is used to control N=N+1 and control the target GPU to enter the stabilization stage and perform the computation and communication operations corresponding to the stabilization stage, and the micro-batch order number to be processed by the warm-up stage is determined according to the number of micro-batches processed by each of the plurality of pipeline stages. The warm-up phase cycle operation is executed cyclically until the target GPU enters a stable phase, and a calculation and communication operation corresponding to the stable phase is executed.
3. The method of claim 2, wherein, According to the sequence relationship between the first micro-batch order number and the micro-batch order number to be processed in the iteration training, a corresponding first warm-up phase operation is started and executed, including: In the case that the first micro-batch order number is not the last micro-batch order number of the iteration training, the second receiving operation and the first calculation operation are started and executed, wherein the second receiving operation and the first calculation operation are executed in parallel; In the case that the first micro-batch order number is the last micro-batch order number of the iteration training, the first calculation operation is started and executed.
4. The method of claim 2, wherein, According to the sequence relationship between the first micro-batch order number and the micro-batch order number to be processed in the warm-up phase, a corresponding second warm-up phase operation is started and executed, including: In the case that the first micro-batch order number is the last micro-batch order number of the warm-up phase, N is controlled to be N+1, and the target GPU is controlled to enter a stable phase, and a calculation and communication operation corresponding to the stable phase is executed; In the case that the first micro-batch order number is not the last micro-batch order number of the warm-up phase, nor the second last micro-batch order number, the first sending operation and the counting operation are started and executed; In the case that the first micro-batch order number is the second last micro-batch order number of the warm-up phase, the first sending operation, the third receiving operation and the counting operation are started and executed, wherein the first sending operation and the third receiving operation are executed in parallel.
5. The method of claim 1, wherein, The calculation and communication operation corresponding to the stable phase is executed, including: The second sending operation and the fourth receiving operation are started and executed, wherein the second sending operation is used to send the current forward propagation result corresponding to the N-1 micro-batches, and the fourth receiving operation is used to receive the upstream backward propagation result corresponding to the M+1 micro-batches; The stable phase cycle operation is executed: In the case that the fifth receiving operation and the sixth receiving operation are both completed, the batch group currently to be processed by the target GPU is determined, wherein the fifth receiving operation is used to receive the upstream forward propagation result corresponding to the N micro-batches, and the sixth receiving operation is used to receive the upstream backward propagation result corresponding to the M micro-batches; According to the group sequence relationship between the batch group and the batch group to be processed in the stable phase, a corresponding stable transceiving operation is started and executed, wherein the stable transceiving operation includes at least one of the third sending operation and the seventh receiving operation, the third sending operation is used to send the current backward propagation result corresponding to the M-1 micro-batches, the seventh receiving operation is used to receive the upstream forward propagation result corresponding to the N+1 micro-batches, and the batch group to be processed in the stable phase is determined according to the number of micro-batches processed by the plurality of pipeline phases respectively. starting and executing a second calculation operation, wherein the second calculation operation is configured to perform a propagation result calculation of an N+M batch group, including determining a current forward propagation result corresponding to an Nth micro batch according to an upstream forward propagation result corresponding to the Nth micro batch, and determining a current backward propagation result corresponding to an Mth micro batch according to an upstream backward propagation result corresponding to the Mth micro batch; determining whether the batch group currently to be processed by the target GPU is the last batch group of the stabilization stage; in the case of not being the last batch group, controlling N=N+1 and M=M+1, and cyclically executing the stabilization stage cyclic operation until being the last batch group, controlling M=M+1, and controlling the target GPU to enter a cooling stage and execute a cooling stage operation corresponding to the cooling stage.
6. The method of claim 5, wherein, starting and executing a corresponding stabilization transceiving operation according to a group order relationship between the batch group and a batch group to be processed in the stabilization stage, including: in the case of the batch group being the first batch group of the stabilization stage, starting and executing the seventh receiving operation; in the case of the batch group not being the first batch group of the stabilization stage and not being the last batch group of the stabilization stage, starting and executing the seventh receiving operation and the third sending operation, wherein the seventh receiving operation and the third sending operation are executed in parallel; in the case of the batch group being the last batch group of the stabilization stage, starting and executing the third sending operation.
7. The method of claim 1, wherein, the execution of the calculation and communication operation corresponding to the cooling stage includes: executing a cooling stage cyclic operation: determining a second micro batch order number currently to be processed by the target GPU; starting and executing a corresponding cooling transceiving operation according to an order number relationship between the second micro batch order number and a micro batch order number to be processed in the cooling stage, wherein the cooling transceiving operation includes at least one of the following: a fourth sending operation, an eighth receiving operation configured to receive a current backward propagation result corresponding to an M+1th micro batch, and the fourth sending operation configured to send an upstream forward propagation result corresponding to an Nth micro batch, and the micro batch order number to be processed in the cooling stage is determined according to the number of micro batches processed by the plurality of pipeline stages respectively; in the case of a ninth receiving operation being completed, starting and executing a fifth sending operation, wherein the ninth receiving operation is configured to receive an upstream backward propagation result corresponding to an Mth micro batch, and the fifth sending operation is configured to send a current backward propagation result corresponding to an M-1th micro batch; starting and executing a third calculation operation, wherein the third calculation operation is configured to determine a current backward propagation result corresponding to the Mth micro batch according to an upstream backward propagation result corresponding to the Mth micro batch; determining whether the micro batch order number currently to be processed by the target GPU is a last micro batch order number of the cooling stage; In the case that the current micro-batch sequence number is not the last micro-batch sequence number, the control M = M + 1, and the cooling phase cycle operation is executed in a loop until the last micro-batch sequence number is reached, and a sixth sending operation is started and executed, wherein the sixth sending operation is used to send the current back propagation result corresponding to the Mth micro-batch in the case that the fifth sending operation is completed.
8. The method of claim 7, wherein, The cooling transceiving operation started and executed according to the sequence number relationship between the second micro-batch sequence number and the micro-batch sequence number to be processed in the cooling phase comprises: In the case that the second micro-batch sequence number is the first micro-batch sequence number of the cooling phase, the fourth sending operation and the eighth receiving operation are started and executed, wherein the fourth sending operation and the eighth receiving operation are executed in parallel; In the case that the second micro-batch sequence number is neither the first micro-batch sequence number nor the last micro-batch sequence number of the cooling phase, the eighth receiving operation is started and executed.
9. A model training system, comprising: The target model comprises a plurality of model layer segments, and the plurality of model layer segments are respectively driven by corresponding GPUs, and the corresponding GPUs are used to train the corresponding model layer segments using the method of claim 1 to update the model parameters of the corresponding model layer segments to obtain the target model.
10. An electronic device, comprising: It comprises: a processor; a memory for storing instructions executable by the processor; wherein the processor is configured to execute the instructions to implement the method of any one of claims 1 to 8.
11. A computer readable storage medium, characterized in that, When the instructions in the computer readable storage medium are executed by the processor of the electronic device, the electronic device can execute the method of any one of claims 1 to 8.