Model training method, device, system and equipment
By splitting the computation graph during hybrid expert model training to perform computation and communication operations in parallel, the problems of communication overhead and GPU memory usage are solved, thereby improving training efficiency and resource utilization.
Patent Information
- Application Number
- CN202511035746.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-25
- Publication Date
- 2025-11-07
- Estimated Expiration
- 2045-07-25
Smart Images

Figure CN120910563A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present disclosure relates to the field of artificial intelligence, and in particular, to a model training method, device, system and equipment. BACKGROUND
[0002] The parallel mode used by the traditional AI infra training framework to train a model (such as a hybrid expert model) is generally divided into pipeline parallelism, tensor parallelism and expert parallelism.
[0003] Pipeline parallelism cuts the model according to layers, and the calculation result of the previous layer needs to be transmitted to the next layer for calculation through P2P (Peer-to-Peer) communication during the calculation process. Tensor parallelism and expert parallelism are similar, both of which cut the model in a layer and distribute it to different units (such as processes) to reduce the use of memory in a single unit through the way of calculation and communication scheduling before and after. These parallel modes can cut the model for large-scale parallel calculation on a cluster, but they need to generate communication between units as a cost.
[0004] Therefore, how to reduce the communication overhead in model training is a problem that needs to be solved at present. SUMMARY
[0005] An object of the present disclosure is to reduce the communication overhead in model training.
[0006] According to a first aspect of the present disclosure, a model training method is provided, comprising: in the execution of at least part of the forward process for the Nth batch of training samples, cutting a dynamically generated computation graph into a plurality of sub-forward processes, and obtaining a plurality of sub-backward processes corresponding to the plurality of sub-forward processes; after the forward process of the Nth batch of training samples is completed, and in the execution of at least part of the forward process of the Pth batch of training samples, in at least one time period, using a computing unit to perform one of the sub-forward processes in the at least part of the forward process of the Pth batch of training samples and one of the plurality of sub-backward processes of the Nth batch of training samples in parallel, and the sub-forward process and the sub-backward process performed in parallel, one of which belongs to a computing operation, and the other belongs to a communication operation.
[0007] Optionally, cutting the dynamically generated computation graph into a plurality of sub-forward processes comprises: identifying a first node in the computation graph that does not need to calculate the gradient based on the attribute information of each node in the computation graph; and cutting the computation graph at the first node as a cutting position to obtain a plurality of sub-forward processes.
[0008] Optionally, splitting the computation graph with the first node as a split position comprises: creating a new output tensor for each output tensor of the first node; and connecting the new output tensor as an input of a next sub-forward process.
[0009] Optionally, the method further comprises: for the sub-forward process, creating a first pseudo tensor with a length of zero and a second pseudo tensor with a length of zero; connecting the first pseudo tensor as an output of the sub-forward process and an output tensor of the sub-forward process; and connecting the second pseudo tensor as an input of a next sub-forward process and the new output tensor.
[0010] Optionally, the model is a hybrid expert model, and a forward process of the hybrid expert model comprises a routing stage, a distribution stage, a multi-layer perceptron (MLP) stage, and an aggregation stage, and the method further comprises: in the forward process, not saving a forward calculation result of the MLP stage, and in the backward process, re-computing for an MLP part other than a last layer in the MLP stage.
[0011] Optionally, the method further comprises: in the forward process, not saving a forward calculation result of the distribution stage, and in the backward process, re-computing for the distribution stage.
[0012] According to a second aspect of the present disclosure, a model training apparatus is further provided, comprising: a splitting module configured to split a dynamically generated computation graph into a plurality of sub-forward processes in performing at least part of a forward process for an Nth batch of training samples, and obtain a plurality of sub-backward processes corresponding to the plurality of sub-forward processes; and a parallel execution module configured to, after the forward process of the Nth batch of training samples is completed, and in performing at least part of a forward process of a Pth batch of training samples, utilize a computing unit to parallelly execute one of the sub-forward processes in the at least part of the forward process of the Pth batch of training samples and one of the sub-backward processes of the Nth batch of training samples in at least one time period, and the parallelly executed sub-forward process and sub-backward process, one of which belongs to a computing operation and the other of which belongs to a communication operation.
[0013] According to a third aspect of the present disclosure, there is also provided a distributed training system of a model, comprising a plurality of computing units, wherein the computing units split a dynamically generated computation graph into a plurality of sub-forward processes and obtain a plurality of sub-backward processes corresponding to the plurality of sub-forward processes in performing at least part of a forward process for an Nth batch of training samples; the computing units perform one of the sub-forward processes in performing at least part of a forward process for a Pth batch of training samples and one of the sub-backward processes of the Nth batch of training samples in parallel in at least one time period after the forward process of the Nth batch of training samples is completed and in performing at least part of the forward process of the Pth batch of training samples, and the sub-forward process and the sub-backward process performed in parallel, one of which belongs to a computing operation and the other belongs to a communication operation.
[0014] According to a fourth aspect of the present disclosure, there is provided a computing device, comprising a processor; and a memory having stored thereon executable code that, when executed by the processor, causes the processor to perform the method according to the first aspect.
[0015] According to a fifth aspect of the present disclosure, there is provided a computer program product, comprising executable code that, when executed by a processor of an electronic device, causes the processor to perform the method according to the first aspect.
[0016] According to a sixth aspect of the present disclosure, there is provided a non-transitory machine-readable storage medium having stored thereon executable code that, when executed by a processor of an electronic device, causes the processor to perform the method according to the first aspect.
[0017] The present disclosure splits a dynamically generated computation graph into a plurality of sub-forward processes and obtains a plurality of sub-backward processes corresponding to the plurality of sub-forward processes in performing at least part of a forward process for an Nth batch of training samples; the computing units perform one of the sub-forward processes in performing at least part of a forward process for a Pth batch of training samples and one of the sub-backward processes of the Nth batch of training samples in parallel in at least one time period after the forward process of the Nth batch of training samples is completed and in performing at least part of the forward process of the Pth batch of training samples, and the sub-forward process and the sub-backward process performed in parallel, one of which belongs to a computing operation and the other belongs to a communication operation. Thereby, at least part of the communication cost overhead can be masked in the model training process, thereby improving the efficiency of model training and improving the utilization rate of the computing units (such as a cluster). BRIEF DESCRIPTION OF DRAWINGS
[0018] The above and other objects, features and advantages of the present disclosure will become more apparent from the following detailed description when taken in conjunction with the accompanying drawings, in which like reference characters refer to like parts throughout the figures, and wherein:
[0019] Figure 1 Two mainstream structure diagrams of large language models are shown.
[0020] Figure 2 A schematic flowchart of a model training method according to an embodiment of the present disclosure is shown.
[0021] Figure 3 A schematic diagram of forward and backward processes in model training is shown.
[0022] Figure 4A A structure diagram of a computational graph is shown.
[0023] Figure 4B A schematic diagram of splitting the computational graph shown in Figure 4A is shown.
[0024] Figure 4C A schematic diagram of optimizing the computational graph shown in Figure 4B is shown.
[0025] Figure 5 A schematic flowchart of parallel execution of forward and backward processes is shown.
[0026] Figure 6 A schematic diagram of re-computation in a hybrid expert model is shown.
[0027] Figure 7 A structure diagram of a model training apparatus according to an embodiment of the present disclosure is shown.
[0028] Figure 8 A structure diagram of a computing device according to an embodiment of the present disclosure is shown. DETAILED DESCRIPTION
[0029] Preferred embodiments of the present disclosure will be described herein below with reference to the accompanying drawings. While the preferred embodiments of the present disclosure are shown in the drawings, it is to be understood that the present disclosure can be implemented in various forms and should not be limited by the embodiments set forth herein. Rather, these embodiments are provided so that this disclosure will be thorough and complete, and will fully convey the scope of the present disclosure to those skilled in the art.
[0030] The following describes a mixture of experts (MoE) model as an example. It should be understood that the present disclosure can be used in communication overhead generated in various model training processes based on model splitting, and is not limited to the mixture of experts model.
[0031] The present disclosure does not have special requirements for model splitting methods (i.e., model parallel technology). It can be pipeline parallelism (PP) or tensor parallelism (TP). Pipeline parallelism refers to horizontally splitting a model into multiple sub-modules according to layers or stages, and each computing unit is responsible for running one of the sub-modules. Tensor parallelism refers to vertically splitting a parameter matrix of a single layer (splitting according to rows or columns), each device saves part of the parameters, and performs the same layer structure but different parameter calculations. In tensor parallelism, the model structure does not change, but each parameter is only 1 / tp_size (device parallel number) in size, and each computing unit runs the entire process, which requires synchronization of the results.
[0032] Current artificial intelligence is widely used, and large language model services have been integrated into various aspects of life, such as daily search queries and online shopping, which are accompanied by the emergence of large model AI assistants.
[0033] Currently, there are two main structures of large language models (LLM), one is dense, and the other is a mixture of experts. For how to solve a mixed task containing multiple domain knowledge, the simplest way is to bring together the experts in these fields to jointly process this problem, and then aggregate the results to obtain a high-quality answer. The mixture of experts model is designed based on this way. Compared with the dense model, the biggest difference in structure of the mixture of experts model is the use of a component called a gating network and a multilayer perceptron (MLP) containing multiple specialized sub-model structures.
[0034] Figure 1 The following shows two main structure diagrams of large language models.
[0035] Figure 1 The left side shows the basic structure of a large language model, including a self-attention layer, a first normalization layer, a feedforward neural network (FFN), and a second normalization layer. Figure 1The right side respectively shows a schematic diagram of the feedforward neural network adopting a dense model structure and a hybrid expert model structure. As shown in the left side of FIG. 1, the dense model includes a plurality of layers of perceptrons, and each layer of perceptrons includes a plurality of perceptrons. Figure 1 As shown in the right side of FIG. 1, the dense model only includes one FFN, and the hybrid expert model includes a plurality of FFNs, each of which can be regarded as an expert.
[0036] The gating network determines which specific expert to which the current task to be processed is assigned according to a specific router algorithm, and then the specific expert in the multi-layer perceptron processes the specific professional problem. Therefore, a significant advantage of the hybrid expert model compared with the dense model is that the hybrid expert model can be more effectively pre-trained with less resources, thereby obtaining better model effect. However, due to the special structure of the hybrid expert model, the current traditional AI infra training framework cannot effectively train the model. Therefore, how to efficiently train the hybrid expert model on a large-scale cluster has been one of the hot issues in the industry.
[0037] Currently, one of the challenges of efficiently training the traditional AI infra training framework on a large-scale cluster is how to sufficiently mask the overhead of communication cost and improve the overall training efficiency of the cluster.
[0038] In order to enable the hybrid expert model to exert the computing performance of the cluster itself on a large-scale cluster, the expert parallel mode is adopted to store a plurality of expert models on different units of the cluster. After each time passing through the gating network, an allgather (operation of collecting distributed data to all nodes) or alltoall (operation of each participating node sending data to all other nodes while also receiving data from all other nodes) collective communication operation is required to send the task to be processed to the unit where the expert model specified by the gating circuit is located. This process will result in a large number of collective communication operations during the training process. The collective communication overhead will continue to rise with the increase of the expert parallel scale. For the current mainstream large-scale hybrid expert model training, the communication overhead will account for 30%-50% of the overall training time. Therefore, the collective communication often becomes the bottleneck of the training of the hybrid expert model on a large-scale cluster.
[0039] In order to reduce the collective communication overhead, a problem existing in the implementation needs to be solved first.
[0040] The following is described by taking Pytorch as an example. It should be understood that the technical solutions of the present disclosure can also be applied to other similar deep learning frameworks. A conventional AI infra framework is generally developed based on Pytorch. The forward and backward propagation of each component needs to be implemented between the components in the framework, and the order between the components needs to be adjusted to hand over to the underlying Pytorch. Then, the dynamic computation graph is constructed by Pytorch to normally run the training of the model and complete the forward and backward calculations.
[0041] To ensure the correctness of the calculation, the computation graph generated by Pytorch can only generate the corresponding backward calculation process after completing the forward calculation process according to the established process on each process. Therefore, it is very difficult to control the execution timing of the current training stage when implementing the training model on the existing framework.
[0042] To this end, the present disclosure proposes an interrupt switching strategy. The "interrupt" in the interrupt switching strategy refers to the splitting processing of the computation graph dynamically generated in the single execution of the forward process of the computing unit, to obtain multiple sub-forward processes and corresponding multiple sub-backward processes. The "switching" in the interrupt switching strategy refers to the parallel execution of a sub-forward process of a later training batch and a sub-backward process of an earlier training batch by the computing unit. Moreover, the parallel execution of the sub-forward process and the sub-backward process, one of which belongs to calculation, and the other of which belongs to communication. In this way, the communication overhead in the model training process can be hidden, thereby improving the efficiency of the model training and improving the utilization rate of the computing unit (such as a cluster).
[0043] Figure 2 A schematic flowchart of a model training method according to an embodiment of the present disclosure is shown.
[0044] Figure 2 The method shown can be performed by any one of the computing units in a cluster for implementing distributed training of a model. The computing units in the cluster can also be referred to as computing nodes, worker nodes, or cluster nodes. The computing unit can be a physical or virtual independent computing unit. For example, the computing unit can be a GPU (Graphics Processing Unit). For another example, the computing unit can also be a process. It should be understood that, due to the large size of the training samples, the entire training sample is usually divided into multiple batches during training. The training sample can be a combination of one or more of text data, image data, voice data, or video data. Before being used for training, these data can be appropriately preprocessed to convert them into a plurality of tokens suitable for training, and the preprocessing operations include but are not limited to data cleaning, standardization, or structuring, etc.
[0045] Referring toFigure 2 At step S210, the dynamically generated computation graph is split into multiple sub-forward processes in performing at least part of the forward process for the Nth (N is a positive integer) batch of training samples, and multiple sub-backward processes corresponding to the multiple sub-forward processes are obtained.
[0046] The present disclosure first adopts a way of splitting the Pytorch dynamically generated computation graph to generate multiple sub-forward processes, which can include multiple calculation stages (such as attention calculation stages) and communication stages (such as dispatch communication stages). In a forward process, when a certain sub-forward process is completed, the backward input and the computation graph of the sub-forward process can be saved.
[0047] Figure 3 The forward process (i.e., the forward propagation process) and the backward process (i.e., the backward propagation process) in the model training process are shown. It should be understood that the forward propagation process is the process of calculating the prediction value of the model. Given the input data, a series of calculations are performed through the layers and activation functions in the model to obtain the output of the model. The backward propagation process is the process of calculating the gradient of the loss with respect to the model parameters. Specifically, the gradient of the loss function with respect to each layer parameter needs to be calculated, and the gradient is passed back to the optimizer for updating. Through forward propagation, the prediction value of the model can be calculated; through backward propagation, the gradient of the loss with respect to the model parameters can be calculated, and the optimizer is used for updating. By repeating these two processes multiple times, the performance of the model can be gradually optimized, thereby realizing the training of the deep learning model.
[0048] Figure 3 The forward process in the model training process is shown in the middle. Figure 3 The right side is the backward stage automatically generated by Pytorch. In order to reasonably control the appropriate timing of the execution of the calculation and communication stages in the model training process, the original computation graph of Pytorch can be directly split. Since the computation graph is destroyed, the corresponding backward calculation stage will not be generated. Therefore, the forward calculation stage can be saved. For example, as shown in Figure 3 As shown on the left side, when a certain sub-forward process (i.e., the communication or calculation of a certain stage in the forward process) is completed, the corresponding sub-backward process of the sub-forward process can be saved in the form of a stack. Until the entire forward process is completed, the backward processes are executed in turn according to the execution timing.
[0049] In the model training process, Pytorch automatically generates the corresponding Autograd computation graph for the forward process through the automatic differentiation mechanism of Autograd (automatic differentiation engine), so that automatic differentiation can be completed during the reverse process, ensuring the normal training of the model. The Autograd computation graph of Pytorch records and generates the forward computation flowchart through the hybrid execution technology, and then generates the reverse computation graph. However, for the use scenario of parallel execution of the forward process and the reverse process and overlapping calculation and communication, the single forward process needs to be cut into multiple sub-forward processes, and the corresponding multiple sub-reverse processes are obtained, but Pytorch itself does not support accurate division of the forward and reverse processes, and manual division has the problems of inconvenient use, easy to cause graph division error and difficult to debug, multiple storage of reverse results leading to doubling of the memory, and the like.
[0050] Therefore, the present disclosure proposes a way of automatically dividing the computation graph. An example of the automatic division principle is as follows.
[0051] In PyTorch, the computation graph is a directed acyclic graph composed of nodes and edges. The nodes in the computation graph represent Tensors or Functions, and the edges represent the dependency relationship between Tensors and Functions. Each Tensor object contains a pointer grad_fn pointing to the generating Function node, and the pointer grad_fn represents which function operation generates the tensor. In PyTorch, requires_grad is a very important attribute, which is used to identify whether a Tensor needs to retain gradient information in calculation, that is, whether it needs to calculate the gradient. When using PyTorch for deep learning model training, it is usually necessary to set requires_grad = True for the trainable parameters (i.e. trainable Tensors) in the model, so as to calculate the gradient during backpropagation. requires_grad is a unique attribute of the Tensor node, and the Function node does not need to calculate the gradient and does not have the requires_grad attribute.
[0052] Therefore, based on the attribute information of each node in the computation graph, a first node in the computation graph that does not need to calculate a gradient can be identified, the computation graph is split at the first node as a split position, and a plurality of sub-forward processes are obtained. The first node may be, for example, a Function node that does not have a requires_grad attribute. In the present disclosure, the operation represented by the Function node can include a calculation operation and a communication operation. For example, nodes corresponding to collective communication operations such as dispatch and combine in the training process of a hybrid expert model can all be regarded as Function nodes. Splitting the computation graph at the first node as a split position can mean splitting at the output of the first node (i.e., the Tensor node connected to the first node) as a breakpoint, and cutting the edge between the output of the first node and the subsequent node. Through the splitting, a plurality of sub-forward processes can be obtained, each sub-forward process corresponding to an execution stage, and the Function involved in the execution stage can be a calculation operation or a communication operation.
[0053] In PyTorch, each Function node has two methods, forward() and backward(), which are used for forward propagation to calculate loss and backward propagation to calculate gradient in the computation graph. When performing forward calculation, Autograd dynamically creates corresponding nodes by overloading various operations of Tensor (such as addition, multiplication, matrix multiplication, etc.), and connects them into a directed acyclic graph (DAG). At this time, the structure of the computation graph depends on the specific calculation path. When calling loss.backward(), the Autograd engine traverses the nodes from the end to the beginning according to the computation graph, and executes the backward() function of each node in turn to calculate the gradient. The gradient is passed through the chain rule, and automatic differentiation is achieved. Therefore, after splitting the computation graph into a plurality of sub-forward processes, a plurality of sub-backward processes corresponding to the plurality of sub-forward processes can be automatically obtained.
[0054] The output tensor of each sub-forward process is used as the input tensor of the next sub-forward process. In order to maintain the integrity of the sub-forward process, when splitting the computation graph, a new output tensor can be created for each output tensor of the first node. Each output tensor of the first node is the output of the sub-forward process corresponding to the first node, and the created new output tensor is the input of the next sub-forward process.
[0055] The following is an exemplary code for splitting the computation graph.
[0056] args_a_old: [tensor_1_old, int_2, tensor_nograd_3]
[0057] args_b_old:{'a':tensor_4_old,'b':float_5,'c':tensor_6_old}
[0058] args_a_new,args_b_new=break_graph(args_a_old,args_b_old)
[0059] args_a_new:[tensor_1_new,int_2,tensor_nograd_3]
[0060] args_b_new:{'a':tensor_4_new,'b':float_5,'c':tensor_6_new}
[0061] The above code can be described as follows: args_a_old and args_b_old are cut off because the former is calculated and the latter is communicated to these two args. args_a and args_b are output tensors, but the form can be more complex, such as a container structure like list or dict, and the objects in the list or dict need to be extracted. If it is a tensor with requires_grad=True (requires gradient calculation), a new tensor needs to be established through the detach operation. The new tensor has no connection with the previous part of the calculation graph. Then the new tensor is replaced in the original format to the position in args, so that the new args_a_new has no connection with the previous graph.
[0062] Figure 4A A structural diagram of a calculation graph is shown. Figure 4A The node a and the node b in the calculation graph are tensors, which correspond to args_a_old and args_b_old in the above code that need to be cut off.
[0063] Figure 4B An effect diagram of the calculation graph shown in Figure 4A after being cut off is shown.
[0064] Referring to Figure 4B, node a1 is equivalent to node a, i.e. args_a_old in the above code. Node b1 is equivalent to node b, i.e. args_b_old. Node a2 is the new tensor after detach operation on node a, i.e. args_a_new in the above code. Node b2 is the new tensor after detach operation on node b, i.e. args_b_new in the above code. The detach operation is used to return a new Tensor which shares the same memory space with the original Tensor, but will not be tracked by the computation graph, i.e. it will not participate in backpropagation. It should be noted that, Figure 4B Only a schematic diagram representing a logically desired effect is shown.
[0065] In order to reduce storage overhead, the disclosure proposes that, for a sub-forward process, a first fake tensor with a length of zero and a second fake tensor with a length of zero can be created. The first fake tensor is connected with the output tensor of the sub-forward process as the output of the sub-forward process; and the second fake tensor is connected with the new output tensor as the input of the next sub-forward process. There is no edge between the first fake tensor and the second fake tensor, so the computation graph can be broken, achieving the purpose of splitting the computation graph. In addition, the first fake tensor with a length of zero can be used as the starting point of the call of the corresponding sub-backward process of the sub-forward process, and there is no need to store the real tensor, so the storage (video memory) consumption can also be reduced.
[0066] Figure 4C A schematic diagram of the optimization of the computation graph shown in Figure 4B is shown.
[0067] Referring to Figure 4C , fake_1 corresponds to the first fake tensor mentioned above, and fake_2 corresponds to the second fake tensor mentioned above. fake_1 is the starting point of the sub-backward process corresponding to the sub-forward process represented in the upper half of Figure 4C . fake_2 is the end point of the sub-backward process corresponding to the sub-forward process represented in the lower half of Figure 4C . It should be noted that, Figure 4C The fake_1 at the end of the first graph needs to be stored for backpropagation, and the fake_2 of the second graph does not need to be stored, only the gradients of a2 and b2 need to be stored in the process of backpropagation from a2 and b2 to fake_2. Then when executing fake_1, the gradients of a2 and b2 need to be obtained from a2 and b2, and sent back to the edges from fake_1 to a1 and b1, so as to start Figure 4CThe sub-forward process represented in the upper half corresponds to the reverse process. That is, the two tensor nodes a1 and b1 are connected with fake_1, which makes fake_1 able to pass the gradient of a2 and b2 to a1 and b1 in the corresponding reverse process. Figure 4C In the sub-forward process represented in the lower half, fake_2 is connected with the two tensor nodes a2 and b2, which makes a2 and b2 able to pass their own gradients to fake_2 based on the connection relationship, and fake_2 stores the gradients of a2 and b2 for use in the reverse process from fake_1 to a1 and b1. There is no connection between fake_1 and fake_2, so the gradient saved by fake_2 can be passed to fake_1 by manual or another executor.
[0068] In step S220, after the forward process of the Nth batch of training samples is executed, and in the execution of at least part of the forward process of the Pth batch of training samples, at least one time period, the computing unit executes one of the sub-forward processes in the at least part of the forward process of the Pth batch of training samples and one of the sub-reverse processes in the multiple sub-reverse processes of the Nth batch of training samples in parallel, and the sub-forward process and the sub-reverse process executed in parallel, one of which belongs to a computing operation, and the other belongs to a communication operation.
[0069] The at least part of the forward process refers to the forward process that needs to be executed for the computing unit. In the model training process after the configuration is completed, the forward process and the corresponding reverse process executed by the same computing unit are determined. Moreover, the same computing unit can execute only part of the forward process configured for it, or can execute the complete forward process, which depends on the model parallel mode adopted. For example, in the pipeline parallel scenario, a single computing unit only executes part of the forward process, and in the tensor parallel scenario, a single computing unit can execute the entire forward process.
[0070] The execution of the at least part of the forward process of the Nth batch of training samples is prior to the execution of the at least part of the forward process of the Pth batch of training samples. The Nth batch of training samples and the Pth batch of training samples can be separated by 1 batch, 2 batches or more batches of training samples. Exemplarily, P = N + 1.
[0071] The reverse process is essentially an automatic differentiation process based on the chain rule, that is, the different sub-reverse processes in the reverse process of the same training batch have a sequence in execution. Therefore, although the sub-forward processes and the sub-reverse processes of different batches in the same computing unit are executed in parallel, the sub-reverse processes within the same batch are strictly executed in sequence, and the calculation logic will not be wrong.
[0072] For example, the at least partial forward process responsible by the computing unit can be divided into a computation phase and a communication phase alternately, and correspondingly, the backward process corresponding to the partial forward process can also be divided into a communication phase and a computation phase alternately. The computing unit can perform one computation phase and one communication phase in parallel in each step, wherein one is from the at least partial forward process for the P-th batch of training samples, and the other is from the backward process corresponding to the at least partial forward process for the N-th batch of training samples. Although the execution time of each of the computation phase and the communication phase performed in parallel in each step can not be completely consistent, the complete overlap of computation and communication in the training process can be achieved as a whole. Moreover, the problem that the execution time of each of the computation phase and the communication phase performed in parallel in each step can not be completely consistent can also be solved by corresponding means (for example, adjusting hyperparameters) to make the execution time of each of the computation phase and the communication phase performed in parallel in each step as consistent as possible.
[0073] It should be understood that the computing unit can also perform the above-mentioned computation graph splitting operation in performing the at least partial forward process of the P-th batch of training samples, so that the backward process corresponding to the at least partial forward process of the P-th batch of training samples can be performed in parallel with the at least partial forward process of other batches of training samples.
[0074] Figure 5 A flowchart of parallel execution of forward process and backward process is shown.
[0075] As Figure 5 shown, the mixed expert model can be split into five parts, which are a self-attention stage, a pre-processing stage before MLP, a dispatch communication stage, an MLP computation stage, and a combine communication stage. Each of these stages has a forward process and a backward process. After being split into two models (i.e., model 1 and model 2 in Figure 5 , the two models will run simultaneously. It should be understood that the simultaneous running of the two models refers to the simultaneous running for different training batches (i.e., different data), rather than the simultaneous running for the same training batch. For example, in the first time period, the forward process represented by model 1 can be performed for data a, and the backward process represented by model 2 can be performed for data b; in the subsequent second time period, the backward process represented by model 2 can be performed for data a, and the forward process represented by model 1 can be performed for data c; in the subsequent third time period, the forward process represented by model 1 can be performed for data d, and the backward process represented by model 2 can be performed for data c.
[0076] Taking an example of executing the forward process represented by model 1 for data a and synchronously executing the reverse process represented by model 2 for data b, the parallel running process of model 1 and model 2 can be divided into four steps. Step one: when model 1 (i.e. the forward process) runs, it first enters the preprocessing calculation stage before attention and MLP, while model 2 (i.e. the reverse process) first enters the reverse communication stage of combine. Step two: when model 1 completes the calculation, it enters the forward communication stage of dispatch, while model 2 enters the reverse calculation stage of MLP after completing the communication; step three: when model 1 completes the communication, it enters the forward calculation stage of MLP, while model 2 enters the reverse dispatch communication stage; step four: when model 1 completes the calculation, it enters the forward communication stage of combine, and model 2 enters the reverse MLP preprocessing and attention calculation stage after completing the communication.
[0077] In each step of the above implementation process, the two models will alternately be in a calculation stage and a communication stage, realizing complete overlap of calculation and communication and improving the use efficiency of the entire cluster computing resources and communication resources. Although, in each calculation and communication overlapping stage, the calculation time and the communication time may not be completely consistent, but by adjusting the hyperparameters (expert parallel number) and the number of resources allocated to communication and calculation resources, the complete overlap of calculation and communication time can be realized as much as possible, thereby reducing the overall time required for training, improving the training efficiency and the utilization rate of the computing unit.
[0078] So far, in combination with Figures 1 to 5 An exemplary description is made of how to fully mask the communication overhead to improve the efficiency of model training.
[0079] Another challenge for the traditional AI Infra training framework to efficiently train on a large-scale cluster is how to balance the use of video memory and the cost of calculation to improve the use efficiency of the cluster.
[0080] With the development of the era of large models, the size of the model is becoming larger and larger, although the memory of the new graphics cards launched by computing vendors is also increasing, but the growth of the memory of the graphics card is far from keeping up with the growth rate of the parameters of the large model. The excessive size of the model parameters will lead to the need to use more graphics cards for parallel computing, resulting in additional communication overhead in the training process. The traditional way is to use the re-computation strategy. In the forward process of large model training, the intermediate calculated activation values are saved, and then these activation values are used to calculate the gradient of the weight in the reverse process, so a large amount of memory fragments will be generated. Re-computation is not to save these activation values, but to re-compute them in the reverse process. Although such a way will lead to a decrease in computing efficiency, it can save memory overhead and reduce communication overhead. Therefore, how to balance the use of memory and computing cost is also one of the important challenges of the AI Infra training framework on large-scale clusters.
[0081] For the mixed expert model scenario, the present disclosure models and analyzes the memory usage and computing amount between the components in the mixed expert model, and accordingly proposes a more reasonable re-computation method to balance the use of memory and computing cost.
[0082] In the entire training process, the forward calculation result will be used in the gradient update operation of the weight matrix in the reverse direction, so a large number of forward calculation results need to be saved before and after calculation, resulting in a large memory overhead. The common way to solve this problem is to perform re-computation operation on the model, without saving the forward calculation result, but instead re-computing a forward calculation to generate the required output. However, such an operation will result in a lot of computational redundancy.
[0083] Figure 6 A re-computation part in the mixed expert model is shown.
[0084] Referring to Figure 6 The forward process of the mixed expert model includes a routing stage, a distribution stage, an MLP stage (i.e., a collective communication dispatch stage), and an aggregation stage (i.e., a collective communication combine stage).
[0085] For the MLP layer of the mixed expert model, first, it needs to be determined according to the gating network which expert needs to be assigned to a specific expert, and then sent to the corresponding process for calculation through the collective communication dispatch stage. Each expert will go through three steps of the first layer of fully connected matrix multiplication fcl (the first fully connected layer), activation function calculation, and the second layer of fully connected matrix multiplication fc2 (the second fully connected layer), and then the results are summarized through the collective communication combine stage.
[0086] Recompiling the entire model directly is equivalent to re-executing the entire forward pass, resulting in significant computational and communication costs. If the MLP is recomputed directly, it would lead to matrix multiplication by fc1 in the first fully connected layer, requiring the activation function to be preserved. Furthermore, two dispatch and combine phases are needed during the backward pass.
[0087] Considering the reverse process, the last layer in the MLP (corresponding to Figure 5 The output of the second fully connected layer in the MLP is useless for the ensemble communication combine phase. Therefore, for the last layer in the MLP, only the input tensor of that layer is needed for the gradient calculation. Thus, during the forward pass, the forward computation results of the MLP phase can be omitted, and during the backward pass, the MLP parts other than the last layer in the MLP phase are recomputed.
[0088] In other words, although the computation results of the MLP are not saved, when recompiling the MLP, it is considered that the computation result of the last layer in the MLP is meaningless. Therefore, only the part of the MLP excluding the last fully connected layer is recomputed. This reduces the recompiling overhead.
[0089] Generally, in the forward computation process, the dispatch phase is more time-consuming than the matrix multiplication fc1 and fc2 calculations. In addition, since the communication generates more GPU memory overhead, the results after the dispatch phase can be saved on each process, and the matrix multiplication fc1 and activation function can be recalculated to calculate the input for the gradient calculation of matrix multiplication fc2. The dispatch phase can be omitted in the backward process to reduce one communication operation. Although it will slightly increase the GPU memory consumption, it can reduce the amount of computation that needs to be recalculated and improve the overall computational efficiency of the training framework.
[0090] In some exemplary embodiments, the forward computation results of the distribution phase may not be saved during the forward process, and the distribution phase may be recomputed during the reverse process. That is, the distribution phase may also be recomputed.
[0091] It should be understood that whether to recompile the distribution stage can be determined by analyzing the relationship between the recompiling overhead of the distribution stage and the memory overhead of saving the calculation results of the distribution stage (i.e., the distribution results) based on the actual situation. When it is determined that the recompiling overhead of the distribution stage is less than or significantly less than the memory overhead, the distribution stage should be recomputed.
[0092] For example, if the number of selected experts in the MLP stage is greater than or equal to a first threshold, it indicates that the number of selected experts is large, and each expert needs to generate a memory overhead through the dispatch communication, and the memory overhead is large. In this case, the re-computation can be performed for the dispatch stage. Correspondingly, if the number of selected experts in the MLP stage is less than the first threshold, it indicates that the number of selected experts is small, and the memory overhead is small, and the re-computation does not need to be performed for the dispatch stage. For example, the first threshold can be set to 2.
[0093] The model training method of the present disclosure can also be implemented as a model training device. Figure 7 The structure diagram of the model training device according to an embodiment of the present disclosure is shown. The functional units of the model training device can be realized by hardware, software or a combination of hardware and software that implements the principles of the present disclosure. Those skilled in the art can understand that, Figure 7 The described functional units can be combined or divided into sub-units to realize the principles of the above-mentioned application. Therefore, the description herein can support any possible combination of the functional units described herein, or division, or further limitation. The following briefly describes the functional units that the reading device can have and the operations that each functional unit can perform. For the details involved, please refer to the related description above, which will not be repeated here.
[0094] Referring to Figure 7 The model training device 700 can include a segmentation module 710 and a parallel execution module 720.
[0095] The segmentation module 710 is configured to segment the dynamically generated computation graph into a plurality of sub-forward processes during the execution of at least part of the forward process for the Nth batch of training samples, and obtain a plurality of sub-backward processes corresponding to the plurality of sub-forward processes.
[0096] The parallel execution module 720 is configured to, after the forward process of the Nth batch of training samples is completed, and during the execution of at least part of the forward process of the Pth batch of training samples, utilize the computing unit to perform, in at least one time period, one of the sub-forward processes in the at least part of the forward process of the Pth batch of training samples and one of the sub-backward processes in the plurality of sub-backward processes of the Nth batch of training samples in parallel, and the sub-forward process and the sub-backward process performed in parallel, one of which belongs to a computing operation and the other of which belongs to a communication operation.
[0097] In some example embodiments, the segmentation module 710 can identify a first node in the computation graph that does not need to calculate the gradient based on the attribute information of each node in the computation graph, segment the computation graph at the first node as a segmentation position, and obtain a plurality of sub-forward processes.
[0098] In some further example embodiments, the splitting module 710 creates a new output tensor for each output tensor of the first node, each output tensor of the first node being an output of a sub-forward process corresponding to the first node; and uses the new output tensor as an input of a next sub-forward process.
[0099] In some further example embodiments, the splitting module 710 can connect the first pseudo tensor as an output of the sub-forward process with an output tensor of the sub-forward process; and use the second pseudo tensor as an input of a next sub-forward process.
[0100] In some example embodiments, the model is a hybrid expert model, the forward process of the hybrid expert model including a routing stage, a distribution stage, an MLP stage, and an aggregation stage, the model training apparatus 700 does not save the forward calculation result of the MLP stage in the forward process, and in the backward process, re-computes for the MLP part other than the last layer in the MLP stage.
[0101] In some further example embodiments, in the forward process, the model training apparatus 700 does not save the forward calculation result of the distribution stage, and in the backward process, re-computes for the distribution stage.
[0102] The present disclosure also proposes a distributed training system of a model, including a plurality of computing units, wherein the computing units split a dynamically generated computation graph into a plurality of sub-forward processes and obtain a plurality of sub-backward processes corresponding to the plurality of sub-forward processes in performing at least part of a forward process for an Nth batch of training samples; the computing units perform one of the sub-forward processes in the at least part of the forward process of a Pth batch of training samples and one of the sub-backward processes of the Nth batch of training samples in parallel in at least one time period after the forward process of the Nth batch of training samples is completed and in performing the at least part of the forward process of the Pth batch of training samples, and the sub-forward process and the sub-backward process performed in parallel, one of which belongs to a computing operation and the other of which belongs to a communication operation. Details of operations that can be performed by the computing units can be referred to the above related description.
[0103] Figure 8 A structural schematic diagram of a computing device according to an embodiment of the present disclosure is shown, which can be used to implement the above model training method.
[0104] Referring to Figure 8 , the computing device 800 includes a memory 810 and a processor 820.
[0105] The processor 820 can be a single core processor or a multiple core processor. In some embodiments, the processor 820 can include a general purpose processor and one or more special purpose processors, such as graphics processors (GPUs), digital signal processors (DSPs), etc. In some embodiments, the processor 820 can be implemented using custom circuitry, such as an application specific integrated circuit (ASIC) or a field programmable gate array (FPGA).
[0106] The memory 810 can include various types of storage units, such as a system memory, a read-only memory (ROM), and a permanent storage device. The ROM can store static data or instructions required by the processor 820 or other modules of the computer. The permanent storage device can be a read and write memory device. The permanent storage device can be a non-volatile storage device that does not lose stored instructions and data even when the computer is powered off. In some embodiments, the permanent storage device uses a mass storage device (e.g., a magnetic or optical disk, a flash memory) as a permanent storage device. In some other embodiments, the permanent storage device can be a removable storage device (e.g., a floppy disk, an optical disk). The system memory can be a read and write memory device or a volatile read and write memory device, such as a dynamic random access memory. The system memory can store some or all of the instructions and data required by the processor during runtime. In addition, the memory 810 can include a combination of any computer readable storage media, including various types of semiconductor memory chips (DRAM, SRAM, SDRAM, flash memory, programmable read-only memory), magnetic disks and / or optical disks. In some embodiments, the memory 810 can include a read and / or write removable storage device, such as a compact disc (CD), a read-only digital versatile disc (e.g., DVD-ROM, dual-layer DVD-ROM), a read-only Blu-ray disc, an ultra density disc, a flash memory card (e.g., SD card, min SD card, Micro-SD card, etc.), a magnetic floppy disk, etc. The computer readable storage media does not include a carrier wave and a transient electronic signal transmitted through a wire or wireless transmission.
[0107] The memory 810 stores executable code, which, when processed by the processor 820, can cause the processor 820 to perform the model training method described above.
[0108] The model training method, device, system, and apparatus according to the present disclosure have been described in detail above with reference to the accompanying drawings.
[0109] Furthermore, the method according to the present application can also be implemented as a computer program or computer program product comprising computer program code instructions for executing the above steps defined in the above method of the present application.
[0110] Alternatively, the present application can also be implemented as a non-transitory machine readable storage medium (or computer readable storage medium, or machine readable storage medium) having stored thereon executable code (or computer program, or computer instruction code) which, when executed by a processor of an electronic device (or computing device, server, etc.), causes the processor to carry out the steps of the above method according to the present application.
[0111] Those skilled in the art will further appreciate that the various illustrative logical blocks, modules, circuits, and algorithm steps described in connection with the disclosure herein can be implemented as electronic hardware, computer software, or combinations of both.
[0112] The flow diagrams and block diagrams in the drawings are representative of the architecture, functionality, and operation of possible implementations of systems and methods according to the present application. In this regard, each block in the flow diagrams and block diagrams can represent a module, segment, or portion of code, which comprises one or more executable instructions for implementing the specified logical functions. It should also be noted that in some alternative implementations, the functions noted in the blocks can occur out of the order noted in the figures. For example, two blocks shown in succession may, in fact, be executed substantially concurrently or the blocks can sometimes be executed in the reverse order, depending upon the functionality involved. It will also be noted that each block of the block diagrams and / or flowchart illustrations, and combinations thereof, can be implemented by special purpose hardware-based systems that perform the specified functions or operations, or combinations of special purpose hardware and computer instructions.
[0113] Embodiments of the present application have been described above with the aid of example illustrations and are not limited to the embodiments disclosed. The description above is illustrative and not restrictive. Many modifications and changes can become apparent to the skilled person. The present application is not limited to the disclosed embodiments but encompasses every novel solution falling within the scope of the appended claims. The choice of terms is intended to best explain the principle of the embodiments, the practical application or improvement of the technology in the market, or to enable other ordinary skilled persons in the art to understand the embodiments disclosed herein.
Claims
1. A model training method, comprising: in executing at least part of a forward process for an Nth batch of training samples, splitting a dynamically generated computation graph into a plurality of sub-forward processes, and obtaining a plurality of sub-backward processes corresponding to the plurality of sub-forward processes; after the forward process of the Nth batch of training samples is executed, and in executing the at least part of the forward process for a Pth batch of training samples, using a computing unit to execute one of the at least part of the forward process for the Pth batch of training samples and one of the plurality of sub-backward processes for the Nth batch of training samples in parallel in at least one time period, and the sub-forward process and the sub-backward process executed in parallel, one of which belongs to a computing operation and the other belongs to a communication operation.
2. The method of claim 1, wherein, splitting the dynamically generated computation graph into the plurality of sub-forward processes, comprising: identifying a first node in the computation graph that does not need to compute a gradient based on attribute information of each node in the computation graph; splitting the computation graph at the first node to obtain the plurality of sub-forward processes. 3.The method of claim 2, wherein splitting the computation graph at the first node comprises: creating a new output tensor for each output tensor of the first node; and using the new output tensor as an input of a next sub-forward process. 4.The method of claim 3, further comprising: for the sub-forward process, creating a first pseudo tensor with a length of zero and a second pseudo tensor with a length of zero; using the first pseudo tensor as an output of the sub-forward process and connecting the output tensor of the sub-forward process; and using the second pseudo tensor as an input of a next sub-forward process and connecting the new output tensor.
5. The method of claim 1, wherein, The model is a hybrid expert model, and a forward process of the hybrid expert model comprises a routing stage, a distribution stage, a multi-layer perceptron (MLP) stage, and an aggregation stage, and the method further comprises: in the forward process, not saving a forward calculation result of the MLP stage, and in the backward process, re-computing an MLP part other than a last layer in the MLP stage.
6. The method of claim 5, wherein, The method further comprises: in the forward process, not saving a forward calculation result of the distribution stage, and in the backward process, re-computing the distribution stage. 7.A model training apparatus, comprising: a splitting module configured to split a dynamically generated computation graph into a plurality of sub-forward processes in executing at least part of a forward process for an Nth batch of training samples, and obtain a plurality of sub-backward processes corresponding to the plurality of sub-forward processes; a parallel execution module configured to, after the forward process of the Nth batch of training samples is executed, and in executing the at least part of the forward process for a Pth batch of training samples, use a computing unit to execute one of the at least part of the forward process for the Pth batch of training samples and one of the plurality of sub-backward processes for the Nth batch of training samples in parallel in at least one time period, and the sub-forward process and the sub-backward process executed in parallel, one of which belongs to a computing operation and the other belongs to a communication operation.
8. A distributed training system of a model, comprising a plurality of computing units, wherein, the computing unit splits a dynamically generated computation graph into a plurality of sub-forward processes and obtains a plurality of sub-backward processes corresponding to the plurality of sub-forward processes in performing at least part of a forward process for an Nth batch of training samples; the computing unit, after the forward process of the Nth batch of training samples is completed, and in performing at least part of a forward process of a Pth batch of training samples, performs one of the at least part of the forward process of the Pth batch of training samples and one of the plurality of sub-backward processes of the Nth batch of training samples in parallel in at least one time period, and the one of the sub-forward process and the sub-backward process, one of which belongs to a computing operation and the other of which belongs to a communication operation, are performed in parallel.
9. A computing device, comprising: a processor; and a memory having stored thereon executable code that, when executed by the processor, causes the processor to perform the method of any one of claims 1 to 6.
10. A computer program product, comprising executable code that, when executed by a processor of an electronic device, causes the processor to perform the method of any one of claims 1 to 6.
11. A non-transitory machine-readable storage medium having stored thereon executable code that, when executed by a processor of an electronic device, causes the processor to perform the method of any one of claims 1 to 6.
Citation Information
Patent Citations
Model training method, device and equipment based on pipeline parallelism
CN113177632A
Large model assembly line parallel training method and system based on gradient sensing parameter freezing
CN118568499A
Large model training method and system supporting parallel hot switching
CN119558371A
Method, electronic device, and computer program product for scheduling computing resources
US20230297420A1
Forward-style Gradient GeMMs
US20230385077A1
Cited By
Parallel training strategy online switching method and deep learning model training system
CN122021969A
Information processing device, information processing method, and program.
JP7901949B1