Assembly line parallel training method for multi-modal large model
Through taboo search and simulated annealing algorithm, the pipeline parallel training of multimodal large models is optimized, and the recomputation layer is selected in combination with the greedy method, which solves the problem of unbalanced computing load and video memory usage, and realizes an efficient training process.
Patent Information
- Application Number
- CN202510639216.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-19
- Publication Date
- 2025-06-20
- Estimated Expiration
- 2045-05-19
AI Technical Summary
The existing pipeline parallel training method is difficult to balance the computing load and video memory usage in multimodal model training, resulting in inefficient training.
The taboo search algorithm is used combined with simulated annealing algorithm to optimize the pipeline stage division of the model, and select the recomputation layer based on the greedy method to achieve the balance between computing load and video memory usage.
Through the optimization phase division and recomputation strategies, training efficiency is significantly improved, unnecessary computing overhead is reduced, and video memory usage is effectively controlled.
Smart Images

Figure CN120179416A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to model training technology, and specifically to a pipeline parallel training method for multi-modal large models. Background Art
[0002] Multi-modal models can integrate data of multiple modalities such as images and texts, and have achieved remarkable results in the field of artificial intelligence. However, with the increase in the scale of the model and the sequence of data, the demand for video memory and computing further increases, and an efficient distributed training method is the key to realizing large-scale model training.
[0003] Currently, the commonly used distributed training methods are mainly data parallelism and model parallelism. Single data parallelism cannot be used when the scale of model parameters exceeds the video memory capacity of a single device, and model parallelism needs to be combined. Model parallelism includes pipeline parallelism and tensor parallelism. Pipeline parallelism calculates by splitting the layers of the model onto different devices. In the mainstream pipeline parallel scheduling method 1F1B, the devices in the i-th stage of the p stages of the pipeline will retain the intermediate calculation results of p-i micro-batches, resulting in unbalanced video memory occupancy and computing load between devices. For multi-modal models composed of different modalities, this imbalance problem is more prominent.
[0004] The current pipeline parallel splitting technologies of mainstream frameworks and technologies are usually implemented based on 1F1B scheduling. In the splitting algorithms of these methods, only computing time or a simple parameter equalization strategy is mainly considered, and the choice for recomputation is only coarse-grained full-scale recomputation, which will greatly increase the training time; it is difficult to fully adapt to the differences in computing resources and video memory requirements between different modality sub-modules in multi-modal model training, does not consider the differences between different modality modules, and cannot efficiently train the model. Summary of the Invention
[0005] Object of the Invention: Aiming at the above problems, the object of the present invention is to provide a pipeline parallel training method for multi-modal large models, which determines the pipeline stage division of the model through a tabu search algorithm, and applies a recomputation strategy based on a greedy method to each segmentation point to determine the selection of the pipeline splitting stage and the recomputation layer, so as to achieve balanced optimization of computing load and video memory usage.
[0006] Technical Solution: A pipeline parallel training method for a multi-modal large model of the present invention includes the following steps: Calculate the execution time and video memory occupancy data of each layer in the multi-modal large model; According to the execution time and video memory occupancy data of each layer, calculate the weight assigned to each layer in the multi-modal large model, calculate the cumulative value based on the weight, and evenly distribute it according to the number of devices to obtain the initial stage division result; Optimize the initial stage partitioning result using an improved tabu search method to obtain an optimized solution, and use this solution as the final partitioning result of the multi-modal large model. Use this final partitioning result to perform pipeline parallel training on the multi-modal large model.
[0007] Further, the process of obtaining the initial stage partitioning result includes: Calculate the weight assigned to each layer of the multi-modal large model. The formula is: , In the formula, is the weight of the k-th layer, and represent the adjustment coefficients of execution time and video memory occupancy data respectively, and ; represents the fixed video memory occupancy of the k-th layer, represents the total execution time of the k-th layer; Accumulate the weights of each layer in order of the layers , when the accumulated weight first exceeds of the total weight, , divide the first layer to the current layer into the first stage , and so on, divide into p stages, denoted as , and assign a GPU device to each stage.
[0008] Further, the steps of optimizing the initial stage partitioning result using the improved tabu search method include: Step 301, take the initial stage partitioning result as the current solution, clear the tabu list, and set the maximum adjustment step size ; Step 302, generate neighborhood solutions: For the current solution, adjust the split point b between adjacent stages S and S i and S i+1 with step size to generate a set of neighborhood solutions; Step 303, for each neighborhood solution in the set of neighborhood solutions, calculate the video memory occupancy of each GPU device in the corresponding partitioning scheme; if it does not exceed the upper limit of the GPU device, jump to Step 304; otherwise, use the greedy method to recalculate and optimize this neighborhood solution; Step 304, evaluate the solution quality, and select the optimal solution according to the cost; Step 305, select solutions: Calculate the cost of the current neighborhood solution. If the cost of the current neighborhood solution is lower than the cost of the optimal solution, take the current neighborhood solution as the new optimal solution and put the current neighborhood solution into the tabu list; if the cost of no neighborhood solution is lower than the cost of the current optimal solution, then according to the probability Select the sub-optimal solution and then put the sub-optimal solution into the taboo list; where P is a randomly generated value between 0 and 1, and T is the current temperature parameter; Step 306, update the taboo list: If no better solution is found after several consecutive iterations, trigger the diversification operation, raise the temperature T to , clear half of the taboo list, and randomly generate the initial solution; otherwise, reduce the temperature to according to the exponential cooling strategy; where, is the preset upper limit of the temperature, is the temperature increase coefficient, is the temperature decrease coefficient; Step 307, repeat the iterative steps from Step 302 to Step 306 until the termination condition is met to obtain the optimized stage division scheme.
[0009] Furthermore, Step 302 includes: Update the step size according to the temperature T , and use the updated step size to move the splitting point b to the right to , or move it to the left to to generate a new division scheme and add it to the neighborhood solution set; where the update formula of the step size is: .
[0010] Furthermore, in Step 303, the steps of re-calculating and optimizing the neighborhood solution by using the greedy method include: Calculate the video memory benefit and the re-calculation cost of each layer, that is, the time of the forward propagation of the k-th layer. For the k-th layer, the video memory benefit is defined as the activated value video memory saved after enabling re-calculation, and the re-calculation cost is the time of re-executing the forward propagation; Calculate the ratio of the video memory benefit to the re-calculation cost, and the formula is: ; Sort all the layers of the current stage in descending order according to the MCR value, and preferentially select the layers with a high ratio to enable re-calculation; Finally, output the list of re-calculation layer indexes of each stage.
[0011] Furthermore, in Step 303, the video memory calculation formula is as follows: , In the formula, and are the starting layer index and the ending layer index of the stage respectively, k is the stage The layer index in satisfies , represents the fixed video memory occupancy, is the recomputation variable, 1 means enabling recomputation, and 0 means disabling it.
[0012] Furthermore, step 304 includes: Construct a cost function Cost with the goal of minimizing the maximum stage training time, and the formula is: , In the formula, is the total training time of stage S i , which consists of the basic calculation time and the additional time introduced by recomputation, and the formula is: , In the formula, represents the total execution time of the k-th layer; Calculate the cost of each neighborhood solution, and select the neighborhood solution corresponding to the lowest cost as the optimal solution.
[0013] Furthermore, in step 1, the formula for calculating the execution time of each layer is: , In the formula, respectively represent the execution time of the k-th layer of the multi-modal large model during backpropagation.
[0014] Beneficial effects: Compared with the prior art, the significant advantages of the present invention are: 1. The present invention adopts a tabu search combined with a simulated annealing algorithm. Through the search optimization strategy, it ensures to find a better balance between video memory usage and computing load, thereby improving the overall training efficiency; 2. The present invention dynamically selects the layers that need to be recomputed based on video memory, and preferentially selects the layers with high video memory benefits and low computing costs. Compared with the coarse-grained full-scale recomputation method, the fine-grained recomputation strategy of the present invention significantly reduces unnecessary computing overhead, effectively controls video memory occupancy, not only improves the training efficiency, but also optimizes resource utilization. Description of the Drawings
[0015] Figure 1 is a flowchart of a pipeline parallel training method for a multi-modal large model; Figure 2 is a schematic diagram of a pipeline splitting framework; Figure 3 is the video memory occupancy of each stage of the CLIP model with a batch size of 4; Figure 4 is the video memory occupancy of each stage of the CLIP model with a batch size of 8; Figure 5 For the acceleration of the experimental model when the batch size is 4; Figure 6 For the acceleration of the experimental model when the batch size is 8. Specific implementation manner
[0016] In order to make the objectives, technical solutions and advantages of the present application clearer and more understandable, the present application will be further described in detail below with reference to the accompanying drawings and embodiments.
[0017] A pipeline parallel training method for a multimodal large model described in this embodiment, the flowchart is as Figure 1 shown, and the method includes the following steps: Step 1, calculate the execution time and video memory occupancy data of each layer in the multimodal large model.
[0018] In the preprocessing stage, analyze the levels of the multimodal large model to obtain the calculation time and video memory requirement data of each layer, providing a basis for subsequent pipeline splitting and video memory optimization. First, perform time analysis to quantify the computational overhead of each pipeline stage. The execution time of each layer in the multimodal large model can be measured through the Profiling and hook functions of the framework tool, and the basic computational time of the stage can be calculated.
[0019] Furthermore, the formula for calculating the execution time of each layer is: , wherein, and respectively represent the execution time of the k-th layer of the multimodal large model in the forward propagation and backward propagation processes.
[0020] Video memory requirement analysis is used to evaluate the video memory occupancy of each stage to evaluate the video memory occupancy of each stage. The purpose of video memory requirement analysis is to provide data support for subsequent pipeline splitting and recomputation strategies, ensuring that the training process runs efficiently within the device video memory limit.
[0021] Step 2, according to the execution time and video memory occupancy data of each layer, calculate the weight assigned to each layer of the multimodal large model, calculate the cumulative value based on the weight, and evenly distribute it according to the number of devices to obtain the initial stage division result.
[0022] In the initialization stage of the algorithm, it is necessary to generate a reasonable division scheme as the starting point for searching. For this purpose, a hybrid cost allocation strategy is designed. There are differences in the calculation time and video memory requirements of the layers of the multimodal large model, so the weighted sum of calculation and memory occupancy is used to calculate the weight assigned to each layer of the model. Based on the calculated weight, calculate the cumulative value and evenly distribute it according to the number of devices to determine the layers divided in each stage and generate the initial division result .
[0023] Further, the process of obtaining the initial stage division result includes: Calculating the weights assigned to each layer of the multimodal large model, with the formula: , In the formula, is the weight of the k-th layer, and respectively represent the adjustment coefficients of the execution time and the video memory occupancy data, and ; represents the fixed video memory occupancy of the k-th layer, and the fixed video memory occupancy consists of model parameters, model gradients, and optimizer parameters, represents the total execution time of the k-th layer; Accumulate the weights of each layer in sequence according to the layer order , when the accumulated weight first exceeds of the total weight, , divide the first layer to the current layer into the first stage , and so on, divide into p stages, denoted as , and assign a GPU device to each stage.
[0024] Step 3: Use the improved tabu search method to optimize the initial stage division result to obtain an optimized solution, and use this solution as the final division result of the multimodal large model, and use this final division result to perform pipelined parallel training on the multimodal large model.
[0025] To achieve efficient training model splitting of the multimodal large model, an optimized tabu search method is selected. Tabu search can find an optimized stage division result in the complex layer assignment through the search mechanism. After the initial division scheme is generated, the stage division assignment is performed by finding the neighborhood solution, and the neighborhood solution is a new scheme obtained by adjusting the pipeline division stage. In order to optimize the search space and avoid the search range always staying at the local optimal solution, in this embodiment, the idea of simulated annealing is combined in the tabu search, and the temperature parameter T is introduced to enhance the search range. The temperature is used to control the number of candidate solutions and the adjustment range. When the temperature is high, the search range is larger, and there will be more candidate solutions, which is beneficial to finding the global optimal solution and avoiding being unable to jump out of the local optimal solution too early. When the temperature is low, the number of candidate solutions will be reduced, mainly focusing on the local optimization of the current solution. Adjust the split points of adjacent stages to generate a new scheme. This method improves the flexibility of the tabu search and can find a better division scheme.
[0026] Further, the steps of using the improved tabu search method to optimize the initial stage division result include: Step 301: Take the initial stage division result as the current solution, clear the taboo list, and set the maximum adjustment step size. ; Step 302: Generate neighborhood solutions. For the current solution, adjust the splitting point b between adjacent stages S and S i and S i+1 with the step size to generate a set of neighborhood solutions. Step 303: For each neighborhood solution in the set of neighborhood solutions, calculate the video memory occupancy of each GPU device in its corresponding division scheme. If it does not exceed the upper limit of the GPU device, jump to Step 304; otherwise, re-calculate and optimize the neighborhood solution using a greedy method. Step 304: Evaluate the solution quality and select the optimal solution according to the cost. Step 305: Screen the solution. Calculate the cost of the current neighborhood solution. If the cost of the current neighborhood solution is lower than the cost of the optimal solution, take the current neighborhood solution as the new optimal solution and put the current neighborhood solution into the taboo list; if the cost of no neighborhood solution is lower than the cost of the current optimal solution, select the sub-optimal solution according to the probability P < T, and then put the sub-optimal solution into the taboo list; where P is a randomly generated value between 0 and 1, and T is the current temperature parameter. Step 306: Update the taboo list. If no better solution is found after several consecutive iterations, trigger the diversification operation, increase the temperature T to , clear half of the taboo list, and randomly generate an initial solution; otherwise, decrease the temperature according to the exponential cooling strategy ; where is the preset upper limit of the temperature, is the temperature increase coefficient, such as 3.0; is the temperature decrease coefficient, such as 0.97. Step 307: Repeat Steps 302 to 306 until the termination condition is met to obtain the optimized stage division scheme.
[0027] To prevent searching for duplicate solutions and getting stuck in a local loop due to repeated searches, a taboo list is used to record a previous segment of the executed solution. The initial value of the taboo list is set to to balance the previous search results and the exploration range, and dynamically adjust the length of the taboo list according to the search status. If a better solution is found, it indicates that the current search direction is effective, and the length of the list is shortened to accelerate local optimization. On the contrary, if no better solution is found for a long time, it is necessary to jump out of the current area and increase the length of the list to increase the search range.
[0028] Furthermore, Step 302 includes: Update the step size according to the temperature T, and use the updated step size to move the splitting point b to the right to , or move left to , generate a new partitioning scheme and add it to the neighborhood solution set; where the step size is updated according to the formula: .
[0029] To ensure that the video memory occupancy in each stage meets the device limit, a greedy method is adopted in this example for recomputation optimization. This strategy is called before evaluating each candidate solution to dynamically adjust the recomputation layer allocation in each stage, thereby optimizing video memory usage and maintaining the balance of the computational load. By analyzing the video memory gain and recomputation cost of each layer, the algorithm can minimize the additional computational overhead as much as possible while meeting the video memory upper limit, achieving efficient resource allocation. If the initial video memory does not exceed the device upper limit, no recomputation is required. Otherwise, analyze the recomputation value of each layer. For the k-th layer, the video memory gain is defined as the video memory saved for the activation values after enabling recomputation. In the 1F1B scheduling, the activation values increase with the stage depth . The recomputation cost is the time for re-executing the forward propagation, that is, the ratio of the video memory gain to the recomputation cost.
[0030] Furthermore, in step 303, the steps of adopting a greedy method to perform recomputation optimization on the neighborhood solution include: Calculate the video memory gain of each layer and the recomputation cost , that is, the time for the forward propagation of the k-th layer. For the k-th layer, the video memory gain is defined as the video memory saved for the activation values after enabling recomputation. In the 1F1B scheduling, the activation values increase with the stage depth . The formula is: , where is the size of the intermediate result of the k-th layer, including the output tensor and the internal tensor, p represents the total number of pipeline stages, and i represents the stage in which the k-th layer is partitioned; the recomputation cost is the time for re-executing the forward propagation; Calculate the ratio of the video memory gain to the recomputation cost, and the formula is: ; Sort all the layers in the current stage in descending order according to the MCR value, and preferentially select the layers with a high ratio to enable recomputation; Finally, output the recomputation layer index list for each stage. The recomputation layer index list specifies which layers need to generate intermediate results through recomputation during the training process and is used for video memory calculation.
[0031] Furthermore, in step 303, the video memory calculation formula is as follows: , In the formula, and are the starting layer index and the ending layer index of stage respectively, k is the layer index in stage , satisfying , represents the fixed video memory occupancy, which consists of model parameters, model gradients, and optimizer parameters. is the recomputation variable, 1 means enabling recomputation, and 0 means not enabling.
[0032] After generating candidate solutions, it is necessary to evaluate the quality of the solutions to select the current optimal solution. Before evaluating the current solution, it is necessary to first call the video memory-aware recomputation strategy, that is, to ensure that the solution meets the training constraints before evaluating the solution. The optimization goal is to minimize the training iteration time. In the 1F1B scheduling of pipeline parallelism, the training time is determined by the calculation time of the longest stage, that is, the optimization goal can be transformed into minimizing the maximum calculation time of each stage. Then, according to the cost, select the solution with the lowest cost, and accept the sub-optimal solution with a temperature-related probability to expand the search space and avoid premature convergence.
[0033] Furthermore, step 304 includes: Construct a cost function Cost with the goal of minimizing the maximum stage training time. The formula is: , In the formula, is the total training time of stage S i , which consists of the basic calculation time and the additional time introduced by recomputation. The formula is: , In the formula, represents the total execution time of the k-th layer; Calculate the cost of each neighborhood solution, and filter out the neighborhood solution corresponding to the lowest cost as the optimal solution.
[0034] After evaluating the cost of the solution, enter the iterative optimization stage, and find a better partitioning solution through continuous search. Set the maximum number of iterations during the search process, which can not only ensure sufficient exploration of the solution space but also limit the longest search time to speed up the algorithm. Then introduce a termination condition, that is, when no better solution is found after the number of iterations reaches the preset number, the search is terminated in advance to avoid ineffective search. At the start of the search, set the temperature parameter to T. The T for finding the neighborhood solution previously is passed in here, and the initial value of T is set to 1.0, which is used to control the entire iterative process. By continuously reducing the temperature value, the algorithm has the ability to explore a wide range of solution spaces in the initial stage and gradually performs local search optimization as the iteration progresses. Filter solutions based on the cost calculation of the neighborhood solution. If the cost of the current solution is lower than the current best solution, update the current partitioning scheme and update the tabu list according to the improvement amplitude. If there is no solution with a cost lower than the current one, accept the sub-optimal solution with a probability P < T. P is a randomly generated probability. This mechanism dynamically adjusts the acceptance rate using temperature and the current value to ensure search diversity in the initial stage. To avoid falling into local optima, when the continuous iteration exceeds a certain number of times, trigger the set diversity operation, increase the temperature, clear half of the tabu list, and then generate an initial solution. Through this iterative optimization mechanism, the algorithm can gradually approach the global optimal solution in the complex layer allocation of the multi-modal model, ensuring continuous improvement in the balance of computational load and video memory occupancy.
[0035] The splitting results can be evaluated, and the evaluation metrics include the single-iteration time, the peak video memory of each stage, the proportion of the computing time of each stage, and the GPU utilization rate. The single-iteration time is defined as the average time taken to complete one training batch; the peak video memory of each stage is the maximum video memory occupancy of each pipeline stage; the proportion of the computing time of each stage is the percentage of the computing time of this stage in the total training time. The GPU utilization rate is the average usage rate of the GPU during training. The calculation formulas for each metric are as follows: , , , where IT represents the single-iteration time, Total Time is the total duration of completing all batch iterations, is the total number of batches; represents the proportion of the computing time of the i-th stage, is the sum of the computing times of all stages, GPU Utilization represents the GPU utilization rate, Model FLOPs is the number of floating-point operations executed by the model, FLOPS is the number of floating-point operations per second of the GPU, and Time is the total training time.
[0036] Figure 2This is the overall framework and optimization process for the pipeline parallel training of multimodal large models in the present invention, including the technical process in the upper part and the sample diagram in the lower part. The upper part starts from the allocation of the model and hardware configuration, and through analyzing performance data including computing time and video memory occupancy, pipeline splitting optimization, and recomputation strategy, finally outputs the optimized stage division plan; the lower part shows that after the multimodal data is input, the multimodal model is divided into several stages, such as stage 0 to stage 3, and allocated to devices for model training.
[0037] The present invention relates to the field of distributed training of multimodal large models. Specifically, for large deep learning models that need to process multiple modal data such as images and texts simultaneously, an efficient pipeline parallel training method is proposed. It is applied to the distributed training process of multimodal large models on a multi-GPU cluster. The specific scenario is to train a model with more than ten billion parameters and process multiple types of data simultaneously, such as image and text data. The trained multimodal large model can be used to process images and texts.
[0038] To further illustrate the effectiveness and superiority of the pipeline parallel training method for a multimodal large model described in the present invention, the following examples are used for illustration. In the examples, two representative multimodal large models, CLIP and COCA, are selected as experimental models. These two experimental models are divided by the method described in this example. Among them, CLIP is an image-text retrieval model, and COCA is used for generative image-text tasks. The model size is about 20 billion parameters. The specific parameter configurations are shown in Table 1 below.
[0039] Table 1
[0040] This example is implemented on a cluster consisting of 2 machines. Each machine has 8 H800 80G graphics cards. The CPU is 2*Intel(R)Xeon(R)Platinum 8468v. The GPUs are connected through NVLinks. The software environment is CUDA 11.8 and PyTorch 2.4.1. PipeDream is used as the baseline. PipeDream is a generalized pipeline parallelism system for deep neural network training, aiming to improve training efficiency through pipeline parallel processing. PipeDream uses dynamic programming to divide the model and adds equal-split and full recomputation settings for comparison. The iteration time of the two models, CLIP and COCA, and the video memory occupancy of each stage are measured to evaluate the video memory utilization rate and acceleration.
[0041] By Figure 3 and Figure 4The video memory occupancy of each stage of the CLIP model with a batch size of 4 and 8, which is 19.7B, is shown. The black line represents the upper limit of the GPU video memory. The stages exceeding the upper limit of the video memory are the evaluation values. "even" is the parameter equal division provided by the distributed framework by default. "even-ful" is based on "even" plus full recomputation. "PipeDream-ful" is PipeDream plus full recomputation. The method described in the present invention is denoted as Mpipe. It can be seen from the figure that the GPU video memory utilization is significantly improved compared with the baseline. The single-node GPU utilization rate in the 4-batch case exceeds 95% at most. The video memory utilization rate in the 8-batch case exceeds the baseline by more than 15GB on average. Since more activation values need to be saved in the previous stages, the video memory in the previous stages is significantly higher than that in the subsequent stages.
[0042] Figures 5 to 6 The iteration time speedup ratios of the 19.7B CLIP and COCA models indicate that the method described in the present invention can achieve an acceleration of up to 1.12 times and 1.17 times in the training time in different batches. In the figure, OOM represents out-of-memory. Based on the method described in the present invention, the video memory and calculation can be balanced by combining the video memory optimization according to the calculation cost. The calculation efficiency is better than equal division and PipeDream, indicating that the iteration time of the method proposed in the present invention is better than the baseline.
Claims
1. A pipeline parallel training method for a multimodal large model, characterized in that: The steps include: Calculate the execution time and memory usage of each layer in a large multimodal model; According to the execution time and memory usage data of each layer, the weight assigned to each layer of the multimodal large model is calculated, the cumulative value is calculated based on the weight, and evenly distributed according to the number of devices to obtain the initial stage division result; An improved taboo search method is used to optimize the initial stage division results to obtain an optimized solution, which is used as the final division result of the multimodal large model. The final division result is used to perform pipeline parallel training on the multimodal large model.
2. The pipeline parallel training method for a multimodal large model according to claim 1, characterized in that: The process of obtaining the initial stage division results includes: Calculate the weights assigned to each layer of the multimodal large model using the formula: , In the formula, is the weight of the kth layer, and Respectively represent the adjustment coefficients of execution time and video memory usage data, and ; Indicates the fixed video memory occupancy of the kth layer, represents the total execution time of the kth layer; Accumulate the weights of each layer in order , when the accumulated weight exceeds the total weight for the first time hour, , the first layer to the current layer is divided into the first stage , and so on, it is divided into p stages, recorded as , each stage is assigned a GPU device.
3. The pipeline parallel training method for a multimodal large model according to claim 2, characterized in that: The steps of optimizing the initial stage partitioning results using the improved tabu search method include: Step 301: Use the initial phase division result as the current solution, clear the taboo table, and set the maximum adjustment step length. ; Step 302, generate neighborhood solutions: for the current solution, use a step size of Adjust the adjacent stage S i and S i+1 The segmentation point b generates a neighborhood solution set; Step 303, for each domain solution in the neighborhood solution set, calculate the video memory occupancy of each GPU device in the corresponding partitioning scheme; if it does not exceed the upper limit of the GPU device, jump to step 304; otherwise, use a greedy method to recalculate and optimize the neighborhood solution; Step 304, evaluate the solution quality and select the optimal solution according to the cost; Step 305, filter solutions: calculate the cost of the current neighborhood solution. If the cost of the current neighborhood solution is lower than the cost of the optimal solution, take the current neighborhood solution as the new optimal solution and put the current neighborhood solution into the taboo table. If the cost of no neighborhood solution is lower than the cost of the current optimal solution, select the new neighborhood solution according to the probability. Select the suboptimal solution and then put it into the taboo table; where P is a randomly generated value between 0 and 1, and T is the current temperature parameter; Step 306, update the taboo table: if no better solution is found after several consecutive iterations, trigger the diversification operation and increase the temperature T to , clear half of the taboo table and randomly generate an initial solution; otherwise, use the exponential cooling strategy to reduce the temperature to ;in, is the preset upper temperature limit, is the temperature rise coefficient, is the temperature reduction coefficient; Step 307, repeating steps 302 to 306 until the termination condition is met, and obtaining an optimized stage division scheme.
4. The pipeline parallel training method for a multimodal large model according to claim 3, characterized in that: Step 302 includes: Update the step size according to the temperature T , using the updated step size , move the split point b to the right , or move left to , generate a new partitioning scheme and add it to the neighborhood solution set; The step length The update formula is: 。 5. A pipeline parallel training method for a multimodal large model according to claim 3 or 4, characterized in that: In step 303, the steps of recalculating and optimizing the neighborhood solution using a greedy method include: Calculate the memory benefit of each layer and recalculation cost , which is the time for the k-th layer forward propagation. For the k-th layer, the memory gain Defined as the activation value memory saved after enabling recalculation, the recalculation cost is the time to re-execute the forward propagation; Calculate the ratio of video memory benefit to recalculation cost, the formula is: ; Sort all layers in the current stage by MCR value from high to low, and give priority to recalculating the layers with high ratios; Finally, a list of recomputed layer indices for each stage is output.
6. The pipeline parallel training method for a multimodal large model according to claim 5, characterized in that: In step 303, the video memory calculation formula is as follows: , In the formula, and The stages The starting layer index and the ending layer index of The layer index in , satisfies , Indicates fixed video memory usage. For the recalculation variable, 1 means recalculation is enabled, 0 means it is not enabled.
7. The pipeline parallel training method for a multimodal large model according to claim 6, characterized in that: Step 304 includes: The cost function Cost is constructed with the goal of minimizing the maximum stage training time. The formula is: , In the formula, For stage S i The total training time is composed of the basic calculation time and the additional time introduced by recalculation, and the formula is: , In the formula, represents the total execution time of the kth layer; Calculate the cost of each neighborhood solution and select the neighborhood solution corresponding to the lowest cost as the optimal solution.
8. The pipeline parallel training method for a multimodal large model according to claim 7, characterized in that: In step 1, the formula for calculating the execution time of each layer is: , In the formula, They respectively represent the execution time of the kth layer of the multimodal large model during the back propagation process.
Citation Information
Patent Citations
Heterogeneous environment-oriented large model hybrid parallel training method and system
CN117633527A
Assembly line parallel division and memory optimization method for large-scale model training
CN119336489A
Pipeline parallel distributed training method, device and system for deep neural network
CN119987999A
Highly performant pipeline parallel deep neural network training
US20190362227A1