Automatic Search Method for Pipeline Allocation Based on Intelligent Agents
Through the automatic search method of pipeline allocation of intelligent agents, task allocation in pipeline parallel training is optimized, the problems of idle equipment and slow solution speed are solved, and efficient pipeline parallel training is realized to adapt to different model scales.
Patent Information
- Application Number
- CN202411466156.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-10-21
- Publication Date
- 2025-07-04
- Estimated Expiration
- 2044-10-21
AI Technical Summary
The existing pipeline parallel technology causes equipment to be idle when task allocation is unbalanced, affecting training efficiency, and the existing technology solves slowly and cannot directly adapt to the new model.
The automatic search method of pipeline allocation based on intelligent agents is adopted, and the execution time and communication time of the neural network layer are obtained through the performance analysis module, a large language model is built, and iterative search pipeline allocation strategies are automatically generated, and an experience database is used to perform automatic search to optimize task allocation.
It significantly shortens the search time, improves the efficiency of parallel training, ensures the search effect, has good versatility and robustness, and is suitable for different model sizes.
Smart Images

Figure CN119416865B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of distributed training, and particularly relates to an automatic search method for pipeline allocation based on intelligent agents. Background Art
[0002] With the continuous increase in the scale of deep learning models and datasets, the single-machine mode has become difficult to meet the needs of training ultra-large models. Therefore, distributed training and its parallel technologies have become the mainstream methods for optimizing training efficiency.
[0003] In distributed training, pipeline parallelism is an important parallelization strategy that allows the training process of a deep learning model to be divided into multiple stages, with each stage executed in parallel on different devices or nodes. This parallelization method can evenly distribute the computational load across multiple devices to improve the overall training efficiency, and is thus particularly suitable for models with a large number of layers and a large amount of computation.
[0004] However, pipeline parallelism also faces difficulties and challenges: The pipeline requires careful planning and scheduling to avoid processor idleness. Since the workload is carried out in stages, unbalanced task allocation will cause some stages to wait for other stages to complete, resulting in device idleness and thus affecting the overall training efficiency. Summary of the Invention
[0005] Aiming at the problems existing in the existing pipeline parallel technology, the present invention provides an automatic search method for pipeline allocation based on intelligent agents, which can improve the parallel training efficiency of the target large model, significantly shorten the search time while ensuring the search effect, and has good generality and robustness.
[0006] To achieve the above object, the technical method adopted by the present invention is as follows:
[0007] The automatic search method for pipeline allocation based on intelligent agents includes the following steps:
[0008] Step 1: Use a performance analysis module to obtain the performance analysis results of a target large model containing N neural network layers, including the execution time of the i-th, i = 1, 2,..., N neural network layers, as the execution time of the target large model on the cluster device;
[0009] Step 2: Assume that the global batch size of the target large model is 1, the user sets the number of pipeline stages M and the micro-batch size, and calculate the activation memory v of the i-th neural network layer according to the network structure, sequence length, hidden layer, and micro-batch size of the target large model i , i = 1, 2,..., N;
[0010] Furthermore, calculate the communication time e of the i-th neural network layer i, i = 1, 2, ..., N:
[0011]
[0012] Wherein, b is the bandwidth between any two devices in the cluster device;
[0013] Step 3, construct a large language model, and set the output data format of the large language model to [p1, p2, ..., p M , where p j , j = 1, 2, ..., M is the number of neural network layers included in the j-th, j = 1, 2, ..., M pipeline stage;
[0014] Step 4, automatically generate an iterative search pipeline allocation strategy, specifically including:
[0015] Step 4.1, let the iteration number k = 1; splice the number of neural network layers, the number of pipeline stages, the mini-batch size of the target large model, and the execution time and communication time of the i-th neural network layer as the input data for the first iteration; construct an experience database, and let the experience database for the first iteration be empty;
[0016] Step 4.2, in the k-th iteration, input the input data into the large language model, and the output is used as the pipeline allocation strategy A k , and splice the execution time and communication time of the N neural network layers in the input data corresponding to the pipeline allocation strategy A k to obtain the input O k for the k-th iteration;
[0017] Step 4.3, evaluate the pipeline allocation strategy A k , and determine the evaluation result Q k obtained in the k-th iteration according to the evaluation formula;
[0018] Step 4.4, construct the experience obtained in the k-th iteration as (O k , A k , Q k ), retrieve the experience closest to (O k , A k , Q k ) in the experience database, and use the retrieved experience and (O k , A k , Q k ) together as the input data for the k + 1-th iteration. If no experience closest to (O k , A k , Q k ) is retrieved in the experience database, then only (O k , A k , Qk ) as the input data for the (k + 1)-th iteration; meanwhile, add (O k , A k , Q k ) to the experience database;
[0019] Step 4.5: Determine whether the iteration number k reaches the preset upper limit of the iteration number, or whether the pipeline allocation strategies obtained from consecutive multiple iterations remain unchanged. If neither condition is satisfied, let k = k + 1 and return to Step 4.2; otherwise, output the pipeline allocation strategy A k obtained in the k-th iteration as the optimal pipeline allocation strategy.
[0020] Furthermore, the specific method for obtaining the execution time of each neural network layer in Step 1 is as follows:
[0021] First, discard the first n rounds of training data of the target large model; then, count the operator duration between the first occurrence of the operator of the i-th (i = 1, 2,..., N) neural network layer and the first occurrence of the operator of the (i + 1)-th neural network layer as the initial execution time of the i-th neural network layer; for the last m rounds of training data, calculate the average value of the initial execution time of the i-th neural network layer as the execution time of the i-th neural network layer.
[0022] Furthermore, the large language model described in Step 3 is constructed through pre-instruction fine-tuning.
[0023] Furthermore, the evaluation formula in Step 4.3 is:
[0024]
[0025] In the formula, B is the micro-batch size after pipeline parallel processing; max{·} represents taking the maximum value; e j is the execution time of the j-th pipeline stage, that is, the sum of the execution times of the p j layers of neural network layers included in the j-th pipeline stage; c j is the communication time between the j-th pipeline stage and the (j + 1)-th pipeline stage.
[0026] Furthermore, the method for determining the experience closest to (O k , A k , Q k ) in Step 4.4 is as follows:
[0027] Calculate the Euclidean distance between the input O k of (O k , A k ) and the input O of each experience in the experience database respectively, and find the experience corresponding to the minimum Euclidean distance as the experience closest to (O k k ,A k ,Q k ) The closest experience.
[0028] Further, the maximum value of N is generally taken as 128; M is enumerated as a power of 2 and its value does not exceed N; n generally takes a value of 20; the value range of m is 20 to 30.
[0029] Compared with the prior art, the beneficial effects of the present invention are as follows:
[0030] The present invention proposes an automatic search method for pipeline allocation based on intelligent agents, which does not require model training and can be directly applied for pipeline parallelism, thus solving the problems of slow solution speed, online training requirement and inability to directly adapt to new models in the prior art; the present invention can improve the parallel training efficiency of the target large model, significantly shorten the search time while ensuring the search effect, and the search time only increases with the expansion of pp (micro-batch size) and is hardly affected by the total number of neural network layers of the target large model, having good generality and robustness. BRIEF DESCRIPTION OF THE DRAWINGS
[0031] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the following drawings are only some embodiments of the present invention, and those of ordinary skill in the art can also obtain other drawings without creative efforts based on these drawings.
[0032] Figure 1 is a schematic flowchart of the automatic search method for pipeline allocation based on intelligent agents proposed in Embodiment 1;
[0033] Figure 2 is a schematic principle diagram of the automatic search method for pipeline allocation based on intelligent agents proposed in Embodiment 1;
[0034] Figure 3 is a schematic diagram of pipeline parallel processing proposed in Embodiment 1. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0035] In order to further understand the present invention, the following describes the preferred implementation schemes of the present invention in combination with embodiments. However, it should be understood that these descriptions are only for further explaining the features and advantages of the present invention and not for limiting the claims of the invention.
[0036] Embodiment 1
[0037] In pipeline parallelism, the neural network layers of the target large model are distributed among multiple devices (such as graphics cards), and the training batch is subdivided into smaller micro-batches. Forward and backward propagation are carried out in a pipelined manner among these micro-batches. However, unbalanced task allocation can cause some stages to wait for other stages to complete, resulting in device idleness and thus affecting the overall training efficiency.
[0038] Based on the above problems, this embodiment proposes an automatic search method for pipeline allocation based on intelligent agents. The process is as Figure 1 shown, and the principle is as Figure 2 shown, including the following steps:
[0039] Step 1: Use the performance analysis module to obtain the performance analysis results of the target large model containing 24 neural network layers, including the execution time of the i-th neural network layer, where i = 1, 2,..., 24, as the execution time of the target large model on the cluster devices; in this embodiment, the performance analysis module is specifically the Profiler module.
[0040] Among them, the method for obtaining the execution time of the i-th neural network layer is specifically as follows:
[0041] First, discard the first 10 rounds of training data of the target large model; then, count the operator duration between the first appearance of the operator of the i-th neural network layer, where i = 1, 2,..., 24, and the first appearance of the operator of the (i + 1)-th neural network layer as the initial execution time of the i-th neural network layer; for the last 20 rounds of training data, calculate the average value of the initial execution time of the i-th neural network layer as the execution time of the i-th neural network layer.
[0042] Step 2: Assume that the global batch size of the target large model is 1, the user sets the number of pipeline stages M = 4 and the micro-batch size, and calculate the activation memory v i , i = 1, 2,..., 24;
[0043] Furthermore, calculate the communication time e i , i = 1, 2,..., 24:
[0044]
[0045] In the formula, when the cluster devices are homogeneous, b is the bandwidth between any two devices in the cluster devices;
[0046] Step 3: Build a large language model for the intelligent agent, specifically by fine-tuning Llama-7b based on the alpaca dataset, and set the output data format of the large language model to [p1, p2, p3, p4], where p j , j = 1, 2, 3, 4 is the number of neural network layers included in the j-th, j = 1, 2, 3, 4 pipeline stage.
[0047] In the pipeline allocation problem, only consecutive neural network layers can be allocated to the same pipeline stage. In particular, the computation of the i-th neural network layer can be reused with the computation of the i+1-th neural network layer. For example, the first and second neural network layers can be allocated to one stage, and their execution times can be summed because they can form an effective pipeline stage; however, it is impossible to sum the first and third neural network layers because they are not consecutive and cannot be allocated to the same pipeline stage.
[0048] Example: For a target large model containing N = 24 neural network layers, the output data of the large language model can be [5, 6, 7, 6], where the first pipeline stage includes 5 neural network layers, namely the first to fifth neural network layers; the second pipeline stage includes 6 neural network layers, namely the sixth to eleventh neural network layers; the third pipeline stage includes 7 neural network layers, namely the twelfth to eighteenth neural network layers; the fourth pipeline stage includes 6 neural network layers, namely the nineteenth to twenty-fourth neural network layers.
[0049] Step 4: Automatically generate an iterative search pipeline allocation strategy. The large language model of the intelligent agent continuously learns how to generate an optimal pipeline allocation strategy through proposal and feedback, specifically including:
[0050] Step 4.1: Let the iteration number k = 1; splice the number of neural network layers, the number of pipeline stages, the mini-batch size of the target large model, and the execution time and communication time of the i-th neural network layer as the input data for the first iteration; build an experience database and set the experience database for the first iteration to be empty.
[0051] Step 4.2: In the k-th iteration, input the input data into the large language model, and the output is used as the pipeline allocation strategy A k , and splice the execution time and communication time of the N neural network layers corresponding to the pipeline allocation strategy A k in the input data to obtain the input O k for the k-th iteration.
[0052] Step 4.3: Evaluate the pipeline allocation strategy A k , and determine the evaluation result Q obtained in the k-th iteration according to the evaluation formulak ; Among them, the evaluation formula is:
[0053]
[0054] In the formula, B is the micro-batch size after adopting pipeline parallel processing; max{·} represents taking the maximum value; e j is the execution time of the j-th pipeline stage, that is, the sum of the execution times of the p j layers of neural network layers included in the j-th pipeline stage; c j is the communication time between the j-th pipeline stage and the (j + 1)-th pipeline stage.
[0055] For different settings of the micro-batch size, the device usually has different utilization rates. Intuitively, the first term of the evaluation formula reflects the laggard effect, and this effect will be amplified by the size of the micro-batch B. If it is assumed that B is 1, the above formula can be simplified to This is consistent with the actual plan: a single batch only performs a single forward pass and a single backward pass, and sends or receives gradients or activations.
[0056] Step 4.4, construct the experience obtained in the k-th iteration as (O k , A k , Q k ). Retrieve the experience closest to (O k , A k , Q k ) in the experience database. Use the retrieved experience and (O k , A k , Q k ) together as the input data for the (k + 1)-th iteration. If no experience closest to (O k , A k , Q k ) is retrieved in the experience database, then only use (O k , A k , Q k ) as the input data for the (k + 1)-th iteration; meanwhile, add (O k , A k , Q k ) to the experience database;
[0057] Among them, the method for determining the experience closest to (O k , A k , Q k ) is:
[0058] Calculate the input O of (O k , A k , Q k ) respectively kThe Euclidean distance between each experience in the experience database and the input O, and find the experience corresponding to the minimum Euclidean distance as the experience closest to (O k , A k , Q k ).
[0059] Step 4.5: Determine whether the iteration count k reaches the preset upper limit of the iteration count, or whether the pipeline allocation strategies obtained from three consecutive iterations remain unchanged. If neither condition is met, set k = k + 1 and return to Step 4.2; otherwise, output the pipeline allocation strategy A obtained in the k-th iteration k , as the optimal pipeline allocation strategy.
[0060] If it is assumed that all neural network layers have the same execution cost and the cluster is equipped with the same device communication bandwidth, then the optimal pipeline allocation strategy is to balance the number of layers.
[0061] However, when there are differences in the execution time and communication time of each pipeline stage, it will cause relatively large pipeline bubbles, as shown in Figure 3 the upper part. The ideal situation is to achieve what is shown in the lower part through a reasonable pipeline allocation algorithm, by reducing the pipeline bubbles during training, thereby achieving an optimized effect in model training. Figure 3
[0062] In traditional pipeline allocation, the optimal pipeline allocation strategy is usually manually designed based on knowledge and experience, but it does not utilize the method of automatic search. This leads to challenges in design when the model scale expands. In traditional search methods, the global space is searched based on dynamic programming or depth-first search to obtain the optimal pipeline allocation strategy, but it wastes the previous search experience and increases the scope of the search space. Therefore, in this embodiment, by setting up an experience database and automatically searching for experiences, the search time can be significantly shortened.
[0063] Taking the AMP algorithm (the technology "AMP: Automatically Finding Model Parallel Strategies with Heterogeneity Awareness" disclosed in 2022) as the comparative example of this embodiment, it solves the pipeline allocation strategy based on the dynamic programming algorithm.
[0064] By adjusting the total number of neural network layers N of pp and the target large model, and comparing the time-consuming of the automatic search pipeline allocation strategy of the AMP algorithm and the automatic search method for pipeline allocation based on intelligent agents proposed in this embodiment, as shown in Table 1, it can be seen that the time-consuming of the automatic search pipeline allocation strategy of the AMP algorithm increases sharply with the expansion of the input data scale (i.e., pp and N). However, the time-consuming of the automatic search pipeline allocation strategy in this embodiment only increases the search time with the expansion of the pp scale, and is hardly affected by the expansion of the N scale. This embodiment has obvious advantages for large-scale target large models. For example, when N = 96, the AMP algorithm takes 27,838 seconds, about 7.7 hours, while this embodiment only takes 37 seconds.
[0065] Table 1
[0066] N=24 N=32 N=48 N=64 N=96 This embodiment pp = 2 1.29 4.27 32.45 160.99 1659.89 2.07 pp = 4 2.55 12.48 91.45 472.85 4795.99 20.38 pp = 8 3.44 26.32 212.49 1027.25 9571.54 30.92 pp = 16 4.99 43.97 402.61 1957.95 27838.39 37.99
[0067] A large number of examples are used to compare the quality of the automatic search pipeline allocation strategy of the AMP algorithm and the automatic search method for pipeline allocation based on intelligent agents proposed in this embodiment. As shown in Table 2, it can be seen that this embodiment makes a trade-off between search time and search effect compared with the AMP algorithm. The error of the pipeline allocation strategy in this embodiment is concentrated at about 1%, indicating that the quality of the automatic search pipeline allocation strategy in this embodiment is already very close to that of the AMP algorithm using dynamic programming, that is, the search time is significantly shortened while ensuring the search effect.
[0068] Table 2
[0069]
[0070] The prior art uses dynamic programming algorithms or online training reinforcement learning algorithms to solve the pipeline allocation strategy, facing the problems of relatively long solution time and inability to transfer the trained strategy model; this embodiment uses an automatic search method for pipeline allocation based on intelligent agents, without model training, and directly applies it for pipeline parallelism, thus solving the problems of slow solution speed, the need for online training, and inability to directly adapt to new models in the prior art, and improving the efficiency of pipeline parallelism.
[0071] The above embodiments are for better understanding of the present invention, and are not limited to the best implementation manner, and do not constitute a limitation to the content and protection scope of the present invention. Any product identical or similar to the present invention obtained by anyone under the inspiration of the present invention or by combining the features of the present invention with other prior art features falls within the protection scope of the present invention.
Claims
1. An automatic search method for pipeline allocation based on intelligent agents, characterized in that Including the following steps: Step 1: Use the performance analysis module to obtain the performance analysis results of the target large model containing N neural network layers, including the execution time of the i-th neural network layer, where i = 1, 2, …, N, as the execution time of the target large model on the cluster device; Step 2: Assume that the global batch size of the target large model is 1. The user sets the number of pipeline stages M and the micro-batch size. According to the network structure, sequence length, hidden layer, and micro-batch size of the target large model, calculate the activation storage amount v of the i-th neural network layer i , and then calculate the communication time e of the i-th neural network layer i : Step 3: Construct a large language model and set the output data format of the large language model to [p1, p2,..., p M , where p j is the number of neural network layers included in the j-th pipeline stage, and j = 1, 2,..., M; Step 4: Automatically generate an iterative search pipeline allocation strategy, specifically including: Step 4.1: Let the iteration number k = 1; splice the number of neural network layers, the number of pipeline stages, the mini-batch size of the target large model, and the execution time and communication time of the i-th neural network layer as the input data for the first iteration; construct an empirical database and set the empirical database for the first iteration to be empty; Step 4.2: In the k-th iteration, input the input data into the large language model, and the output is used as the pipeline allocation policy A k , and splice the execution time and communication time of the N neural network layers in the input data corresponding to the pipeline allocation policy A k to obtain the input O of the k-th iteration k ; Step 4.3: Evaluate the pipeline allocation strategy A k and determine the evaluation result Q obtained in the k-th iteration according to the evaluation formula k ; Step 4.
4. Construct the experience obtained in the k-th iteration as (O k , A k , Q k ). Retrieve the experience closest to (O k , A k , Q k ) in the experience database. Use the retrieved experience and (O k , A k , Q k ) together as the input data for the (k + 1)-th iteration. If no experience closest to (O k , A k , Q k ) is retrieved in the experience database, then use only (O k , A k , Q k ) as the input data for the (k + 1)-th iteration. At the same time, add (O k , A k , Q k ) to the experience database; Step 4.
5. Determine whether the iteration number k reaches the preset upper limit of the iteration number, or whether the pipeline allocation strategies obtained by consecutive multiple iterations remain unchanged. If neither condition is satisfied, let k = k + 1 and return to Step 4.2; otherwise, output the pipeline allocation strategy A obtained in the k-th iteration. k , as the optimal pipeline allocation strategy.
2. The automatic search method for pipeline allocation based on intelligent agents according to claim 1, wherein The specific method for obtaining the execution time of each neural network layer in Step 1 is as follows: First, discard the first n rounds of training data of the target large model; then, count the operator duration between the first appearance of the operator of the i-th neural network layer and the first appearance of the operator of the i + 1-th neural network layer as the initial execution time of the i-th neural network layer; For the subsequent m rounds of training data, calculate the average value of the initial execution time of the i-th neural network layer as the execution time of the i-th neural network layer.
3. The automatic search method for pipeline allocation based on intelligent agent according to claim 1, characterized in that The large language model described in Step 3 is constructed through pre-instruction fine-tuning.
4. The automatic search method for pipeline allocation based on intelligent agent according to claim 1, characterized in that The evaluation formula in Step 4.3 is: where B is the micro-batch size after pipelined parallel processing; max{} represents taking the maximum value; e j is the execution time of the j-th pipelining stage, that is, the sum of the execution times of the p j layers of neural network layers included in the j-th pipelining stage; c j is the communication time between the j-th pipelining stage and the (j + 1)-th pipelining stage.
5. The automatic search method for pipeline allocation based on intelligent agents according to claim 1, wherein In step 4.4, the method for determining the experience closest to (O k , A k , Q k ) is as follows: Calculate (O k , A k , Q k )'s input O k and the Euclidean distance between the input O corresponding to each experience in the experience database, find the experience corresponding to the minimum Euclidean distance, and use it as the experience closest to (O k , A k , Q k ).
Citation Information
Patent Citations
Neural network pipeline parallel training method for optimizing model division
CN116167436A
Model parallel training method and device
CN118410859A