Parallel strategy optimal selection method, and neural network solver training method and apparatus
Through the combination of neural network solvers and heuristic strategies, the parallel strategies of quickly obtaining and selecting large-scale deep learning models is solved, and the problems of low training efficiency and high technical difficulty in the existing technology are achieved, and fast and efficient parallel strategy acquisition and model training are achieved.
Patent Information
- Application Number
- PCT/CN2024/126172
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2023-12-01
- Filing Date
- 2024-10-21
- Publication Date
- 2025-06-05
AI Technical Summary
In the distributed training of large-scale deep learning models, a single parallel mode is difficult to maximize the hardware computing power, and a hybrid parallel mode requires complex algorithm development, resulting in inefficient training and high technical difficulty.
A parallel strategy optimization method is provided, using neural network solvers and heuristic strategies to quickly obtain multiple candidate mixed parallel strategies, predict the overhead value of candidate strategies, and quickly determine the theoretical optimal strategy through the optimization process. The neural network solver is saved after self-game training, avoiding repeated training and improving migration.
It significantly shortens the time cost of parallel strategy acquisition, improves training efficiency, solves the problems of slow solution speed and online training in the existing technology, and can quickly adapt to the new model.
Smart Images

Figure CN2024126172_05062025_PF_FP_ABST
Abstract
Description
Parallel strategy optimization method and neural network solver training method and device
[0001] This application claims priority to the Chinese patent application filed with the State Intellectual Property Office on December 1, 2023, with application number 202311653104.8 and invention name “Parallel Strategy Optimization Method and Neural Network Solver Training Method and Device”, the entire contents of which are incorporated by reference into this application. Technical Field
[0002] The present application relates to the field of artificial intelligence, and in particular to a parallel strategy optimization method, electronic equipment, computer storage medium, computer program product and chip, and a neural network solver training method and device. Background Art
[0003] As deep learning models and datasets continue to grow in size, the training process becomes computationally intensive. Traditional single-machine models are no longer sufficient for training extremely large models. To address this issue, distributed training parallelism has become a mainstream approach for optimizing the training efficiency of large models. Currently, there are two basic approaches for distributed training of large models: data parallelism and model parallelism. Data parallelism divides the training data across different compute nodes. The model replica on each node processes a local subset of the data, and parameter updates are aggregated through network communication. Model parallelism approaches are categorized into two types based on how the model is partitioned: pipeline parallelism and tensor parallelism. Pipeline parallelism divides the model into multiple parts along the forward propagation direction, assigning each part to a node for execution, and running them in parallel like a pipeline. Tensor parallelism splits large tensors in the model (such as hidden layer weights) across different compute nodes, allowing each node to store and update only a portion of the tensor, reducing the memory usage of each node.
[0004] A single parallel model struggles to maximize hardware computing power. However, a hybrid parallel model requires algorithm developers to call the parallel sharding API provided by the AI framework based on the characteristics of the algorithm model. This approach increases the technical difficulty of distributed training of AI algorithms. Furthermore, due to algorithm developers' lack of understanding of the characteristics of AI frameworks and computing devices, model parallel training efficiency is low. The specific distributed tuning work increases the difficulty of algorithm development and reduces the efficiency of algorithm research. As the target large model scales, how to quickly acquire parallel strategies and reduce the time cost of acquisition becomes an urgent problem.
[0005] Summary of the Invention
[0006] In view of the problems and shortcomings of the above-mentioned prior art, the present invention provides a method for optimizing a parallel strategy, an electronic device, a computer storage medium, and a method and device for training a neural network solver. The optimization method provided in the present application quickly obtains a variety of candidate hybrid parallel strategies based on a neural network solver and a heuristic strategy, predicts the cost values of the candidate hybrid parallel strategies, and optimizes the candidate hybrid parallel strategies based on the predicted cost values, and quickly determines the hybrid parallel strategies. The training method for the neural network solver provided in the present application constructs an offline training device, and saves the self-game training of the neural network solver for direct application without repeated training, and has excellent transferability, further reducing the time cost of obtaining the hybrid parallel strategy.
[0007] In a first aspect, the present invention provides an optimization method for a parallel strategy.
[0008] In a second aspect, the present invention provides a method for training a neural network solver, for training the neural network solver described in the first aspect.
[0009] In a third aspect, the present invention provides an offline training device for training the neural network solver described in either the first aspect or the second aspect.
[0010] In a fourth aspect, the present invention provides an electronic device comprising a memory and a processor, wherein the memory is used to store a computer program, and the processor is used to run the computer program to implement the steps of the preferred method described in the first aspect.
[0011] In a fifth aspect, the present invention provides a computer storage medium storing program instructions, which, when executed, implement the steps of the preferred method described in the first aspect.
[0012] In a sixth aspect, the present invention provides a computer program product, comprising code, characterized in that when the code is executed, it is used to implement the preferred method as described in any one of claims 1-6.
[0013] In a seventh aspect, the present invention provides a chip comprising a processor, characterized in that the processor is used to run the preferred method as described in any one of claims 1-6.
[0014] In a first aspect, the present invention provides an optimization method for a parallel strategy.
[0015] In an implementation of the first aspect above, the preferred method includes:
[0016] Acquire basic information of the target large model based on cluster information of the target large model, wherein the basic information of the target large model includes candidate tensor parallelism;
[0017] Based on the candidate tensor parallelism, determine candidate data parallelism and candidate pipeline parallelism through a heuristic candidate strategy to obtain a first parallel strategy;
[0018] Calculating the first parallel strategy through a neural network solver to obtain a split point strategy, adding the split point strategy to the first parallel strategy to obtain a second parallel strategy, and predicting a simulation cost value of the second parallel strategy;
[0019] The second parallel strategies are screened based on the simulated overhead values of the second parallel strategies to obtain a theoretically optimal strategy.
[0020] In the above implementation, the cluster information includes the number of GPUs of each device, the basic information of the target large model also includes the layer time overhead, and the target large model includes at least one neural network layer. The basic information of the target large model is obtained based on the cluster information of the target large model. The specific implementation steps are:
[0021] Obtaining the candidate tensor parallelism based on the number of GPUs of each device, where the value of the candidate tensor parallelism is less than or equal to the number of GPUs of each device;
[0022] Under the candidate tensor parallelism, the time overhead of each neural network layer of the at least one neural network layer is collected as the layer-level time overhead.
[0023] The candidate tensor parallelism can be a set of values less than or equal to the number of GPUs per device, or one of these values.
[0024] As a preferred embodiment, the candidate tensor parallelism is a set of all values less than or equal to the number of GPUs of each device, because the richer the values of the candidate tensor parallelism, the greater the number of second parallel strategies participating in the optimization, and the better the optimization effect.
[0025] In the above implementation, the cluster information also includes the number of devices. Based on the candidate tensor parallelism, the candidate data parallelism and the candidate pipeline parallelism are determined by a heuristic candidate strategy to obtain a first parallel strategy. The specific implementation steps are:
[0026] Traversing the candidate tensor parallelisms to perform data parallelism calculations to obtain candidate data parallelism, where the candidate data parallelism is calculated based on the candidate tensor parallelism, the number of devices, and the number of GPUs of each device;
[0027] Where factor is the factor acquisition function, dp is the candidate data parallelism, N is the number of devices, M is the number of GPUs per device, and tmp is the candidate tensor parallelism.
[0028] The pipeline parallelism is calculated based on the candidate tensor parallelism, the candidate data parallelism, the number of devices, and the number of GPUs of each device to obtain the candidate pipeline parallelism.
[0029] Among them, pp is the candidate data parallelism.
[0030] Based on the heuristic candidate strategy, the tensor parallelism is traversed to obtain all possible candidate data parallelism values, and the pipeline parallelism is solved one by one based on the traversal results. The heuristic candidate strategy can obtain the parameters of the first parallel strategy simply, efficiently and to the maximum extent.
[0031] In the above implementation, the basic information of the target large model includes hierarchical time overhead, the target large model includes at least one neural network layer, the first parallel strategy is calculated by the neural network solver to obtain the split point strategy, the split point strategy is added to the first parallel strategy to obtain the second parallel strategy, and the simulation overhead value of the second parallel strategy is predicted. The specific implementation steps are:
[0032] The neural network solver is used to calculate the parameters of the first parallel strategy to obtain a split point strategy, wherein the parameters of the first parallel strategy include candidate tensor parallelism, candidate pipeline parallelism and candidate data parallelism.
[0033] The split point strategy is added to the first parallel strategy to obtain the second parallel strategy.
[0034] Based on the parameters and hierarchical time overhead of the second parallel strategy, a simulation overhead value of the second parallel strategy is predicted, where the parameters of the second parallel strategy include candidate tensor parallelism, candidate pipeline parallelism, candidate data parallelism, and a split point strategy.
[0035] In the above implementation, the neural network solver includes an encoder and a decoder. The neural network solver calculates the parameters of the first parallel strategy to obtain the split point strategy. The parameters of the first parallel strategy include candidate tensor parallelism, candidate pipeline parallelism, and candidate data parallelism. The specific implementation steps are:
[0036] Get candidate tensor parallelism and candidate pipeline parallelism;
[0037] Characterize the layer time overhead under the candidate tensor parallelism and the candidate pipeline parallelism to obtain a feature matrix;
[0038] Input the feature matrix into the encoder for encoding to obtain the segmentation point code sequence;
[0039] The segmentation point coding sequence is input into the decoder for decoding to obtain the segmentation point strategy.
[0040] In the above implementation, the hierarchical time overhead under the candidate tensor parallelism and the candidate pipeline parallelism are characterized to obtain a feature matrix. The specific implementation steps are:
[0041] Perform a full estimation based on the candidate pipeline parallelism to obtain a first communication overhead matrix;
[0042] Obtain a first computational cost vector based on the hierarchical time cost under the candidate tensor parallelism, and perform matrix transformation on the first computational cost vector to obtain a first computational cost matrix;
[0043] The first communication overhead matrix and the first computation overhead matrix are concatenated column by column to obtain a feature matrix.
[0044] Through the characterization process, the first communication overhead matrix and the first computation overhead matrix are aligned based on columns, so as to facilitate further splicing of the two.
[0045] The neural network solver constructed in the present invention is a reusable general strategy neural network. After the network is self-game trained, it is saved and applied directly without the need for online retraining and learning every time. A neural network solver is obtained that is decoupled from the application scenario, which greatly improves the solution speed of the split point strategy, solves the problems of the existing technology that the solution speed is slow, requires online training, and cannot be directly adapted to the new model, and further quickly obtains the second parallel strategy.
[0046] In the above implementation, the simulation overhead value includes data overhead and model overhead. The simulation overhead value of the second parallel strategy is predicted based on the parameters and level time overhead of the second parallel strategy. The specific implementation steps are:
[0047] Predict the model overhead of the second parallelization strategy based on the candidate tensor parallelism, candidate pipeline parallelism, split point strategy, and inter-layer time overhead;
[0048] Predicting a data overhead of the second parallel strategy based on the candidate data parallelism;
[0049] Based on the sum of the data overhead and the model overhead, a simulation overhead value of the second parallel strategy is calculated.
[0050] In the above implementation, the model overhead of the second parallel strategy is predicted based on the candidate data parallelism, and the specific implementation steps are:
[0051] The stage overhead of the second parallel strategy is calculated based on the split point strategy and the layer time overhead, which includes the stage computation overhead and the stage communication overhead. The neural network layer of each stage is obtained based on the split point strategy, and the layer time overhead of the neural network layer belonging to the same stage is summed. The product of the sum value and the micro-batch size is used as the stage computation overhead of the stage. At the same time, the stage communication overhead is calculated based on the data volume calculation and theoretical bandwidth of each stage.
[0052] The model overhead is calculated based on the stage overhead and tensor parallelism of the second parallel strategy;
[0053] Among them, t F j Calculate the cost for the jth stage of candidate tensor parallelism F, max t F j The maximum value of the stage calculation overhead when the candidate tensor parallelism is F; e F j is the communication cost of the jth stage when the candidate tensor parallelism is F, T p is the model overhead of the pipeline corresponding to the data parallel group p, gbs p is the global batch size of the pipeline corresponding to the data parallel group p, p is the data parallel group, p = 1, 2, ..., G, G is the candidate data parallelism, K is the candidate pipeline parallelism, and mbs is the micro-batch size.
[0054] In the above implementation, the data overhead of the second parallel strategy is predicted, and the specific implementation steps are:
[0055] Based on the candidate data parallelism, the data overhead of the second parallel strategy is predicted.
[0056] Where a is the data communication volume, T dp_sync is the data overhead, minb p is the minimum bandwidth in data parallel group p.
[0057] In a second aspect, the present invention provides a method for training a neural network solver, for training the neural network solver described in the first aspect.
[0058] In an implementation of the second aspect above, the training method includes:
[0059] Generate data based on probability distribution;
[0060] Performing characterization on the generated data to obtain input data;
[0061] Inputting input data into at least two neural network solvers respectively to obtain at least two split point strategies;
[0062] At least two simulation cost values are calculated based on at least two split point strategies;
[0063] Performing numerical conversion based on at least two simulated cost values to obtain at least two reward scores;
[0064] A loss function is calculated based on at least two reward scores and at least two split point strategies, and weights of at least two neural network solvers are updated based on the loss function, wherein the weights are shared between the at least two neural network solvers.
[0065] In the above implementation, the loss function is calculated based on at least two reward scores and at least two split point strategies, including:
[0066] The target large model is divided based on the dividing point strategy to obtain K stages; one of the K stages includes at least one continuous neural network layer.
[0067] The loss function is calculated based on the reward score and the probability value of the last neural network layer of each stage:
[0068] in, For the split point strategy C σ The probability value of the last neural network layer of the jth stage, R σ For the split point strategy C σ Reward score, C σ is the σth split point strategy.
[0069] In any of the above implementations, the probability distribution generation includes uniform distribution generation and random distribution generation.
[0070] The existing technology uses a dynamic programming algorithm to search for strategies online for each model, which cannot reuse existing strategy networks. Each search takes a long time, affecting the efficiency of obtaining parallel strategies. The present invention trains a neural network solver in a self-supervised offline manner, adopts self-game learning and self-supervised learning methods, abstracts specific problems into a general strategy solver part, decouples the training process from the solver application scenario, avoids repeated training, and enables the network to have generalization capabilities, which can be directly applied to different models. This can solve the problems of the existing technology that are slow to solve, require online training, and cannot be directly adapted to new models, thereby improving the efficiency of pipeline segmentation.
[0071] In a third aspect, the present invention provides an offline training device for training the neural network solver described in either the first aspect or the second aspect.
[0072] In an implementation of the third aspect above, the offline training device includes:
[0073] Input data acquisition module: used to obtain input data;
[0074] Neural network solution module: used to solve the split point and simulate the cost prediction based on the input data to obtain the split point strategy and simulation cost value;
[0075] Reward module: used to evaluate the split point strategy based on the simulated cost value and obtain a reward score.
[0076] In the above implementation, the neural network solving module includes at least two neural network solvers, the neural network solver includes an encoder and a decoder, the encoder and the decoder are connected in series, the at least two neural network solvers are connected in parallel, and the input data acquisition module, the neural network solving module and the reward module are connected in series.
[0077] In a fourth aspect, the present invention provides an electronic device comprising a memory and a processor, wherein the memory is used to store a computer program, and the processor is used to run the computer program to implement the steps of the preferred method described in the first aspect.
[0078] In a fifth aspect, the present invention provides a computer storage medium storing program instructions, which, when executed, implement the steps of the preferred method described in the first aspect.
[0079] In a sixth aspect, the present invention provides a computer program product, comprising code, characterized in that when the code is executed, it is used to implement the preferred method as described in any one of claims 1-6.
[0080] In a seventh aspect, the present invention provides a chip comprising a processor, characterized in that the processor is used to run the preferred method as described in any one of claims 1-6. BRIEF DESCRIPTION OF THE DRAWINGS
[0081] FIG1 is a flow chart of an embodiment of a preferred method of a parallel strategy disclosed in the present invention;
[0082] FIG2 is an example diagram of segmentation results of a neural network model using the segmentation point strategy of the present invention;
[0083] FIG3 is a schematic structural diagram of an offline training device according to the present invention. DETAILED DESCRIPTION
[0084] Data parallelism distributes a batch of training data evenly across multiple accelerator devices, with each device maintaining a complete copy of the pipeline. After each training session, gradient synchronization communication is required between devices to ensure that the models stored on all devices have exactly the same parameters. Although data parallelism can accelerate the training of target large models by increasing the batch size, it makes model optimization difficult. In addition, when data parallelism is in operation, gradients need to be transmitted through communication to synchronously update model parameters. The amount of communication is extremely large, and for systems with smaller communication bandwidth, it will seriously slow down the operation speed. Pipeline parallelism divides the model by layer, with each device storing a portion of continuous Transformer layers (defined as stages). Due to the sequential execution order between layers during the forward and reverse calculations of the model, there is a large amount of device idle time during pipeline parallel training. Tensor parallelism adopts a complex intra-layer tensor splitting method, but intra-layer splitting introduces excessive global communication overhead.
[0085] Example 1
[0086] The parallel strategies generated by existing technologies cannot adapt to the cluster environment where the target model is located, resulting in strategy unavailability. This application designs different embedded modules to capture environmental information such as computing resources and communication resources, and generates the optimal strategy that takes into account the target model structure and cluster environment, ensuring the correctness and availability of the strategy.
[0087] Referring to FIG1 , the present application provides a preferred method for a parallel strategy, the method comprising:
[0088] S1 obtains basic information of a target large model, where the basic information of the target large model includes candidate tensor parallelism.
[0089] The basic information of the target large model also includes the hierarchical time overhead under the candidate tensor parallelism.
[0090] S11 obtains the candidate tensor parallelism based on the number of GPUs of each device, wherein the value of the candidate tensor parallelism is an integer value less than or equal to the number of GPUs of each device;
[0091] The candidate tensor parallelism can be a set of values less than or equal to the number of GPUs per device, or one of the values. As a preferred embodiment, the candidate tensor parallelism can be a set of all values less than or equal to the number of GPUs per device, because the richer the candidate tensor parallelism values, the better the optimization effect of the parallel strategy.
[0092] S12 splits the target large model to obtain at least one neural network layer, and collects the time overhead of each neural network layer of at least one neural network layer when the data batch size is 1 under the candidate tensor parallelism as the layer time overhead under the tensor parallelism.
[0093] S2 obtains a first parallel strategy based on the candidate tensor parallelism through heuristic candidate strategy calculation. The first parallel strategy includes data parallelism, tensor parallelism and pipeline parallelism. The parameters of the first parallel strategy include candidate data parallelism, candidate tensor parallelism and pipeline parallelism.
[0094] The first parallel strategy is obtained by calculating the candidate tensor parallelism through a heuristic candidate strategy, including:
[0095] S21 traverses the candidate tensor parallelism to perform data parallelism calculation to obtain candidate data parallelism, wherein the candidate data parallelism is calculated based on the candidate tensor parallelism, the number of devices, and the number of GPUs of each device;
[0096] Where factor is the factor acquisition function, dp is the data parallelism, N is the number of devices, M is the number of GPUs per device, and tmp is the tensor parallelism;
[0097] During the data parallelism calculation process, the number of devices and the number of GPUs in each device are both fixed values, and the candidate data parallelism is calculated by traversing the values of the candidate tensor parallelism.
[0098] S22 calculates pipeline parallelism based on the candidate tensor parallelism and the candidate data parallelism to obtain a candidate pipeline parallelism, wherein the pipeline parallelism is calculated based on the candidate tensor parallelism, the candidate data parallelism, the number of devices, and the number of GPUs of each device:
[0099] S23 obtains a first parallel strategy based on the candidate tensor parallelism, the candidate data parallelism, and the candidate pipeline parallelism.
[0100] The first parallel strategy includes data parallelism, tensor parallelism and pipeline parallelism, wherein the data parallelism includes candidate data parallelism, the tensor parallelism includes candidate tensor parallelism, and the pipeline parallelism includes candidate pipeline parallelism.
[0101] In the first parallel strategy, the pipeline parallelism represents the total number of pipeline parallel stages, that is, the target large model is divided into several stages in the forward propagation direction, but the specific division points cannot be known. Under the same pipeline parallelism, differences in the division point positions will result in different transmission efficiency of the target large model.
[0102] S3 calculates the first parallel strategy based on the neural network solver to obtain the split point strategy, adds the split point strategy to the first parallel strategy to obtain the second parallel strategy, and predicts the simulation overhead value of the second parallel strategy.
[0103] The parameters of the second parallel strategy include the candidate data parallelism, the candidate tensor parallelism, the candidate pipeline parallelism and a split point strategy, where the split point strategy is used to determine the position of the pipeline split point.
[0104] Step S3 includes:
[0105] S31 obtains candidate tensor parallelism and candidate pipeline parallelism of the first parallel strategy;
[0106] S32 performs characterization processing based on the layer time overhead under the candidate tensor parallelism and the candidate pipeline parallelism to obtain a feature matrix, and inputs the feature matrix into a neural network solver to obtain pipeline split points under different pipeline parallelisms;
[0107] The neural network solver includes an encoder and a decoder, and step S32 includes:
[0108] S321 performs characterization processing based on the layer time overhead under the candidate tensor parallelism and the candidate pipeline parallelism to obtain a feature matrix;
[0109] In step S32, characterization processing is performed based on the layer time overhead under the candidate tensor parallelism and the candidate pipeline parallelism, including:
[0110] 1) Perform a full estimate based on the candidate pipeline parallelism to obtain the first communication overhead matrix: Based on the candidate pipeline parallelism, the target large model is divided into K stages. Assuming that each neural network layer of the target large model may be the end of any stage and send data to the next stage, there are K-1 communication overhead estimates for the target large model. Therefore, the first communication overhead matrix is a matrix with K-1 rows and L-1 columns, where K is the number of stages obtained by the pipeline parallel segmentation of the target large model, that is, the pipeline parallelism, and L is the total number of neural network layers contained in the target large model. A stage includes at least one continuous neural network layer.
[0111] 2) Perform matrix conversion based on the hierarchical time overhead under the candidate tensor parallelism to obtain a first computational overhead matrix: Since the target large model includes L neural network layers, corresponding to L hierarchical time overheads, the first computational overhead is a vector of L elements, which cannot be aligned with the first communication overhead matrix. Therefore, a simple left-right summation is used for conversion to align the first computational overhead with the first communication overhead based on columns; the positions between the neural network layers of the target large model are recorded as inter-layer points, and each inter-layer point has two eigenvalues, which are the sum of the computational overheads of all neural network layers to the left of the inter-layer point and the sum of the computational overheads of all neural network layers to the right of the inter-layer point. At this time, the computational overhead data is converted into a matrix with 2 rows and L-1 columns.
[0112] 3) Concatenate the first communication overhead matrix and the first computation overhead matrix by columns to obtain a feature matrix with K+1 rows and L-1 columns.
[0113] S322 inputs the feature matrix into the encoder for encoding to obtain a segmentation point code sequence;
[0114] S323 inputs the segmentation point coding sequence into the decoder for calculation to obtain a segmentation point strategy.
[0115] The target large model is divided into K stages according to the pipeline parallelism, that is, a total of K-1 pipeline splitting points. The splitting point strategy is used to determine the specific splitting positions of these K-1 pipeline splitting points. The pipeline is split with the neural network layer as the smallest unit, and one stage includes at least one continuous neural network layer.
[0116] Through offline training of the neural network solver, the neural network solver learns the optimal pipeline segmentation method under different pipeline parallelisms, and obtains a trained neural network solver. Please refer to Example 2 for the training process. Directly migrate and apply the trained neural network solver to calculate the pipeline segmentation point for the first parallel strategy to obtain the segmentation point strategy. The preferred method of the parallel strategy in this application can directly migrate and apply the trained neural network solver to obtain the segmentation point strategy, without the need to use the target domain data to perform further online training on the trained neural network solver, which solves the problem that the existing technology requires online training and cannot directly adapt to the new model, and achieves the effect of quickly and efficiently obtaining the segmentation point strategy, thereby further improving the acquisition efficiency of the second parallel strategy.
[0117] S33 adds the split point strategy to the first parallel strategy to obtain a second parallel strategy, where parameters of the second parallel strategy include the candidate data parallelism, the candidate tensor parallelism, the candidate pipeline parallelism, and the split point strategy.
[0118] To help understand the role of the split point strategy and pipeline parallelism, please refer to Figure 2, which is a schematic diagram of the pipeline split strategy of a 4-layer neural network model in this embodiment. The neural network in Figure 2 consists of 4 neural network layers. The pipeline parallelism pp is 3, indicating that the neural network is divided into 3 stages, including 2 pipeline split points. However, the specific location of the pipeline split points cannot be confirmed by pipeline parallelism alone, nor can the specific content of each stage be confirmed. According to the split point strategy, the two pipeline split points are determined to be located between the 1st layer and the 2nd layer, and between the 3rd layer and the 4th layer, respectively. The three stages obtained by the split point strategy are the 1st stage 210, the 2nd stage 220 and the 3rd stage 230, respectively. Among them, the 1st stage 210 includes the 1st neural network layer 211, the 2nd stage 220 includes the 2nd neural network layer 221 and the 3rd neural network layer 222, and the 3rd stage 230 includes the 4th neural network layer 231.
[0119] The pipeline parallelism is used to indicate the total number of pipeline parallel stages, and the split point strategy is used to indicate the split position of the pipeline parallel. Based on the split point position, the length of each stage can be obtained at the same time. The length of each stage specifically refers to the number of neural network layers in the stage.
[0120] S34 predicts a simulation cost value of the second parallel strategy based on the layer time cost and the parameters of the second parallel strategy.
[0121] S341 obtains any second parallel strategy, the data parallelism G, tensor parallelism F, and pipeline parallelism K of the second parallel strategy. Each stage obtained based on the split point strategy is represented by j, and there are K stages in total, j = 1, 2, .., K.
[0122] S342 obtains the phase overhead of the second parallel strategy, where the phase overhead includes phase calculation overhead and phase communication overhead.
[0123] S3421 obtains the stage computation overhead based on the split point strategy and the layer time overhead calculation; wherein the n-m+1 consecutive neural network layers from m to n contained in the j-th stage are obtained based on the split point strategy, m≤n≤L.
[0124] Among them, t F,i is the layer-level computational overhead of the i-th neural network layer when the candidate tensor parallelism is F, m is the first neural network layer in the j-th stage, n is the last neural network layer in the j-th stage, mbs is the micro-batch size, Where gbs is the preset global batch size and dp is the candidate data parallelism.
[0125] S3422 calculates the communication overhead of each stage based on the data volume and bandwidth of each stage; the communication overhead of transmitting data from stage j to stage j+1 is recorded as e F j This part of the communication overhead will be estimated based on the micro-batch size, data transmission volume and theoretical bandwidth, and is recorded as e F j ; e j =mbs×v j / b j
[0126] Among them, v j represents the amount of data in the jth stage when the batch is 1, b j Represents the bandwidth from the device in stage j to the device in stage j+1.
[0127] S343 calculates the model cost based on the stage cost of the second parallel strategy:
[0128] Among them, t F j Calculate the cost for the jth stage with tensor parallelism F, max t F j The maximum value of the computation overhead for a stage with tensor parallelism of F; e F j is the communication overhead of the jth stage when the tensor parallelism is F, T p is the model overhead of the pipeline corresponding to the data parallel group p, K is the pipeline parallelism, gbs p is the global batch size of the pipeline corresponding to the data parallel group p, p is the data parallel group, p = 1, 2, ..., G.
[0129] The model overhead is calculated under the pipeline strategy, pipeline parallelism and tensor parallelism, and the model overhead is used to measure the performance of pipeline parallelism and tensor parallelism.
[0130] S344 estimates the communication overhead of data parallelism based on the candidate data parallelism to obtain data overhead.
[0131] Obtaining data parallel groups based on data parallelism, and obtaining the minimum bandwidth in the data parallel groups;
[0132] The data overhead is estimated by dividing the data communication volume by the minimum bandwidth in the data parallel grouping to obtain the data overhead.
[0133] Where a is the data communication volume, T dp_sync is the data overhead, minb pis the minimum bandwidth in the data parallel group, p is the data parallel group, p=1,2,…,G.
[0134] The data overhead is obtained under the data parallelism degree, and the data overhead is used to measure the performance of data parallelism.
[0135] S345 calculates the simulation cost value of the second parallel strategy based on the sum of the data cost and the model cost. dp_sync +T p
[0136] Among them, T dp_sync is the data overhead, T p The model overhead.
[0137] S4 screens the second parallel strategies based on the simulated cost values of the second parallel strategies to obtain a theoretically optimal strategy.
[0138] The second parallel strategies are arranged in ascending order based on the simulated global overhead value, and the first second parallel strategy from the sorting result is selected as the theoretical optimal strategy, wherein the theoretical optimal strategy includes the theoretical optimal data parallelism, the theoretical optimal tensor parallelism, the theoretical optimal pipeline parallelism and the theoretical optimal splitting point strategy.
[0139] Based on pipeline parallelism, the target large model is divided into layers and distributed to different devices, with each device storing one stage;
[0140] Devices are grouped according to data parallelism to obtain multiple pipelines, with each data group obtaining a pipeline.
[0141] According to Table 1, the present invention obtains parallel strategies based on heuristic candidate strategies and neural network solvers, which significantly shortens the time cost compared to the AMP algorithm for obtaining parallel strategies. As the model scale of the target large model increases, the time cost of the AMP algorithm for obtaining parallel strategies increases sharply, requiring about 170 minutes on the GPT2-48 layer model, while the present invention only takes 1.2 seconds to complete the acquisition of parallel strategies on the GPT2-48 layer model. The AMP algorithm, used as a comparison method for the present invention, comes from the 2022 public technology "AMP: Automatically Finding Model Parallel Strategies with Heterogeneity Awareness", which is solved based on a dynamic programming algorithm.
[0142] Table 1
[0143] Example 2
[0144] To obtain the pipeline segmentation capability of the neural network solver and determine the optimal segmentation point strategy, the neural network solver needs to be trained offline. Existing technologies use a dynamic programming algorithm to search for strategies online for each model, making it impossible to reuse existing strategy networks. Each search is time-consuming, affecting the efficiency of acquiring strategies. The strategy network needs to be retrained for each new model, resulting in poor migration. The present invention constructs a reusable universal strategy neural network, which is directly applied after self-game training. This eliminates the need for online retraining and learning each time, improving the generalization capability of the strategy network itself. For new model environments, strategies can be directly generated by simply loading the pre-trained strategy network, eliminating the need for repeated training and significantly improving migration and adaptation efficiency.
[0145] An offline training method for a neural network solver, comprising:
[0146] First, an offline training device is constructed. The offline training device includes at least two neural network solvers and one reward module. The neural network solver includes an encoder and a decoder.
[0147] Referring to FIG3 , the training process of an offline training device using two neural network solvers is illustrated. Offline training device 3000 includes an input data acquisition module 3100, a neural network solver module 3200, and a reward module 3300 connected in series. Neural network solver module 3200 includes two neural network solvers connected in parallel (neural network solver 3210 and neural network solver 3220). Neural network solvers 3210 and 3220 share weights.
[0148] Secondly, obtain input data, which includes obtaining generated data and characterization processing.
[0149] Generate data based on uniform distribution, including generation time overhead, random parameters, and communication overhead of the training process;
[0150] The generation time cost is the layer time cost.
[0151] The generated data is subjected to characterization processing to obtain input data.
[0152] The specific steps of the characterization process have been explained in detail in Example 1 and will not be repeated here.
[0153] Then, the input data is input into two neural network solvers respectively to obtain two candidate split point strategies C1 and C2, and two candidate cost values are obtained based on the two candidate split point strategies. The two neural network solvers share weights;
[0154] Both candidate split point strategies C1 and C2 divide the pipeline into K stages. However, due to the different split positions, the length of each stage obtained by splitting is also different, which intuitively shows that the neural network layers included in each stage are inconsistent.
[0155] Based on the generated data, the stage costs of the two split point strategies are obtained respectively. The stage costs include the stage calculation cost and the stage communication cost.
[0156] The communication cost of each stage of the training process is obtained when the generated data is directly generated based on the uniform distribution. For the communication cost from each stage j to stage j+1, we use e j To express.
[0157] The computational cost of each training stage is calculated based on the generation time cost:
[0158] Among them, t i is the computational cost of generating the i-th neural network layer, m is the first neural network layer in the j-th stage, n is the last neural network layer in the j-th stage, and n ≥ m.
[0159] The simulation cost value is obtained based on the stage cost calculation;
[0160] Among them, t j The stage calculation cost of the jth stage, max t j The maximum value of the cost of calculating the phase; e j is the stage communication overhead between the jth stage and the j+1th stage, which is obtained when the generated data is obtained. T is the simulation overhead value, K is the total number of stages, and B is the random parameter, which is obtained when the generated data is obtained. B≥K-1.
[0161] The reward module 3300 performs numerical conversion based on the simulated cost value to obtain a reward score.
[0162] Generally speaking, the simulated global cost value is negatively correlated with the reward score of the split point strategy. That is, the smaller the simulated global cost value, the higher the reward score of the split point strategy given by the reward network.
[0163] By comparing T1 and T2, we can obtain the corresponding reward scores R1 and R2. The reward score value is 1 or -1. When the simulated cost of the split point strategy is small, the reward score is positive. When the simulated cost of the split point strategy is large, the reward score is negative. Specifically, when T1 < T2, R2 is -1 and R1 is 1; when T2 < T1, R1 is -1 and R2 is 1.
[0164] The loss function is calculated based on the probability value and reward score value of the last neural network layer in each stage.
[0165] in, is the probability value of the last neural network layer in the jth stage of the split point strategy C1, is the probability value of the last neural network layer in the jth stage of the split point strategy C2, and R1 and R2 are the reward scores.
[0166] The last neural network layer in each stage sends data to the next stage, so there are K-1 probability values in the entire pipeline that need to be passed to the next stage, so there are K-1 probability values that constitute the joint probability.
[0167] When the training stop condition is reached, the training is stopped and the trained neural network solver is obtained.
[0168] The training stopping condition includes the rate of change of the loss function being less than a preset threshold.
[0169] The existing technology uses dynamic programming algorithms or online training reinforcement learning algorithms to solve parallel strategies, which cannot solve the problems of long solution time and inability to migrate the trained strategy model; the present invention obtains a neural network solver through the offline training method, and directly migrates the application for pipeline splitting, thereby solving the problems of slow solution speed, need for online training, and inability to directly adapt to new models in the existing technology, thereby improving the efficiency of pipeline splitting.
[0170] In a specific embodiment of the present invention, the time overhead includes: computing overhead and communication overhead; the computing overhead includes one or more of omnidirectional computing time, reverse computing time, and parameter update time; the communication overhead includes computing value transmission time, computing value aggregation time, and model gradient aggregation time.
[0171] The above embodiments are only used to illustrate the technical solutions of the present invention rather than to limit the same. Those skilled in the art may modify or make equivalent substitutions for the technical solutions of the present invention without departing from the principles and scope of the present invention. The scope of protection of the present invention shall be subject to the claims.
Claims
1. A method for optimizing a parallel strategy, characterized in that: The preferred method comprises: Acquire basic information of the target large model based on cluster information of the target large model, wherein the basic information of the target large model includes candidate tensor parallelism; Based on the candidate tensor parallelism, determine candidate data parallelism and candidate pipeline parallelism through a heuristic candidate strategy to obtain a first parallel strategy; Calculating the first parallel strategy through a neural network solver to obtain a split point strategy, adding the split point strategy to the first parallel strategy to obtain a second parallel strategy, and predicting a simulation overhead value of the second parallel strategy; The second parallel strategy is screened based on the simulation cost value of the second parallel strategy to obtain a theoretically optimal strategy.
2. The preferred method according to claim 1, characterized in that: The cluster information of the target large model includes the number of GPUs of each device, the basic information of the target large model also includes layer time overhead, the target large model includes at least one neural network layer, and the basic information of the target large model is obtained based on the cluster information of the target large model, including: Acquire the candidate tensor parallelism based on the number of GPUs of each device, where the value of the candidate tensor parallelism is less than or equal to the number of GPUs of each device; Under the candidate tensor parallelism, the time overhead of each neural network layer of the at least one neural network layer is collected as the layer-level time overhead.
3. The preferred method according to any one of claims 1-2, characterized in that: The first parallel strategy is calculated by a neural network solver to obtain a split point strategy, the split point strategy is added to the first parallel strategy to obtain a second parallel strategy, and a simulation cost value of the second parallel strategy is predicted, including: Calculating the parameters of the first parallel strategy by the neural network solver to obtain a split point strategy, wherein the parameters of the first parallel strategy include the candidate tensor parallelism, the candidate pipeline parallelism and the candidate data parallelism; Adding the split point strategy to the parameters of the first parallel strategy to obtain the parameters of the second parallel strategy, wherein the parameters of the second parallel strategy include the candidate tensor parallelism, the candidate pipeline parallelism, the candidate data parallelism and the split point strategy; Based on the parameters of the second parallel strategy and the level time overhead, a simulation overhead value of the second parallel strategy is predicted.
4. The preferred method according to any one of claims 1 to 3, characterized in that: The neural network solver includes an encoder and a decoder. The neural network solver is used to calculate the parameters of the first parallel strategy to obtain the split point strategy, including: Obtaining the candidate tensor parallelism and the candidate pipeline parallelism; Characterizing the hierarchical time overhead and the candidate pipeline parallelism under the candidate tensor parallelism to obtain a feature matrix; Inputting the feature matrix into the encoder for encoding to obtain a segmentation point code sequence; The segmentation point coding sequence is input into the decoder for decoding to obtain the segmentation point strategy.
5. The preferred method according to any one of claims 1 to 3, characterized in that: The simulation cost value includes data cost and model cost, and the predicting the simulation cost value of the second parallel strategy based on the parameter of the second parallel strategy and the level time cost includes: Predicting a model overhead of the second parallel strategy based on the candidate tensor parallelism, the candidate pipeline parallelism, the split point strategy, and the inter-layer time overhead; Among them, t F j The stage cost for calculating the jth stage with candidate tensor parallelism F is max t F j The maximum value of the stage calculation overhead when the candidate tensor parallelism is F; e F j is the communication cost of the jth stage when the candidate tensor parallelism is F, T p is the model overhead of the pipeline corresponding to the data parallel group p, gbs p is the global batch size of the pipeline corresponding to the data parallel group p, p is the data parallel group, p=1,2,…,G, G is the candidate data parallelism, K is the candidate pipeline parallelism, mbs is the micro-batch size; based on the candidate data parallelism, predict the data overhead of the second parallel strategy; Where a is the data communication volume, T dp_sync is the data overhead, minb p is the minimum bandwidth in data parallel group p; Based on the sum of the data overhead and the model overhead, a simulation overhead value of the second parallel strategy is calculated.
6. The preferred method according to any one of claims 2 to 5, characterized in that: The cluster information also includes the number of devices. Based on the candidate tensor parallelism, the candidate data parallelism and the candidate pipeline parallelism are determined by a heuristic candidate strategy to obtain a first parallel strategy, including: Traversing the candidate tensor parallelism to perform data parallelism calculation to obtain the candidate data parallelism, where the candidate data parallelism is calculated based on the candidate tensor parallelism, the number of devices, and the number of GPUs of each device; Pipeline parallelism is calculated based on the candidate tensor parallelism, the candidate data parallelism, the number of devices, and the number of GPUs of each device to obtain the candidate pipeline parallelism.
7. A training method for a neural network solver, characterized in that: The training method comprises: Generate data based on probability distribution; Performing characterization processing on the generated data to obtain input data; Inputting the input data into at least two neural network solvers respectively to obtain at least two split point strategies; At least two simulation cost values are calculated based on the at least two split point strategies; Performing numerical conversion based on at least two simulated cost values to obtain at least two reward scores; A loss function is calculated based on the at least two reward scores and the at least two split point strategies, and weights of the at least two neural network solvers are updated based on the loss function, and the weights are shared between the at least two neural network solvers.
8. The training method according to claim 7, characterized in that: The calculating of the loss function based on the at least two reward scores and the at least two split point strategies includes: Based on the segmentation point strategy, the target large model is segmented to obtain at least one stage, each of the at least one stage includes at least one continuous neural network layer; A loss function is obtained by calculating based on the reward score and the probability value of the last neural network layer of the stage.
9. The training method according to any one of claims 7-8, characterized in that: The probability distribution generation includes uniform distribution generation and random distribution generation.
10. An offline training device, characterized in that: The offline training device comprises: Input data acquisition module: used to obtain input data; Neural network solution module: used to solve the split point and simulate the cost prediction based on the input data to obtain the split point strategy and simulation cost value; Reward module: used to evaluate the split point strategy based on the simulation cost value to obtain a reward score.
11. The off-line training device according to claim 10, characterized in that: The neural network solving module includes at least two neural network solvers, and the neural network solver includes an encoder and a decoder. The encoder and the decoder are connected in series, and the at least two neural network solvers are connected in parallel. The input data acquisition module, the neural network solving module and the reward module are connected in series.
12. An electronic device, characterized in that: The method comprises a memory and a processor, wherein the memory is used to store a computer program, and the processor is used to run the computer program to implement the steps of the preferred method according to any one of claims 1 to 6.
13. A computer storage medium storing program instructions, characterized in that: When the program instructions are executed, the steps of the preferred method according to any one of claims 1 to 6 are implemented.
14. A computer program product comprising code, characterized in that When the code is executed, it is used to implement the preferred method according to any one of claims 1 to 6.
15. A chip, comprising a processor, characterized in that: The processor is configured to execute the preferred method as described in any one of claims 1-6.
Citation Information
Patent Citations
Automatic parallel strategy searching method based on network-level simulation, medium and equipment
CN115879529A
Edge-end collaborative deep learning calculation acceleration system and method
CN116341624A
Parallel strategy search method for efficient training of artificial intelligence large model
CN116680301A
End-side hierarchical neural network model training method and device, and computer equipment
CN116958862A
Layered Gradient Accumulation and Modular Pipeline Parallelism for Improved Training of Machine Learning Models
US20220383084A1
Cited By
Assembly line parallel training method suitable for heterogeneous equipment
CN120429090A
Model parameter determination method and device, chip, electronic equipment, storage medium and computer program product
CN120873877A
GPU (Graphics Processing Unit) program optimization method for parallel environment
CN121029422A
Model distributed training optimization method and device, equipment and storage medium
CN121683918A