Automatic parallelization method and apparatus for large model, and storage medium
By establishing a strategic search space with parallel dimensions from top to bottom and pruning, the high cost and inefficiency problem of manually selecting parallel modes and configuration parameters in large model training is solved, and efficient automatic parallel training and accurate cost estimation are achieved.
Patent Information
- Application Number
- PCT/CN2024/103557
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2023-12-15
- Filing Date
- 2024-07-04
- Publication Date
- 2025-06-19
AI Technical Summary
The prior art manually selects parallel mode and configuration parameters in large-scale model training requires a lot of relevant knowledge and trial and error costs, and automatic parallel tools ignore the impact of communication optimization on policy selection, resulting in low search efficiency.
By building a strategic search space through data parallelism, flow parallelism and tensor parallelism dimensions from top to bottom, and using important observations of experimental or theoretical verification as a criterion, the search space is pruned, the search space is reduced, and the search efficiency is improved.
There is no need to build a search space based on fragments at the operator level, which greatly reduces the search space and improves the search efficiency. By considering communication optimization, the accuracy of cost estimation is improved, and the optimal strategy is selected for parallel training.
Smart Images

Figure CN2024103557_19062025_PF_FP_ABST
Abstract
Description
Large model automatic parallel method, device and storage medium Technical Field
[0001] The present invention relates to the field of large model training, and in particular to a large model automatic parallelization method, device and storage medium. Background Art
[0002] Parallel training is one of the important means to increase the scale of large models and accelerate training. Common parallel methods include data parallelism, tensor parallelism, and pipeline parallelism.
[0003] Since each parallel mode has a large number of configuration parameters and different parallel modes can be mixed, given the available computing cluster and the model structure to be trained, manually selecting the appropriate parallel mode and configuration parameters requires a lot of relevant knowledge and trial and error costs.
[0004] Therefore, some existing technologies provide automatic parallelization tools, such as ColossalAI, but they ignore the impact of communication optimization on strategy selection, and their parallel strategies are based on the combination of sharding strategies for each layer of parameters, resulting in a very large strategy search space and low search efficiency.
[0005] Summary of the Invention
[0006] The purpose of the present invention is to provide a large-scale model automatic parallel method, device and storage medium. It does not need to construct an automatic parallel search space based on operator-level sharding, but instead establishes a strategy search space from top to bottom through several parallel dimensions, and further prunes the search space based on important observations verified experimentally or theoretically as criteria, thereby greatly reducing the search space and improving search efficiency.
[0007] The purpose of the present invention can be achieved by the following technical solutions:
[0008] A large model automatic parallel method, comprising:
[0009] Obtain the model structure of the large model to be trained and the cluster information used for training, wherein the cluster information includes at least the number of nodes and the number of computing cards contained in each node;
[0010] Generate a search space based on the product of the number of nodes and the number of computing cards contained in each node, wherein the search space consists of multiple candidate parallel solutions, and the product of the data parallel dimension, pipeline parallel maintenance, and tensor parallel dimension in each candidate parallel solution is equal to the product of the number of nodes and the number of computing cards contained in each node;
[0011] Based on the model structure and cluster information of the large model and pre-configured rules, the search space is pruned to remove some candidate parallel solutions;
[0012] Calculate the computation time and communication time of all remaining candidate parallel solutions in the search space, and obtain the total time consumption of each candidate parallel solution based on the computation time and communication time of each candidate parallel solution;
[0013] A candidate parallel solution with the shortest total time consumption is selected to perform parallel training on the large model.
[0014] The candidate parallel solutions in the search space are obtained by exhaustive enumeration.
[0015] The model structure includes at least the number of layers of the model, and the preconfigured rules include at least:
[0016] The dimension of pipeline parallelism is less than or equal to the number of layers of the large model.
[0017] The pre-configured rules also include:
[0018] The tensor parallel dimension is less than or equal to the number of computing cards contained in each node.
[0019] The cluster information also includes the video memory of a single computing card.
[0020] The pre-configured rules also include:
[0021] The required video memory of the candidate parallel solution is less than or equal to the video memory of a single computing card, wherein the required video memory is obtained based on the model structure.
[0022] The calculation time is obtained by fitting based on the calculation time of a single sample.
[0023] The total time consumption of each candidate parallel solution is obtained based on the computation time and communication time of each candidate parallel solution, specifically including:
[0024] Determine whether there is masking in the communication and calculation process. If so, the larger of the communication time and the calculation time is used as the total time. Otherwise, the sum of the communication time and the calculation time is used as the total time.
[0025] The large model is trained in parallel based on PyTorch.
[0026] A large model automatic parallel device includes a memory, a processor, and a program stored in the memory, wherein the processor implements the above method when executing the program.
[0027] A storage medium stores a program thereon, wherein the program implements the above method when executed.
[0028] Compared with the prior art, the present invention has the following beneficial effects:
[0029] 1. Instead of constructing an automatically parallel search space based on operator-level sharding, the strategic search space is established top-down through several parallel dimensions. Furthermore, the search space is pruned using important observations verified experimentally or theoretically as guidelines, which greatly reduces the search space and improves search efficiency.
[0030] 2. When making cost estimates, the impact of communication optimization is taken into account, making the cost estimate more accurate and enabling the optimal strategy to be accurately selected. BRIEF DESCRIPTION OF THE DRAWINGS
[0031] FIG1 is a schematic flow chart of the main steps of the method of the present invention;
[0032] FIG2 is a schematic diagram of the technical route of the present invention. DETAILED DESCRIPTION
[0033] The present invention is described in detail below with reference to the accompanying drawings and specific embodiments. This embodiment is implemented based on the technical solution of the present invention, and provides a detailed implementation method and specific operation process, but the protection scope of the present invention is not limited to the following embodiments.
[0034] Overall, this application provides a solution for automated parallel training of large models. Given the model structure and cluster information, we establish a policy search space along three dimensions: data parallelism, pipeline parallelism, and tensor parallelism. We then prune this search space using several key criteria. For each constructed policy, we perform a cost estimate while taking communication optimization into account and select the optimal policy. Finally, we automatically deploy and execute the optimal policy on the cluster using PyTorch.
[0035] Specifically, a large model automatic parallel method, as shown in Figures 1 and 2, includes:
[0036] Obtain the model structure of the large model to be trained and the cluster information used for training, where the cluster information includes at least the number of nodes and the number of computing cards contained in each node;
[0037] Generate a search space based on the product of the number of nodes and the number of computing cards contained in each node, where the search space consists of multiple candidate parallel solutions, and the product of the data parallel dimension, pipeline parallel maintenance, and tensor parallel dimension in each candidate parallel solution is equal to the product of the number of nodes and the number of computing cards contained in each node;
[0038] Based on the model structure and cluster information of the large model and pre-configured rules, the search space is pruned to remove some candidate parallel solutions;
[0039] Calculate the computation time and communication time of all remaining candidate parallel solutions in the search space, and obtain the total time consumption of each candidate parallel solution based on the computation time and communication time of each candidate parallel solution;
[0040] Select the candidate parallel solution with the shortest total time to train the large model in parallel.
[0041] There is no need to construct an automatically parallel search space based on operator-level sharding. Instead, a top-down strategy search space is established through several parallel dimensions. The search space is further pruned using important observations verified experimentally or theoretically as guidelines, which greatly reduces the search space and improves search efficiency.
[0042] In some embodiments, candidate parallel solutions in the search space are obtained by exhaustive enumeration, if a more comprehensive search space can be obtained. However, in some embodiments, candidate parallel solutions can also be obtained based on some empirical criteria, thereby reducing the total processing time of pruning and subsequent calculations.
[0043] When the number of nodes is greater than one, inter-node communication is very expensive, and pipeline parallelism requires less communication than other approaches, so pipeline parallelism should be prioritized. However, for large models, which are typically partitioned by transformer blocks, the pipeline parallelism dimension cannot exceed the number of transformer blocks. Therefore, in most implementations, the model structure includes at least the number of model layers, and the preconfigured rule includes at least: the pipeline parallelism dimension must be less than or equal to the number of layers in the large model.
[0044] At the same time, since the communication volume of tensor parallelism is very large, tensor parallelism should not be considered between nodes, that is, the tensor parallel dimension should be less than or equal to the number of computing accelerators in each node. Therefore, in most embodiments, the preconfigured rules also include: the tensor parallel dimension is less than or equal to the number of computing cards contained in each node.
[0045] Moreover, in most embodiments, the cluster information also includes the video memory of a single computing card, and the pre-configured rules also include: the video memory required by the candidate parallel solution is less than or equal to the video memory of a single computing card, wherein the required video memory is obtained based on the model structure. Generally, the model structure includes the number of layers of the model, the size of the hidden features, the sequence length, and the batch size of the training data. Based on this information, the video memory size required for each candidate parallel solution can be calculated. If the video memory size required by a candidate parallel solution exceeds the video memory of a single computing card, this candidate parallel solution needs to be eliminated.
[0046] Furthermore, when a parallel dimension is 1, it indicates that the parallel strategy is not used. Only when the dimension is greater than 1 does it indicate that the strategy is used. Based on this, we can construct practical strategies based on the parallel dimension in the strategy search space. These include individual strategies such as data parallelism and pipeline parallelism, as well as combined strategies such as data + pipeline parallelism, data + tensor parallelism, pipeline + tensor parallelism, and pipeline + tensor + data parallelism.
[0047] Finally, in determining the total time consumption, for computation time, this embodiment measures the sample-by-sample computation time on a single device and estimates the overall computation time using a fitting function. For communication time, we use cluster information to determine the intra-node and inter-node bandwidth, theoretically calculate the communication volume, and divide this by the bandwidth to obtain the estimated communication time. Furthermore, parallel training involves significant overlap between communication and computation to accelerate training. For each parallel strategy, either individually or in combination, when estimating the actual cost, this application uses the larger of the communication and computation times for the overlapped portion, rather than the sum of the two, to achieve a more accurate cost estimate.
[0048] The identification of mutually masked parts can be done by setting up a judgment library based on experience, and then obtaining them through table lookup. For example, gradient synchronous communication and reverse calculation in data parallelism, and P2P communication and forward and reverse calculation at different stages in stream parallelism all have mutually masked parts.
[0049] If the above functions are implemented in the form of software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, or the part that contributes to the prior art, or the part of the technical solution, can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes several instructions for enabling a computer device (which can be a personal computer, server, or network device, etc.) to execute all or part of the steps of the method described in each embodiment of the present invention. The aforementioned storage medium includes: various media that can store program codes, such as a USB flash drive, a mobile hard disk, a read-only memory (ROM), a random access memory (RAM), a magnetic disk, or an optical disk.
Claims
1. A large model automatic parallel method, characterized in that: include: Obtaining the model structure of the large model to be trained and the cluster information used for training, wherein the cluster information includes at least the number of nodes and the number of computing cards contained in each node; Generate a search space based on the product of the number of nodes and the number of computing cards contained in each node, wherein the search space consists of multiple candidate parallel schemes, and the product of the data parallel dimension, pipeline parallel maintenance and tensor parallel dimension in each candidate parallel scheme is equal to the product of the number of nodes and the number of computing cards contained in each node; Based on the model structure and cluster information of the large model and pre-configured rules, the search space is pruned to remove some candidate parallel solutions; Calculate the computation time and communication time of all remaining candidate parallel solutions in the search space, and obtain the total time consumption of each candidate parallel solution based on the computation time and communication time of each candidate parallel solution; A candidate parallel solution with the shortest total time consumption is selected to perform parallel training on the large model.
2. The large model automatic parallel method according to claim 1, characterized in that: The candidate parallel solutions in the search space are obtained by exhaustive enumeration.
3. The large model automatic parallel method according to claim 1, characterized in that: The model structure at least includes the number of layers of the model, and the preconfigured rules at least include: The dimension of pipeline parallelism is less than or equal to the number of layers of the large model.
4. The large model automatic parallel method according to claim 3, characterized in that: The pre-configured rules also include: The tensor parallel dimension is less than or equal to the number of computing cards contained in each node.
5. A large model automatic parallel method according to claim 3 or 4, characterized in that: The cluster information also includes the video memory of a single computing card. The pre-configured rules also include: The required video memory of the candidate parallel solution is less than or equal to the video memory of a single computing card, wherein the required video memory is obtained based on the model structure.
6. The large model automatic parallel method according to claim 1, characterized in that: The calculation time is obtained by fitting based on the calculation time of a single sample.
7. The large model automatic parallel method according to claim 1, characterized in that: The total time consumption of each candidate parallel solution is obtained based on the calculation time and communication time of each candidate parallel solution, specifically including: Determine whether there is concealment in the communication and calculation process. If so, the larger one of the communication time and the calculation time is taken as the total time consumption. Otherwise, the sum of the communication time and the calculation time is taken as the total time consumption.
8. The large model automatic parallel method according to claim 1, characterized in that: The large model is trained in parallel based on PyTorch.
9. A large model automatic parallel device, comprising a memory, a processor, and a program stored in the memory, characterized in that: When the processor executes the program, the method according to any one of claims 1 to 8 is implemented.
10. A storage medium having a program stored thereon, characterized in that: When the program is executed, the method according to any one of claims 1 to 8 is implemented.
Citation Information
Patent Citations
Automatic parallel strategy searching method based on network-level simulation, medium and equipment
CN115879529A
Scheduling strategy determination method and system for pipeline parallel training
CN116450312A
Large model automatic parallelization method and device and storage medium
CN117892787A
Layered Gradient Accumulation and Modular Pipeline Parallelism for Improved Training of Machine Learning Models
US20220383084A1
Method, device, and system for determining distributed training algorithm framework configuration
WO2023123275A1