Interleaved Pipeline Parallel Training Method, Apparatus, Device, Storage Medium, and Program Product
By determining the number of layers contained in each block based on the calculation unit number, block number and total number of layers in large model training, the interleaved pipeline parallel training method is used to solve the problem of memory capacity limitation of the calculation unit, and a wider training efficiency improvement and applicable scenario expansion are achieved.
Patent Information
- Application Number
- CN202411968005.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-30
- Publication Date
- 2025-07-29
- Estimated Expiration
- 2044-12-30
AI Technical Summary
In the prior art, the memory capacity of the computing unit is limited, resulting in the interleaved pipeline parallel training only when the number of layers of the computing unit can be divided by the number of blocks during training, which limits the improvement of training efficiency.
By determining the number of layers contained in each block based on the calculation unit number, the number of blocks and the total number of layers of the large model, the interleaved pipeline parallel training method is adopted to break through the constraint that the number of layers responsible for each calculation unit must be divisible by the number of blocks, and the interleaved pipeline parallel training process of equal division and non-equality division is supported, and block allocation is combined with bubble time and memory consumption as indicators.
It has achieved improvement in training efficiency in more scenarios, expanded the applicability of parallel training of interleaved pipelines, and achieved a good compromise between bubble time and memory consumption.
Smart Images

Figure CN119376796B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of artificial intelligence (AI), and more specifically, to an interleaved pipeline parallel training method, device, equipment, storage medium, and program product. Background Art
[0002] In the field of deep learning, a computing unit (such as a GPU or GPGPU) is usually used to perform model training. The video memory capacity of the computing unit is generally limited. For the training scenario of a large model, the parameter scale of the large model is usually huge (for example, the number of parameters ranges from billions to trillions). Pipeline Parallelism (PP) refers to splitting the model by layer, deploying different layers of the model to different computing units for pipeline scheduling, and each computing unit is called a different PP stage.
[0003] Pipeline parallel training usually includes non-interleaved and interleaved. Currently, interleaved pipeline parallel training can only be performed when the number of layers of the computing unit can be divided evenly by the number of chunks, which is not conducive to improving the training efficiency. Summary of the Invention
[0004] The present invention provides an interleaved pipeline parallel training method, device, equipment, storage medium, and program product, which helps to improve the training efficiency.
[0005] The technical solutions of the embodiments of the present invention are as follows:
[0006] An interleaved pipeline parallel training method, comprising:
[0007] Based on the number of computing units, the number of chunks, and the total number of layers of the large model, determine the number of layers included in each chunk, where the number of layers included in each chunk is configurable; the determining the number of layers included in each chunk based on the number of computing units, the number of chunks, and the total number of layers of the large model includes one of the following: determining the number of layers included in each chunk with bubble time as an indicator based on the number of computing units, the number of chunks, and the total number of layers; determining the number of layers included in each chunk with video memory consumption as an indicator based on the number of computing units, the number of chunks, and the total number of layers; determining the number of layers included in each chunk with bubble time and video memory consumption as indicators based on the number of computing units, the number of chunks, and the total number of layers;
[0008] Based on the number of layers included in each chunk, perform interleaved pipeline parallel training on the large model.
[0009] In one embodiment, determining the number of layers included in each block with the bubble time as an indicator based on the number of computing units, the number of blocks, and the total number of layers includes:
[0010] Based on the number of computing units and the number of blocks, determine multiple candidate solutions for allocating the total number of layers to each block, where each candidate solution includes the respective number of layers included in each block;
[0011] Determine the bubble time of each candidate solution;
[0012] Select the candidate solution with the least bubble time from the multiple candidate solutions;
[0013] Based on the candidate solution with the least bubble time, determine the number of layers included in each block.
[0014] In one embodiment, determining the bubble time of each candidate solution includes:
[0015] Determine the forward time for each block of each candidate solution to perform forward calculation on a unit micro-batch;
[0016] Determine the maximum forward time from the multiple forward times of the multiple blocks of each candidate solution;
[0017] Determine the backward time for each block of each candidate solution to perform backward calculation on a unit micro-batch;
[0018] Determine the maximum backward time from the multiple backward times of the multiple blocks of each candidate solution;
[0019] Based on the maximum forward time, the maximum backward time, and the number of computing units, determine the bubble time of each candidate solution.
[0020] In one embodiment, determining the number of layers included in each block with the video memory consumption as an indicator based on the number of computing units, the number of blocks, and the total number of layers includes:
[0021] Based on the number of computing units and the number of blocks, determine multiple candidate solutions for allocating the total number of layers to each block, where each candidate solution includes the respective number of layers included in each block;
[0022] Determine the video memory consumption of the computing unit that first performs the calculation in each candidate solution;
[0023] Select the candidate solution with the least video memory consumption from the multiple candidate solutions;
[0024] Based on the candidate solution with the least video memory consumption, determine the number of layers included in each block.
[0025] In one embodiment, determining the video memory consumption of the computing unit that first performs calculations in each candidate solution includes:
[0026] Determining the first activation value of all blocks of the computing unit that first performs calculations in each candidate solution for performing calculations on a unit micro-batch;
[0027] Determining the second activation value of the first block of the computing unit that first performs calculations in each candidate solution for performing calculations on a unit micro-batch;
[0028] Based on the first activation value, the second activation value, and the number of computing units, determining the video memory consumption of the computing unit that first performs calculations in each candidate solution.
[0029] In one embodiment, determining the number of layers included in each block with bubble time and video memory consumption as metrics based on the number of computing units, the number of blocks, and the total number of layers includes:
[0030] Based on the number of computing units and the number of blocks, determining multiple candidate solutions for allocating the total number of layers to each block, where each candidate solution includes the respective number of layers included in each block;
[0031] Determining the video memory consumption of the computing unit that first performs calculations in each candidate solution;
[0032] Selecting at least one candidate solution from the multiple candidate solutions whose video memory consumption is lower than a first predetermined threshold;
[0033] Determining the bubble time of the at least one candidate solution;
[0034] Selecting the candidate solution with the least bubble time from the at least one candidate solution;
[0035] Based on the candidate solution with the least bubble time, determining the number of layers included in each block.
[0036] In one embodiment, determining the number of layers included in each block with bubble time and video memory consumption as metrics based on the number of computing units, the number of blocks, and the total number of layers includes:
[0037] Based on the number of computing units and the number of blocks, determining multiple candidate solutions for allocating the total number of layers to each block, where each candidate solution includes the respective number of layers included in each block;
[0038] Determining the bubble time of each candidate solution;
[0039] Selecting at least one candidate solution from the multiple candidate solutions whose bubble time is lower than a second predetermined threshold;
[0040] Determine the video memory consumption of the computing unit that first performs calculations among the at least one candidate solution;
[0041] Select the candidate solution with the minimum video memory consumption from the at least one candidate solution;
[0042] Based on the candidate solution with the minimum video memory consumption, determine the number of layers included in each block.
[0043] In one embodiment, the determining the number of layers included in each block based on the number of computing units, the number of blocks, and the total number of layers, with bubble time and video memory consumption as metrics includes:
[0044] Based on the number of computing units and the number of blocks, determine multiple candidate solutions for distributing the total number of layers to each block, and each candidate solution includes the respective number of layers included in each block;
[0045] Determine the bubble time of each candidate solution and the video memory consumption of the computing unit that first performs calculations in each candidate solution;
[0046] Select at least one candidate solution from the multiple candidate solutions whose video memory consumption is lower than the first predetermined threshold and whose bubble time is lower than the second predetermined threshold;
[0047] Select the candidate solution with the minimum video memory consumption or the candidate solution with the least bubble time from the at least one candidate solution;
[0048] Based on the selected candidate solution, determine the number of layers included in each block.
[0049] In one embodiment, before determining the number of layers included in each block based on the number of computing units, the number of blocks, and the total number of layers of the large model, the method further includes:
[0050] Based on the total number of layers and the number of computing units, determine the number of layers of each computing unit;
[0051] The determining the number of layers included in each block based on the number of computing units, the number of blocks, and the total number of layers of the large model includes:
[0052] When the number of layers of each computing unit is not divisible by the number of blocks, the number of layers included in at least two blocks is different;
[0053] When the number of layers of each computing unit is divisible by the number of blocks, the number of layers included in each block is the same.
[0054] An interleaved pipelined parallel training device, applied to any of the above interleaved pipelined parallel training methods, includes:
[0055] A determination module, configured to determine the number of layers included in each block based on the number of computing units, the number of blocks, and the total number of layers of the large model, where the number of layers included in each block is configurable;
[0056] A training module, configured to perform interleaved pipeline parallel training on the large model based on the number of layers included in each block; the determination module is configured to perform one of the following:
[0057] Determine the number of layers included in each block with bubble time as an indicator based on the number of computing units, the number of blocks, and the total number of layers;
[0058] Determine the number of layers included in each block with video memory consumption as an indicator based on the number of computing units, the number of blocks, and the total number of layers;
[0059] Determine the number of layers included in each block with bubble time and video memory consumption as indicators based on the number of computing units, the number of blocks, and the total number of layers.
[0060] In one embodiment, the determination module is configured to: based on the number of computing units and the number of blocks, determine multiple candidate schemes for allocating the total number of layers to each block, where each candidate scheme includes the respective number of layers included in each block; determine the bubble time of each candidate scheme; select the candidate scheme with the least bubble time from the multiple candidate schemes; and based on the candidate scheme with the least bubble time, determine the number of layers included in each block.
[0061] In one embodiment, the determination module is configured to: based on the number of computing units and the number of blocks, determine multiple candidate schemes for allocating the total number of layers to each block, where each candidate scheme includes the respective number of layers included in each block; determine the video memory consumption of the computing unit that performs the calculation first in each candidate scheme; select the candidate scheme with the minimum video memory consumption from the multiple candidate schemes; and based on the candidate scheme with the minimum video memory consumption, determine the number of layers included in each block.
[0062] In one embodiment, the determination module is configured to perform one of the following:
[0063] Based on the number of computing units and the number of blocks, determine multiple candidate schemes for allocating the total number of layers to each block, where each candidate scheme includes the respective number of layers included in each block; determine the video memory consumption of the computing unit that performs the calculation first in each candidate scheme; select at least one candidate scheme whose video memory consumption is lower than a first predetermined threshold from the multiple candidate schemes; determine the bubble time of the at least one candidate scheme; select the candidate scheme with the least bubble time from the at least one candidate scheme; and based on the candidate scheme with the least bubble time, determine the number of layers included in each block;
[0064] Based on the number of computing units and the number of partitions, determine multiple candidate solutions for allocating the total number of layers to each partition, where each candidate solution includes the respective number of layers included in each partition; determine the bubble time of each candidate solution; select at least one candidate solution from the multiple candidate solutions whose bubble time is lower than a second predetermined threshold; determine the video memory consumption of the computing unit that first performs the calculation in the at least one candidate solution; select the candidate solution with the minimum video memory consumption from the at least one candidate solution; based on the candidate solution with the minimum video memory consumption, determine the number of layers included in each partition; based on the number of computing units and the number of partitions, determine multiple candidate solutions for allocating the total number of layers to each partition, where each candidate solution includes the respective number of layers included in each partition; determine the bubble time of each candidate solution and the video memory consumption of the computing unit that first performs the calculation in each candidate solution; select at least one candidate solution from the multiple candidate solutions whose video memory consumption is lower than a first predetermined threshold and whose bubble time is lower than a second predetermined threshold; select the candidate solution with the minimum video memory consumption or the candidate solution with the least bubble time from the at least one candidate solution; based on the selected candidate solution, determine the number of layers included in each partition.
[0065] An electronic device, comprising:
[0066] A memory;
[0067] A processor;
[0068] Wherein an application program executable by the processor is stored in the memory, and is used to cause the processor to execute the interleaved pipeline parallel training method described in any one of the above.
[0069] A computer-readable storage medium, on which computer-readable instructions are stored, and when the computer-readable instructions are executed by a processor, the processor is caused to execute the interleaved pipeline parallel training method described in any one of the above.
[0070] A program product, comprising a computer program, and when the computer program is executed by a processor, the interleaved pipeline parallel training method described in any one of the above is implemented.
[0071] As can be seen from the above technical solution, in the embodiment of the present invention, based on the number of computing units, the number of partitions, and the total number of layers of the large model, the number of layers included in each partition is determined, where the number of layers included in each partition is configurable. Based on the number of computing units, the number of partitions, and the total number of layers of the large model, determining the number of layers included in each partition includes one of the following: determining the number of layers included in each partition with the bubble time as an indicator based on the number of computing units, the number of partitions, and the total number of layers; determining the number of layers included in each partition with the video memory consumption as an indicator based on the number of computing units, the number of partitions, and the total number of layers; determining the number of layers included in each partition with the bubble time and the video memory consumption as indicators based on the number of computing units, the number of partitions, and the total number of layers; performing interleaved pipelined parallel training on the large model based on the number of layers included in each partition. It can be seen that by breaking through the constraint that the number of layers responsible for each computing unit must be divisible by the number of partitions, an interleaved pipelined parallel training process that supports both equal and unequal partitioning is realized, improving the training efficiency and expanding the applicable scenarios of the interleaved pipelined parallel training. In addition, the number of layers included in the partition can also be allocated based on the bubble time or the video memory consumption, and a good compromise can also be achieved between the bubble time and the video memory consumption. Description of the Drawings
[0072] Figure 1 FIG. is a schematic diagram of a non-interleaved one-forward-one-backward (1F1B) pipelined parallel training process according to the related art.
[0073] Figure 2 FIG. is a schematic diagram of an interleaved pipelined parallel training process according to the related art.
[0074] Figure 3 FIG. is a schematic flowchart of an interleaved pipelined parallel training method according to an embodiment of the present invention.
[0075] Figure 4 FIG. is a schematic diagram of an interleaved pipelined parallel training process according to an embodiment of the present invention.
[0076] Figure 5 FIG. is a schematic structural diagram of an interleaved pipelined parallel training device according to an embodiment of the present invention.
[0077] Figure 6 FIG. is a schematic structural diagram of an electronic device according to an embodiment of the present invention. Detailed Embodiments
[0078] To make the objectives, technical solutions, and advantages of the present invention clearer, the present invention will be further described in detail below with reference to the accompanying drawings.
[0079] For the sake of brevity and intuitiveness in description, the solutions of the present invention will be elaborated below by describing several representative embodiments. A large number of details in the embodiments are only used to help understand the solutions of the present invention. However, it is obvious that the technical solutions of the present invention can be implemented without being limited to these details. In order to avoid unnecessarily obscuring the solutions of the present invention, some embodiments are not described in detail, but only the frameworks are given. Hereinafter, "including" means "including but not limited to", and "according to..." means "at least according to..., but not limited to only according to...". Due to the language habits of Chinese, when the quantity of a component is not specifically indicated hereinafter, it means that the component can be one or more, or can be understood as at least one.
[0080] In pipelined parallel training (usually 1F1B), each layer of the large model is separately assigned to its own computing unit, and each training batch is divided into multiple micro-batches. Each micro-batch performs forward propagation and backward propagation in sequence among different computing units, which can maximize the utilization rate of computing devices and improve training efficiency.
[0081] Figure 1 It is a schematic diagram showing an exemplary non-interleaved 1F1B pipelined parallel training process according to the related art. The core idea of the 1F1B strategy is that after one micro-batch forward pass of the large model, a backward pass is immediately performed, thereby reducing memory occupancy and parallel bubbles, improving training efficiency and throughput. This strategy is particularly effective when dealing with long sequences and can significantly improve the efficiency and stability of large model training. Under this scheduling mechanism, the computing devices at each stage start to process and generate outputs immediately after receiving the outputs from the previous stage, and then pass them to the computing devices at the next stage. However, this scheduling mechanism may lead to low memory usage efficiency or the generation of pipeline bubbles (i.e., idle time).
[0082] In Figure 1 taking four computing units A to D as an example for description. As Figure 1 shown, it is assumed that the number of micro-batches that can be processed in one non-interleaved scheduling is 8, that is, the accumulation count is 8. The computing units A to D execute their respective computing processes in parallel. Among them:
[0083] For computing unit A, it sequentially executes: (1) forward operations for micro-batch numbers 1 to 4 ( Figure 1 identified by the dark blue boxes with numbers 1 to 4 in Figure 1 ); (2) wait (i.e., bubble time, Figure 1is identified by a light green square with the number 1); (4) the forward operation of micro-batch number 5 ( Figure 1 is identified by a dark blue square with the number 5); (5) the backward operation of micro-batch number 2 ( Figure 1 is identified by a light green square with the number 2)… and so on.
[0084] For computing unit B, execute in sequence: (1) Wait; (2) The forward operations of micro-batch numbers 1 to 4; (3) Wait; (4) The backward operation of micro-batch number 1; (5) Wait; (6) The backward operation of micro-batch number 2… and so on.
[0085] For computing unit C, execute in sequence: (1) Wait; (2) The forward operations of micro-batch numbers 1 to 2; (3) Wait; (4) The backward operation of micro-batch number 1; (5) The forward operation of micro-batch number 3; (6) The backward operation of micro-batch number 2… and so on.
[0086] For computing unit D, execute in sequence: (1) Wait; (2) The forward operation of micro-batch number 1; (3) The backward operation of micro-batch number 1; (4) The forward operation of micro-batch number 2; (5) The backward operation of micro-batch number 2; (6) The forward operation of micro-batch number 3… and so on.
[0087] It can be seen that Figure 1 there are a large number of bubble times (i.e., a large number of white squares) in the pipelined parallelism of
[0088] Currently, an interleaved 1F1B pipeline scheduling is adopted to try to solve the technical problem of the too long bubble time. By allocating multiple chunks to computing units, the idle time is reduced and the overall throughput is improved. In the interleaved 1F1B pipeline scheduling, each computing unit can calculate several layers in the model.
[0089] Figure 2 is a schematic diagram showing an exemplary interleaved pipeline parallel training process according to related technologies. In the interleaved scheduling, each computing unit is allocated multiple chunks (exemplarily 2 chunks in Figure 2 ), allowing interleaved execution of forward and backward between multiple chunks at different stages. For example, for interleaved 1F1B, on the basis of non-interleaved 1F1B, a finer-grained division is performed. All the layers originally responsible for each computing unit are further divided according to the number of chunks, and multiple chunks are called interleaved, which can reduce the bubble time.
[0090] In Figure 2In this example, four computing units A to D are used for description. Assume that the number of micro-batches that can be processed in one interleaved scheduling is 8 (i.e., the accumulation count is 8), and the total number of layers of the large model is 40. Then the number of layers responsible for calculation by each computing unit is 40 / 4 = 10 layers. Assume that the number of chunks is 2, divided into chunk 1 and chunk 2, and the number of layers in each chunk is 10 / 2 = 5. Assume that chunk 1 contains layers 1 to 5, and chunk 2 includes layers 6 to 10.
[0091] Taking computing unit A as an example, the following operations are performed in sequence: (1) Forward operations on chunk 1 (containing layers 1 to 5) of micro-batch sequence numbers 1 to 4 ( Figure 2 identified by the dark blue boxes with numbers 1 to 4 in Figure 2 ); (2) Forward operations on chunk 2 (containing layers 6 to 10) of micro-batch sequence numbers 1 to 4 ( Figure 2 identified by the light blue boxes with numbers 1 to 4 in Figure 2 ); (3) Forward operations on chunk 1 (containing layers 1 to 5) of micro-batch sequence numbers 5 to 6 ( Figure 2 identified by the dark blue boxes with numbers 5 to 6 in Figure 2 ); (4) Wait (i.e., bubble time, Figure 2 identified by the white boxes in Figure 2 ); (5) Forward operations on chunk 1 (containing layers 1 to 5) of micro-batch sequence number 7 ( Figure 2 identified by the dark blue box with number 7 in Figure 2 ); (6) Backward operations on chunk 2 (containing layers 6 to 10) of micro-batch sequence number 1 ( Figure 2 identified by the light green box with number 1 in Figure 2 ); (7) Forward operations on chunk 1 (containing layers 1 to 5) of micro-batch sequence number 8 ( Figure 2 identified by the dark blue box with number 8 in Figure 2identified by a dark green square with the number 1);... and so on.
[0092] Similarly, computing units B to D are assigned 2 chunks, allowing for interleaved execution of forward and backward passes between different stages.
[0093] It can be seen that compared with Figure 1 , Figure 2 the waiting time (i.e., bubbles) in the pipeline parallelism of Figure 2 is significantly reduced. However, in the related art taking Figure 2 as an example, only when the number of layers (e.g., 10) responsible for by a computing unit can be divided evenly by the number of chunks (e.g., 2), can interleaved pipeline parallelism be executed. Otherwise, when the number of layers responsible for by a computing unit cannot be divided evenly by the number of chunks, it is necessary to fallback to execute a non-interleaved pipeline parallel training process, which is not conducive to improving the training efficiency. For example, when Figure 2 the number of layers responsible for by the computing unit in Figure 1 is 11 and the number of chunks is 2, then interleaved pipeline parallelism is not executed, and a non-interleaved pipeline parallel training process as shown in Figure 1 is executed.
[0094] In other words, in the commonly used distributed frameworks in the industry currently, only the interleaved pipeline parallel training process with evenly divisible layers is supported. Assume that the total number of layers of a large model is N, the number of computing units is T, and the number of layers responsible for by each computing unit is M (i.e., M = N / T). In the interleaved pipeline parallel training process with evenly divisible layers, it is required that M must be divisible by the number of chunks (i.e., the result of dividing M by the number of chunks is an integer without a remainder), and the number of layers responsible for by each chunk needs to be the same. If the number of layers M responsible for by a computing device cannot be divided evenly by the number of chunks, then the interleaved pipeline parallel training cannot be used, and it can only fallback to the non-interleaved pipeline parallel training, thus limiting the applicable scenarios of the interleaved pipeline parallel training.
[0095] Considering that the factors affecting the training process mainly include: (1) bubble time, where finer-grained chunking can reduce the bubble time generated by pipeline parallelism; (2) video memory, the number of chunks affects the video memory consumption of each computing unit. In the embodiments of the present invention, the constraint that the number of layers responsible for by each computing unit must be divisible by the number of chunks is broken, and an interleaved pipeline parallel training process supporting both evenly divided chunks and non-evenly divided chunks is realized. In the embodiments of the present invention, according to the user's configuration, combined with the parameters of the current large model (such as the total number of model layers and the number of computing units, etc.) and information such as the video memory capacity on the computing unit, etc., and considering factors such as the video memory consumption and bubble time of the current model, the chunk size (i.e., the number of layers included in each chunk) in the interleaved pipeline parallel training process that meets the user's expectations can be automatically allocated.
[0096] The above disclosure details the technical defects existing in the related art, the reasons for these technical defects, and the thought analysis process for overcoming these technical defects. In fact, the recognition of the above technical defects is not common knowledge in this field, but a novel discovery by the inventor during research. Additionally, the tracing of the reasons for the technical defects and the thought analysis process for overcoming these technical defects are also the step-by-step analysis results of the inventor during the actual research process, and are not common knowledge in this field.
[0097] Figure 3 FIG. is a schematic flowchart of an interleaved pipeline parallel training method according to an embodiment of the present invention.
[0098] As Figure 3 shown, the method includes:
[0099] Step 101: Determine the number of layers included in each block based on the number of computing units, the number of blocks, and the total number of layers of the large model, where the number of layers included in each block is configurable. Determining the number of layers included in each block based on the number of computing units, the number of blocks, and the total number of layers of the large model includes one of the following: determining the number of layers included in each block with the bubble time as an indicator based on the number of computing units, the number of blocks, and the total number of layers; determining the number of layers included in each block with the video memory consumption as an indicator based on the number of computing units, the number of blocks, and the total number of layers; determining the number of layers included in each block with the bubble time and the video memory consumption as indicators based on the number of computing units, the number of blocks, and the total number of layers.
[0100] Here, the number of computing units and the total number of layers of the large model are known parameters. Among them: the number of blocks can be specified by the user or can be a specified default value (the default value can be changed). Among them: the number of layers included in each block is configurable. That is to say: the number of layers included in each block can be the same, or at least two blocks can have different numbers of layers.
[0101] In one embodiment, before step 101, the method further includes: determining the number of layers of each computing unit based on the total number of layers and the number of computing units; determining the number of layers included in each block based on the number of computing units, the number of blocks, and the total number of layers of the large model includes: when the number of layers of each computing unit is not divisible by the number of blocks, at least two blocks have different numbers of layers; when the number of layers of each computing unit is divisible by the number of blocks, each block has the same number of layers.
[0102] For example: Assume that the total number of layers of the large model is N, the number of computing units is T, and the number of layers responsible for each computing unit is M (that is, M = N / T). When the number of layers M of each computing unit is not divisible by the number of blocks, at least two blocks have different numbers of layers; when the number of layers of each computing unit is divisible by the number of blocks, each block has the same number of layers.
[0103] Therefore, in the embodiments of the present invention, the constraint that the number of layers responsible for each computing unit must be divisible by the number of chunks is broken, and the number of layers included in each chunk is made configurable, so that an interleaved pipeline parallel training process that supports both equal-sized chunks and non-equal-sized chunks can be supported simultaneously.
[0104] When the number of layers M of each computing unit is not divisible by the number of chunks, the number of layers included in at least two chunks is different. At this time, the number of layers included in each chunk can be specifically determined based on the bubble time or the video memory consumption.
[0105] In one embodiment, step 101 includes one of the following:
[0106] Method (1): Based on the number of computing units, the number of chunks, and the total number of layers, determine the number of layers included in each chunk with the bubble time as an indicator.
[0107] In method (1), determine the number of layers included in each chunk with the bubble time as an indicator. Usually, after determining the number of layers included in the chunk, the bubble time is minimized.
[0108] In one embodiment, method (1) specifically includes: Based on the number of computing units and the number of chunks, determine multiple candidate schemes for distributing the total number of layers to each chunk, and each candidate scheme includes the respective number of layers included in each chunk; determine the bubble time of each candidate scheme; select the candidate scheme with the least bubble time from multiple candidate schemes; based on the candidate scheme with the least bubble time, determine the number of layers included in each chunk.
[0109] In one embodiment, determining the bubble time of each candidate scheme includes: determining the forward time for each chunk of each candidate scheme to perform forward calculation on a unit micro-batch; determining the maximum forward time from the multiple forward times of multiple chunks of each candidate scheme; determining the backward time for each chunk of each candidate scheme to perform backward calculation on a unit micro-batch; determining the maximum backward time from the multiple backward times of multiple chunks of each candidate scheme; based on the maximum forward time, the maximum backward time, and the number of computing units, determine the bubble time of each candidate scheme. Therefore, based on the maximum forward time, the maximum backward time, and the number of computing units, the quantization process of the bubble time can be realized.
[0110] The following exemplary description discusses the typical process of quantifying and evaluating the bubble time of candidate schemes.
[0111] If there is no restriction on the number of layers contained in each block to achieve uneven partitioning, the size of the bubble time will be determined by the block with the largest number of layers. The total forward time of all layers responsible for by the computing unit is equal to the sum of the forward times of all blocks, and the total backward time of all layers is equal to the sum of the backward times of all blocks. For the interleaved pipelining scheduling of uneven partitioning, the total bubble time will be determined by the block with the largest time consumption (i.e., the block with the largest number of layers).
[0112] Therefore, the total bubble time ( T ) of any -th computing unit is related to the summation result of the longest forward time ( ) among all blocks and the longest backward time ( ) among all blocks.
[0113] ; Formula (1)
[0114] Wherein: is the forward time of a micro-batch of the T -th block of the i -th computing unit; is the backward time of a micro-batch of the T -th block of the i -th computing unit; is the number of computing units; max is the maximum value function; i is the serial number of the block.
[0115] Then, calculate the summation value of the total bubble time of all computing units in each candidate scheme, which is the bubble time of each candidate scheme.
[0116] Example: Assume that the total number of layers of the large model is 20, the number of computing units is 4, and the number of layers responsible for calculation by each computing unit is 20 / 4 = 5 layers. Assume that the number of blocks is 2, called Block 1 and Block 2. Determine multiple candidate schemes for allocating the number of layers (5) responsible for calculation by each computing unit to Block 1 and Block 2.
[0117] For example, Candidate Scheme 1: Block 1 contains 1 layer and Block 2 contains 4 layers; Candidate Scheme 2: Block 1 contains 2 layers and Block 2 contains 3 layers; Candidate Scheme 3: Block 1 contains 3 layers and Block 2 contains 2 layers; Candidate Scheme 4: Block 1 contains 4 layers and Block 2 contains 1 layer, and so on.
[0118] Then, calculate the respective bubble times of the above 5 candidate schemes based on Formula (1), and select the candidate scheme with the least bubble time from the above 5 candidate schemes. Assume that the bubble time of Candidate Scheme 2 is the least, then it can be determined that Block 1 contains 2 layers and Block 2 contains 3 layers.
[0119] Method (2): Based on the number of computing units, the number of blocks, and the total number of layers, the number of layers in each block is determined using the memory consumption as an indicator.
[0120] In method (2), the number of layers in each block is determined based on the memory consumption. Generally, after determining the number of layers in the block, the memory consumption is minimized.
[0121] In one embodiment, method (2) specifically includes: determining multiple candidate schemes for allocating the total number of layers to each block based on the number of computing units and the number of blocks, each candidate scheme including the number of layers included in each block; determining the video memory consumption of the computing unit that first performs calculations in each candidate scheme; selecting the candidate scheme with the lowest video memory consumption from the multiple candidate schemes; and determining the number of layers included in each block based on the candidate scheme with the lowest video memory consumption.
[0122] In one embodiment, determining the video memory consumption of the computational unit that first performs calculations in each candidate solution includes: determining a first activation value for all blocks of the computational unit that first performs calculations in each candidate solution for a unit micro-batch; determining a second activation value for the first block of the computational unit that first performs calculations in each candidate solution for a unit micro-batch; and determining the video memory consumption of the computational unit that first performs calculations in each candidate solution based on the first activation value, the second activation value, and the number of computational units. Therefore, a process of quantifying video memory consumption can be implemented based on the first activation value, the second activation value, and the number of computational units.
[0123] The following exemplary description discusses how to quantitatively evaluate the memory consumption of candidate solutions.
[0124] Large model memory is usually divided into static memory and dynamic memory. Usually, static memory is the memory of the resident computing unit during the training process, including the parameters that need to be trained, the optimizer state required by the optimizer, etc. The consumption of dynamic memory is usually characterized by activation values, which are dynamically requested and released as the training progresses. Generally speaking, the forward activation value needs to be held until the reverse is consumed. For pipeline parallelism, in order to minimize bubbles as much as possible, the reverse calculation of the first micro-batch and the forward calculation of the first micro-batch usually include multiple micro-batch forward calculations. Therefore, in the absence of recalculation, starting from the warm-up phase, a large number of activation values will need to be held until the start of the first reverse calculation.
[0125] Moreover, the memory usage of activation values is usually different for different computational units, with the first computational unit usually occupying more.
[0126] For interleaved pipeline parallel training, due to finer-grained partitioning and the pursuit of fewer bubbles, the video memory occupancy of each computing unit is more than that of ordinary 1F1B.
[0127] Based on formula (2), calculate the video memory consumption of the computing unit that first performs calculations (such as Figure 1 and Figure 2 device A in) in each candidate solution.
[0128] ; formula (2)
[0129] Where: is the video memory consumption of the computing unit that first performs calculations (that is, the first computing unit); is the first activation value of all blocks of the computing unit that first performs calculations for a unit micro-batch; The first block that first performs calculations of the computing unit that first performs calculations (that is, the first computing unit) (i.e., chunk 1 ) the second activation value of the unit micro-batch; is the number of computing units.
[0130] Example: Assume that the total number of layers of the large model is 20, the number of computing units is 4, and the number of layers responsible for calculation by each computing unit is 20 / 4 = 5 layers. Assume that the number of blocks is 2, called block 1 and block 2. Determine multiple candidate solutions for allocating the number of layers (5) responsible for calculation by each computing unit to block 1 and block 2.
[0131] For example, candidate solution 1: block 1 contains 1 layer, block 2 contains 4 layers; candidate solution 2: block 1 contains 2 layers, block 2 contains 3 layers; candidate solution 3: block 1 contains 3 layers, block 2 contains 2 layers; candidate solution 4: block 1 contains 4 layers, block 2 contains 1 layer, and so on.
[0132] Then, based on formula (2), calculate the video memory consumption of the computing unit that first performs calculations for the above 5 candidate solutions, and select the candidate solution with the least video memory consumption from the above 5 candidate solutions. Assume that the video memory consumption of candidate solution 1 is the least, then determine that block 1 contains 1 layer and block 2 contains 4 layers.
[0133] Method (3): Based on the number of computing units, the number of blocks, and the total number of layers, determine the number of layers contained in each block with bubble time and video memory consumption as indicators.
[0134] In one embodiment, method (3) specifically includes: based on the number of computing units and the number of blocks, determining multiple candidate solutions for distributing the total number of layers to each block, where each candidate solution includes the respective number of layers included in each block; determining the video memory consumption of the computing unit that first performs the calculation in each candidate solution; selecting at least one candidate solution from the multiple candidate solutions whose video memory consumption is lower than the first predetermined threshold; determining the bubble time of at least one candidate solution; selecting the candidate solution with the least bubble time from at least one candidate solution; and determining the number of layers included in each block based on the candidate solution with the least bubble time.
[0135] It can be seen that, based on the video memory consumption as the first filtering criterion, at least one candidate solution whose video memory consumption meets the expectation is screened out, and then, with the bubble time as the selection criterion, the candidate solution with the least bubble time is selected from at least one candidate solution.
[0136] Example: Assume that the total number of layers of the large model is 20, the number of computing units is 4, and the number of layers responsible for calculation by each computing unit is 20 / 4 = 5 layers. Assume that the number of blocks is 2, called block 1 and block 2. Determine multiple candidate solutions for distributing the number of layers (5) responsible for calculation by each computing unit to block 1 and block 2. For example, candidate solution 1: block 1 includes 1 layer and block 2 includes 4 layers; candidate solution 2: block 1 includes 2 layers and block 2 includes 3 layers; candidate solution 3: block 1 includes 3 layers and block 2 includes 2 layers; candidate solution 4: block 1 includes 4 layers and block 2 includes 1 layer, and so on. Calculate the video memory consumption of the computing unit that first performs the calculation for the above 5 candidate solutions based on formula (2), and select at least one candidate solution whose video memory consumption meets the expectation from the above 5 candidate solutions. Assume that candidate solution 1, candidate solution 2, and candidate solution 3 are selected. Calculate the bubble time of candidate solution 1, candidate solution 2, and candidate solution 3 based on formula (1), and then select the candidate solution with the least bubble time from candidate solution 1, candidate solution 2, and candidate solution 3. Assume it is candidate solution 2, then determine that block 1 includes 2 layers and block 2 includes 3 layers.
[0137] In one embodiment, method (3) specifically includes: based on the number of computing units and the number of blocks, determining multiple candidate solutions for distributing the total number of layers to each block, where each candidate solution includes the respective number of layers included in each block; determining the bubble time of each candidate solution; selecting at least one candidate solution from the multiple candidate solutions whose bubble time is lower than the second predetermined threshold; determining the video memory consumption of the computing unit that first performs the calculation in at least one candidate solution; selecting the candidate solution with the minimum video memory consumption from at least one candidate solution; and determining the number of layers included in each block based on the candidate solution with the minimum video memory consumption.
[0138] Example: Assume that the total number of layers of the large model is 20, the number of computing units is 4, and the number of layers responsible for calculation by each computing unit is 20 / 4 = 5 layers. Assume that the number of blocks is 2, called block 1 and block 2. Determine multiple candidate solutions for allocating the number of layers (5) responsible for calculation by each computing unit to block 1 and block 2. For example, candidate solution 1: block 1 contains 1 layer and block 2 contains 4 layers; candidate solution 2: block 1 contains 2 layers and block 2 contains 3 layers; candidate solution 3: block 1 contains 3 layers and block 2 contains 2 layers; candidate solution 4: block 1 contains 4 layers and block 2 contains 1 layer, and so on. Calculate the bubble time of the above 5 candidate solutions based on formula (1), and select at least one candidate solution whose bubble time meets the expectation from the above 5 candidate solutions. Assume that candidate solution 2, candidate solution 3, and candidate solution 4 are selected. Then, calculate the video memory consumption of the computing unit that performs the calculation first in candidate solution 2, candidate solution 3, and candidate solution 4 based on formula (2), and select the candidate solution with the least video memory consumption of the computing unit that performs the calculation first from candidate solution 2, candidate solution 3, and candidate solution 4. Assume it is candidate solution 3, then determine that block 1 contains 3 layers and block 2 contains 2 layers.
[0139] It can be seen that based on the bubble time as the first filtering index, at least one candidate solution whose bubble time meets the expectation is selected, and then based on the video memory consumption as the selection index, the candidate solution with the least video memory consumption is selected from at least one candidate solution.
[0140] In one implementation, method (3) specifically includes: based on the number of computing units and the number of blocks, determine multiple candidate solutions for allocating the total number of layers to each block, and each candidate solution includes the respective number of layers contained in each block; determine the bubble time of each candidate solution and the video memory consumption of the computing unit that performs the calculation first in each candidate solution; select at least one candidate solution whose video memory consumption is lower than the first predetermined threshold and whose bubble time is lower than the second predetermined threshold from multiple candidate solutions; select the candidate solution with the least video memory consumption or the least bubble time from at least one candidate solution; based on the selected candidate solution, determine the number of layers contained in each block.
[0141] It can be seen that based on the bubble time and the video memory consumption as the selection indexes, the candidate solutions whose bubble time and video memory consumption meet the expectation are selected.
[0142] Example: Assume that the total number of layers of the large model is 20, the number of computing units is 4, and the number of layers each computing unit is responsible for calculating is 20 / 4 = 5 layers. Assume that the number of chunks is 2, called chunk 1 and chunk 2. Determine multiple candidate schemes for distributing the number of layers (5) that each computing unit is responsible for calculating to chunk 1 and chunk 2. For example, candidate scheme 1: chunk 1 contains 1 layer and chunk 2 contains 4 layers; candidate scheme 2: chunk 1 contains 2 layers and chunk 2 contains 3 layers; candidate scheme 3: chunk 1 contains 3 layers and chunk 2 contains 2 layers; candidate scheme 4: chunk 1 contains 4 layers and chunk 2 contains 1 layer, and so on. Calculate the bubble time of the above 5 candidate schemes based on formula (1), calculate the video memory consumption of the computing unit that first performs the calculation for the above 5 candidate schemes based on formula (2), and select a candidate scheme whose bubble time and video memory consumption both meet the expectations from the above 5 candidate schemes. Assume that candidate scheme 3 and candidate scheme 4 are selected. Then, based on the minimum bubble time, select candidate scheme 3 from candidate scheme 3 and candidate scheme 4, and it is determined that chunk 1 contains 3 layers and chunk 2 contains 2 layers.
[0143] The above demonstration describes a typical process for determining the number of layers contained in each chunk based on the number of computing units, the number of chunks, and the total number of layers of the large model. Those skilled in the art can realize that this description is only exemplary and is not used to limit the protection scope of the embodiments of the present invention.
[0144] Step 102: Perform interleaved pipelined parallel training on the large model based on the number of layers contained in each chunk.
[0145] Here, based on the number of layers contained in each chunk determined in step 101, perform interleaved pipelined parallel training on the large model. Each computing unit in the pipeline is assigned multiple pipeline stages. Compared with non-interleaved, the amount of calculation in each pipeline stage is less.
[0146] Figure 4 It is a schematic diagram showing an example of the interleaved pipelined parallel training process according to an embodiment of the present invention.
[0147] In Figure 4 it is described by taking 4 computing units A to D as an example. Assume that the number of micro-batches that can be processed in one interleaved scheduling method is 8 (i.e., the number of accumulation times is 8), and the total number of layers of the large model is 40, then the number of layers each computing unit is responsible for calculating is 40 / 4 = 10 layers. Assume that the number of chunks is 2, called chunk 1 and chunk 2. Assume that based on Figure 3The method shown determines that block 1 contains 3 layers and block 2 contains 7 layers, with the bubble time or video memory consumption as the metric. It can be seen that the number of layers in block 1 and block 2 is not the same. Determine the layers each of computing devices A - D is responsible for. Among them: Based on the order of the computing devices and in ascending order of layer numbers, in the earlier - executed block (block 1) of each computing device, allocate the respective first 3 layers in sequence (since block 1 contains 3 layers), and in the later - executed block (block 2), allocate the respective subsequent 7 layers in sequence (since block 2 contains 7 layers). For example, for computing device A, block 1 is responsible for layers 1 - 3, and block 2 is responsible for layers 13 - 19; for computing device B, block 1 is responsible for layers 4 - 6, and block 2 is responsible for layers 20 - 26; for computing device C, block 1 is responsible for layers 7 - 9, and block 2 is responsible for layers 27 - 33; for computing device D, block 1 is responsible for layers 10 - 12, and block 2 is responsible for layers 34 - 40.
[0148] Taking computing unit A as an example, computing unit A should be responsible for the calculations of layers 1 - 3 and layers 13 - 19. Execute in sequence: (1) The forward operations of block 1 (containing layers 1 - 3) for micro - batch sequence numbers 1 - 4 ( Figure 4 identified by the dark - blue boxes with numbers 1 - 4); (2) The forward operations of block 2 (containing layers 13 - 19) for micro - batch sequence numbers 1 - 4 ( Figure 4 identified by the light - blue boxes with numbers 1 - 4); (3) The forward operations of block 1 (containing layers 1 - 3) for micro - batch sequence numbers 5 - 6 ( Figure 4 identified by the dark - blue boxes with numbers 5 - 6); (4) Wait (i.e., bubble time, Figure 4 identified by the white boxes); (5) The forward operations of block 1 (containing layers 1 - 3) for micro - batch sequence number 7 ( Figure 4 identified by the dark - blue box with number 7); (6) The backward operations of block 2 (containing layers 13 - 19) for micro - batch sequence number 1 ( Figure 4 identified by the light - green box with number 1); (7) The forward operations of block 1 (containing layers 1 - 3) for micro - batch sequence number 8 ( Figure 4 identified by the dark - blue box with number 8); (8): Wait (i.e., bubble time, Figure 4 identified by the white boxes); (9) The backward operations of block 2 (containing layers 13 - 19) for micro - batch sequence number 2 ( Figure 4 identified by the light - green box with number 2); (10) The forward operations of block 2 (containing layers 13 - 19) for micro - batch sequence number 5 ( Figure 4 identified by the light - blue box with number 5); (11) The backward operations of block 2 (containing layers 13 - 19) for micro - batch sequence number 3 ( Figure 4identified by a light green square with the number 3; (12) forward operation of block 2 of micro-batch sequence number 6 (including layers 13 to 19) ( Figure 4 identified by a light blue square with the number 6; (13) backward operation of block 2 of micro-batch sequence number 4 (including layers 13 to 19) ( Figure 4 identified by a light green square with the number 4; (14) forward operation of block 2 of micro-batch sequence number 7 (including layers 13 to 19) ( Figure 4 identified by a light blue square with the number 7; (15) backward operation of block 1 of micro-batch sequence number 1 (including layers 1 to 3) ( Figure 4 identified by a dark green square with the number 1);... and so on.
[0149] Taking computing unit B as an example, computing unit B should be responsible for the calculations of layers 4 to 6 and layers 20 to 26. Execute in sequence: (1) Wait (i.e., bubble time, Figure 4 identified by a white square); (2) Forward operation of block 1 of micro-batch sequence numbers 1 to 4 (including layers 4 to 6) ( Figure 4 identified by dark blue squares with the numbers 1 to 4); (3) Wait (i.e., bubble time, Figure 4 identified by a white square); (4) Forward operation of block 2 of micro-batch sequence numbers 1 to 4 (including layers 20 to 26) ( Figure 4 identified by light blue squares with the numbers 1 to 4); (5) Wait (i.e., bubble time, Figure 4 identified by a white square); (6) Forward operation of block 1 of micro-batch sequence number 5 (including layers 4 to 6) ( Figure 4 identified by a light blue square with the number 5);... and so on.
[0150] Similarly, computing units C to D are assigned 2 non-uniformly divided blocks and are responsible for the calculations of their respective layers, thereby allowing for interleaved execution of forward and backward between different stages.
[0151] It can be seen that interleaved pipelined parallel training is implemented based on non-uniformly divided blocks without having to fall back to non-interleaved pipelined parallel training, thus improving the training efficiency.
[0152] Figure 5 It is a schematic diagram of an interleaved pipelined parallel training device according to an embodiment of the present invention. As Figure 5As shown, the interleaved pipeline parallel training device 200 includes: a determination module 201, configured to determine the number of layers included in each block based on the number of computing units, the number of blocks, and the total number of layers of the large model, where the number of layers included in each block is configurable; a training module 202, configured to perform interleaved pipeline parallel training on the large model based on the number of layers included in each block; the determination module 201, configured to perform at least one of the following: determining the number of layers included in each block with the bubble time as an indicator based on the number of computing units, the number of blocks, and the total number of layers; determining the number of layers included in each block with the video memory consumption as an indicator based on the number of computing units, the number of blocks, and the total number of layers; determining the number of layers included in each block with the bubble time and the video memory consumption as indicators based on the number of computing units, the number of blocks, and the total number of layers.
[0153] In one embodiment, the determination module 201 is configured to determine multiple candidate solutions for allocating the total number of layers to each block based on the number of computing units and the number of blocks, where each candidate solution includes the respective number of layers included in each block; determine the bubble time of each candidate solution; select the candidate solution with the least bubble time from the multiple candidate solutions; and determine the number of layers included in each block based on the candidate solution with the least bubble time.
[0154] In one embodiment, the determination module 201 is configured to determine multiple candidate solutions for allocating the total number of layers to each block based on the number of computing units and the number of blocks, where each candidate solution includes the respective number of layers included in each block; determine the video memory consumption of the computing unit that first performs the calculation in each candidate solution; select the candidate solution with the least video memory consumption from the multiple candidate solutions; and determine the number of layers included in each block based on the candidate solution with the least video memory consumption.
[0155] In one embodiment, the determination module 201 is configured to perform one of the following:
[0156] (1) Determine multiple candidate solutions for allocating the total number of layers to each block based on the number of computing units and the number of blocks, where each candidate solution includes the respective number of layers included in each block; determine the video memory consumption of the computing unit that first performs the calculation in each candidate solution; select at least one candidate solution with a video memory consumption lower than a first predetermined threshold from the multiple candidate solutions; determine the bubble time of the at least one candidate solution; select the candidate solution with the least bubble time from the at least one candidate solution; and determine the number of layers included in each block based on the candidate solution with the least bubble time.
[0157] (2) Based on the number of computing units and the number of chunks, determine multiple candidate solutions for allocating the total number of layers to each chunk, where each candidate solution includes the respective number of layers included in each chunk; determine the bubble time of each candidate solution; select at least one candidate solution from the multiple candidate solutions whose bubble time is lower than the second predetermined threshold; determine the video memory consumption of the computing unit that first performs calculations in at least one candidate solution; select the candidate solution with the minimum video memory consumption from at least one candidate solution; based on the candidate solution with the minimum video memory consumption, determine the number of layers included in each chunk.
[0158] (3) Based on the number of computing units and the number of chunks, determine multiple candidate solutions for allocating the total number of layers to each chunk, where each candidate solution includes the respective number of layers included in each chunk; determine the bubble time of each candidate solution and the video memory consumption of the computing unit that first performs calculations in each candidate solution; select at least one candidate solution from the multiple candidate solutions whose video memory consumption is lower than the first predetermined threshold and whose bubble time is lower than the second predetermined threshold; select the candidate solution with the minimum video memory consumption or the candidate solution with the least bubble time from at least one candidate solution; based on the selected candidate solution, determine the number of layers included in each chunk.
[0159] In summary, in the embodiment of the present invention, based on the number of computing units, the number of chunks, and the total number of layers of the large model, determine the number of layers included in each chunk, where the number of layers included in each chunk is configurable; based on the number of layers included in each chunk, perform interleaved pipeline parallel training on the large model. It can be seen that the constraint that the number of layers responsible for each computing unit must be divisible by the number of chunks is broken, realizing a non-interleaved pipeline parallel training process that supports both equal division and non-equal division, improving the training efficiency, and also expanding the applicable scenarios of interleaved pipeline parallel training. In addition, a good compromise can also be achieved between bubble time and video memory consumption. The embodiment of the present invention also proposes an electronic device with a processor-memory architecture. Figure 6 is a structural diagram of an electronic device according to an embodiment of the present invention. As Figure 6 shown, the electronic device includes a processor 301, a memory 302, and a computer program stored on the memory 302 and executable on the processor 301. When the computer program is executed by the processor 301, it implements any one of the above interleaved pipeline parallel training methods. Among them, the memory 302 can be specifically implemented as various storage media such as electrically erasable programmable read-only memory (EEPROM), flash memory, programmable read-only memory (PROM), etc. The processor 301 can be implemented as including one or more central processing units or one or more field programmable gate arrays, where the field programmable gate array integrates one or more central processing unit cores. Specifically, the central processing unit or the central processing unit core can be implemented as a CPU, GPU, GPGPU, MCU, or DSP, etc.
[0160] It should be noted that not all steps and modules in the above processes and structure diagrams are necessary, and some steps or modules can be ignored according to actual needs. The execution order of each step is not fixed and can be adjusted as required. The division of each module is only for the convenience of description in terms of function. In actual implementation, a module can be implemented by multiple modules, and the functions of multiple modules can also be implemented by the same module. These modules can be located in the same device or in different devices.
[0161] The hardware modules in each embodiment can be implemented mechanically or electronically. For example, a hardware module can include specially designed permanent circuits or logic devices (such as dedicated processors, such as FPGA or ASIC) for performing specific operations. For instance, specific operations can be completed in various types of chips (e.g., artificial intelligence chips). A hardware module can also include programmable logic devices or circuits temporarily configured by software (such as including general-purpose processors or other programmable processors) for executing specific operations. As for whether to specifically adopt a mechanical method, or use dedicated permanent circuits, or use temporarily configured circuits (such as configured by software) to implement the hardware module, it can be determined based on cost and time considerations.
[0162] The present invention also provides a machine-readable storage medium storing instructions for causing a machine to execute the method as described in this application. Specifically, a system or device equipped with a storage medium can be provided, on which software program codes for implementing the functions of any one of the above embodiments are stored, and the computer (or CPU or MPU) of the system or device is caused to read and execute the program codes stored in the storage medium. In addition, based on the instructions of the program codes, the operating system operating on the computer, etc. can be caused to complete part or all of the actual operations. The program codes read from the storage medium can also be written into the memory provided in the expansion board inserted into the computer or into the memory provided in the expansion unit connected to the computer, and then based on the instructions of the program codes, the CPU, etc. installed on the expansion board or expansion unit are caused to execute part and all of the actual operations, thereby implementing the functions of any one of the above embodiments. The storage medium embodiments for providing program codes include floppy disks, hard disks, magneto-optical disks, optical disks (such as CD-ROM, CD-R, CD-RW, DVD-ROM, DVD-RAM, DVD-RW, DVD+RW), magnetic tapes, non-volatile memory cards, and ROM. Optionally, the program codes can be downloaded from a server computer or the cloud via a communication network.
[0163] In this text, "schematic" means "serving as an example, instance or illustration", and any illustration or embodiment described as "schematic" in this text should not be construed as a more preferred or more advantageous technical solution. To make the drawings concise, only the parts related to the present invention are schematically shown in each drawing, and do not represent the actual structure of the product as a whole. In addition, to make the drawings concise and easy to understand, in some drawings, for components with the same structure or function, only one of them is schematically shown, or only one of them is labeled. In this text, "a" does not mean that the quantity of the parts related to the present invention is limited to "only one", and "a" does not exclude the case where the quantity of the parts related to the present invention is "more than one". In this text, "upper", "lower", "front", "rear", "left", "right", "inner", "outer", etc. are only used to represent the relative positional relationship between the relevant parts, rather than limiting the absolute positions of these relevant parts.
[0164] The above is only a preferred embodiment of the present invention and is not intended to limit the protection scope of the present invention. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present invention shall be included within the protection scope of the present invention.
Claims
1. An interleaved pipeline parallel training method, characterized in that, Including: Based on the number of computing units, the number of chunks, and the total number of layers of the large model, determine the number of layers included in each chunk, where the number of layers included in each chunk is configurable. Based on the number of computing units and the number of chunks, determine multiple candidate schemes for allocating the total number of layers to each chunk, and each candidate scheme includes the respective number of layers included in each chunk. The determining the number of layers included in each chunk based on the number of computing units, the number of chunks, and the total number of layers of the large model includes one of the following: determining the number of layers included in each chunk with bubble time as an indicator based on the number of computing units, the number of chunks, and the total number of layers; determining the number of layers included in each chunk with memory consumption as an indicator based on the number of computing units, the number of chunks, and the total number of layers; Determining the number of layers included in each chunk with both bubble time and memory consumption as indicators based on the number of computing units, the number of chunks, and the total number of layers; Based on the number of layers included in each chunk, perform interleaved pipelined parallel training on the large model; Before determining the number of layers included in each chunk based on the number of computing units, the number of chunks, and the total number of layers of the large model, the method further includes: based on the total number of layers and the number of computing units, determine the number of layers of each computing unit; The determining the number of layers included in each chunk based on the number of computing units, the number of chunks, and the total number of layers of the large model includes: when the number of layers of each computing unit is not divisible by the number of chunks, the number of layers included in at least two chunks is different; when the number of layers of each computing unit is divisible by the number of chunks, the number of layers included in each chunk is the same.
2. The method according to claim 1, wherein The determining the number of layers included in each chunk with bubble time as an indicator based on the number of computing units, the number of chunks, and the total number of layers of the large model includes: Determine the bubble time of each candidate scheme; Select the candidate scheme with the least bubble time from the multiple candidate schemes; Based on the candidate scheme with the least bubble time, determine the number of layers included in each chunk.
3. The method according to claim 2, wherein The determining the bubble time of each candidate scheme includes: Determine the forward time for each chunk of each candidate scheme to perform forward calculation on a unit micro-batch; Determine the maximum forward time from the multiple forward times of the multiple chunks of each candidate scheme; Determine the backward time for each chunk of each candidate scheme to perform backward calculation on a unit micro-batch; Determine the maximum backward time from the multiple backward times of the multiple chunks of each candidate scheme; Based on the maximum forward time, the maximum backward time, and the number of computing units, determine the bubble time of each candidate scheme.
4. The method according to claim 1, wherein The determining the number of layers included in each chunk with memory consumption as an indicator based on the number of computing units, the number of chunks, and the total number of layers of the large model includes: Determine the memory consumption of the computing unit that first performs calculation in each candidate scheme; Select the candidate scheme with the least memory consumption from the multiple candidate schemes; Based on the candidate scheme with the least memory consumption, determine the number of layers included in each chunk.
5. The method according to claim 4, characterized in that, The determining the memory consumption of the computing unit that first performs calculation in each candidate scheme includes: Determine the first activation value for performing calculations on all blocks of the computing unit that first performs calculations in each candidate solution for a unit micro-batch; Determine the second activation value for performing calculations on the first block that first performs calculations of the computing unit that first performs calculations in each candidate solution for a unit micro-batch; Based on the first activation value, the second activation value, and the number of computing units, determine the video memory consumption of the computing unit that first performs calculations in each candidate solution.
6. The method according to claim 1, characterized in that, The determining the number of layers included in each block with bubble time and video memory consumption as metrics based on the number of computing units, the number of blocks, and the total number of layers includes: Determine the video memory consumption of the computing unit that first performs calculations in each candidate solution; Select at least one candidate solution from the multiple candidate solutions whose video memory consumption is lower than a first predetermined threshold; Determine the bubble time of the at least one candidate solution; Select the candidate solution with the least bubble time from the at least one candidate solution; Based on the candidate solution with the least bubble time, determine the number of layers included in each block.
7. The method according to claim 1, characterized in that The determining the number of layers included in each block with bubble time and video memory consumption as metrics based on the number of computing units, the number of blocks, and the total number of layers includes: Determine the bubble time of each candidate solution; Select at least one candidate solution from the multiple candidate solutions whose bubble time is lower than a second predetermined threshold; Determine the video memory consumption of the computing unit that first performs calculations in the at least one candidate solution; Select the candidate solution with the minimum video memory consumption from the at least one candidate solution; Based on the candidate solution with the minimum video memory consumption, determine the number of layers included in each block.
8. The method according to claim 1, wherein The determining the number of layers included in each block with bubble time and video memory consumption as metrics based on the number of computing units, the number of blocks, and the total number of layers includes: Determine the bubble time of each candidate solution and the video memory consumption of the computing unit that first performs calculations in each candidate solution; Select at least one candidate solution from the multiple candidate solutions whose video memory consumption is lower than a first predetermined threshold and whose bubble time is lower than a second predetermined threshold; Select the candidate solution with the minimum video memory consumption or the candidate solution with the least bubble time from the at least one candidate solution; Based on the selected candidate solution, determine the number of layers included in each block.
9. An interleaved pipeline parallel training device, applied to the interleaved pipeline parallel training method according to any one of claims 1-8, characterized in that, Includes: A determination module for determining the number of layers included in each block based on the number of computing units, the number of blocks, and the total number of layers of the large model, where the number of layers included in each block is configurable, and based on the number of computing units and the number of blocks, determine multiple candidate solutions for distributing the total number of layers to each block, and each candidate solution includes the respective number of layers included in each block; where before determining the number of layers included in each block based on the number of computing units, the number of blocks, and the total number of layers of the large model, determine the number of layers of each computing unit based on the total number of layers and the number of computing units; Determining the number of layers included in each block based on the number of computing units, the number of blocks, and the total number of layers of the large model includes: when the number of layers of each computing unit is not divisible by the number of blocks, the number of layers included in at least two blocks is different; when the number of layers of each computing unit is divisible by the number of blocks, the number of layers included in each block is the same; A training module, configured to perform interleaved pipelining parallel training on the large model based on the number of layers included in each block; The determining module is configured to perform one of the following: Determine the number of layers included in each block with bubble time as an indicator based on the number of computing units, the number of blocks, and the total number of layers; Determine the number of layers included in each block with video memory consumption as an indicator based on the number of computing units, the number of blocks, and the total number of layers; Determine the number of layers included in each block with bubble time and video memory consumption as indicators based on the number of computing units, the number of blocks, and the total number of layers.
10. The apparatus according to claim 9, wherein The determining module is configured to determine the bubble time of each candidate solution; select the candidate solution with the least bubble time from the multiple candidate solutions; and determine the number of layers included in each block based on the candidate solution with the least bubble time.
11. The apparatus according to claim 9, wherein The determining module is configured to determine the video memory consumption of the computing unit that first performs calculations in each candidate solution; select the candidate solution with the minimum video memory consumption from the multiple candidate solutions; Determine the number of layers included in each block based on the candidate solution with the minimum video memory consumption.
12. The apparatus according to claim 9, wherein The determining module is configured to perform one of the following: Determine the video memory consumption of the computing unit that first performs calculations in each candidate solution; select at least one candidate solution with a video memory consumption lower than a first predetermined threshold from the multiple candidate solutions; determine the bubble time of the at least one candidate solution; Select the candidate solution with the least bubble time from the at least one candidate solution; Determine the number of layers included in each block based on the candidate solution with the least bubble time; Determine the bubble time of each candidate solution; Select at least one candidate solution with a bubble time lower than a second predetermined threshold from the multiple candidate solutions; determine the video memory consumption of the computing unit that first performs calculations in the at least one candidate solution; Select the candidate solution with the minimum video memory consumption from the at least one candidate solution; Determine the number of layers included in each block based on the candidate solution with the minimum video memory consumption; Determine the bubble time of each candidate solution and the video memory consumption of the computing unit that first performs calculations in each candidate solution; select at least one candidate solution with a video memory consumption lower than a first predetermined threshold and a bubble time lower than a second predetermined threshold from the multiple candidate solutions; select the candidate solution with the minimum video memory consumption or the least bubble time from the at least one candidate solution; and determine the number of layers included in each block based on the selected candidate solution.
13. An electronic device, characterized in that, including: A memory; A processor; Among them, an application program executable by the processor is stored in the memory, and is used to cause the processor to execute the interleaved pipeline parallel training method according to any one of claims 1-8.
14. A computer-readable storage medium, characterized in that, Computer-readable instructions are stored on the computer-readable storage medium, and when the computer-readable instructions are executed by a processor, the processor is caused to execute the interleaved pipeline parallel training method according to any one of claims 1-8.
15. A program product, comprising a computer program, characterized in that, When the computer program is executed by a processor, the interleaved pipeline parallel training method according to any one of claims 1-8 is implemented.