A Two-Dimensional Sequence Splitting Method and System for Large Model Pipeline Parallel Training

By adopting a two-dimensional sequence splitting method in parallel training of large model pipelines and optimizing redundant storage in combination with time and space dimensions, the problems of high memory usage and long communication time caused by long sequence samples are solved, and more efficient training efficiency is achieved.

CN119883383BActive Publication Date: 2025-07-01ZHEJIANG UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510379220.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-03-28
Publication Date
2025-07-01
Estimated Expiration
2045-03-28

AI Technical Summary

Technical Problem

In parallel training of large model pipelines, long sequence samples lead to high memory usage and long communication time, resulting in low training efficiency, and single-dimensional sequence splitting cannot fully perform the performance of cluster equipment.

Method used

The two-dimensional sequence splitting method is used to split the long sequence in the time and space dimensions, and the redundant sequence length and proportion are optimized by combining the linear integer programming algorithm, the redundant storage space of the GPU and CPU is used to reduce communication overhead, and the training process is accelerated through the two-dimensional sequence parallel operator.

Benefits of technology

The overall training efficiency of parallel training of large model pipelines is improved, and the throughput is increased by about 10% to 15%, which is better than single-dimensional splitting or simple combination solutions.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119883383B_ABST
    Figure CN119883383B_ABST
Patent Text Reader

Abstract

The present invention discloses a two-dimensional sequence splitting method and system under large model pipeline parallel training, belonging to the field of artificial intelligence in computer science. The present invention includes: a data collection module obtains device basic information and model configuration information, including the bandwidth between GPUs, the device video memory size, the device CPU memory size, the bandwidth between GPU and CPU, the model dimension, the number of model layers, and the length of the input data sequence; a decision maker generates an optimal decision according to the obtained data; the decision content includes the redundant sequence length, the proportion of the redundant sequence stored in the GPU, the proportion of the redundant sequence stored in the CPU, and the number of splits in the time dimension; a deep learning training module integrates the optimal decision into the model training process to improve the overall training performance of the system. The present invention combines the idle video memory space and the bandwidth between GPU and CPU to achieve sequence splitting and efficient training in both the time and space dimensions, while maximizing the training efficiency of pipeline parallel training.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of computer science and artificial intelligence, and particularly to a two-dimensional sequence splitting method and system under large model pipeline parallel training. Background Art

[0002] With the development of artificial intelligence technology, the scale and complexity of models have been continuously increasing, and more powerful performance and more accurate prediction results have been obtained. These large models, such as the Transformer model, usually contain hundreds of millions or even billions of parameters. For example, one of the latest open-source Transformer large models, Llama 3.1 - 405B, contains more than 400 billion parameters. At the same time, the proposal of the latest technologies in the field of large models (such as RAG) has also led to a significant increase in the length of input samples included in the large model training process. For the current most advanced open-source large model, the Deepseek-V3 model, its context window length reaches 128,000 tokens. The increase in the parameters of the model itself brings a huge demand for video memory capacity, making it very difficult to train these models on a single GPU or CPU. To solve this problem, pipeline parallel technology is often introduced in the large model training process. The core idea of pipeline parallel technology is to split a huge model into multiple smaller sub-models, and each sub-model can be independently trained and calculated on different hardware devices, forming a pipeline with each other to efficiently utilize the potential of the devices.

[0003] The addition of long sequence samples makes the video memory capacity once again become a severe problem restricting the training of Transformer large models, giving rise to two problems: 1. The video memory occupancy of the activation tensors generated by long sequence samples is relatively high, and 2. The introduction of more device idle time (bubbles) also leads to a decline in the above-mentioned pipeline parallel performance. To solve these problems, sequence chunking (parallelism) is introduced into the training of large models. The core idea of sequence chunking is to split a relatively long sequence into several shorter sequences and redistribute these shorter sequences to multiple devices or different time periods, and obtain the attention outputs of each subsequence through chunked calculation during the attention calculation process. Sequence splitting techniques include two categories: one is to split and parallelize in the spatial dimension (such as the RingAttention Attention parallel computing algorithm), where a long sequence is split into equal-length subsequences distributed to multiple devices for simultaneous calculation. The disadvantage of this type of sequence parallelism is that the parallelism is affected by the number of devices, and it is necessary to ensure that the calculation time can cover the communication time; the other is to split in the time dimension (such as TeraPipe, Seq1F1B technology), where a long sequence is split into subsequences with gradually decreasing lengths according to a uniform computational amount and calculated block by block in sequence on one device. The disadvantage of this type of sequence splitting is that it has a relatively high video memory occupancy for activation tensors, and the effect is not as good as sequence splitting in the spatial dimension.

[0004] The simple idea is to directly combine the two. First, use temporal splitting to split the long sequence into different time periods, and then use spatial splitting to split the subsequences within the corresponding time periods to multiple devices for parallel calculation. However, simply combining the two will create new problems. Since the Attention mechanism needs to calculate pairwise between the tokens in the sequence and all the previous tokens, the communication volume of each chunk will increase linearly with the position of the chunk in the sequence. However, the increase rate of the computational volume in the attention mechanism part is less than the increase rate of the communication volume, and there will be a situation where the communication time is greater than the calculation time during training. At this time, there will be additional communication overhead during the training process, reducing the overall training efficiency. Summary of the Invention

[0005] Aiming at the problem that the performance of long sequence samples is relatively low under the pipeline parallel training of large models in the prior art, and the single-dimensional sequence splitting technology cannot fully utilize the performance of cluster devices, the purpose of the present invention is to provide a two-dimensional sequence splitting parallel acceleration method and system for pipeline parallel training of large models.

[0006] The purpose of the present invention is achieved through the following technical solutions: A two-dimensional sequence splitting method for pipeline parallel training of large models includes the following steps:

[0007] S1. Use the real RingAttention load to test and collect the inter-GPU bandwidth, device video memory size, device CPU memory size, GPU-CPU bandwidth, model dimension, model layer number, and input data sequence length;

[0008] S2. Make an intelligent decision based on the information collected in step S1. The decision content includes the redundant sequence length, the proportion of redundant sequences saved on the GPU, the proportion of redundant sequences saved on the CPU, and the time dimension split number;

[0009] S3. According to the time dimension split number and the length of each subsequence, split the input sequence into the first subsequences with decreasing lengths, further split the subsequences into M equal-length second subsequences, and distribute the M second subsequences to the GPUs in M sequence parallel groups; input the second subsequences into the Transformer large model optimized with a two-dimensional sequence parallel operator for training. During the training process, use the redundant sequence length, the proportion of redundant sequences saved on the GPU, and the proportion of redundant sequences saved on the CPU as parameters and input them into the two-dimensional sequence parallel operator, and during training, execute according to the above parameters: save the redundant sequences to the GPU, transfer the redundant sequences from the GPU to the CPU, and transfer the redundant sequences from the CPU to the GPU.

[0010] Further, step S1 is specifically:

[0011] When the user deploys the training framework to the cluster and inputs the model and training configuration, record the model dimension, model layer number, and input data sequence length, and run the RingAttention operator in the sequence parallel group for several rounds with sequences of arbitrary lengths and dimensions. Calculate the FP16 precision computing power of the GPU device and the communication bandwidth between GPU devices based on the time obtained from the Profile during the operator execution. Then copy the above sequences between the GPU and the CPU for several rounds, and calculate the bandwidth between the GPU and the CPU based on the time obtained from the Profile during the copy.

[0012] Further, when running the RingAttention operator in the sequence parallel group for several rounds with sequences of arbitrary lengths and dimensions, select sequences with a length of 32k and a dimension of 1024.

[0013] Further, step S2 includes:

[0014] S2.1: Initialize the search range [N low , N high ) of the time dimension split number as [1, N max ); where N max is the maximum time dimension split number set by the user, and N lowis the lower bound of the search interval of the current round of binary search, N high The upper bound of the search interval for the current round of binary search;

[0015] S2.2: In the time dimension, the number of splits N is taken as the interval [N low , N high )Intermediate value N mid In the case of , the sequence is split into N segments according to the standard that the amount of calculation in each segment is equal mid subsequences; under the condition that the Attention communication overhead is less than or equal to the Attention computation overhead, linear integer programming is used to solve the shortest redundant sequence length, the corresponding GPU redundant sequence ratio and the CPU redundant sequence ratio;

[0016] S2.3: Calculate the memory overhead of the current decision. If it exceeds the memory budget, update the interval [N low , N high ) is [N low , N mid ), otherwise, update interval [N low , N high ) is [N mid , N high ); If the number of elements in the interval is 1, return the result, otherwise repeat steps S2.2-S2.3.

[0017] Furthermore, the step S2.2 includes: the constraint conditions and optimization objectives for searching the shortest redundant sequence length and the optimal CPU to GPU redundant sequence ratio under the shortest redundant sequence length are shown in the following formula:

[0018] argmin L local

[0019] F a =4dl N L

[0020] p cpu +p gpu =1

[0021]

[0022] 4*L local *d*p cpu ≤M cpu ;

[0023] Among them, L is the total length of the sequence, L local represents the length of redundant sequence, N is the number of time dimension sequence splits, l N is the length of the last subsequence, M is the number of spatial dimension sequence splits, E is the FP16 computing power of a single GPU, and B gpu , Bcpu They are the communication bandwidth between GPUs and the communication bandwidth between GPU and CPU respectively. d is the model dimension, and F a is the computational amount of the attention of the last subsequence, and p cpu and p gpu respectively represent the redundant sequence ratios stored in the CPU memory and the GPU video memory. M cpu is the size of the CPU memory space.

[0024] Furthermore, the step S3 includes: when performing the attention mechanism calculation for each layer of the Transformer model, the following steps are executed:

[0025] S3.1: Calculate the attention of the current Q and the current KV; at the same time, receive the next KV block from the GPU numbered i - 1 and send the current KV to the GPU numbered i + 1. If there is GPU / CPU redundancy, ignore the redundant part of the data during sending / receiving, and only send / receive the non-redundant part; if the CPU redundancy ratio is not 0 at this time, copy the redundant part of the next KV block from the CPU memory to the GPU video memory at the same time;

[0026] S3.2: Update the attention calculation result of the local block;

[0027] S3.3: If according to the decision, the KV in the current time period needs to be redundant, save the KV participating in the calculation in step S3.1 to a tensor;

[0028] S3.4: Update the current KV to the next KV block, and repeat steps S3.1 - S3.4 until the attention mechanism calculation has been performed for the current Q and each KV block;

[0029] S3.5: If according to the decision, the KV in the current time period needs to be redundant, split the tensor storing the KV in step S3.3 into two parts. Among them, one part is stored in the GPU according to the ratio specified by the decision, and the remaining part of the KV is asynchronously copied to the CPU memory and then the tensor in step S3.3 is released;

[0030] S3.6: If the CPU redundancy ratio is not 0, asynchronously prefetch the first CPU redundant KV block of the next layer.

[0031] The present invention also provides a two-dimensional sequence splitting system for large model pipeline parallel training. The system includes:

[0032] A data collection module, which is used to obtain device basic information and model configuration information, including the bandwidth between GPUs, the size of the device video memory, the bandwidth between GPU and CPU, the model dimension, the number of model layers, and the length of the input data sequence;

[0033] A decision maker for generating an optimal decision based on the data obtained by the data collection module; wherein the decision content includes the redundant sequence length, the proportion of the redundant sequence stored in the GPU, the proportion of the redundant sequence stored in the CPU, and the number of splits in the time dimension.

[0034] A deep learning training module for integrating the optimal decision into the model training process to improve the overall training performance of the system. The deep learning training module includes a long sequence sample splitting sub-module and a two-dimensional sequence parallel operator sub-module.

[0035] The present invention also provides an electronic device, including a memory and a processor, the memory is coupled to the processor; wherein, the memory is used to store program data, and the processor is used to execute the program data to implement the two-dimensional sequence splitting method under the large model pipeline parallel training.

[0036] The present invention also provides a computer-readable storage medium, on which a computer program is stored, and when the program is executed by a processor, it implements the two-dimensional sequence splitting method under the large model pipeline parallel training.

[0037] The beneficial effects of the present invention are as follows: The traditional single-dimensional sequence splitting cannot maximize the efficiency of the large model pipeline parallel training under the condition of limited video memory occupancy, but simply combining the two-dimensional sequence splitting will introduce new communication overhead. Different from the traditional single dimension or simple combination, the two-dimensional sequence splitting parallel acceleration scheme proposed by the present invention combines the spatial and temporal dimension splitting and designs corresponding operators, uses the redundant Key / Value Tensor copies and Key / Value Tensor memory-video memory transfer saved on the GPU and CPU, and uses the linear integer programming algorithm to model and realize the dynamic decision of the system, so as to reduce the data transmission volume between GPUs with limited video memory and memory, reduce the overall communication overhead, and thus accelerate the overall training efficiency of long sequence samples under the large model pipeline parallel training. Compared with not using sequence splitting, the two-dimensional sequence splitting can increase the throughput by about 10%, and compared with the single-dimensional sequence splitting or simply combining the two-dimensional splitting, the present invention can increase the throughput by 2-5%. Description of the Drawings

[0038] Figure 1 It is the system architecture diagram of the present invention;

[0039] Figure 2 It is the experimental result diagram of the two-dimensional sequence splitting compared with no splitting, single-dimensional splitting, and simple combination splitting under the pipeline parallel training. Detailed Embodiments

[0040] The terms used in this invention are for the purpose of describing specific embodiments only and are not intended to limit the invention. The singular forms "a", "the", and "said" used in this invention and the appended claims are also intended to include the plural forms unless the context clearly dictates otherwise. It should also be understood that the term "and / or" used herein refers to and encompasses any and all possible combinations of one or more of the associated listed items.

[0041] It should be understood that although the terms first, second, third, etc. may be used in this invention to describe various information, such information should not be limited to these terms. These terms are only used to distinguish the same type of information from each other. For example, without departing from the scope of this invention, the first information may also be referred to as the second information, and similarly, the second information may also be referred to as the first information. Depending on the context, the word "if" as used herein may be interpreted as "when" or "while" or "in response to a determination".

[0042] The present invention will be described in detail below with reference to the accompanying drawings. Without conflict, the features in the following embodiments and implementation manners can be combined with each other.

[0043] As Figure 1 shown, an embodiment of the present invention provides a two-dimensional sequence splitting system under large model pipeline parallel training. The system includes:

[0044] A data collection module for obtaining device basic information and model configuration information, including the bandwidth between GPUs, the device video memory size, the bandwidth between GPU and CPU, the model dimension, the number of model layers, and the length of the input data sequence.

[0045] A decision maker for generating an optimal decision according to the data obtained by the data collection module; wherein the decision content includes the redundant sequence length, the proportion of the redundant sequence saved in the GPU, the proportion of the redundant sequence saved in the CPU, and the number of splits in the time dimension.

[0046] A deep learning training module for integrating the optimal decision into the model training process to improve the overall training performance of the system. The deep learning training module includes a long sequence sample splitting sub-module and a two-dimensional sequence parallel operator sub-module.

[0047] Based on the above system, an embodiment of the present invention also provides a two-dimensional sequence splitting method under large model pipeline parallel training. The method includes:

[0048] Step 1: When the user deploys the training framework to the cluster and inputs the model and training configuration, the data collection module records the model dimension, the number of model layers, and the length of the input data sequence. Then, it runs the RingAttention operator several rounds within the sequence parallel group using a sequence with a length of 32k and a dimension of 1024. The FP16 precision computing power of the GPU device and the communication bandwidth between GPU devices are calculated by estimating the time obtained from the Profile during the operator execution. Next, the above sequence is copied between the GPU and the CPU several rounds, and the bandwidth between the GPU and the CPU is calculated by estimating the time obtained from the Profile during the copy operation.

[0049] Step 2: The decision-making module makes dynamic decisions based on the collected information. The specific approach can be divided into the following sub-steps:

[0050] Step 2.1: Initialize the search interval [N low , N high ) for the time dimension split number as [1, N max ), where N max is the maximum time dimension split number set by the user, N low is the lower bound of the search interval for the current round of binary search, and N high is the upper bound of the search interval for the current round of binary search.

[0051] Step 2.2 (Redundancy ratio optimization): When the time dimension split number N takes the middle value N low , N high ) of the interval [N mid , the sequence is split into N mid subsequences according to the standard of equal computational load per segment. Then, linear integer programming is used to solve for the shortest redundant sequence length to be saved, as well as the corresponding GPU redundant sequence ratio and CPU redundant sequence ratio, under the condition that the Attention communication overhead is less than or equal to the Attention computational overhead (i.e., the communication time can be masked by the computational time).

[0052] The specific constraint conditions and optimization objectives can be expressed by the following formula:

[0053] argmin L local

[0054] F a = 4dl N L

[0055] p cpu + P gpu = 1

[0056]

[0057] 4 * Llocal *d*p cpu ≤M cpu ;

[0058] Among them, L is the total length of the sequence, L local represents the length of the redundant sequence, N is the number of splits of the time dimension sequence, l N is the length of the last subsequence, M is the number of splits of the space dimension sequence (RingAttention parallelism), E is the single GPU FP16 computing power (Flops / s), B gpu 、B cpu are the inter-GPU communication bandwidth and the GPU-CPU communication bandwidth (bytes / s) respectively, d is the model dimension, F a is the Attention computation amount (Flops) of the last subsequence, p cpu 、p gpu represent the proportion of the redundant sequence saved in the CPU memory and the GPU video memory respectively, M cpu is the CPU memory space size.

[0059] Step 2.3: Calculate the video memory overhead of the current policy. If it exceeds the video memory budget, update the interval [N low , N high ) to [N low , N mid ). Otherwise, update the interval [N low , N high ) to [N mid , N high ). If the number of elements in the interval is 1, return the result. Otherwise, repeat steps 2.2 - 2.3.

[0060] Step Three: The deep learning training module first uses the sequence splitting sub-module according to the number of splits N of the time dimension obtained in Step Two, splits the input sequence into the first subsequences with decreasing lengths according to the standard that the computation amount of each sub-sequence is equal, and then further splits each first subsequence into M equal-length second subsequences and distributes the second subsequences to the GPUs in M sequence parallel groups; input the obtained second subsequences into the Transformer large model optimized by the two-dimensional sequence parallel operator for training. When each layer of the Transformer model executes the attention mechanism calculation, the following steps are performed:

[0061] Step 3.1: Calculate the attention of the current Q and the current KV; simultaneously receive the next KV block from the GPU numbered i - 1 and send the current KV to the GPU numbered i + 1. If there is GPU / CPU redundancy, ignore the redundant data during sending / receiving and only send / receive the non-redundant part; at this time, if the CPU redundancy ratio is not 0, simultaneously copy the redundant part of the next KV block from the CPU memory to the GPU video memory.

[0062] Step 3.2: Update the attention calculation result of the local block.

[0063] Step 3.3: If, according to the policy, the KV in the current time period needs to be redundant, save the KV participating in the calculation in step 3.1 to a tensor.

[0064] Step 3.4: Update the current KV to the next KV block, and repeat steps 3.1, 3.2, 3.3, and 3.4 until the attention mechanism calculation is performed for the current Q and each KV block.

[0065] Step 3.5: If, according to the policy, the KV in the current time period needs to be redundant, split the tensor storing the KV in step 3.3 into two parts. One part is stored in the GPU according to the ratio specified by the policy, and the remaining part of the KV is asynchronously copied to the CPU memory and then the tensor in step 3.3 is released.

[0066] Step 3.6: Finally, if the CPU redundancy ratio is not 0, asynchronously prefetch the first CPU redundant KV block of the next layer.

[0067] An embodiment of the present invention also provides an electronic device, including a memory and a processor, the memory is coupled to the processor; wherein, the memory is used to store program data, and the processor is used to execute the program data to implement the two-dimensional sequence splitting method under large model pipeline parallel training.

[0068] An embodiment of the present invention also provides a computer-readable storage medium, on which a computer program is stored, and when the program is executed by a processor, it implements the two-dimensional sequence splitting method under large model pipeline parallel training.

[0069] Embodiment 1

[0070] To verify the advantages of the present invention compared with traditional methods, the present invention provides the following specific embodiments. Among them, Megatron-LM is one of the most widely used pipeline parallel training frameworks at present (without sequence splitting optimization technology), RingAttention (spatial dimension sequence splitting) is one of the most advanced sequence parallel technologies, and Seq1F1B is one of the most advanced technologies for splitting sequences in the time dimension. The specific experiments are as follows:

[0071] Experimental configuration:

[0072] (1) Operating system: Ubuntu 22.04 LTS;

[0073] (2) CPU: Model is 128-core AMD EPYC 7763 CPU, equipped with 1TB DRAM;

[0074] (3) GPU: 4 * NVIDIA A800 with 80GB video memory.

[0075] Model configuration:

[0076] (1) Model: llama2-7B;

[0077] (2) Dataset: LongAlpaca, the average length of the test samples is 8k;

[0078] (3) Parallel configuration: PP (pipeline parallelism) = 2; CP (RingAttention sequence parallelism) = 2;

[0079] (4) Training configuration: BatchSize = 4, MicrobatchSize = 1;

[0080] (5) Memory limit: The CPU memory usage limit for a single card is 20GB.

[0081] Test metrics:

[0082] Average throughput per GPU (tokens / s).

[0083] Final test results:

[0084] The final test results are as Figure 2 shown: Without using any sequence parallel methods, the average throughput of Megatron-LM is 1414 tokens / s; when only using the RingAttention algorithm in the training pipeline, the throughput is 1528 tokens / s; when only using Seq1F1B in the training pipeline, the throughput is 1498 tokens / s; when using both simply combined, the throughput is 1465 tokens / s; the average throughput of the method of the present invention can reach 1556 tokens / s. Compared with no splitting, two-dimensional sequence splitting parallelism improves the throughput by about 10%, and improves the throughput by 2 - 5% compared with other single-dimension or simple combination schemes.

[0085] The above embodiments are only used to illustrate the design concept and features of the present invention, and the purpose is to enable those skilled in the art to understand the content of the present invention and implement it accordingly. The protection scope of the present invention is not limited to the above embodiments. Therefore, all equivalent changes or modifications made according to the principles and design concepts disclosed by the present invention are within the protection scope of the present invention.

Claims

1. A two-dimensional sequence splitting method under large model pipeline parallel training, characterized in that: The steps include: S1. Use real RingAttention load test and collect inter-GPU bandwidth, device video memory size, device CPU memory size, GPU-CPU bandwidth, model dimension, number of model layers, and input data sequence length; S2. Making intelligent decisions based on the information collected in step S1, wherein the decision content includes the length of the redundant sequence, the proportion of the redundant sequence stored in the GPU, the proportion of the redundant sequence stored in the CPU, and the number of time dimension splits; S3. Split the input sequence into first subsequences of decreasing length according to the number of splits in the time dimension and the length of each subsequence, further split the subsequence into M second subsequences of equal length, and distribute the M second subsequences to the GPUs in the M sequence parallel groups; input the second subsequences into the Transformer large model optimized by the two-dimensional sequence parallel operator for training; wherein, during the training process, the redundant sequence length, the proportion of redundant sequences stored in the GPU, and the proportion of redundant sequences stored in the CPU are used as parameters to be input into the two-dimensional sequence parallel operator, and during training, the following are guided by the above parameters: saving redundant sequences to the GPU, transferring redundant sequences from the GPU to the CPU, and transferring redundant sequences from the CPU to the GPU.

2. According to claim 1, a two-dimensional sequence splitting method under large model pipeline parallel training is characterized in that: The step S1 is specifically as follows: When the user deploys the training framework to the cluster and inputs the model and training configuration, the model dimension, number of model layers, and length of the input data sequence are recorded, and a sequence of arbitrary length and dimension is used to run the RingAttention operator for several rounds in the sequence parallel group. The FP16 precision computing capability of the GPU device and the communication bandwidth between GPU devices are calculated based on the time obtained from the profile when the operator is executed. The above sequence is then used to copy several rounds between the GPU and CPU, and the bandwidth between the GPU and CPU is calculated based on the time obtained from the profile when copying.

3. According to claim 2, a two-dimensional sequence splitting method under large model pipeline parallel training is characterized in that: The RingAttention operator is run for several rounds in the sequence parallel group using sequences of arbitrary length and dimension, and a sequence with a length of 32k and a dimension of 1024 is selected.

4. The two-dimensional sequence splitting method under large model pipeline parallel training according to claim 1 is characterized in that: The step S2 comprises: S2.1: Initialize the time dimension split number search interval [N low , N high ) is [1, N max ), where N max The maximum number of time dimension splits set by the user, N low is the lower bound of the search interval of the current round of binary search, N high The upper bound of the search interval for the current round of binary search; S2.2: In the time dimension, the number of splits N is taken as the interval [N low , N high )Intermediate value N mid In the case of , the sequence is split into N segments according to the standard that the amount of calculation in each segment is equal mid subsequences; under the condition that the Attention communication overhead is less than or equal to the Attention computation overhead, linear integer programming is used to solve the shortest redundant sequence length, the corresponding GPU redundant sequence ratio and the CPU redundant sequence ratio; S2.3: Calculate the memory overhead of the current decision. If it exceeds the memory budget, update the interval [N low , N high ) is [N low , N mid ), otherwise, update interval [N low , N high ) is [N mid , N high ); If the number of elements in the interval is 1, return the result, otherwise repeat steps S2.2-S2.

3.

5. According to claim 4, a two-dimensional sequence splitting method under large model pipeline parallel training is characterized in that: The step S2.2 includes: the constraints and optimization targets for searching the shortest redundant sequence length and the optimal CPU to GPU redundant sequence ratio under the shortest redundant sequence length are shown in the following formula: Among them, L is the total length of the sequence, L local represents the length of redundant sequence, N is the number of time dimension sequence splits, l N is the length of the last subsequence, M is the number of spatial dimension sequence splits, E is the FP16 computing power of a single GPU, and B gpu , B cpu are the inter-GPU communication bandwidth and the GPU-CPU communication bandwidth respectively, d is the model dimension, and F a The Attention calculation amount for the last subsequence, p cpu 、p gpu Respectively represent the proportion of redundant sequences stored in CPU memory and GPU memory, M cpu The size of the CPU memory space.

6. The two-dimensional sequence splitting method under large model pipeline parallel training according to claim 1 is characterized in that: The step S3 includes: performing the following steps when each layer of the Transformer model performs attention mechanism calculation: S3.1: Calculate the attention of the current Q and the current KV; at the same time, receive the next KV block from the GPU numbered i-1 and send the current KV to the GPU numbered i+1. If there is GPU / CPU redundancy, ignore the redundant part of the data when sending / receiving, and only send / receive the non-redundant part; if the CPU redundancy ratio is not 0 at this time, copy the redundant part of the next KV block from the CPU memory to the GPU memory at the same time; S3.2: Update the attention calculation results of the local block; S3.3: If according to the decision, the KV of the current time period needs to be redundant, the KV involved in the calculation in step S3.1 is saved into a tensor; S3.4: Update the current KV to the next KV block, and repeat steps S3.1-S3.4 until the current Q and each KV block have performed attention mechanism calculations; S3.5: If, according to the decision, the KV of the current time period needs to be redundant, the tensor storing the KV in step S3.3 is split into two parts, wherein one part is stored in the GPU according to the ratio specified by the decision, and the remaining part of the KV is asynchronously copied to the CPU memory and then the tensor in step S3.3 is released; S3.6: If the CPU redundancy ratio is not 0, asynchronously pre-fetch the first CPU redundancy KV block of the next layer.

7. A two-dimensional sequence segmentation system under large model pipeline parallel training, characterized in that: The system comprises: The data collection module is used to obtain basic device information and model configuration information, including inter-GPU bandwidth, device memory size, GPU-CPU bandwidth, model dimension, number of model layers, and input data sequence length; A decision maker, used to generate an optimal decision based on the data acquired by the data collection module; wherein the decision content includes the length of redundant sequences, the proportion of redundant sequences stored in the GPU, the proportion of redundant sequences stored in the CPU, and the number of time dimension splits; A deep learning training module is used to integrate the optimal decision into the model training process to improve the overall training performance of the system. The deep learning training module includes a long sequence sample splitting submodule and a two-dimensional sequence parallel operator submodule. The long sequence sample splitting submodule is used to split the input sequence into first subsequences of decreasing length according to the standard that the computational amount of each subsequence is equal, and then further split each first subsequence into M second subsequences of equal length and distribute the second subsequences to the GPUs in the M sequence parallel groups. The two-dimensional sequence parallel operator submodule is used to optimize the Transformer large model using the two-dimensional sequence parallel operator.

8. An electronic device, comprising a memory and a processor, characterized in that: The memory is coupled to the processor; wherein the memory is used to store program data, and the processor is used to execute the program data to implement a two-dimensional sequence splitting method under large model pipeline parallel training as described in any one of claims 1-6.

9. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the program is executed by a processor, a two-dimensional sequence splitting method under large model pipeline parallel training as described in any one of claims 1 to 6 is implemented.

Citation Information

Patent Citations

  • Attention mechanism calculation and model reasoning method and device, equipment and medium

    CN118798263A

  • Estimating Resource Costs for Computing Tasks for a Reconfigurable Dataflow Computing System

    US20240086235A1