Large model distributed training adaptive method and system based on sequence length
By testing the best training strategies for sequence data of different lengths and using the first adaptation decreasing algorithm for data splicing and grouping, the training strategy is dynamically adjusted, and the problem of waste of computing resources in the long-tail distribution data set and the high complexity of the sequence splicing algorithm is solved, and efficient long-sequence data processing and training efficiency are achieved.
Patent Information
- Application Number
- CN202510093769.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-21
- Publication Date
- 2025-06-10
AI Technical Summary
When the prior art processes natural language processing data sets with long-tail distribution, computing resources are seriously wasted, and the existing sequence splicing algorithm is complex, making it difficult to efficiently process long sequence data.
By testing the best training strategies for sequence data of different lengths, the short sequence data are spliced in the pre-processing of training data, and the spliced data is divided into different microbatches according to the length. During the training process, the training strategy is dynamically adjusted according to the length of the training data in the microbatches, and the first adaptation decreasing algorithm is used for splicing and grouping.
It significantly improves the training efficiency on the long-tail distribution dataset, reduces the use of fill symbols, and realizes efficient processing of long-sequence data.
Smart Images

Figure CN120124713A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of information technology, and particularly relates to an adaptive method and system for distributed training of large models based on sequence length. Background Art
[0002] The Zero technology aims to solve the problem of video memory limitation caused by storing the entire model in data parallel training. Its video memory consumption is mainly divided into two parts: model state and other states. The model state includes model parameters, gradients, and optimizer states (such as the first-order momentum and second-order momentum in the Adam optimizer). Other states cover intermediate results, temporary buffers, and unused fragmented spaces during model training. Zero is divided into three levels: Zero1, Zero2, and Zero3, which respectively split the optimizer state, gradients, and model parameters in data parallelism. During training, only the complete data is obtained through communication when a certain part of the parameters is needed, and the unnecessary part is immediately discarded after use, thus significantly reducing the video memory occupancy. This mechanism enables data parallelism to support larger-scale model training.
[0003] DeepSpeed Ulysses is a parallel strategy designed for long sequences. By splitting the input tensor along the sequence dimension and using an all-to-all collective communication to transfer the split dimension to the embedding layer dimension before attention calculation and restoring the split dimension to the sequence dimension again through all-to-all communication after attention calculation, it effectively improves the ability to process long sequences and the training speed.
[0004] In terms of sequence optimization of the Transformer model, PackedBert proposed a sequence packing method aiming to solve a common problem in natural language processing: in order to unify the sequence length, a large number of padding tokens usually need to be added, which leads to waste of computing resources. This method combines short sequences into long sequences through an efficient algorithm, significantly reducing the use of padding tokens. Specifically, it includes two algorithms: the shortest-pack-first histogram packing algorithm, which packs sequences from long to short according to sequence length and limits the packing depth; and the non-negative least squares histogram packing algorithm, which transforms the packing problem into a weighted non-negative least squares problem. In addition, in order to maintain the mathematical consistency of the model, corresponding adjustments need to be made to position embeddings, attention masks, the loss of each sequence, and hyperparameters.
[0005] The disadvantages of the existing technical solutions are as follows:
[0006] Most natural language processing datasets, such as large-scale datasets like CommonCrawl, Wikipedia, and GitHub, typically exhibit long-tailed distribution characteristics. This means that short sequence data accounts for a relatively high proportion in the dataset, while long sequence data is less. Therefore, maintaining a high DeepSpeed Ulysses sequence parallelism hybrid parallel strategy for a small number of long sequences throughout the training process will cause significant computational resource waste. For example: the longest data length in the dataset is 64k. To train 64k-length training data, a hybrid parallel strategy with a DeepSpeed Ulysses sequence parallelism of 8 and a zero data parallelism of 2 needs to be run. If a hybrid parallel strategy with a lower DeepSpeed Ulysses sequence parallelism is adopted, such as a DeepSpeed Ulysses sequence parallelism of 4 and a zero data parallelism of 4, there will be a situation of out-of-memory. However, there are a large number of short sequences in this dataset, and short sequences are more computationally efficient in a hybrid parallel strategy with a high zero data parallelism and a low DeepSpeed-Ulysses parallelism.
[0007] The two sequence concatenation schemes used by PackedBert - the shortest-pack-first histogram packing algorithm and the non-negative least squares histogram packing algorithm - have complexities of and where n is the number of samples in the dataset and s m is the longest sequence length. These algorithms have a high complexity when dealing with long sequences, and the shortest-pack-first algorithm only supports the concatenation of a small number of sequences. Summary of the Invention
[0008] In view of the above problems, the present invention provides an adaptive method and system for distributed training of large models based on sequence length.
[0009] The technical solution adopted by the present invention is as follows:
[0010] An adaptive method for distributed training of large models based on sequence length, comprising the following steps:
[0011] Testing the best training strategy for sequence data of different lengths;
[0012] Concatenating short sequence data during the preprocessing of training data, and dividing the concatenated data into different micro-batches according to the length;
[0013] Dynamically adjusting the training strategy according to the length of the training data in the micro-batch during the training process.
[0014] Further, the testing of the best training strategy for sequence data of different lengths includes:
[0015] Under the condition of ensuring that the total sequence length of the training samples in the global batch is equal, the running time of training data of different lengths under different hybrid parallel training strategies is tested;
[0016] For the training data of each length, a hybrid parallel training strategy suitable for the training data of this length is obtained through testing, as well as the number of samples in the micro-batch that maximizes the training efficiency under this hybrid parallel training strategy.
[0017] Furthermore, the hybrid parallel training strategy refers to a hybrid parallel training strategy composed of data parallel zero and sequence parallel DeepSpeedUlysses.
[0018] Furthermore, the first-fit decreasing algorithm is used for the splicing and the training data is divided into different micro-batches; the first-fit decreasing algorithm constructs micro-batches according to the global batch, where the global batch contains several sequences of unequal lengths, the micro-batch contains several sequences of equal lengths, one sequence in the micro-batch is spliced by one or more sequences in the global batch, and one piece of data in the global batch can only exist uniquely in one micro-batch.
[0019] Furthermore, the first-fit decreasing algorithm includes the following steps:
[0020] Sort the training data in the global batch from longest to shortest sequence length;
[0021] When there are sequences in the global batch that have not been added to a certain micro-batch, try to construct a micro-batch until all sequences in the global batch are added to a certain micro-batch;
[0022] Obtain the longest sequence among the remaining sequences in the global batch. The remaining sequences refer to the sequences that have not been added to a certain micro-batch. According to the longest remaining sequence, determine the number of sequences and the length of the sequences in the currently constructed micro-batch: round up the length of the longest remaining sequence as the sequence length of the micro-batch; through the determined sequence length of the micro-batch, query the results in the test phase to obtain the most efficient hybrid parallel strategy and the micro-batch size as the hybrid parallel strategy and the micro-batch size of the current micro-batch;
[0023] After determining the size and sequence length of the micro-batch, each spliced sequence in the micro-batch is equivalent to a bucket, and the size of each bucket is the sequence length of the micro-batch. Try to put the remaining sequences in the global batch into the buckets as much as possible so that the total remaining space in the buckets is minimized. The sequences put into the same bucket are finally spliced into one sequence;
[0024] Traverse the remaining sequences in the global batch from long to short. For the traversed sequences, traverse all the buckets from front to back and put them into the first bucket that can accommodate the sequence. If no suitable bucket is found for the sequence, skip the sequence.
[0025] Further, the formula for rounding up is as follows:
[0026]
[0027] where s represents the sequence length.
[0028] Further, dynamically adjusting the training strategy according to the length of the training data in the micro-batch during the training process includes:
[0029] The hybrid parallel strategy adopted for different micro-batches is determined by querying the results in the test phase according to the sequence length in the micro-batch. When switching different hybrid parallel strategies, keep the communication group required for zero data parallelism unchanged, including all devices. The communication group for DeepSpeed Ulysses sequence parallelism changes dynamically with the size of the sequence parallelism.
[0030] An adaptive system for distributed training of large models based on sequence length, which includes:
[0031] A test module for testing the best training strategy for sequence data of different lengths;
[0032] A splicing and grouping module for splicing short sequence data in the preprocessing of training data and dividing the spliced training data into different micro-batches according to the length;
[0033] A dynamic adjustment module for dynamically adjusting the training strategy according to the length of the training data in the micro-batch during the training process.
[0034] The beneficial effects of the present invention are as follows:
[0035] 1) By pre-testing the training strategies suitable for sequences of different lengths, splicing short sequence data in the preprocessing of training data, dividing the spliced sequences into micro-batches according to the length, and dynamically adjusting the training strategy according to the length of the training data during the training process, the training efficiency in most real datasets (with long-tail distribution) can be improved.
[0036] 2) In order to efficiently implement the splicing of long sequences, the present invention adopts the first-fit decreasing algorithm, which can not only efficiently process the splicing of long sequences, but also greatly reduce the use of padding symbols.
[0037] 3) The present invention uses GPT2 and Llama with 7B parameters to achieve a 109%, 78%, 128%, and 76% increase in training speed on the Wikipedia and GitHub datasets respectively. BRIEF DESCRIPTION OF THE DRAWINGS
[0038] Figure 1 is a schematic flowchart of an adaptive method for distributed training of large models based on sequence length.
[0039] Figure 2 is a schematic diagram of communication group splitting under different hybrid parallel strategies. Figure 3 is a schematic diagram of communication group division in the embodiment. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0040] To make the above objects, features, and advantages of the present invention more obvious and understandable, the present invention will be further described in detail below through specific embodiments and the accompanying drawings.
[0041] With the rapid development of machine learning, especially deep learning, large language models have become the key to improving the performance of intelligent tasks. With excellent expressiveness and flexibility, these models have achieved remarkable results in the fields of natural language processing, image recognition, etc. However, some more complex tasks have put forward higher requirements for the long sequence processing ability of large language models. For example, in the text summarization task, the model needs to efficiently extract and integrate the core information of long texts. Therefore, it has become particularly important to introduce long sequence data during the training process to enhance the processing ability of the model. However, training long sequence data requires more computing resources. How to efficiently train long sequence data in large language models has become a hot issue of concern in the academic and industrial communities.
[0042] DeepSpeed Ulysses is a sequence parallel strategy designed specifically for long sequence training. Compared with data parallelism (zero, a data parallel strategy), it can handle longer sequences, but it is less efficient when dealing with short sequences. In real training scenarios, the training data has an obvious long-tail distribution, that is, only a small number of data sequences are long, and the vast majority of data sequences are short. Therefore, the present invention proposes a solution that can dynamically adjust the training strategy of large language models according to the sequence length of the training data. The technical key points of the present invention are as follows:
[0043] 1) During the training process of large language models, through pre-testing the training strategies suitable for sequences of different lengths, concatenating short sequence data during the preprocessing of training data and dividing the concatenated sequences into micro-batches according to length, and dynamically adjusting the training strategy according to the length of the training data during the training process, the large language model can efficiently process the long-tail distribution dataset in three stages.
[0044] 2) Phase of pre - testing training strategies suitable for sequences of different lengths: First, test training data of different lengths, such as 1k, 2k, 4k, 8k, etc. Then, for each length of training data, obtain the hybrid parallel training strategy suitable for this length of training data through testing, as well as the micro - batch size that maximizes the training efficiency under a specific strategy, that is, the number of samples in the micro - batch.
[0045] 3) During the pre - processing of training data, splice short - sequence data and group the spliced sequences by length: During the training process of large - language models, training data is usually divided into several global batches, and further divided into several micro - batches based on the global batches. After all micro - batches of a global batch are iterated, a model parameter update is performed. Guided by the analysis data in the testing phase, splice the training data within the global batch through the first - fit decreasing algorithm and divide it into different micro - batches.
[0046] 4) During the training process, switch different hybrid parallel training strategies according to the length of the training data in the micro - batch for training: During this process, always keep the parameters, gradients, optimizer states, etc. split among all devices, while the DeepSpeed Ulysses communication group changes dynamically with the length of the training data in the batch.
[0047] The present invention consists of three processes. In the first process, the present invention tests the training environment suitable for training data of different lengths. Here, the training environment includes hybrid parallel strategies and the number of samples in the micro - batch that runs most efficiently under a specific hybrid parallel strategy, that is, the size of the micro - batch. The second link attempts to splice sequences of different lengths together within the global batch and divide them into several micro - batches based on the test data obtained in the first process. In the third stage, during operation, different hybrid parallel strategies are executed according to the sequence length of each micro - batch. The flow chart of the distributed training adaptive method for large models based on sequence length is as Figure 1 shown, and the testing process in the first stage is not shown in the figure.
[0048] 1. Test and obtain training strategies suitable for training data of different lengths
[0049] First, ensure that the total sequence length of the training samples in the global batch is equal, and test the running time of sequences of different lengths under different hybrid parallel strategies. Here, the hybrid parallel strategy refers to the hybrid parallel strategy composed of data parallel zero3 (taking zero3 as an example for analysis, the method of the present invention is also applicable to zero1 and zero2) and sequence parallel DeepSpeed Ulysses. For example, set the data parallelism of zero3 to 2 and the sequence parallelism of DeepSpeed Ulysses to 8 (dp2sp8), or set the data parallelism of zero3 to 4 and the sequence parallelism of DeepSpeed Ulysses to 4 (dp4sp4).
[0050] Table 1 shows the test results in an environment with two nodes, each node containing 8 40G A100 GPUs. For example, the meaning of the second row in the table is to test the training efficiency of 128k-length training data under different hybrid parallel strategies. For 128k-length training data, it can only run under the hybrid parallel strategy with the data parallelism of zero3 being 1 and the sequence parallelism of DeepSpeed Ulysses being 16 (dp1sp16). Reducing the sequence parallelism of DeepSpeed Ulysses will result in out-of-memory (OOM). The meaning of the last row in the table is to test the training efficiency of 1k-length training data under different hybrid parallel strategies. It can be found that the running efficiency is the highest when running under the hybrid parallel strategy with the data parallelism of zero3 being 16 and the sequence parallelism of DeepSpeed Ulysses being 1 (dp16sp1). That is, for shorter sequences, it is more suitable to run under the hybrid parallel strategy with a high data parallelism of zero3 and a low sequence parallelism of DeepSpeed Ulysses. The situations of the other rows can be analyzed similarly.
[0051] Table 2 shows the micro-batch size (the number of samples in the micro-batch) with the highest running efficiency of training data of different lengths under different hybrid parallel strategies. For example, the meaning of the second row in the table is that for 64k-length training data, under the hybrid parallel strategy with the data parallelism of zero3 being 16 and the sequence parallelism of DeepSpeed Ulysses being 1 (dp16sp1), the micro-batch size with the highest running efficiency is 2, while under the dp16sp1 hybrid parallel strategy, the micro-batch size with the highest running efficiency is 1. The statistics of this information are to guide the subsequent sequence splicing in the second stage and divide the spliced sequences into different micro-batches.
[0052] Table 1 Running Speeds of Different Sequence Lengths at Different Sequence Parallelisms
[0053]
[0054] Table 2 Sizes of the most efficient micro - batches for different sequence lengths at different degrees of sequence parallelism
[0055]
[0056] 2. During the pre - processing of training data, splice and group short - sequence data
[0057] In this process, the present invention splices the sequences within a global batch and divides them into several micro - batches. Specifically, the global batch contains several sequences with unequal lengths, and the micro - batch contains several sequences with equal lengths. One sequence in the micro - batch is spliced from one or more sequences in the global batch, and one piece of data in the global batch can only exist in one micro - batch uniquely. To make the sequences in the micro - batch have equal lengths, padding characters may be added at the end of the sequence. The present invention hopes that this process is efficient and the additional padding characters added are few. Therefore, the present invention uses the first - fit decreasing algorithm, and the process of this algorithm is as follows:
[0058] 1) First, sort the training data in the global batch from the longest to the shortest sequence length.
[0059] 2) When there are sequences in the global batch that have not been added to a certain micro - batch, try to construct a micro - batch until all sequences in the global batch are added to a certain micro - batch.
[0060] 2.1) Obtain the longest sequence among the remaining sequences in the global batch. The remaining sequences refer to the sequences that have not been added to a certain micro - batch. According to the longest remaining sequence, determine the number of sequences and the length of the sequences in the currently constructed micro - batch.
[0061] 2.1.1) Determine the sequence length of the micro - batch: Round up the length of the longest sequence among the remaining sequences using the following formula, where s represents the length of the sequence. Here, it is considered that if the rounding - up range is too small, it will be difficult to use short sequences for splicing, and if the rounding - up range is too large, a large number of unnecessary padding characters are required. Therefore, the present invention uses a dynamic rounding - up range.
[0062]
[0063] 2.1.2) Determine the number of sequences in a micro-batch (i.e., the micro-batch size): Based on the determined sequence length of the micro-batch, round up to the sequence lengths in Table 1 and Table 2 and query. For example, if the sequence length of the micro-batch is 3k, round up to 4k. Take the hybrid parallel strategy and micro-batch size corresponding to the 4k-length sequences in Table 1 and Table 2 as the hybrid parallel strategy and micro-batch size of the current micro-batch. It should be noted that the final micro-batch size needs to be multiplied by the data parallelism of the hybrid parallel strategy because during training, the data will be evenly divided among different data parallel groups for processing.
[0064] 2.2) When the number of sequences and the sequence length in the micro-batch are determined, each concatenated sequence in the micro-batch is equivalent to a bucket, and the size of each bucket is the sequence length of the micro-batch. We hope to put the remaining sequences in the global batch into the buckets as much as possible to minimize the total remaining space in the buckets. Sequences put into the same bucket will eventually be concatenated into one sequence.
[0065] 2.3) Traverse the remaining sequences in the global batch from longest to shortest. For the traversed sequence, traverse all the buckets from front to back and put it into the first bucket that can accommodate the sequence, that is, the sum of the sequence lengths in the bucket plus the length of this sequence does not exceed the size of the bucket. If the sequence does not find a suitable bucket, skip this sequence.
[0066] The core idea of this algorithm is as Figure 2 shown, where GlobalBatch represents the global batch, Micro Batch represents the micro-batch, sequence length represents the sequence length, and Micro batch size represents the number of sequences in the micro-batch. It should be noted that the present invention uses a balanced tree when implementing the first-fit decreasing algorithm, reducing the complexity of finding and deleting a sequence to O(logn). Therefore, the overall complexity is O(nlogm). Here, n is the number of samples in the entire dataset, and m is the size of the global batch. Because in most cases, the size of the global batch is in the order of magnitude of 10 1 and 10 2 Therefore, the overall complexity is close to the size of the dataset.
[0067] 3. Dynamically adjust the parallel strategy during training
[0068] The hybrid parallel strategy adopted by different micro-batches is determined by querying the results of the test phase according to the sequence length in the micro-batch. For example, if the sequence length of a micro-batch is 2k, and the training data of length 2k obtained in the test phase is most suitable for training under the hybrid parallel strategy where the zero data parallelism is 16 and the DeepSpeed Ulysses sequence parallelism is 1, then this micro-batch is trained using the above hybrid parallel strategy. The hybrid parallel training composed of Zero3 and DeepSpeed Ulysses has a natural advantage in strategy switching. Zero3 can not only be used complementarily with DeepSpeed Ulysses, but also extend its function of parameter state splitting to the parallel groups of DeepSpeed Ulysses. Specifically, it can perform parameter state splitting in the sequence and data parallel groups. When performing All Gather and Reduce Scatter operations, the communication groups involved are also extended to the sequence and data parallel groups. As Figure 3 , when the zero3 data parallelism (dp) and the DeepSpeed Ulysses sequence parallelism (sp) are both 2, when performing operations such as All Gather and Reduce Scatter in Zero3, these operations will be carried out across four devices. While when performing the AllToAll communication in DeepSpeed Ulysses, the communication is limited between two devices within the same sequence parallel group. In the case where the zero3 data parallelism is 1 and the DeepSpeed Ulysses sequence parallelism is 4, when performing operations such as All Gather and Reduce Scatter in Zero3, the corresponding communication group remains unchanged, while when performing the AllToAll communication in DeepSpeed Ulysses, the communication group needs to be dynamically changed, and at this time, the AllToAll needs to be carried out among four devices.
[0069] The present invention implements the switching of the hybrid parallel strategy in the Megatron-DeepSpeed framework (a popular distributed training framework). The initialization of the data parallel communication group and the sequence parallel communication group is carried out in the Parallel_state file in the Megatron-DeepSpeed framework and saved in the mpu object. Therefore, by modifying the initialization code, the creation of all communication groups involved in sp16 (DeepSpeed Ulysses sequence parallelism is 16), sp8, sp4, and sp2 can be completed at the start of training. In addition, the mpu object is passed into the DeepSpeed framework. When performing the AllToAll communication in DeepSpeedUlysses in the DeepSpeed framework, the sequence parallel communication group in the mpu will be read. Therefore, a feasible modification strategy is to switch the sequence parallel group in the mpu according to the sequence length of the training data before each micro-batch training and re-pass it into the DeepSpeed framework to overwrite the previous mpu state.
[0070] In summary, the technical solution of the present invention can effectively improve the training efficiency of long-tail distribution datasets. The present invention customizes a dynamic training strategy through three stages to adapt to the characteristics of long-tail distribution datasets: first, test the best training strategy for sequences of different lengths; then splice short sequence data during data preprocessing and divide them into different training batches; finally, dynamically adjust the strategy according to the data length in each batch during training. When the training data is long, a hybrid parallel strategy of low-parallelism Zero data parallel and high-parallelism DeepSpeed Ulysses is adopted; when the data is short, a hybrid parallel strategy of high-parallelism Zero data parallel and low-parallelism DeepSpeed Ulysses is adopted. At the same time, in order to efficiently implement the splicing of long sequence data, the present invention adopts a new sequence splicing algorithm, which can not only efficiently process the splicing of long sequences, but also significantly reduce the use of padding symbols. Without changing the training paradigm, the training efficiency is significantly improved.
[0071] The present invention can be used in the field of natural language processing. For example, during the pre-training of large language models, most datasets show a long-tail distribution. The method of the present invention can be used to dynamically adjust the training parallel strategy according to the sequence length during training, so as to achieve the effect of improving the training speed.
[0072] Another embodiment of the present invention provides a large model distributed training self-adaptive system based on sequence length, which includes:
[0073] A test module for testing the best training strategy for sequence data of different lengths;
[0074] The splicing grouping module is used to splice short sequence data in the preprocessing of training data and divide the spliced training data into different micro-batches according to the length.
[0075] The dynamic adjustment module is used to dynamically adjust the training strategy according to the length of the training data in the micro-batch during the training process.
[0076] The division of the above modules is only for illustrative purposes. In actual applications, the above functions can be assigned to different functional modules according to needs to complete all or part of the functions described in the foregoing method. The specific working processes of the above modules can refer to the corresponding processes in the foregoing method embodiments and will not be elaborated herein.
[0077] Another embodiment of the present invention provides a computer device (such as a computer, a server, a smart phone, etc.), which includes a memory and a processor. The memory stores a computer program, and the computer program is configured to be executed by the processor. The computer program includes instructions for executing each step in the method of the present invention.
[0078] Another embodiment of the present invention provides a computer-readable storage medium (such as ROM / RAM, a disk, an optical disc). The computer-readable storage medium stores a computer program, and when the computer program is executed by a computer, each step of the method of the present invention is implemented.
[0079] The specific embodiments of the present invention disclosed above are intended to help understand the content of the present invention and implement it accordingly. Those of ordinary skill in the art can understand that various substitutions, changes, and modifications are possible without departing from the spirit and scope of the present invention. The present invention should not be limited to the content disclosed in the embodiments of this specification, and the protection scope of the present invention is subject to the scope defined by the claims.
Claims
1. A large model distributed training adaptive method based on sequence length, characterized in that: The following steps are involved: Testing the best training strategy for sequence data of different lengths; In the training data preprocessing, short sequence data are concatenated and the concatenated data are divided into different micro-batches according to their lengths; The training strategy is dynamically adjusted during training according to the length of the training data in the mini-batch.
2. The method according to claim 1, characterized in that The optimal training strategy for testing sequence data of different lengths includes: While ensuring that the total length of the sequences of training samples in the global batch is equal, test the running time of training data of different lengths under different hybrid parallel training strategies; For each length of training data, a hybrid parallel training strategy suitable for the length of training data is obtained through testing, as well as the number of samples in the micro-batch that makes the training efficiency highest under the hybrid parallel training strategy.
3. The method according to claim 2, characterized in that The hybrid parallel training strategy refers to a hybrid parallel training strategy consisting of data parallel zero and sequence parallel DeepSpeed Ulysses.
4. The method according to claim 1, characterized in that: A first-fit decreasing algorithm is used to perform the splicing and divide the training data into different micro-batches; the first-fit decreasing algorithm constructs micro-batches according to the global batch, wherein the global batch contains a number of sequences of unequal lengths, the micro-batch contains a number of sequences of equal lengths, a sequence of the micro-batch is spliced by one or more sequences in the global batch, and a piece of data in the global batch can only exist in one micro-batch.
5. The method according to claim 4, characterized in that The first adaptation decreasing algorithm comprises the following steps: Sort the training data in the global batch from long to short according to sequence length; When there are sequences in the global batch that have not been added to a micro-batch, try to construct a micro-batch until all sequences in the global batch are added to a micro-batch; Obtain the longest sequence among the remaining sequences in the global batch. The remaining sequence refers to the sequence that has not been added to a micro-batch. According to the longest remaining sequence, determine the number of sequences and the length of the sequence in the currently constructed micro-batch: round up the length of the longest sequence among the remaining sequences as the sequence length of the micro-batch; query the results of the test phase through the determined sequence length of the micro-batch, obtain the most efficient hybrid parallel strategy and micro-batch size, and use them as the hybrid parallel strategy and micro-batch size of the current micro-batch; After determining the size of the micro-batch and the sequence length, each concatenated sequence in the micro-batch is equivalent to a bucket. The size of each bucket is the sequence length of the micro-batch. Try to put the remaining sequences in the global batch into the bucket so that the total remaining space in the bucket is minimized. The sequences put into the same bucket are finally concatenated into one sequence. Traverse the remaining sequences in the global batch from long to short. For the traversed sequences, traverse all buckets from front to back and put them into the first bucket that can accommodate the sequence; if the sequence does not find a suitable bucket, skip the sequence.
6. The method according to claim 5, characterized in that The formula used for rounding up is as follows: Where s represents the sequence length.
7. The method according to claim 1, characterized in that The training strategy is dynamically adjusted according to the length of the training data in the micro-batch during the training process, including: the hybrid parallel strategy adopted by different micro-batches is determined according to the result of the sequence length query test phase in the micro-batch; when different hybrid parallel strategies are switched, the communication group required for zero data parallelism is kept unchanged, including all devices, and the communication group of DeepSpeed Ulysses sequence parallelism changes dynamically with the size of the sequence parallelism.
8. A large model distributed training adaptive system based on sequence length, characterized in that: include: The testing module is used to test the best training strategy for sequence data of different lengths; The splicing and grouping module is used to splice short sequence data in training data preprocessing and divide the spliced training data into different micro-batches according to their length; The dynamic adjustment module is used to dynamically adjust the training strategy according to the length of the training data in the micro-batch during the training process.
9. A computer device, characterized in that: The method comprises a memory and a processor, wherein the memory stores a computer program, the computer program is configured to be executed by the processor, and the computer program comprises instructions for executing the method according to any one of claims 1 to 7.
10. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores a computer program, and when the computer program is executed by a computer, the method according to any one of claims 1 to 7 is implemented.