Language Task Processing Method, System, Device, Storage Medium and Program Product

By determining the word batch length based on resource configuration information during the pre-filling and decoding stage of the language task processing model, and building multiple pipelines to process word batches in parallel in the decoding stage, the problem of resource utilization unsaturation in the decoding stage of the language task processing model is solved, and higher resource utilization and more stable delay performance are achieved.

CN120068846BActive Publication Date: 2025-07-01SHANDONG HAILIANG INFORMATION TECH RES INST
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510526403.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-04-25
Publication Date
2025-07-01
Estimated Expiration
2045-04-25

AI Technical Summary

Technical Problem

When the language task processing model executes language tasks, there is a problem of resource utilization unsaturation in the decoding stage, resulting in low resource utilization, reduced throughput and increased task processing delay.

Method used

By determining the word batch length according to the resource configuration information of the language task processing model in the pre-filling stage and the decoding stage, the word batch length is determined, and the target request segment is processed in parallel in the pre-filling stage to generate the word batch. During the decoding stage, multiple pipelines are built to process word batches in parallel to improve resource utilization.

Benefits of technology

It effectively improves the resource utilization rate during language task execution, reduces resource coupling, improves the stability of pre-filling delay, and reduces the problem of resource utilization unsaturation in the decoding stage.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120068846B_ABST
    Figure CN120068846B_ABST
Patent Text Reader

Abstract

The present invention discloses a language task processing method, system, device, storage medium and program product, relating to the field of artificial intelligence technology. Among them, the method includes determining resource configuration information in the pre-filling stage and the decoding stage according to the resource requirement information of the language task processing model during the execution of the language task. Obtain a matching number of target request segments from the current request batch request, and perform pre-filling parallel processing on them to generate the current token batch. By obtaining the next token of each token in the newly generated token batch to form a new token batch, multiple new token batches are generated to meet the condition of merging the batch to the token batch length. Decode each token batch in parallel through multiple pipelines, and obtain the corresponding language task processing result according to the decoding results of all request segments of each task processing request. The present invention can solve the problem of unsaturated resource utilization in the related technology when executing language tasks, and can effectively improve resource utilization rate.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of artificial intelligence technology, and in particular, to a method, system, electronic device, computer-readable storage medium, and computer program product for processing language tasks. Background Art

[0002] When a language model executes a language task, its inference generation process includes a prefill stage and a decoding stage with different computational requirements. To avoid the problem of reduced inference latency caused by a single device mixing the processing of the two stages, related technologies split and deploy them to different computing devices for processing according to the different performance requirements of the prefill stage and the decoding stage. As the length of the input context text increases, the performance requirement difference between the two stages will continuously increase, resulting in the underutilization of resources in the decoding stage. It can be seen that related technologies cannot significantly improve the problem of low resource utilization during the execution of language tasks. Summary of the Invention

[0003] The present invention provides a method, system, electronic device, computer-readable storage medium, and computer program product for processing language tasks, which can solve the problem of underutilized resources in the decoding stage when a language task processing model executes a language task, and can effectively improve the resource utilization rate during the execution of language tasks.

[0004] To solve the above technical problems, the present invention provides the following technical solutions:

[0005] The present invention provides a method for processing language tasks, including:

[0006] According to the resource requirement information of the language task processing model during the execution of the language task, determine the resource configuration information of the language task processing model in the prefill stage and the decoding stage respectively, and determine the token batch length according to the prefill resource configuration information and the decoding resource configuration information; from each task processing request in the current request batch, obtain the target request segments with the number matching the prefill resource configuration information, and perform parallel prefill processing on each target request segment to generate the current token batch; in the current request batch, generate multiple new token batches by obtaining the next token of each token in the latest generated token batch at the current moment and forming a new token batch with each next token, so as to meet the condition that the combined length of the current token batch and each new token batch reaches the token batch length; the lengths of each target request segment are the same and are part of the corresponding task processing request; construct multiple pipelines using the decoding resource configuration information, perform parallel decoding processing on each token batch through the multiple pipelines, and obtain the corresponding language task processing result according to the decoding results of all request segments of each task processing request.

[0007] The present invention also provides an electronic device, including a memory and a processor. When the processor executes the computer program stored in the memory, the steps of any of the above language task processing methods are implemented.

[0008] The present invention also provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the steps of any of the above language task processing methods are implemented.

[0009] The present invention also provides a computer program product, including computer program / instructions. When the computer program / instructions are executed by a processor, the steps of any of the above language task processing methods are implemented.

[0010] Finally, the present invention also provides a language task processing system, which at least includes a first processor and a second processor, and the first processor is connected to the second processor; the resource configuration information of the first processor and the second processor is determined according to the resource configuration information of the language task processing model in the pre-filling stage and the decoding stage; wherein, the first processor obtains, from each task processing request in the current request batch, a target request segment with a number matching the pre-filling resource configuration information, and sends the target request segment as the current request batch to the second processor, and the second processor performs pre-filling processing on each target request segment in parallel to generate a current token batch, and sends it to the first processor; the lengths of each target request segment are the same, and are partial contents of the corresponding task processing request; the first processor obtains the next token of each token in the current token batch from the current request batch to generate a new current token batch, and obtains multiple current token batches until the combined length of the current token batches reaches the token batch length, and combines each current token batch into a token batch and sends it to the second processor; the second processor constructs multiple pipelines by using the decoding resource configuration information, and performs decoding processing on each token batch in parallel through the multiple pipelines, and sends the decoding processing result to the first processor; the first processor generates corresponding language task processing results according to the decoding results of all request segments of each task processing request.

[0011] The advantages of the technical solution provided by the present invention are as follows: deploying the decoding stage and the pre-filling stage in the process of executing a language task of a language task processing model on different devices not only helps improve resource utilization rate, but also can reduce resource coupling and improve the stability of pre-filling latency. In the pre-filling stage, multiple task processing requests are merged into a request batch, and multiple target request segments in the request batch are pre-filled in parallel. This can not only increase the total number of tokens output per second in the pre-filling stage, improve the throughput of the pre-filling stage, and improve the resource utilization rate of the pre-filling stage, but also effectively increase the data processing volume in the decoding stage by increasing the throughput of the pre-filling stage, thereby helping to reduce the problem of under-utilized resources in the decoding stage, and can also effectively reduce the pipeline bubbles generated in the pre-filling stage. Further, in the decoding stage, the token batch generated in the pre-filling stage and multiple known token batches are merged, and these multiple token batches are executed concurrently in a pipeline, increasing the number of batches and the batch size that can be processed in the decoding stage, so as to effectively improve the resource utilization rate of the decoding stage and improve the language task processing performance without significantly increasing the memory. In addition, the present invention also provides a corresponding implementation system, electronic device, computer-readable storage medium and computer program product for the language task processing method, further making the method more practical, and the system, electronic device, computer-readable storage medium and computer program product have corresponding advantages. BRIEF DESCRIPTION OF THE DRAWINGS

[0012] In order to more clearly illustrate the technical solutions of the present invention or related technologies, the following will briefly introduce the drawings required for use in the description of the embodiments or related technologies. Obviously, the drawings in the following description are only some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.

[0013] Figure 1 Schematic diagram of the hardware composition framework applicable to the language task processing method provided by the present invention;

[0014] Figure 2 Schematic diagram of the process flow of a language task processing method provided by the present invention;

[0015] Figure 3 Schematic diagram of the inference process flow of the language task processing model provided by the present invention in an exemplary example;

[0016] Figure 4 Schematic diagram of the data processing flow of the language task processing model provided by the present invention;

[0017] Figure 5 Schematic diagram of the token batch merging process flow in an exemplary example provided by the present invention;

[0018] Figure 6 A schematic diagram of the process of pipelined parallel execution of a token batch in an exemplary example provided by the present invention;

[0019] Figure 7 A schematic diagram of the process of pipelined parallel execution of a token batch in another exemplary example provided by the present invention;

[0020] Figure 8 A schematic diagram of the process of pipelined parallel execution in an exemplary example provided by the present invention;

[0021] Figure 9 A structural framework diagram of a language task processing device provided by the present invention under an exemplary embodiment;

[0022] Figure 10 A structural framework diagram of a language task processing system provided by the present invention under an exemplary embodiment;

[0023] Figure 11 A schematic diagram of the architecture of a language task processing system provided by the present invention in an exemplary example. Detailed implementation manners

[0024] In order to enable those skilled in the art to better understand the technical solutions of the present invention, the present invention will be further described in detail below with reference to the accompanying drawings and specific implementation manners. Among them, the terms "first", "second", "third", "fourth", etc. in the specification and the above-mentioned drawings are used to distinguish different objects, rather than to describe a specific order. In addition, the terms "include" and "have" and any variations thereof are intended to cover non-exclusive inclusion. The term "exemplary" means "serving as an example, embodiment, or illustration". Any embodiment described herein as "exemplary" need not be construed as superior or better than other embodiments.

[0025] With the development of artificial intelligence technology, language models are widely applied to the processing of various artificial intelligence tasks, such as question-and-answer tasks, translation tasks, and code generation tasks. With the increase in the parameter scale and data of language models, the resources required for language models to execute related language tasks are also increasing.

[0026] When a language model executes a language task, its inference generation process includes a prefill stage and a decoding stage with different computational requirements. In order to reduce the repeated key-value calculations of tokens, the language model performs key-value caching during the inference process. The decoding process needs to save the key-value caches of all previous tokens until the decoding ends. With the increase in the length of the generated context sequence, this leads to an unlimited increase in the memory occupancy of the key-value cache, further exacerbating the resource requirements of the language model. In order to meet the memory resources and computational resources requirements of the language model, a distributed parallel computing system is used to run the language model.

[0027] Taking into account the data transmission overhead, computing resource utilization, memory capacity, cost and energy consumption in distributed parallel computing systems, the language model performs language tasks through parallel computing. For example, the calculation process in the reasoning stage can combine tensor parallelism, pipeline parallelism and pre-filling-decoding stage split parallelism. Among them, the relevant technology usually splits the pre-filling stage and the decoding stage to different computing devices for processing. For example, a related technology places pre-filling and decoding on different systems to reduce the pressure of single-device processing reasoning, determines the pre-filling distributed system according to the characteristics of the pre-filling calculation, and determines the decoding system according to the characteristics of the decoding calculation, so that each system meets the needs of the corresponding reasoning process, balances the efficiency differences of different computing stages, reduces resource waste, and improves the efficiency of the entire reasoning process. Another related technology deploys the hint calculation and word unit calculation stages to different models of GPUs (Graphics Processing Unit), optimizes the hardware resource management of each stage separately, makes the hardware used in each stage most suitable, and optimizes the communication overhead of key-value cache between GPUs; another related technology assigns pre-filling and decoding calculations to different GPUs to eliminate pre-decoding interference, jointly optimizes resource allocation and parallel strategies for each stage, and also places the two stages according to the bandwidth of the service cluster to minimize the communication caused by decomposition.

[0028] Although the related technology can reduce the inference delay caused by the resource coupling between the pre-filling stage and the decoding stage calculation on a single computing device by splitting the language task in parallel during the pre-filling-decoding stage, the related technology analyzes the computing performance difference between the pre-filling stage and the decoding stage, and matches the GPU with corresponding performance according to the performance difference of the two stages, so as to maximize the utilization of the GPU resources of each stage, thereby solving the problem of unsaturated computing resource utilization in the decoding stage. However, based on the premise that the computing performance and memory requirements of the pre-filling stage and the decoding stage themselves are already unbalanced, as the context length of the user input continues to increase, the imbalance in the requirements of the two stages will increase accordingly, which will lead to the problem of unsaturated computing resource utilization in the decoding stage, reducing the throughput of the language model in the process of executing language tasks and increasing the latency of task processing.

[0029] In view of this, in order to improve the problem of low GPU resource utilization rate in the decoding stage in related technologies, different batch request processing methods are adopted for the pre-filling stage and the decoding stage split on different processors. For the pre-filling stage, length truncation and static batch processing technologies are adopted to improve the resource utilization rate of the pre-filling stage. For the decoding stage, batch processing based on token batch merging and pipeline concurrent execution is adopted to increase the batch size and batch dimension processed in the decoding stage, so as to improve the resource utilization rate of the decoding stage, such as the GPU resource utilization rate, without significantly increasing the memory. Combining with the specific application environment architecture or specific hardware architecture on which the execution of the language task processing method depends, the specific application environment architecture or specific hardware architecture is described herein. The following combines Figure 1 Some possible application scenarios related to the technical solution of the present invention are introduced by way of example, which may include the following content:

[0030] Such as Figure 1 As shown, the hardware composition framework may include a server 11 and a GPU group 12. The total number of GPUs included in the GPU group 12 and the number of CPUs (Central Processing Unit) included in the server 11 are determined according to the resource requirement information in the process of executing the language task by the language task processing model. Each CPU of the server 11 is connected to each other and is connected to each GPU of the GPU group 12 through PCIe (peripheral component interconnect express, high-speed serial computer expansion bus). When the total amount of twice the language task processing model memory, intermediate activation memory, and key-value cache memory corresponding to the maximum generation length in the language task generation scenario is greater than or equal to the total amount of GPU memory on a single CPU, then each GPU of the GPU group 12 is deployed to different CPUs, that is, the pre-filling stage and the decoding stage are split to the GPUs on different CPUs. When the total amount of twice the language task processing model memory, intermediate activation memory, and key-value cache memory corresponding to the maximum generation length in the language task generation scenario is less than the total amount of GPU memory on a single CPU, then each GPU of the GPU group 12 is deployed to the same CPU, that is, the pre-filling stage and the decoding stage are split to the GPUs on the same CPU. A part of the GPUs in the GPU group 12 is used to process the pre-filling stage of the language task processing model, which is defined as the pre-filling GPU group. Such as Figure 1 Taking 4 GPUs as an example, each GPU can be defined as G0-P, G1-P, G2-P, G3-P respectively, where P represents the pre-filling stage and G represents the GPU. A part of the GPUs is used to process the decoding stage of the language task processing model, which is defined as the decoding GPU group. Such as Figure 1Taking 4 GPUs as an example, each GPU can be defined as G4-T, G5-T, G6-T, and G7-T respectively, where T represents the decoding stage. The server 11 is also used to provide a human-computer interaction interface, which can be the interface of the corresponding application software or the interface opened in the browser through a specified URL (Uniform Resource Locator).

[0031] In this application scenario, the user inputs a task processing request through the human-computer interaction interface. The human-computer interaction interface sends the task processing request to the CPU through the network. The CPU combines the task processing requests of multiple users or multiple task processing requests of the same user into one batch. For the sake of convenience of description, it is defined as a request batch. The CPU uniformly truncates all the task processing requests in the request batch to the minimum length to form a prompt batch. The prompt batch contains multiple prompt contents in the form of text characters. Convert all the text characters of the prompt batch into a digital sequence to obtain a prompt batch sequence. Input the prompt batch sequence to the pre-fill GPU group for parallel pre-fill calculation. The parallel method in the pre-fill stage is determined according to the memory requirements of the language task processing model and the number of GPUs in the GPU group 12. After the pre-fill stage is completed, each prompt content in the prompt batch predicts and outputs its next token and the key-value cache of all tokens. Each token constitutes the current token batch. The pre-fill GPU group sends the current token batch to the CPU and sends the key-value cache to the decoding GPU group. The CPU obtains the next token of each token in the current token batch from the request batch, generates a new current token batch, and repeats continuously until the combined length can reach the token batch length, and combines each current token batch into one batch and sends it to the decoding GPU group. The decoding GPU group performs decoding calculations on the key-value cache generated in the previous pre-fill stage and each current token batch in a pipelined concurrent manner. Each token in the batch outputs the next token and the key-value cache after decoding calculation. Continuously repeat the token batch scheduling and decoding calculation of the above steps until all tokens in the request batch are decoded. Finally, send all the tokens generated by the decoding GPU group to the CPU. The CPU converts them into text characters, that is, obtains the language task processing result corresponding to the task processing request, and displays the language task processing result to the user through the human-computer interaction interface.

[0032] It should be noted that the above application scenario is only shown for the convenience of understanding the idea and principle of the present invention, and the embodiments of the present invention are not limited in this regard. On the contrary, the embodiments of the present invention can be applied to any applicable scenario. After introducing the technical solution of the present invention, the various non-limiting embodiments of the present invention will be described in detail below with reference to the drawings and specific embodiments. First, please refer to Figure 2 , the language task processing method provided in this embodiment may include the following content:

[0033] S201: Determine the resource configuration information of the language task processing model in the pre-filling stage and the decoding stage respectively according to the resource requirement information during the execution of the language task by the language task processing model, and determine the token batch length according to the pre-filling resource configuration information and the decoding resource configuration information.

[0034] Among them, the language task is a task that needs to be executed using the language task processing model, including but not limited to question-answering tasks, code generation tasks, text generation tasks. The language task can be input in any way, such as text, table, graph, video, voice, and it at least includes a prompt. The prompt is a prompt that tells the language task processing model how to execute the language task and can be flexibly set according to the actual application scenario. The language task processing model is a pre-trained language model that can process language tasks. When allowing the language task to be input in multiple ways, corresponding encoding modules can be set in the input layer. If voice input is supported, a network model for converting speech to text or converting speech to images also needs to be added to the input layer of the task processing model. If visual or text information input is supported, an image encoder and a text encoder need to be added to the input layer of the language task processing model. The neural network algorithm structure of the language task processing model can be CNN (Convolutional Neural Network), RNN (Recurrent Neural Network), LSTM (Long Short-Term Memory), etc., or it can be a model constructed with an attention network, such as transformer (transformer network model), bert (Bidirectional Encoder Representations from Transformers), GPT (Generative Pre-trained Transformer), Clip (Contrastive Language–Image Pre-training), etc. The present invention does not limit this here. Among them, the attention network refers to a network model trained using the attention mechanism. This model assigns different weights to each part of the input sequence, thereby extracting more important feature information from the input sequence, so that the language task processing model finally obtains a more accurate task processing result.

[0035] In the present invention, in order to reduce resource coupling and improve the stability of pre-filling latency, the pre-filling stage and the decoding stage of a language task such as a text inference task are split and processed on different computing devices. Computing devices such as GPUs, FPGAs, etc. can balance the differences between the computing requirements and memory requirements in the pre-filling stage and the decoding stage, and the parallel granularity is larger than that of model parallelism and tensor parallelism. As Figure 3As shown, when the user inputs a task processing request containing a prompt, the prompt is predicted to generate the next text character, called a token, after inference calculation by the language task processing model. The prompt calculation stage is called the prefill stage. Then, the prompt and the previous token are input into the language task processing model, and the next token is generated after inference calculation. This process is repeated until all tokens are generated. The token generation calculation process is called decoding, that is, the decoding stage calculates and generates the next token based on the token at the previous position and the key-value cache of all previous tokens. The former focuses on compute-intensive requirements, and the latter focuses on memory-intensive requirements. Since the generation of each token requires the input of all previous tokens, then each token generation needs to repeat the calculation of all previous tokens. To reduce the repeated key-value calculation of tokens, a key-value cache optimization method is adopted, that is, the key-value cache data of all previous tokens is retained during each token generation process, and there is no need to repeat the calculation. The input for token generation is the key-value cache of all previous tokens and the previous token. Correspondingly, the prefill stage generates the first predicted token based on all tokens in the prompt context input by the user and retains the key-value (KV) cache of all tokens. The decoding process needs to save the key-value cache of all previous tokens until the decoding ends. As the length of the generated context sequence increases, the memory occupancy of the key-value cache also increases without limit. Coupled with the increase in model capacity, a distributed parallel computing system is used to run the resource requirements of the language task processing model during the execution of language tasks. The distributed parallel computing system is composed of different types of processors, such as CPU+GPU, or CPU+FPGA. The CPU is used to receive task processing requests from several users in the application layer, perform corresponding text-sequence conversion on the task processing requests and the tokens generated in the decoding stage, and perform scheduling and forwarding of batch processing on the task processing requests and tokens. The GPU or FPGA is used to run the prefill calculation and decoding calculation in the inference process of the language task processing model. The resource requirement information includes memory requirement information and compute requirement information. For memory requirements, in the present invention, the prefill stage and the decoding stage of the language task processing model are split onto different computing devices. The calculation processes of these two stages each require a copy of the model parameters of the language task processing model. If deployed on a single CPU node, it is required that the total GPU memory on the single CPU node is at least greater than twice the model memory. In addition, the intermediate activation memory and key-value cache memory occupancy during model inference need to be considered. The intermediate activation memory depends on the memory occupancy of the network layer of the language task processing model, such as depending on the GPU memory occupancy of a single Transformer module layer. The memory of the key-value cache is the memory occupancy of the key matrix and value matrix corresponding to all tokens in the prefill and decoding.For computing requirements, the input for prefill computing is a prompt batch, which is part or all of the request batch merged from multiple task processing requests. It is highly computationally intensive. The input batch size and the number of batches can usually ensure full utilization of prefill resources. However, for decoding computing, the input for a single decoding request is a single token text. To improve the resource utilization rate during the decoding phase, it is necessary to increase the batch size and the number of batches as much as possible. Since the decoding phase needs to store the key-value cache of all tokens, the memory requirement of the computing device used in the decoding phase is higher than that of the computing device used in the prefill phase. Therefore, the batch size and the number of batches processed by the computing device used in the decoding phase are limited. To increase the batch size and the number of batches in the decoding phase, it is necessary to improve the throughput of the prefill phase as much as possible, that is, to increase the number of tokens output per second by the prefill computing. Considering the high computational intensity of the prefill phase, to improve its throughput, it is necessary to increase the number of computing devices used in the prefill phase. Based on the above memory resource requirement information and computing resource requirement information, when the resource utilization rate is the highest, the resource configuration information of the language task processing model in the prefill phase and the decoding phase is determined. For the convenience of description, the resource configuration information in the prefill phase and the decoding phase is defined as prefill resource configuration information and decoding resource configuration information respectively. The prefill resource configuration information and the decoding resource configuration information at least include the number of computing devices used respectively and whether each computing device is deployed across nodes or on the same node. The token batch length is the batch size processed at one time in the decoding phase. Based on the limitation of the key-value memory, the batch length when the resource utilization rate in the decoding phase is the highest can be determined according to the prefill resource configuration information and the decoding resource configuration information.

[0036] S202: From each task processing request in the current request batch, obtain the target request segments with the number matching the prefill resource configuration information, and perform parallel prefill processing on each target request segment to generate the current token batch; in the current request batch, by obtaining the next token of each token in the most recently generated token batch at the current moment and forming a new token batch with each next token, multiple new token batches are generated to meet the condition that the combined length of the current token batch and each new token batch reaches the token batch length.

[0037] In the present invention, when a task processing request input by a user is received, the task processing requests of multiple users or multiple task processing requests of the same user are combined into one batch, which is defined as a request batch for the sake of convenient description. The currently processing request batch is defined as the current request batch. Before prefill calculation is performed on the current request batch, each task processing request of the current request batch passes through the encoding function of the tokenizer of the language task processing model to convert each text character into a corresponding number, obtaining a sequence of numbers. These sequences of numbers are combined into one batch, which can be defined as a prompt batch. Prefill calculation is performed using the language task processing model to predict the next text character of each number in the prompt batch. The predicted text character is a token, which is defined as a prompt token for the sake of distinction. To improve the resource utilization rate in the prefill stage, the present invention is implemented in a model parallel manner, that is, the language task processing model is sliced into several sub-models according to the number of model layers, and each sub-model is deployed to each computing device used in the prefill stage. The memory capacity of each computing device satisfies the memory occupancy of each sub-model. This method can reduce the video memory occupancy of the model weights on each computing device, but may cause pipeline bubble problems. To reduce the pipeline bubbles generated in the prefill stage, the present invention obtains content of the same length from all task processing requests in the current request batch each time to form a prompt batch. The selected content is defined as the target request segment, that is, the lengths of the target request segments are the same and are part of the corresponding task processing requests. Then, the prompt batch is scheduled and processed in a static batch processing manner. The so-called static batch processing requires that all requests in the batch end after all calculations are completed. After prefill calculation is performed on the prompt batch, corresponding tokens are generated for each prompt prediction of the prompt batch, and the tokens of multiple prompts form a token batch. For the sake of convenient description, it is defined as the current token batch. To improve the resource utilization rate in the decoding stage, the batch size and / or the number of batches processed at one time in the decoding stage are increased. After determining the token batch length in S201, if the length of the current token batch does not reach the token batch length, multiple known tokens are selected from the request batch for combination to reach the token batch length. The selection method of the known tokens is as follows: whether the next token of each prompt token in the current token batch is in the current request batch. If it is, it is selected to form a new token batch at the current moment. If the lengths of these two token batches still do not reach the token batch length, the new token batch is used as the newly generated token batch at the current moment, and the above steps are repeated until multiple token batches that meet the token batch length are obtained.

[0038] S203: Construct multiple pipelines using the decoding resource configuration information, perform decoding processing on each current token batch in parallel through the multiple pipelines, and obtain the corresponding language task processing result according to the decoding results of all request segments of each task processing request.

[0039] Batch-combine the multiple tokens obtained in the previous step into one batch. The current token batch refers to the current token batch whose combined token batch length can reach the token batch length determined by S201 and the newly generated token batch through the method of S202. Input these token batches into the language task processing model for decoding calculation. During the decoding calculation process, determine the number of pipelines according to the decoding resource configuration information. Considering the synchronization overhead required for each part of the model executed by different token batches, the number of pipelines cannot be too large, and its maximum value can be the same as the token batch length. In the decoding stage of the present invention, the model parallelism method is adopted, that is, the language task processing model is split into multiple sub-models, and each sub-model is deployed to each computing device used in the decoding stage. In order to achieve pipeline concurrency in the decoding stage, each sub-model is further split based on the number of model layers again, which can be defined as a secondary sub-model. The so-called pipeline parallelism means that the execution of each secondary sub-model on the same pipeline is serial, and the execution of each secondary sub-model on different pipelines is asynchronous but synchronization is required between the same secondary sub-models. After decoding and calculating each current token batch and the key-value cache, the next token will be predicted and generated for each current token batch, and this process is continuously repeated until all tokens are generated. Convert these tokens into corresponding text characters, and these text characters are the task processing results corresponding to the corresponding task processing requests.

[0040] In the technical solution provided in this embodiment, deploying the language task processing model on different devices in the decoding stage and the pre-filling stage during the execution of the language task is not only beneficial to improving resource utilization, but also can reduce resource coupling and improve the stability of the pre-filling delay. In the pre-filling stage, multiple task processing requests are combined into one request batch, and multiple target request segments in the request batch are pre-filled in parallel, which can not only increase the total number of tokens output per second in the pre-filling stage, improve the throughput of the pre-filling stage, improve the resource utilization rate of the pre-filling stage, effectively increase the data processing volume in the decoding stage by increasing the throughput of the pre-filling stage, thereby facilitating the reduction of the problem of unsaturated resource utilization in the decoding stage, but also can effectively reduce the pipeline bubbles generated in the pre-filling stage. Further, in the decoding stage, the token batch generated in the pre-filling stage and multiple known token batches are combined, and these multiple token batches are executed in pipeline concurrency, increasing the number of batches and batch sizes that can be processed in the decoding stage, so as to effectively improve the resource utilization rate of the decoding stage and improve the language task processing performance without significantly increasing the memory.

[0041] After the above embodiments complete the scheduling, token batch combination and scheduling of the request batch through S201 - S203, there is no limitation on the inference calculation of the language task processing model in the pre-filling stage and the decoding stage. The present invention also provides an implementation method for the inference calculation of the language task processing model, which may include the following content:

[0042] The language task processing model at least includes an encoder, an embedding layer, multiple feature extraction layers, an output layer, and a decoder; the encoder is used to convert the prompt batch into numbers through the encoding function of the tokenizer and output a sequence of numbers; the embedding layer is used to receive the sequence of numbers after the conversion of the prompt batch and convert the sequence of numbers into floating-point vectors. The feature extraction layer includes multiple multi-head attention layers and a two-layer fully-connected layer structure, and the outputs of both are connected with a residual structure to capture the relationships between words at different positions in the floating-point vectors for the modeling and processing of the sequence of numbers. The output layer is used to perform linear transformation and normalization processing on the output of the feature extraction layer, and it includes a linear layer and a softmax layer. The decoder is used to convert the normalized data sequence into text characters.

[0043] In this embodiment, as Figure 4 shown, first, the text characters in the prompt batch formed from each task processing request are converted into numbers through the encoding function of the tokenizer in the encoder, and the generated sequence of numbers is converted into floating-point vectors through the operation of the embedding layer. The floating-point vectors are input to the feature extraction layer. The feature extraction layer can adopt a transformer network model, for example. Multiple feature extraction layers perform calculations in parallel. The output of the last feature extraction layer is subjected to linear layer transformation and softmax normalization activation calculation to obtain a probability vector. Each dimension of the vector represents the probability value of predicting each vocabulary in the tokenizer. The subscript of the maximum value of the probability vector is the predicted vocabulary value, and the predicted value is input to the decoding function of the tokenizer to output the predicted next character. That is, the predicted vocabulary value generated by the feature extraction layer is converted into text characters through the decoding scheme of the tokenizer in the decoder.

[0044] Among them, for a language task processing model including N feature extraction layers, the input of the first feature extraction layer is the output of the embedding layer, and the dimension is , where represents the batch size of the input text (the number of requests), that is, the number of prompt batches, represents the number of tokens in the prompt, represents the dimension of the output vector of the embedding layer, and the output and input dimensions of each feature extraction layer are the same. As Figure 4 shown, the data processing flow of each feature extraction layer includes: first, calculate the key (K) matrix, value (V) matrix, and query (Q) matrix through three linear layer transformations, and then perform multi-head dimension splitting on the key matrix, value, and query matrix, and convert the dimension to , represents the dimension of the head. Multiply the key matrix and the value matrix, add the mask matrix, and based on the dimension Perform softmax normalization to obtain a score matrix; then, multiply the score matrix by the query matrix, and finally transform the output dimension through a linear layer transformation to Finally, perform a linear transformation on the output of the multi-head attention layer through two fully connected layers.

[0045] As can be seen from the above, the language task processing model in this embodiment can extract richer features by adopting the multi-head attention mechanism during the feature extraction process, improve the semantic understanding ability of the language task processing model, improve the processing accuracy of the language task, and is conducive to improving the language task processing efficiency through the parallel computing ability of the language task processing model.

[0046] The above embodiment deploys the language task processing model through prefill-decoding split parallelism and model parallelism. This embodiment also gives an exemplary determination method for prefill resource configuration information and decoding resource configuration information, which may include the following content:

[0047] When the sum of the double memory occupancy demand, activation memory occupancy, and the memory occupancy demand of the key-value cache corresponding to the maximum generation length of the language task processing model is greater than or equal to the total memory resources of a single host node, different host nodes are used to process the prefill stage and the decoding stage; when the sum of the double memory occupancy demand, activation memory occupancy, and the memory occupancy demand of the key-value cache corresponding to the maximum generation length of the language task processing model is less than the total memory resources of a single host node, the language task processing model runs on a single host node.

[0048] In this embodiment, whether the total occupancy of the inference memory of the language task processing model, and the computing devices used in the prefill stage and the decoding stage are deployed within a single node or different nodes, the total memory of the computing devices used in the decoding stage and the occupancy of the key-value cache will affect the maximum number of token batches supported by the decoding process, and the number of computing devices used in the prefill stage affects the number of token batches output by the prefill calculation, which in turn affects the resource utilization of the decoding stage. When the types of computing devices used in the prefill stage and the decoding stage are clear, the total memory of a single computing device can be known. Then, the total occupancy of twice the memory occupancy of the language task processing model, the memory occupancy of intermediate activation calculations, and the occupancy of the key-value cache corresponding to the maximum generation length involved in various language task processing scenarios can be calculated, and the total memory of a single node can be calculated. After obtaining these information, compare the sum of the double memory occupancy demand, activation memory occupancy, and the memory occupancy demand of the key-value cache corresponding to the maximum generation length of the language task processing model with the total memory resources of a single host node, and accordingly determine whether it is a single-node deployment or a cross-node deployment.

[0049] After determining the deployment architecture of the language task processing model, it is also necessary to determine the computing devices used in the pre-filling stage and the decoding stage respectively: obtain the pre-filling throughput in the pre-filling stage under different pre-filling resource configuration information, and the decoding throughput in the decoding stage under different decoding resource configuration information; when the pre-filling throughput and the decoding throughput meet the preset same or similar conditions, the corresponding target pre-filling resource configuration information and target decoding resource configuration information are used as the optimal resource configuration information of the language task processing model in the pre-filling stage and the decoding stage; when the pre-filling throughput and the decoding throughput cannot meet the preset same or similar conditions, and the pre-filling throughput is less than the decoding throughput, the total amount of pre-filling resources configured in the pre-filling resource configuration information should be greater than the total amount of decoding resources configured in the decoding resource configuration information, but it is necessary to meet the condition that the sum of the total amount of pre-filling resources and the total amount of decoding resources is less than or equal to the maximum value of the total resources.

[0050] In this embodiment, in order to determine the optimal resource configuration for the pre-filling stage and the decoding stage, different numbers of computing devices can be configured for the pre-filling stage in advance according to the total number of computing devices that can be deployed in parallel in the language task processing model, and the pre-filling maximum throughput corresponding to the different computing devices deployed in the pre-filling stage is analyzed, that is, the maximum number of word units output per second. Similarly, different numbers of computing devices can be configured for the decoding stage according to the total number of computing devices that can be deployed in parallel in the language task processing model, and the decoding maximum throughput corresponding to the different computing devices deployed in the decoding stage is analyzed, that is, the maximum number of word units output per second. Taking the computing device as a GPU, and the resource configuration information of the pre-filling stage and the decoding stage as the number of GPUs as an example, according to the number of GPUs that can be deployed in parallel in the model, different numbers of GPUs are configured for the pre-filling stage, and the pre-filling maximum throughput corresponding to the different numbers of GPUs deployed in the pre-filling stage is analyzed; according to the number of GPUs that can be deployed in parallel in the model, different numbers of GPUs are configured for the decoding stage, and the decoding maximum throughput corresponding to the different numbers of GPUs deployed in the decoding stage is analyzed. Through analysis, it can be determined that when the throughput of the pre-filling stage, that is, the pre-filling throughput, is close to the throughput of the decoding stage, that is, the decoding throughput, it is optimal. At this time, the resource utilization rate of the decoding stage is the highest, and this configuration is the optimal configuration, that is, the target pre-filling resource configuration information and the target decoding resource configuration information are obtained. For example, when the pre-filling throughput is close to the decoding throughput, the number of configured pre-filling GPUs and decoding GPUs is optimal, and the optimal configuration can improve the resource utilization rate of the GPU in the decoding stage. Among them, the preset same similarity condition of this embodiment is used to indicate that the pre-filling throughput is close to the decoding throughput. The condition can be flexibly set according to the actual situation. If the difference between the two is less than or equal to 10, it is considered to be close. Due to the limitations of the total number of computing devices and the performance configuration of the computing devices in the actual deployment, the number of computing devices deployed in the two stages is difficult to meet the optimal conditions under which the pre-filling throughput is close to the decoding throughput. Further, it can also be determined in combination with the total number of computing devices and the throughput of the two stages: when the throughput of the pre-filling stage is less than the throughput of the decoding stage, it will cause unsaturated resource utilization in the decoding stage, and the number of computing devices such as GPU deployment used in the pre-filling stage should be increased. Due to the memory limitation of the key-value cache in the decoding stage, the throughput of the pre-filling stage is greater than the throughput of the decoding stage, which will also lead to unsaturated resource utilization in the decoding stage. The word batches in the decoding stage can be scheduled through S203, and the word batches can be appropriately increased to increase the batch size processed in the decoding stage.

[0051] In the process of determining the resource allocation for the pre-filling stage and the decoding stage, not only the maximum and minimum numbers of deployable computing devices such as GPUs in the actual deployment need to be considered, but also the minimum number of computing devices used in each of the pre-filling stage and the decoding stage needs to be determined to ensure the smooth execution of the task. In this embodiment, the computing devices used in the pre-filling stage are defined as pre-filling processors, and the computing devices used in the decoding stage are defined as decoding processors. According to the memory occupancy requirement of the language task processing model and the memory occupancy requirement of the key-value cache corresponding to the maximum generation length, the minimum memory occupancy is determined; according to the minimum memory occupancy and the total memory of the pre-filling processors, the minimum number of pre-filling processors used by the language task processing model in the pre-filling stage is determined; according to the minimum memory occupancy and the total memory of the decoding processors, the minimum number of decoding processors used by the language task processing model in the decoding stage is determined. In this embodiment, according to the total memory of a single computing device and the memory occupancy of a single model, the minimum number of model parallel splits is determined, that is, the minimum number of computing devices used. The minimum number of processors required for each of the pre-filling and decoding stages is: the minimum number of model parallel splits should be at least greater than the total memory occupancy of a single model and the memory occupancy of the key-value cache corresponding to the maximum generation length divided by the total memory of a single computing device such as a GPU.

[0052] As can be seen from the above, before deploying the inference calculation of the language task processing model in the pre-filling-decoding split parallel and model parallel modes, by determining the number of splits for model parallel deployment, the number of computing devices deployed in the pre-filling stage, the number of computing devices deployed in the decoding stage, and the selection of single-node or cross-node deployment modes for each of the pre-filling and decoding stages as resource configuration information, the resource utilization rate in the inference calculation process of the language task processing model can be effectively improved, and the problems of GPU resource utilization and inference latency can be effectively improved.

[0053] The above embodiments do not make any limitations on how to form a prompt batch. The present invention also provides a schematic implementation manner, which may include: obtaining the number of tokens of each task processing request, and selecting the smallest number of tokens therefrom as the truncation length; respectively obtaining the first n target tokens from each task processing request as the target request segment of each task processing request; the lengths of the target tokens of the same task processing request are all the same as the truncation length. The task processing request is a request containing a prompt, the prompt is composed of multiple tokens, and multiple task processing requests form a batch, that is, a request batch. Based on the truncation length, the token lengths of each task processing request in the request batch are truncated, and the truncation length is the minimum value of the number of tokens of each task processing request in the request batch. Taking Figure 5 as an example, the request batch includes 4 task processing requests, and each task processing request can be respectively expressed as: , , , , where a, b, c, and d respectively represent the tokens of each task processing request. The number of tokens of the first task processing request is the smallest, and the truncation length can be set to 4. The first 4 tokens are obtained from each task processing request to form a prompt batch. That is, the prompt batch can be expressed as: .

[0054] Based on the above embodiments, the present invention also provides a more convenient batch request scheduling method. The batch request scheduling includes constructing a prompt batch in the request batch and scheduling it to the computing device corresponding to the pre-filling stage for pre-filling calculation, and also includes merging the token batches and scheduling them to the computing device corresponding to the decoding stage for decoding calculation. It may include the following contents:

[0055] A batch processing queue is constructed in advance. The batch processing queue includes at least an undecoded batch queue and a decoded batch queue. In an exemplary implementation, the batch processing queue may further include a decoded batch queue. When the current request batch is received, each task processing request of the current request batch is moved into the undecoded batch queue. The first batch of target request segments with the number matching the pre-filling resource configuration information is obtained from the undecoded batch queue and moved into the decoded batch queue to perform pre-filling processing on each first batch of target request segments in the decoded batch queue in parallel. When the first batch of token batches corresponding to the first batch of target request segments is generated, if the first batch of token batches includes the first type of target tokens with end identification information, the target task processing requests corresponding to the first type of target tokens are completed. For the second type of target tokens without end identification information, when the second type of target tokens do not exist in the undecoded batch queue and the decoded batch queue, they are added to the corresponding request batch in the undecoded batch queue. If the task processing requests of the current request batch exist in the undecoded batch queue, the second batch of target request segments with the number matching the pre-filling resource configuration information is obtained from the undecoded batch queue again and moved into the decoded batch queue. If the task processing requests of the current request batch do not exist in the undecoded batch queue, the token batches of the current request batch are merged based on the token batch length. When the batch processing queue further includes a decoded batch queue, correspondingly, when the first batch of token batches corresponding to the first batch of target request segments is generated, the first batch of target request segments is moved into the decoded batch queue. If the first batch of token batches includes the first type of target tokens with end identification information, the request segments of the task processing requests corresponding to the first type of target tokens are deleted from the decoded batch queue.

[0056] In this embodiment, based on the undecoded batch queue, the decoding batch queue, and the decoded batch queue, the task processing requests for the client, the token batch requests output in the pre-filling stage and the decoding stage are managed and scheduled to the corresponding computing devices for calculation. Among them, the undecoded batch queue is used to store the batch data information that has not been scheduled, the decoding queue is used to store the batch data information that has been scheduled but the inference calculation has not been completed, and the decoded batch queue is used to store the batch data information that has completed the inference calculation. In order to manage various batch data through the batch queue, the data included in the request batch, the token batch, and the prompt batch are defined as request information, and each request information can use a queue with a triple data structure Describe, represent each request information as , Represents each token in the request information, Is the subscript of each token in the request information, that is, the token position. Use the queue Describe each batch information of the request batch, the token batch, and the prompt batch, and represent the batch as , Is each request information in the batch, Represents the subscript of each request information in the batch. Use the queue To describe all batch information queues, the batch information queue can be represented as , Represents the subscript of each batch request. The undecoded batch queue, the decoding batch queue, and the decoded batch queue can be represented as , , .

[0057] When a task processing request is received, move the task processing request And the existing request batch Into the undecoded batch queue ; Calculate the minimum value of the number of tokens of each request in the request batch , and move the first Tokens of each task processing request in the request batch from the undecoded batch queue Into the decoding batch queue . All requests in the decoding batch queue Wait to be scheduled for execution, Each batch request of is scheduled and executed in a static batch processing manner, that is, it is required to end after all requests in the batch are completed, which is applicable to batch requests with the same length of all requests in the batch. When the first prompt batch is input, send the prompt batch information in the decoding batch queue To the computing device corresponding to the pre-filling stage for pre-filling calculation, such as Figure 5 ​​As shown, after the pre-filling calculation is completed, the current token batch is obtained . After the pre-filling calculation of the prompt batch or the decoding calculation of the token batch is completed, token batch information will be generated. The tokens in the input batch request corresponding to the token batch are moved from the decoding middle batch queue to the decoded batch queue ; then, it is judged whether there is an EOS end symbol in the tokens of the token batch. If so, the input batch request corresponding to the token batch is removed from the decoded batch queue , and a feedback message indicating the completion of the task is sent to the client. If the token is not an EOS end symbol, the token is searched in the undecoded batch queue and the decoding middle batch queue . If it is not found in both, it is added to the request of the corresponding batch in the undecoded batch queue . Each batch in the undecoded batch queue is traversed to judge whether the current batch exists in the decoding middle batch queue . If not, the tokens in the undecoded batch queue are scheduled. If it exists, the next round of request batch scheduling starts in the above manner

[0058] The above embodiments do not make any limitations on the merging of the current token batch. The present invention also provides an exemplary implementation method: if the next token of the target prompt token of the current token batch is a known token in the current request batch, a first token batch is generated according to the known token corresponding to the target prompt token; if the merged length of the first token batch and the current token batch is less than the token batch length, the known token corresponding to the first token batch in the current request batch is obtained, and a second token batch is generated; if the merged length of the first token batch, the current token batch and the second token batch is the token batch length, the first token batch, the current token batch and the second token batch are decoded and processed in parallel using multiple pipelines. Taking Figure 5 as an example, the current token batch is , the token batch length is 3, that is, three current token batches need to be merged. The next token of the first target prompt token of the current token batch exists as a known token in the request batch, the next token of the second target prompt token exists as a known token in the request batch, and the next token of the third target prompt token exists as a known token in the request batch, and the first token batch is generated. The merged length of the first token batch and the current token batch is 2, so the first token batch is used as the current token batch to repeat the above process, and the second token batch , the combined length of the first word batch, the current word batch and the second word batch is 3, then the whole process ends. If the combined length of the first word batch, the current word batch and the second word batch is not equal to the word batch length, then the second word batch is used. Repeat the above process for the current word batch.

[0059] Based on the batch queue constructed above, this embodiment also provides a scheduling implementation method for implementing word batches based on the batch queue, which may include the following contents:

[0060] If the current request batch is in the undecoded batch queue and decoding batch queue , for example, can traverse the undecoded batch queue Each batch of the current request batch Whether to batch queue during decoding If it does not exist, the current batch will be scheduled. The word unit scheduling process includes: extracting the undecoded batch queue The current batch corresponds to the first token of all requests in the request queue, forming a token batch, which can be expressed as , and the corresponding word is not decoded from the batch queue Move into the decoding batch queue Since the number of tokens in different requests is different, the next token of some requests is known. The above token batch merging method is used to continue to schedule the known next token in each request of the current batch to form a token batch. For example, the next token can be continuously scheduled in the undecoded batch queue. Extract the first token of all requests in the current batch to form a merged token batch, which can be expressed as , and the corresponding word is sent to the undecoded batch queue Move into the decoding batch queue Repeat this process to generate merged word batches until the number of merged word batches reaches the word batch length or all requested words in the current batch are empty. If the number of merged word batches in the previous step does not meet the set value, continue to traverse the undecoded batch queue Repeat the scheduling process of the previous step until the number of merged word batches meets the set value or all requested word units are empty, then exit the word unit batch request scheduling; finally, after completing the batch request scheduling, the decoding batch queue The word batches in are sent to the decoding GPU for decoding calculation. Figure 5 Examples include Figure 6 and Figure 7For the scheduling of each token batch request, the token batch scheduling of batch1 (batch 1) generates three token batches, and schedules the three token batches to three pipelines in the decoding stage, namely stream 1, stream 2, and stream 3. Each token batch is executed on one stream, and the decoding operations of the token batches on the three streams are executed concurrently through the pipelines, indirectly increasing the number of token batches. Moreover, the token batches of different prompt batches, such as Figure 6 the remaining part of batch 1 and the first token of batch 2 can also be combined into one token batch. Figure 7 Batch 1, batch 2, and batch 3 of Figure 7 are combined into one token batch, thereby improving the resource utilization in the decoding stage.

[0061] As can be seen from the above, in this embodiment, the scheduling of request batches and token batches is realized through the batch queue, which can simplify the processing complexity of batch requests, improve the scheduling efficiency of request batches and token batches, and is beneficial to improving the processing efficiency of language tasks. Further, by combining the token batches of different batches, the resource utilization rate can be further improved.

[0062] The above embodiment does not make any limitation on how to perform decoding calculations in parallel through multiple pipelines. Based on the above embodiment, the present invention also gives an exemplary implementation manner, which may include the following content:

[0063] According to the number of model layers of the language task processing model, each sub-model in the decoding stage where the language task processing model is deployed is respectively divided into multiple sub-sub-models, and the number of sub-sub-models of each sub-model is the same as the value of the token batch length; construct multiple pipelines, and each pipeline serially executes each sub-sub-model; the number of pipelines is the same as the value of the token batch length; through the first pipeline, each sub-sub-model is used in turn to decode the first current token batch; when the first pipeline completes the decoding process of the first current token batch using the first sub-sub-model, during the process of the first pipeline using the second sub-sub-model to decode the first current token batch, the second pipeline uses the target sub-sub-model corresponding to the first sub-sub-model to decode the second current token batch; when the second pipeline completes the decoding process of the second current token batch using the target sub-sub-model, during the process of the second pipeline using the next sub-sub-model to decode the second current token batch, the third pipeline uses the sub-sub-model corresponding to the target sub-sub-model to decode the third current token batch. Of course, if the constructed pipelines also include the fourth pipeline, the fifth pipeline, the sixth pipeline, etc., the fourth pipeline, the fifth pipeline, and the sixth pipeline execute the corresponding decoding process in the same way that the execution of each sub-sub-model on the same stream is serial, the execution of sub-sub-models on different streams is asynchronous but synchronization is required between the same sub-sub-models.

[0064] It is understandable that the lengths of the original request batches are different. After the current token batch generated by the pre-padding calculation, when merging tokens, the next tokens of some tokens in the token batch are known and do not need to be calculated and generated by the language task processing model. Multiple merged batches can be generated through the token batch merging scheduling method, and then concurrent pipeline execution of multiple token batches can be initiated in the decoding stage. However, the key-value cache data of the next token still needs to be obtained through the calculation of the previous token batch. Since the key-value cache of tokens is obtained through layer-by-layer calculation of the model, considering the key-value data dependency relationship between the tokens in the two consecutive merged token batches, the model is divided into several parts based on the model layer. According to the key-value cache dependency relationship between the previous and subsequent token batches, the decoding calculations of each token batch are sent sequentially. Different token batches execute different parts of the model at the same time, enabling each token batch to execute concurrently in a pipeline. To ensure the correct dependency of the previous and subsequent token batches in each part of the calculation model, synchronization is required when the previous and subsequent token batches calculate each part of the model. The number of model partitions and the number of merged token batches, that is, the length of the token batch, are the same. Since the number of model partitions is the same as the number of token batches, and the number of token batches is related to the token length in the batch request and the model partition, considering the synchronization overhead required for each token batch to execute each part of the model, the length of the token batch, that is, the number of pipeline stages, cannot be too large.

[0065] As Figure 8 shown, taking the computing device as a GPU and the token batch length as 3 as an example, after splitting the language task processing model into multiple sub-models using model parallelism, in order to achieve pipeline concurrency on the decoding GPU, the sub-models on the GPU are split into 3 sub-sub-models based on the number of model layers. For ease of description, the current token batch generated in the pre-padding stage is defined as the original token batch, and the token batch generated by merging known tokens is defined as the merged token batch. The original token batch is started on the first stream to serially execute the three sub-sub-models, the first merged token batch is started on the second stream to serially execute the three sub-sub-models, and the second merged token batch is started on the third stream to serially execute the three sub-sub-models. However, the execution of each sub-sub-model on each stream needs to be synchronized with the execution of the corresponding sub-sub-model on the previous stream. That is, the execution of the first, second, and third sub-sub-models on the second stream needs to wait for the execution of the first, second, and third sub-sub-models on the first stream to complete respectively, and the execution of the first, second, and third sub-sub-models on the third stream needs to wait for the execution of the first, second, and third sub-sub-models on the second stream to complete respectively. That is to say, the execution of each sub-sub-model on the same stream is serial, and the execution of sub-sub-models on different streams is asynchronous but synchronization is required between the same sub-sub-models.

[0066] As can be seen from the above, in this embodiment, by making the number of pipelines consistent with the token batch length and further dividing the sub-model in the decoding stage into multiple sub-sub-models with the same number as the token batch length, the correct dependencies of the front and back token batches on each part of the computing model are ensured, which is conducive to improving the resource utilization rate in the decoding stage.

[0067] Through the description of the above embodiments, those skilled in the art can clearly understand that the method according to the above embodiments can be implemented by means of software plus a necessary general hardware platform. Of course, it can also be implemented by hardware, but in many cases, the former is a better implementation method. The present invention also provides a corresponding device for the language task processing method, further making the method more practical. Among them, the device can be described from the perspective of functional modules and the perspective of hardware respectively. The language task processing device provided by the present invention will be introduced below, and the device is used to implement the language task processing method provided by the present invention. The description of the features in the corresponding embodiments of the language task processing device can refer to the relevant descriptions of the corresponding embodiments of the language task processing method, and the device embodiments will not be elaborated one by one.

[0068] First, from the perspective of functional modules, please refer to Figure 9 , Figure 9 which is the structural diagram of the language task processing device provided in this embodiment in a specific implementation manner. The device may include:

[0069] A resource configuration module 901, configured to determine the resource configuration information of the language task processing model in the pre-filling stage and the decoding stage respectively according to the resource requirement information of the language task processing model during the execution of the language task, and determine the token batch length according to the pre-filling resource configuration information and the decoding resource configuration information.

[0070] A pre-filling stage processing module 902, configured to obtain target request segments with a number matching the pre-filling resource configuration information from each task processing request in the current request batch, and perform pre-filling processing on each target request segment in parallel to generate the current token batch.

[0071] A decoding stage processing module 903, configured to generate multiple new token batches in the current request batch by obtaining the next token of each token in the latest generated token batch at the current moment and forming a new token batch with each next token, so as to meet the condition that the combined length of the current token batch and each new token batch reaches the token batch length; the lengths of each target request segment are the same and are part of the corresponding task processing request; multiple pipelines are constructed by using the decoding resource configuration information, and each token batch is decoded in parallel through the multiple pipelines, and the corresponding language task processing results are obtained according to the decoding results of all request segments of each task processing request.

[0072] Exemplarily, in some embodiments of this embodiment, the above resource configuration module 901 may further be configured to: when the sum of the double memory occupancy requirement of the language task processing model, the active memory occupancy, and the memory occupancy requirement of the key-value cache corresponding to the maximum generation length is greater than or equal to the total memory resources of a single host node, use different host nodes to process the prefill stage and the decoding stage; when the sum of the double memory occupancy requirement of the language task processing model, the active memory occupancy, and the memory occupancy requirement of the key-value cache corresponding to the maximum generation length is less than the total memory resources of a single host node, the language task processing model runs on a single host node.

[0073] Exemplarily, in some other embodiments of this embodiment, the above resource configuration module 901 may further be configured to: obtain the prefill throughput of the prefill stage under different prefill resource configuration information, and the decoding throughput of the decoding stage under different decoding resource configuration information; use the target prefill resource configuration information and the target decoding resource configuration information respectively corresponding when the prefill throughput and the decoding throughput meet the preset same / similar conditions as the optimal resource configuration information of the language task processing model in the prefill stage and the decoding stage; when the prefill throughput and the decoding throughput cannot meet the preset same / similar conditions, and the prefill throughput is less than the decoding throughput, the total prefill resources configured for the prefill resource configuration information are greater than the total decoding resources configured for the decoding resource configuration information; the sum of the total prefill resources and the total decoding resources is less than or equal to the maximum value of the total resources.

[0074] Exemplarily, in some other embodiments of this embodiment, the above resource configuration module 901 may further be configured to: determine the minimum memory occupancy according to the memory occupancy requirement of the language task processing model and the memory occupancy requirement of the key-value cache corresponding to the maximum generation length; determine the minimum number of prefill processors used by the language task processing model in the prefill stage according to the minimum memory occupancy and the total memory owned by the prefill processors; determine the minimum number of decoding processors used by the language task processing model in the decoding stage according to the minimum memory occupancy and the total memory owned by the decoding processors.

[0075] Exemplarily, in some other embodiments of this embodiment, the above pre-filling stage processing module 902 may also be used to: obtain the number of tokens of each task processing request, and select the smallest number of tokens therefrom as the truncation length; respectively obtain the first n target tokens from each task processing request as the target request segment of each task processing request; the lengths formed by the respective target tokens of the same task processing request are all the same as the truncation length. Exemplarily, in some other embodiments of this embodiment, the above pre-filling stage processing module 902 may also be used to: pre-construct a batch processing queue, the batch processing queue including an undecoded batch queue and a decoded batch queue; when receiving the current request batch, move the task processing requests of the current request batch into the undecoded batch queue, obtain the first batch of target request segments in a number matching the pre-filling resource configuration information from the undecoded batch queue, and move the first batch of target request segments into the decoded batch queue to perform pre-filling processing on the respective first batch of target request segments in the decoded batch queue in parallel; when generating the first batch of token batches corresponding to the first batch of target request segments, if the first batch of token batches includes the first type of target tokens with end identification information, the target task processing requests corresponding to the first type of target tokens are completed, and for the second type of target tokens without end identification information, when the second type of target tokens do not exist in the undecoded batch queue and the decoded batch queue, add them to the corresponding request batch in the undecoded batch queue; if there are task processing requests of the current request batch in the undecoded batch queue, obtain the second batch of target request segments in a number matching the pre-filling resource configuration information from the undecoded batch queue again, and move the second batch of target request segments into the decoded batch queue; if there are no task processing requests of the current request batch in the undecoded batch queue, merge the respective token batches of the current request batch based on the token batch length.

[0076] As an exemplary implementation manner of the above embodiment, the above pre-filling stage processing module 902 may also be used to: the batch processing queue further includes a decoded batch queue; when generating the first batch of token batches corresponding to the first batch of target request segments, move the first batch of target request segments into the decoded batch queue, and if the first batch of token batches includes the first type of target tokens with end identification information, delete the request segments of the task processing requests corresponding to the first type of target tokens from the decoded batch queue.

[0077] Exemplarily, in some other embodiments of the present embodiment, the above decoding stage processing module 903 may further be configured to: if the next token of the target prompt token in the current token batch is a known token in the current request batch, generate a first token batch according to the known token corresponding to the target prompt token; if the combined length of the first token batch and the current token batch is less than the token batch length, obtain the known tokens corresponding to the first token batch in the current request batch, and generate a second token batch; if the combined length of the first token batch, the current token batch, and the second token batch is the token batch length, perform decoding processing on the first token batch, the current token batch, and the second token batch in parallel using multiple pipelines.

[0078] Exemplarily, in some other embodiments of the present embodiment, the above decoding stage processing module 903 may further be configured to: according to the number of model layers of the language task processing model, respectively divide each sub-model in the decoding stage where the language task processing model is deployed into multiple sub-sub-models, and the number of sub-sub-models of each sub-model is the same as the value of the token batch length; construct multiple pipelines, and each pipeline serially executes each sub-sub-model; the number of pipelines is the same as the value of the token batch length; through the first pipeline, sequentially use each sub-sub-model to perform decoding processing on the first current token batch; when the first pipeline completes the decoding processing of the first current token batch using the first sub-sub-model, during the process of the first pipeline using the second sub-sub-model to decode the first current token batch, use the second pipeline to perform decoding processing on the second current token batch using the target sub-sub-model corresponding to the first sub-sub-model; when the second pipeline completes the decoding processing of the second current token batch using the target sub-sub-model, during the process of the second pipeline using the next sub-sub-model to decode the second current token batch, use the third pipeline to perform decoding processing on the third current token batch using the sub-sub-model corresponding to the target sub-sub-model.

[0079] The language task processing device mentioned above is described from the perspective of functional modules. Further, the present invention also provides an electronic device, which is described from the perspective of hardware. The electronic device includes a memory and a processor. A computer program is stored in the memory, and the processor is configured to run the computer program to execute the steps in any one of the above embodiments of the language task processing method.

[0080] The embodiments of the present application also provide a computer-readable storage medium. A computer program is stored in the computer-readable storage medium, where the computer program is configured to execute the steps in any one of the above embodiments of the language task processing method when running. In an exemplary embodiment, the above computer-readable storage medium may include, but is not limited to: USB flash drive, ROM (Read-Only Memory), RAM (Random Access Memory), mobile hard disk, magnetic disk, or optical disc, etc., various media that can store computer programs.

[0081] An embodiment of the present application also provides a computer program product. The computer program product includes a computer program, and when the computer program is executed by a processor, the steps in any of the above-described method embodiments for processing language tasks are implemented.

[0082] An embodiment of the present application also provides another computer program product, including a non-volatile computer-readable storage medium. The non-volatile computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the steps in any of the above-described method embodiments for processing language tasks are implemented.

[0083] Finally, the present invention also provides a language task processing system. Refer to Figure 10 , the language task processing system may at least include a first processor 101 and a second processor 102, and the first processor 101 is connected to the second processor 102; the resource configuration information of the first processor 101 and the second processor 102 is determined according to the resource configuration information of the language task processing model in the pre-filling stage and the decoding stage. The so-called resource configuration information is used to determine whether the second processor 102 is deployed on the same node or different nodes. The node is the processor included in the first processor 101, and to determine the number of processors included in the second processor 102. The first processor 101 obtains a target request segment with a number matching the pre-filling resource configuration information from each task processing request in the current request batch, and sends the target request segment as the current request batch to the second processor 102. The second processor 102 performs parallel pre-filling processing on each target request segment to generate the current token batch, and sends it to the first processor 101; the lengths of each target request segment are the same and are part of the corresponding task processing request; the first processor 101 obtains the next token of each token in the current token batch from the current request batch, and generates a new current token batch in this way to obtain multiple current token batches until the combined length of the current token batches reaches the token batch length, and combines each current token batch into a token batch and sends it to the second processor 102; the second processor 102 constructs multiple pipelines using the decoding resource configuration information, and performs parallel decoding processing on each token batch through the multiple pipelines, and sends the decoding processing result to the first processor 101; the first processor 101 generates a corresponding language task processing result according to the decoding results of all request segments of each task processing request.

[0084] Exemplarily, if the total memory of the second processor 102 is greater than or equal to the sum of twice the memory occupancy requirement of the language task processing model, the activation memory occupancy, and the memory occupancy requirement of the key-value cache corresponding to the maximum generation length, then the second processor 102 is deployed within the same node. That is, there is only one first processor 101. For ease of description, the processor that performs pre-filling calculations is defined as the pre-filling processor group, and the processor that performs decoding calculations is defined as the decoding processor group. The second processor 102 includes a pre-filling processor group and a decoding processor group; the pre-filling processor group and the decoding processor group are respectively connected to the first processor 101; the pre-filling processor group includes multiple pre-filling processors, and the pre-filling processors are connected in sequence; the decoding processor group includes multiple decoding processors, and the decoding processors are connected in sequence; among them, the minimum value of the total number of pre-filling processors and decoding processors is determined according to the total memory of a single processor and the memory occupancy requirement of the language task processing model; the total memory of a single processor is the maximum value of the total memory of the pre-filling processors and the total memory of the decoding processors; the optimal values of the number of pre-filling processors and the number of decoding processors respectively are the number of pre-filling processors included in the pre-filling processor group and the number of decoding processors included in the decoding processor group when the throughput of the pre-filling stage and the throughput of the decoding stage satisfy the preset same or similar conditions; when the throughput of the pre-filling stage is less than the throughput of the decoding stage, the number of pre-filling processors is greater than the number of decoding processors, and the sum of the number of pre-filling processors and the number of decoder processors is greater than or equal to the minimum value of the total number.

[0085] Exemplarily, if the total memory of the second processor 102 is less than the sum of twice the memory occupancy requirement of the language task processing model, the activation memory occupancy, and the memory occupancy requirement of the key-value cache corresponding to the maximum generation length, then the second processor 102 is deployed in different nodes, that is, the first processor 101 includes at least 2 nodes. For ease of description, the first processor 101 may include a first host node and a second host node. The first host node is connected to the second processor 102, the second host node is connected to a third processor, and the third processor is connected to the second processor 102; the second processor 102 includes a pre-filling processor group, and the third processor includes a decoding processor group. The total memory of the pre-filling processor group is greater than twice the memory occupancy requirement of the language task processing model, and the total memory of the decoding processor group is greater than twice the memory occupancy requirement of the language task processing model; the pre-filling processor group includes multiple pre-filling processors, and each pre-filling processor is connected in sequence; the decoding processor group includes multiple decoding processors, and each decoding processor is connected in sequence; each pre-filling processor in the pre-filling processor group communicates through a point-to-point connection method, such as through NVlink (a point-to-point connection method), and each decoding processor in the decoding processor group communicates through a point-to-point connection method; each pre-filling processor in the pre-filling processor group is interconnected and communicates with each decoding processor in the decoding processor group through the remote direct access method.

[0086] To make the technical solution of the present invention clearer and more understandable to those skilled in the art, the present invention also describes the language task processing system with the first processor being a CPU and the second processor being a GPU as an example, such as Figure 11As shown, the language task processing system includes multiple CPUs and multiple GPUs. If the memory capacity of a single GPU cannot meet the memory requirements of the language task processing model, the model parallelism method can be used to split the model inference operations across multiple GPUs. Correspondingly, the prefill calculation and the decoding calculation are each deployed across multiple GPUs. The CPUs and GPUs can be connected via the PCIE method. The GPUs within the same CPU node can communicate with each other via the NVlink method, and the GPUs across CPU nodes are interconnected and communicate via the Remote Direct Memory Access method. Since the bandwidth of NVlink is higher than that of the direct memory access method, single-node internal deployment is preferentially selected when the deployment conditions are met. Among them, the functions of receiving user prompt requests, scheduling and forwarding for batch processing of prompt requests and tokens, and text-sequence conversion on the CPU are encapsulated as functional modules by computer programs that implement these functions and are solidified on the CPU. Correspondingly, the CPU includes a request scheduling module and a text-sequence conversion module. The text-sequence conversion module completes the conversion between the character text and the numerical sequence of each token in the generation request; initially, the input prompt text is converted from text characters to numerical values through the encode function of the tokenizer module, and the numerical values are used as model inputs for prefill and decoding inference calculations. After all decoding is completed, the predicted numerical values are input into the decode function of the tokenizer module to convert the numbers into text characters, completing text generation. The request scheduling module schedules the prompt batch requests from the user side and the token batch requests output from the prefill or decoding stage. The scheduling module schedules the generation inference calculations of the large language model to each GPU by managing three queues. The GPUs are divided into prefill GPUs and decoding GPUs according to the inference stage of the language task processing model. Both the prefill GPUs and the decoding GPUs can include multiple GPUs.

[0087] Based on Figure 11For the language task processing system, when the prompt requests of several users at the application layer are input into the request scheduling module, the request scheduling module uniformly truncates all the prompts in a batch to the minimum length and adopts a static batch processing mechanism to schedule and execute the prompt batch. The size of the prompt batch is determined according to the computing power of the pre-filled GPU. The prompt batch is input into the text-sequence conversion module to convert all the text characters of the prompt into a digital sequence. According to the distributed parallel manner in the pre-filled computing stage, the prompt batch sequence is input into the pre-filled GPU for pre-filled computing. After the pre-filled stage, each prompt predicts and outputs a token and the key-value cache of all tokens. The parallel manner in the pre-filled stage depends on the model memory and the number of GPUs, and model parallelism and tensor parallelism can be adopted. When the token batch generated by the pre-filled computing is input into the request scheduling module, considering the problem of under-utilized GPU resources in the decoding stage, a batch processing scheduling method of token batch merging and pipeline concurrent execution is adopted to increase the batch size processed by the decoding GPU, thereby improving the resource utilization rate of the decoding GPU. Before the token batch performs decoding computing, the key-value cache of the prompt tokens calculated by the previous pre-filled GPU is transferred to the decoding GPU, and the scheduled merged token batch is input into the decoding GPU to perform decoding computing in a pipeline concurrent manner. Each token in the batch outputs the next token and the key-value cache after decoding computing; repeat the token batch scheduling and decoding computing in the previous two steps until all the tokens in the batch complete the decoding computing. Finally, all the tokens in the batch are input into the text-sequence conversion module to convert the digital sequence into text characters, obtaining the processing result of the task processing request and completing the inference computing of the language task.

[0088] As can be seen from the above, this embodiment uses a distributed computing platform composed of CPU + GPU to run the language task processing model, and splits the pre-filled stage and the decoding stage to different GPUs for execution, reducing resource coupling and improving the pre-filled latency stability. The pre-filled stage adopts sequence truncation and static batch processing technologies to reduce pipeline bubbles. The decoding stage adopts a batch processing scheduling method based on token batch merging and pipeline concurrent execution to increase the batch size processed by the GPU, thereby improving the GPU resource utilization rate in the decoding stage without significantly increasing the memory.

[0089] The above has introduced in detail a language task processing method, system, electronic device, computer-readable storage medium, and computer program product provided by the present invention. Each embodiment in this specification is described in a progressive manner. The key point of each embodiment is the difference from other embodiments. For the same or similar parts between the embodiments, reference can be made to each other. Whether the units and algorithm steps of each example described in the disclosed embodiments are executed in the form of electronic hardware or computer software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods for each specific application to implement the described functions, and such implementation should not be considered to exceed the scope of the present invention. Without departing from the principle of the present invention, several improvements and modifications can also be made to the present invention, and these improvements and modifications also fall within the protection scope of the present invention.

Claims

1. A language task processing method, characterized in that: include: Determining resource configuration information of the language task processing model in a pre-filling stage and a decoding stage respectively according to resource demand information of the language task processing model in the process of executing the language task, and determining the word unit batch length according to the pre-filling resource configuration information and the decoding resource configuration information; From each task processing request of the current request batch, obtain target request segments of a number matching the pre-filled resource configuration information, perform pre-filling processing on each target request segment in parallel, and generate the current word unit batch; In the current request batch, multiple new word-unit batches are generated by obtaining the next word-unit of each word-unit in the word-unit batch that is most recently generated at the current moment, and each next word-unit constitutes a new word-unit batch, so as to meet the condition that the length of the current word-unit batch and each new word-unit batch combined reaches the length of the word-unit batch; each target request segment has the same length and is a partial content of the corresponding task processing request; Multiple pipelines are constructed using the decoding resource configuration information, and each word batch is decoded in parallel through the multiple pipelines, and the corresponding language task processing results are obtained according to the decoding results of all request segments of each task processing request.

2. The language task processing method according to claim 1, characterized in that: Determining resource configuration information of the language task processing model in a pre-filling phase and a decoding phase, respectively, according to resource requirement information of the language task processing model in a process of executing the language task, including: When the sum of the double memory usage requirement, the activation memory usage, and the memory usage requirement of the key-value cache corresponding to the maximum generation length of the language task processing model is greater than or equal to the total memory resources of a single host node, different host nodes are used to process the pre-filling stage and the decoding stage; When the sum of the double memory usage requirement, the activation memory usage and the memory usage requirement of the key-value cache corresponding to the maximum generation length of the language task processing model is less than the total memory resources of a single host node, the language task processing model runs on a single host node.

3. The language task processing method according to claim 1, characterized in that: Determining resource configuration information of the language task processing model in a pre-filling phase and a decoding phase, respectively, according to resource requirement information of the language task processing model in a process of executing the language task, including: Acquire the prefilling throughput of the prefilling stage under different prefilling resource configuration information, and the decoding throughput of the decoding stage under different decoding resource configuration information; When the pre-filling throughput and the decoding throughput meet the preset same or similar conditions, the corresponding target pre-filling resource configuration information and the target decoding resource configuration information are used as the optimal resource configuration information of the language task processing model in the pre-filling stage and the decoding stage; When the pre-filling throughput and the decoding throughput cannot satisfy preset identical or similar conditions, and the pre-filling throughput is less than the decoding throughput, the total amount of pre-filling resources configured for the pre-filling resource configuration information is greater than the total amount of decoding resources configured for the decoding resource configuration information; and the sum of the total amount of pre-filling resources and the total amount of decoding resources is less than or equal to the maximum total amount of resources.

4. The language task processing method according to claim 1, characterized in that: Determining resource configuration information of the language task processing model in a pre-filling phase and a decoding phase, respectively, according to resource requirement information of the language task processing model in a process of executing the language task, including: Determine a minimum memory usage value according to the memory usage requirement of the language task processing model and the memory usage requirement of the key-value cache corresponding to the maximum generation length; Determining the minimum number of pre-filled processors used by the language task processing model in the pre-filling phase according to the minimum memory usage and the total memory owned by the pre-filled processors; The minimum number of decoding processors used by the language task processing model in the decoding stage is determined according to the minimum memory occupancy value and the total memory owned by the decoding processor.

5. The language task processing method according to claim 1, characterized in that: From each task processing request in the current request batch, obtain the target request segments that match the pre-filled resource configuration information, including: Obtain the number of tokens in each task processing request, and select the minimum number of tokens as the truncation length; The first n target word units are obtained from each task processing request respectively as the target request segment of each task processing request to form a prompt batch; the length of each target word unit of the same task processing request is the same as the truncation length.

6. The language task processing method according to claim 1, characterized in that: From each task processing request in the current request batch, obtain the target request segments that match the pre-filled resource configuration information, including: Pre-building a batch processing queue, wherein the batch processing queue includes an undecoded batch queue and a decoded batch queue; When a current request batch is received, each task processing request of the current request batch is moved into the undecoded batch queue, a first batch of target request segments of a number matching the pre-filled resource configuration information is obtained from the undecoded batch queue, and the first batch of target request segments is moved into the decoded batch queue, so as to perform pre-filling processing on each first batch of target request segments in the decoded batch queue in parallel; When the first batch of word units corresponding to the first batch of target request segments is generated, if the first batch of word units includes the first type of target word units with end identification information, the target task processing request corresponding to the first type of target word units has been completed, and for the second type of target word units without end identification information, when the second type of target word units do not exist in the undecoded batch queue and the decoded batch queue, they are added to the corresponding request batch of the undecoded batch queue; If there are task processing requests for the current request batch in the undecoded batch queue, a second batch of target request segments whose number matches the pre-filled resource configuration information is obtained from the undecoded batch queue again, and the second batch of target request segments is moved into the decoded batch queue; if there are no task processing requests for the current request batch in the undecoded batch queue, the word-meta batches of the current request batch are merged based on the word-meta batch length.

7. The language task processing method according to claim 6, characterized in that: The batch processing queue also includes a decoded batch queue; When the first batch of word units corresponding to the first batch of target request segments are generated, the corresponding first batch of target request segments are moved into the decoded batch queue. If the first batch of word units includes a first type of target word units with end identification information, the request segment of the task processing request corresponding to the first type of target word units is deleted from the decoded batch queue.

8. The language task processing method according to claim 1, characterized in that: In the current request batch, multiple new word unit batches are generated by obtaining the next word unit of each word unit in the word unit batch that is most recently generated at the current moment, and forming each next word unit into a new word unit batch, including: If the next word of the target prompt word of the current word batch is a known word in the current request batch, generating a first word batch according to the known word corresponding to the target prompt word; If the combined length of the first word unit batch and the current word unit batch is less than the word unit batch length, obtaining the known word units corresponding to the first word unit batch in the current request batch, and generating a second word unit batch; If the combined length of the first word unit batch, the current word unit batch and the second word unit batch is the word unit batch length, multiple pipelines are used to decode the first word unit batch, the current word unit batch and the second word unit batch in parallel.

9. The language task processing method according to any one of claims 1 to 8, characterized in that: Multiple pipelines are constructed using the decoding resource configuration information, and each word batch is decoded in parallel through multiple pipelines, including: According to the number of model layers of the language task processing model, each sub-model deployed in the decoding stage of the language task processing model is divided into a plurality of sub-sub-models, and the number of sub-sub-models of each sub-model is the same as the value of the word batch length; Constructing multiple pipelines, each pipeline serially executing each sub-model; the number of pipelines is the same as the value of the word batch length; Through the first pipeline, each sub-model is used in turn to perform decoding processing on the first current word batch; After the first pipeline completes the decoding process of the first current word unit batch using the first sub-model, while the first pipeline is decoding the first current word unit batch using the second sub-model, the second pipeline is decoding the second current word unit batch using the target sub-sub-model corresponding to the first sub-model; After the second pipeline completes the decoding processing of the second current word batch using the target sub-sub-model, while the second pipeline is decoding the second current word batch using the next sub-sub-model, the third pipeline decodes the third current word batch using the sub-sub-model corresponding to the target sub-sub-model.

10. An electronic device, characterized in that: include: Memory for storing computer programs; A processor, configured to implement the steps of the language task processing method according to any one of claims 1 to 9 when executing the computer program.

11. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the steps of the language task processing method according to any one of claims 1 to 9 are implemented.

12. A computer program product comprising a computer program / instructions, characterized in that When the computer program / instructions are executed by a processor, the steps of the language task processing method according to any one of claims 1 to 9 are implemented.

13. A language task processing system, characterized in that: At least comprising a first processor and a second processor, wherein the first processor is connected to the second processor; resource configuration information of the first processor and the second processor is determined according to resource configuration information of the language task processing model in the pre-filling stage and the decoding stage; The first processor obtains target request segments of a number matching the pre-filled resource configuration information from each task processing request of the current request batch, and sends the target request segments as the current request batch to the second processor, and the second processor performs pre-filling processing on each target request segment in parallel to generate a current word unit batch, and sends it to the first processor; each target request segment has the same length and is a part of the corresponding task processing request; The first processor generates a plurality of new word unit batches in the current request batch by acquiring the next word unit of each word unit in the word unit batch most recently generated at the current moment and forming a new word unit batch with each next word unit, so as to meet the condition that the length of the current word unit batch and each new word unit batch is equal to the length of the word unit batch, and merges each current word unit batch into a word unit batch and sends it to the second processor; The second processor constructs a plurality of pipelines using the decoding resource configuration information, performs decoding processing on each word batch in parallel through the plurality of pipelines, and sends the decoding results to the first processor; The first processor generates a corresponding language task processing result according to the decoding results of all request segments of each task processing request.

14. The language task processing system according to claim 13, characterized in that: The total memory of the second processor is greater than or equal to twice the memory requirement of the language task processing model, the activation memory requirement and the sum of the memory requirement of the key-value cache corresponding to the maximum generation length, and the second processor includes a pre-filling processor group and a decoding processor group; The pre-filling processor group and the decoding processor group are respectively connected to the first processor; the pre-filling processor group includes a plurality of pre-filling processors, each of which is connected in sequence; the decoding processor group includes a plurality of decoding processors, each of which is connected in sequence; The minimum value of the total number of pre-filled processors and decoding processors is determined according to the total memory of a single processor and the memory occupation requirement of the language task processing model; the total memory of a single processor is the maximum value of the total memory of the pre-filled processor and the total memory of the decoding processor; The optimal values ​​of the number of pre-filling processors and the number of decoding processors are the number of pre-filling processors included in the pre-filling processor group and the number of decoding processors included in the decoding processor group when the throughput of the pre-filling stage and the throughput of the decoding stage meet the same preset similar conditions; when the throughput of the pre-filling stage is less than the throughput of the decoding stage, the number of pre-filling processors is greater than the number of decoding processors, and the sum of the number of pre-filling processors and the number of decoder processors is greater than or equal to the minimum value of the total number.

15. The language task processing system according to claim 13, characterized in that: The total memory of the second processor is less than twice the memory requirement of the language task processing model, the activation memory requirement, and the sum of the memory requirement of the key-value cache corresponding to the maximum generation length, the first processor includes a first host node and a second host node, the first host node is connected to the second processor, the second host node is connected to the third processor, and the third processor is connected to the second processor; The second processor includes a pre-filling processor group, the third processor includes a decoding processor group, the total memory of the pre-filling processor group is greater than twice the memory occupation requirement of the language task processing model, and the total memory of the decoding processor group is greater than twice the memory occupation requirement of the language task processing model; the pre-filling processor group includes a plurality of pre-filling processors, each of which is connected in sequence; the decoding processor group includes a plurality of decoding processors, each of which is connected in sequence; The pre-filling processors of the pre-filling processor group communicate through a point-to-point connection, and the decoding processors of the decoding processor group communicate through a point-to-point connection; the pre-filling processors of the pre-filling processor group are interconnected and communicated with the decoding processors of the decoding processor group respectively through a remote direct access method.

Citation Information

Patent Citations

  • Large language model reasoning optimization method and device, computer equipment and storage medium

    CN117194056A

  • Large language model reasoning optimization method and device, electronic equipment and storage medium

    CN119150994A