Language task processing method, system and device, storage medium and program product
By determining the word batch length based on resource configuration information during the pre-filling and decoding stage of the language task processing model, and building multiple pipelines in parallel processing word batches in the decoding stage, the problem of resource utilization unsaturation in the decoding stage in the language task processing model is solved, and resource utilization and task processing performance are improved.
Patent Information
- Application Number
- CN202510526403.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-25
- Publication Date
- 2025-05-30
- Estimated Expiration
- 2045-04-25
AI Technical Summary
When the language task processing model executes language tasks, there is a problem of resource utilization unsaturation in the decoding stage, resulting in low resource utilization, reduced throughput and increased task processing delay.
By determining the word batch length according to the resource configuration information of the language task processing model in the pre-filling stage and the decoding stage, the word batch length is determined, and the target request segment is processed in parallel in the pre-filling stage to generate the word batch. During the decoding stage, multiple pipelines are built to process word batches in parallel to improve resource utilization.
It effectively improves the resource utilization rate during language task execution, reduces resource coupling, improves the stability of pre-filling delay, and reduces the resource utilization unsaturation problem in the decoding stage.
Smart Images

Figure CN120068846A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of artificial intelligence technology, and in particular to a method, system, electronic device, computer-readable storage medium, and computer program product for processing language tasks. Background Art
[0002] When a language model executes a language task, its inference generation process includes a prefill stage and a decoding stage with different computational requirements. To avoid the problem of reduced inference latency caused by a single device mixing the processing of the two stages, related technologies split and deploy them to different computing devices for processing according to the different performance requirements of the prefill stage and the decoding stage. As the length of the input context text increases, the performance requirement difference between these two stages will continuously increase, resulting in the underutilization of resources in the decoding stage. It can be seen that related technologies cannot significantly improve the problem of low resource utilization during the execution of language tasks. Summary of the Invention
[0003] The present invention provides a method, system, electronic device, computer-readable storage medium, and computer program product for processing language tasks, which can solve the problem of underutilized resources in the decoding stage when a language task processing model executes a language task, and can effectively improve the resource utilization rate during the execution of language tasks.
[0004] To solve the above technical problems, the present invention provides the following technical solutions: The present invention provides a method for processing language tasks, including: Determine the resource configuration information of the language task processing model in the prefill stage and the decoding stage respectively according to the resource requirement information during the execution of the language task by the language task processing model, and determine the token batch length according to the prefill resource configuration information and the decoding resource configuration information; obtain the target request segments with the number matching the prefill resource configuration information from each task processing request in the current request batch, and perform prefill processing on each target request segment in parallel to generate the current token batch; in the current request batch, generate multiple new token batches by obtaining the next token of each token in the latest generated token batch at the current moment and forming a new token batch with each next token, so as to meet the condition that the combined length of the current token batch and each new token batch reaches the token batch length; the lengths of each target request segment are the same and are part of the corresponding task processing request; construct multiple pipelines using the decoding resource configuration information, perform decoding processing on each token batch in parallel through the multiple pipelines, and obtain the corresponding language task processing result according to the decoding results of all request segments of each task processing request.
[0005] The present invention also provides an electronic device, including a memory and a processor, and the processor is used to implement the steps of any of the above language task processing methods when executing the computer program stored in the memory.
[0006] The present invention also provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the steps of any of the above language task processing methods are implemented.
[0007] The present invention also provides a computer program product, including a computer program / instructions. When the computer program / instructions are executed by a processor, the steps of any of the above language task processing methods are implemented.
[0008] Finally, the present invention also provides a language task processing system, which at least includes a first processor and a second processor, and the first processor is connected to the second processor; the resource configuration information of the first processor and the second processor is determined according to the resource configuration information of the language task processing model in the pre-filling stage and the decoding stage; wherein, the first processor obtains, from each task processing request in the current request batch, a target request segment with a number matching the pre-filling resource configuration information, and sends the target request segment to the second processor as the current request batch, and the second processor performs parallel pre-filling processing on each target request segment to generate a current token batch and sends it to the first processor; the lengths of each target request segment are the same and are partial contents of the corresponding task processing request; the first processor obtains the next token of each token in the current token batch from the current request batch to generate a new current token batch, and obtains multiple current token batches until the combined length of the current token batches reaches the token batch length, and combines each current token batch into a token batch and sends it to the second processor; the second processor constructs multiple pipelines by using the decoding resource configuration information, and performs parallel decoding processing on each token batch through the multiple pipelines, and sends the decoding processing result to the first processor; the first processor generates a corresponding language task processing result according to the decoding results of all request segments of each task processing request.
[0009] The advantages of the technical solution provided by the present invention are as follows: Deploying the decoding stage and the pre-filling stage of the language task processing model during the execution of the language task on different devices not only helps improve resource utilization, but also reduces resource coupling and improves the stability of the pre-filling latency. During the pre-filling stage, multiple task processing requests are merged into one request batch, and multiple target request segments in the request batch are pre-filled in parallel. This can not only increase the total number of tokens output per second in the pre-filling stage, improve the throughput of the pre-filling stage, and improve the resource utilization rate of the pre-filling stage, but also effectively increase the data processing volume in the decoding stage by increasing the throughput of the pre-filling stage, thereby helping to reduce the problem of under-utilized resources in the decoding stage. It can also effectively reduce the pipeline bubbles generated in the pre-filling stage. Further, during the decoding stage, the token batch generated in the pre-filling stage and multiple known token batches are merged, and these multiple token batches are executed concurrently using a pipeline, increasing the number of batches and the batch size that can be processed in the decoding stage. Thus, without significantly increasing the memory, the resource utilization rate of the decoding stage is effectively improved, and the language task processing performance is enhanced. In addition, the present invention also provides a corresponding implementation system, electronic device, computer-readable storage medium, and computer program product for the language task processing method, further making the method more practical, and the system, electronic device, computer-readable storage medium, and computer program product have corresponding advantages. BRIEF DESCRIPTION OF THE DRAWINGS
[0010] To more clearly illustrate the technical solutions of the present invention or related technologies, the following will briefly introduce the drawings required for use in the description of the embodiments or related technologies. Obviously, the drawings in the following description are only some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.
[0011] Figure 1 Schematic diagram of the hardware composition framework applicable to the language task processing method provided by the present invention; Figure 2 Schematic diagram of the process flow of a language task processing method provided by the present invention; Figure 3 Schematic diagram of the inference process flow of the language task processing model provided by the present invention in an exemplary example; Figure 4 Schematic diagram of the data processing flow of the language task processing model provided by the present invention; Figure 5 Schematic diagram of the token batch merging process flow in an exemplary example provided by the present invention; Figure 6 Schematic diagram of the process flow of pipeline parallel execution of token batches in an exemplary example provided by the present invention; Figure 7 A schematic diagram of the process of pipelined parallel execution of token batches in another exemplary example provided by the present invention; Figure 8 A schematic diagram of the pipelined parallel execution process in an exemplary example provided by the present invention; Figure 9 A structural framework diagram of a language task processing device provided by the present invention under an exemplary embodiment; Figure 10 A structural framework diagram of a language task processing system provided by the present invention under an exemplary embodiment; Figure 11 A schematic diagram of the architecture of a language task processing system provided by the present invention in an exemplary example. Detailed implementation manners
[0012] In order to enable those skilled in the art to better understand the technical solutions of the present invention, the present invention will be further described in detail below with reference to the accompanying drawings and specific implementation manners. Among them, the terms "first", "second", "third", "fourth", etc. in the specification and the above drawings are used to distinguish different objects, rather than to describe a specific order. In addition, the terms "include" and "have" and any variations thereof are intended to cover non-exclusive inclusion. The term "exemplary" means "serving as an example, embodiment, or illustration". Any embodiment described herein as "exemplary" is not necessarily to be construed as superior to or better than other embodiments.
[0013] With the development of artificial intelligence technology, language models are widely applied to the processing of various artificial intelligence tasks, such as question-and-answer tasks, translation tasks, and code generation tasks. As the parameter scale and data of the language model increase, the resources required for the language model to execute related language tasks are also increasing.
[0014] When a language model executes a language task, its inference generation process includes a prefill stage and a decoding stage with different computational requirements. In order to reduce the repeated key-value calculations of tokens, the language model performs key-value caching during the inference process. The decoding process needs to save the key-value caches of all previous tokens until the decoding ends. As the length of the generated context sequence increases, this leads to an unlimited increase in the memory occupancy of the key-value cache, further exacerbating the resource requirements of the language model. In order to meet the memory resources and computational resources requirements of the language model, a distributed parallel computing system is used to run the language model.
[0015] Considering issues such as data transmission overhead, computing resource utilization, memory capacity, cost, and energy consumption in distributed parallel computing systems, the language model executes language tasks through parallel computing methods. For example, in the inference phase, tensor parallelism, pipeline parallelism, and prefill-decoding phase split parallelism can be combined. Among them, related technologies usually split the prefill phase and the decoding phase to be processed on different computing devices. For example, a related technology places the prefill and decoding on different systems to reduce the pressure of single-device inference processing. The prefill distributed system is determined according to the characteristics of prefill computing, and the decoding system is determined according to the characteristics of decoding computing, so that each system meets the requirements of the corresponding inference process, balances the efficiency differences of different computing phases, reduces resource waste, and improves the efficiency of the entire inference process. Another related technology deploys the prompt calculation and token calculation phases to different models of GPUs (Graphics Processing Units), individually optimizing the hardware resource management of each phase, making the hardware used in each phase the most suitable, and optimizing the communication overhead of the key-value cache between GPUs. There is also a related technology that allocates prefill and decoding calculations to different GPUs to eliminate pre-decoding interference, jointly optimizes resource allocation and parallel strategies for each phase, and also places these two phases according to the bandwidth of the service cluster to minimize the communication caused by decomposition.
[0016] Although related technologies can reduce the inference latency caused by resource coupling between the prefill phase and the decoding phase calculations on a single computing device by splitting and parallelizing the prefill-decoding phase to execute language tasks, however, since related technologies analyze the computational performance differences between the prefill phase and the decoding phase, and match GPUs with corresponding performance according to the performance differences of the two phases to maximize the utilization of GPU resources in each phase, thereby solving the problem of unsaturated utilization of computing resources in the decoding phase. However, on the premise that the computational performance and memory requirements of the prefill phase and the decoding phase themselves are already unbalanced, as the length of the user input context continues to increase, the imbalance in the requirements of the two phases will correspondingly increase, which will still lead to the problem of unsaturated utilization of computing resources in the decoding phase, reducing the throughput during the execution of language tasks by the language model and increasing the latency of task processing.
[0017] In view of this, in order to improve the problem of low GPU resource utilization rate in the decoding stage in the related technology, different batch request processing methods are adopted for the pre-filling stage and the decoding stage split onto different processors. For the pre-filling stage, length truncation and static batch processing technologies are used to improve the resource utilization rate of the pre-filling stage. For the decoding stage, batch processing based on token batch merging and pipeline concurrent execution is adopted to increase the batch size and batch dimension processed in the decoding stage, thereby improving the resource utilization rate of the decoding stage, such as the GPU resource utilization rate, without significantly increasing the memory. Combining with the specific application environment architecture or specific hardware architecture on which the execution of the language task processing method depends, the specific application environment architecture or specific hardware architecture is described herein. The following combines Figure 1 Some possible application scenarios related to the technical solution of the present invention are introduced by way of example, which may include the following content: Such as Figure 1 As shown, the hardware composition framework may include a server 11 and a GPU group 12. The total number of GPUs included in the GPU group 12 and the number of CPUs (Central Processing Unit) included in the server 11 are determined according to the resource requirement information in the process of executing the language task by the language task processing model. Each CPU of the server 11 is connected to each other and is connected to each GPU of the GPU group 12 through PCIe (peripheral component interconnect express, high-speed serial computer expansion bus). When the total amount of twice the language task processing model memory, intermediate activation memory, and key-value cache memory corresponding to the maximum generation length in the language task generation scenario is greater than or equal to the total GPU memory on a single CPU, then each GPU of the GPU group 12 is deployed to different CPUs, that is, the pre-filling stage and the decoding stage are split onto the GPUs of different CPUs. When the total amount of twice the language task processing model memory, intermediate activation memory, and key-value cache memory corresponding to the maximum generation length in the language task generation scenario is less than the total GPU memory on a single CPU, then each GPU of the GPU group 12 is deployed to the same CPU, that is, the pre-filling stage and the decoding stage are split onto the GPUs of the same CPU. A part of the GPUs in the GPU group 12 is used to process the pre-filling stage of the language task processing model, which is defined as the pre-filling GPU group. For example, Figure 1 Taking 4 GPUs as an example, each GPU can be respectively defined as G0-P, G1-P, G2-P, G3-P, where P represents the pre-filling stage and G represents the GPU. A part of the GPUs is used to process the decoding stage of the language task processing model, which is defined as the decoding GPU group. For example, Figure 1Taking 4 GPUs as an example, each GPU can be respectively defined as G4-T, G5-T, G6-T, and G7-T, where T represents the decoding stage. The server 11 is also used to provide a human-computer interaction interface, which can be the interface of the corresponding application software or the interface opened in a browser through a specified URL (Uniform Resource Locator).
[0018] In this application scenario, the user inputs a task processing request through the human-computer interaction interface. The human-computer interaction interface sends the task processing request to the CPU through the network. The CPU combines the task processing requests of multiple users or multiple task processing requests of the same user into one batch. For the sake of easy description, it is defined as a request batch. The CPU uniformly truncates all the task processing requests in the request batch to the minimum length to form a prompt batch. The prompt batch contains multiple prompt contents in the form of text characters. Convert all the text characters of the prompt batch into a digital sequence to obtain a prompt batch sequence. Input the prompt batch sequence into the pre-fill GPU group for parallel pre-fill calculation. The parallel method in the pre-fill stage is determined according to the memory requirements of the language task processing model and the number of GPUs in the GPU group 12. After the pre-fill stage is completed, each prompt content in the prompt batch predicts and outputs its next token and the key-value cache of all tokens. Each token constitutes the current token batch. The pre-fill GPU group sends the current token batch to the CPU and sends the key-value cache to the decoding GPU group. The CPU obtains the next token of each token in the current token batch from the request batch, generates a new current token batch, and repeats continuously until the combined length can reach the token batch length, and combines each current token batch into one batch and sends it to the decoding GPU group. The decoding GPU group performs decoding calculations on the key-value cache generated in the previous pre-fill stage and each current token batch in a pipelined concurrent manner. Each token in the batch outputs the next token and the key-value cache after decoding calculation. Continuously repeat the token batch scheduling and decoding calculation of the above steps until all tokens in the request batch complete the decoding calculation. Finally, send all the tokens generated by the decoding GPU group to the CPU. The CPU converts them into text characters, that is, obtains the language task processing result corresponding to the task processing request, and displays the language task processing result to the user through the human-computer interaction interface.
[0019] It should be noted that the above application scenario is only shown for the convenience of understanding the idea and principle of the present invention, and the embodiments of the present invention are not restricted in this regard. On the contrary, the embodiments of the present invention can be applied to any applicable scenario. After introducing the technical solution of the present invention, the various non-limiting embodiments of the present invention will be described in detail below with reference to the drawings and specific embodiments. First, please refer to Figure 2 , the language task processing method provided in this embodiment may include the following content: S201: Determine the resource configuration information of the language task processing model in the pre-filling stage and the decoding stage respectively according to the resource requirement information during the execution of the language task by the language task processing model, and determine the token batch length according to the pre-filling resource configuration information and the decoding resource configuration information.
[0020] Among them, the language task is a task that needs to be executed using the language task processing model, including but not limited to question-and-answer tasks, code generation tasks, and text generation tasks. The language task can be input in any way, such as text, table, graph, video, voice, and it includes at least a prompt. The prompt is a prompt that tells the language task processing model how to execute the language task and can be flexibly set according to the actual application scenario. The language task processing model is a pre-trained language model that can process language tasks. When allowing the language task to be input in multiple ways, corresponding encoding modules can be set in the input layer. If voice input is supported, a network model for converting speech to text or converting speech to images also needs to be added to the input layer of the task processing model. If visual or text information input is supported, an image encoder and a text encoder need to be added to the input layer of the language task processing model. The neural network algorithm structure of the language task processing model can be CNN (Convolutional Neural Network), RNN (Recurrent Neural Network), LSTM (Long Short-Term Memory), etc., or it can be a model constructed by an attention network, such as transformer (transformer network model), bert (Bidirectional Encoder Representations from Transformers), GPT (Generative Pre-trained Transformer), Clip (Contrastive Language–Image Pre-training), etc. The present invention does not make any limitations here. Among them, the attention network refers to a network model trained using the attention mechanism. This model assigns different weights to each part of the input sequence, thereby extracting more important feature information from the input sequence, so that the language task processing model finally obtains a more accurate task processing result.
[0021] In the present invention, in order to reduce resource coupling and improve the stability of pre-filling latency, the pre-filling stage and the decoding stage of a language task such as a text inference task are split and processed on different computing devices, such as GPUs, FPGAs, etc., which can balance the differences between the computing requirements and memory requirements of the pre-filling stage and the decoding stage, and the parallel granularity is larger than that of model parallelism and tensor parallelism. For example Figure 3As shown, when the user inputs a task processing request containing a prompt, the prompt is predicted and generates the next text character through the inference calculation of the language task processing model, which is called a token. The prompt calculation stage is called the prefill stage. Then, the prompt and the previous token are input into the language task processing model, and the next token is generated through inference calculation. This process is repeated until all tokens are generated. The token generation calculation process is called decoding. That is, the decoding stage calculates and generates the next token based on the token at the previous position and the key-value cache of all previous tokens. The former focuses on compute-intensive requirements, and the latter focuses on memory-intensive requirements. Since the generation of each token requires the input of all previous tokens, then each token generation needs to repeat the calculation of all previous tokens. To reduce the repeated key-value calculation of tokens, a key-value cache optimization method is adopted, that is, the key-value cache data of all previous tokens is retained during each token generation process, and there is no need to repeat the calculation. The input for token generation is the key-value cache of all previous tokens and the previous token. Correspondingly, the prefill stage generates the first predicted token based on all tokens in the prompt context input by the user and retains the key-value (KV) cache of all tokens. The decoding process needs to save the key-value cache of all previous tokens until the decoding ends. As the length of the generated context sequence increases, the memory occupancy of the key-value cache also increases without limit. Coupled with the increase in model capacity, a distributed parallel computing system is used to run the resource requirements of the language task processing model during the execution of language tasks. The distributed parallel computing system is composed of different types of processors, such as CPU+GPU, or CPU+FPGA. The CPU is used to receive task processing requests from several users at the application layer, perform corresponding text-sequence conversion on the task processing requests and the tokens generated in the decoding stage, and schedule and forward the batch processing of the task processing requests and tokens. The GPU or FPGA is used to run the prefill calculation and decoding calculation in the inference process of the language task processing model. The resource requirement information includes memory requirement information and compute requirement information. For memory requirements, the present invention splits the prefill stage and the decoding stage of the language task processing model onto different computing devices. The calculation processes of these two stages each require a copy of the model parameters of the language task processing model. If deployed on a single CPU node, it is required that the total GPU memory on a single CPU node is at least greater than twice the model memory. In addition, the intermediate activation memory and key-value cache memory occupancy during model inference need to be considered. The intermediate activation memory depends on the memory occupancy of the network layer of the language task processing model, such as depending on the GPU memory occupancy of a single Transformer module layer. The memory of the key-value cache is the memory occupancy of the key matrix and value matrix corresponding to all tokens in the prefill and decoding.For computing requirements, the input for prefill computing is a prompt batch, which is part or all of the request batch merged from multiple task processing requests. It is highly computationally intensive. The input batch size and the number of batches can usually ensure full utilization of prefill resources. However, for decoding computing, the input for a single decoding request is a single token text. To improve the resource utilization rate during the decoding phase, it is necessary to increase the batch size and the number of batches as much as possible. Since the decoding phase needs to store the key-value cache of all tokens, the memory requirement of the computing device used in the decoding phase is higher than that of the computing device used in the prefill phase. Therefore, the batch size and the number of batches processed by the computing device used in the decoding phase are also affected. To increase the batch size and the number of batches during the decoding phase, it is necessary to improve the throughput of the prefill phase as much as possible, that is, to increase the number of tokens output per second by the prefill computing. Considering the high computational intensity of the prefill phase, to improve its throughput, it is necessary to increase the number of computing devices used in the prefill phase. Based on the above memory resource requirement information and computing resource requirement information, when the resource utilization rate is the highest, the resource configuration information of the language task processing model in the prefill phase and the decoding phase is determined. For the convenience of description, the resource configuration information in the prefill phase and the decoding phase are respectively defined as prefill resource configuration information and decoding resource configuration information. The prefill resource configuration information and the decoding resource configuration information at least include the number of computing devices used respectively and whether each computing device is deployed across nodes or on the same node. The token batch length is the batch size processed at one time during the decoding phase. Based on the limitation of the key-value memory, the batch length when the resource utilization rate of the decoding phase is the highest can be determined according to the prefill resource configuration information and the decoding resource configuration information.
[0022] S202: From each task processing request in the current request batch, obtain the target request segments with the number matching the prefill resource configuration information, and perform prefill processing on each target request segment in parallel to generate the current token batch; in the current request batch, generate multiple new token batches by obtaining the next token of each token in the token batch newly generated at the current moment and forming a new token batch with each next token, so as to meet the condition that the combined length of the current token batch and each new token batch reaches the token batch length.
[0023] In the present invention, when a task processing request input by a user is received, task processing requests of multiple users or multiple task processing requests of the same user are combined into one batch, which is defined as a request batch for convenience of description. The currently being processed request batch is defined as the current request batch. Before prefill calculation is performed on the current request batch, each task processing request of the current request batch passes through the encoding function of the tokenizer of the language task processing model to convert each text character into a corresponding number, obtaining a sequence of numbers. These sequences of numbers are combined into one batch, which can be defined as a prompt batch. Prefill calculation is performed using the language task processing model to predict the next text character for each number in the prompt batch. The predicted text character is a token, which is defined as a prompt token for convenience of distinction. To improve the resource utilization rate in the prefill stage, the present invention is implemented in a model parallel manner, that is, the language task processing model is sliced into several sub-models according to the number of model layers, and each sub-model is deployed to each computing device used in the prefill stage. The memory capacity of each computing device meets the memory occupancy of each sub-model. This method can reduce the video memory occupancy of the model weights on each computing device, but may cause pipeline bubble problems. To reduce the pipeline bubbles generated in the prefill stage, in the present invention, the same length of content is obtained from all task processing requests in the current request batch each time to form a prompt batch, and the selected content is defined as the target request segment, that is, the lengths of all target request segments are the same and are partial content of the corresponding task processing requests. Then, the prompt batch is scheduled and processed in a static batch processing manner. The so-called static batch processing requires that all requests in the batch end after all calculations are completed. After prefill calculation is performed on the prompt batch, corresponding tokens are generated for each prompt prediction in the prompt batch, and the tokens of multiple prompts form a token batch, which is defined as the current token batch for convenience of description. To improve the resource utilization rate in the decoding stage, the batch size and / or the number of batches processed at one time in the decoding stage are increased. After determining the length of the token batch in S201, if the length of the current token batch does not reach the length of the token batch, multiple known tokens are selected from the request batch for combination to reach the length of the token batch. The selection method of the known tokens is as follows: whether the next token of each prompt token in the current token batch is in the current request batch. If it is, it is selected to form a new token batch at the current moment. If the lengths of these two token batches still do not reach the length of the token batch, the new token batch is used as the newly generated token batch at the current moment, and the above steps are repeated until multiple token batches that meet the length of the token batch are obtained.
[0024] S203: Construct multiple pipelines using the decoding resource configuration information, perform decoding processing on each current token batch in parallel through the multiple pipelines, and obtain the corresponding language task processing result according to the decoding results of all request segments of each task processing request.
[0025] Batch-combine the multiple tokens obtained in the previous step into one batch. The current token batch refers to the current token batch whose combined token batch length can reach the token batch length determined by S201 and the newly generated token batch through the method of S202. Input these token batches into the language task processing model for decoding calculation. During the decoding calculation process, determine the number of pipelines according to the decoding resource configuration information. Considering the synchronization overhead required for each part of the model to execute different token batches, the number of pipelines cannot be too large, and its maximum value can be the same as the token batch length. In the decoding stage of the present invention, the model parallelism method is adopted, that is, the language task processing model is split into multiple sub-models, and each sub-model is deployed to each computing device used in the decoding stage. In order to achieve pipeline concurrency in the decoding stage, each sub-model is further split based on the number of model layers again, which can be defined as a secondary sub-model. The so-called pipeline parallelism means that the execution of each secondary sub-model on the same pipeline is serial, and the execution of each secondary sub-model on different pipelines is asynchronous, but synchronization is required between the same secondary sub-models. After decoding and calculating each current token batch and the key-value cache, the next token will be predicted and generated for each current token batch. Repeat this process continuously until all tokens are generated, and convert these tokens into corresponding text characters. These text characters are the task processing results corresponding to the corresponding task processing requests.
[0026] In the technical solution provided in this embodiment, deploying the decoding stage and the pre-filling stage of the language task processing model during the execution of the language task on different devices not only helps improve resource utilization, but also reduces resource coupling and improves the stability of the pre-filling delay. In the pre-filling stage, multiple task processing requests are combined into one request batch, and multiple target request segments in the request batch are pre-filled in parallel, which can not only increase the total number of tokens output per second in the pre-filling stage, improve the throughput of the pre-filling stage, and improve the resource utilization rate of the pre-filling stage, but also effectively increase the data processing volume in the decoding stage by increasing the throughput of the pre-filling stage, thereby helping to reduce the problem of under-utilized resources in the decoding stage, and can also effectively reduce the pipeline bubbles generated in the pre-filling stage. Further, in the decoding stage, the token batch generated in the pre-filling stage and multiple known token batches are combined, and these multiple token batches are executed in pipeline concurrency, increasing the number of batches and batch sizes that can be processed in the decoding stage, so as to effectively improve the resource utilization rate of the decoding stage and improve the language task processing performance without significantly increasing the memory.
[0027] After the above embodiments complete the scheduling, token batch combination, and scheduling of the request batch through S201 - S203, there are no specific limitations on the inference calculation of the language task processing model in the pre-filling stage and the decoding stage. The present invention also provides an implementation method for the inference calculation of the language task processing model, which may include the following content: The language task processing model at least includes an encoder, an embedding layer, multiple feature extraction layers, an output layer, and a decoder; the encoder is used to convert the prompt batch into numbers through the encoding function of the tokenizer and output a sequence of numbers; the embedding layer is used to receive the sequence of numbers after the conversion of the prompt batch and convert the sequence of numbers into a floating-point vector. The feature extraction layer includes multiple multi-head attention layers and a two-layer fully connected layer structure, and the outputs of both are connected with a residual structure to capture the relationships between words at different positions in the floating-point vector for the modeling and processing of the sequence of numbers. The output layer is used to perform linear transformation and normalization processing on the output of the feature extraction layer, and it includes a linear layer and a softmax layer. The decoder is used to convert the normalized data sequence into text characters.
[0028] In this embodiment, as Figure 4 shown, first, the text characters in the prompt batch formed from each task processing request are converted into numbers through the encoding function of the tokenizer in the encoder, and the generated sequence of numbers is converted into a floating-point vector through the operation of the embedding layer. The floating-point vector is input to the feature extraction layer. The feature extraction layer can adopt a transformer network model, for example. Multiple feature extraction layers perform calculations in parallel. The output of the last feature extraction layer is subjected to linear layer transformation and softmax normalization activation calculation to obtain a probability vector. Each dimension of the vector represents the probability value of predicting each vocabulary in the tokenizer. The subscript of the maximum value of the probability vector is taken as the predicted vocabulary value, and the predicted value is input to the decoding function of the tokenizer to output the predicted next character. That is, the predicted vocabulary value generated by the feature extraction layer is converted into text characters through the decoding scheme of the tokenizer in the decoder.
[0029] Among them, for a language task processing model including N feature extraction layers, the input of the first feature extraction layer is the output of the embedding layer, and the dimension is , where represents the batch size of the input text (the number of requests), that is, the number of prompt batches, represents the number of tokens in the prompt, represents the dimension of the output vector of the embedding layer, and the output and input dimensions of each feature extraction layer are the same. As Figure 4 shown, the data processing flow of each feature extraction layer includes: first, calculate the key (K) matrix, value (V) matrix, and query (Q) matrix through three linear layer transformations, and then perform multi-head dimensional splitting on the key matrix, value, and query matrix, and convert the dimension to , represents the dimension of the head. Multiply the key matrix and the value matrix, then add the mask matrix, and based on the dimension Perform softmax normalization to obtain a score matrix; then, multiply the score matrix by the query matrix, and finally transform the output dimension through a linear layer transformation to Finally, perform a linear transformation on the output of the multi-head attention layer through two fully connected layers.
[0030] As can be seen from the above, the language task processing model in this embodiment can extract richer features by adopting the multi-head attention mechanism during the feature extraction process, improve the semantic understanding ability of the language task processing model, improve the processing accuracy of the language task, and is conducive to improving the language task processing efficiency through the parallel computing ability of the language task processing model.
[0031] The above embodiment deploys the language task processing model through prefill-decoding split parallelism and model parallelism. This embodiment also gives an exemplary determination method for prefill resource configuration information and decoding resource configuration information, which may include the following content: When the sum of the double memory occupancy demand, activation memory occupancy, and the memory occupancy demand of the key-value cache corresponding to the maximum generation length of the language task processing model is greater than or equal to the total memory resources of a single host node, different host nodes are used to process the prefill stage and the decoding stage; when the sum of the double memory occupancy demand, activation memory occupancy, and the memory occupancy demand of the key-value cache corresponding to the maximum generation length of the language task processing model is less than the total memory resources of a single host node, the language task processing model runs on a single host node.
[0032] In this embodiment, whether the total occupancy of the inference memory of the language task processing model, and the computing devices used in the prefill stage and the decoding stage are deployed within a single node or different nodes, the total memory of the computing devices used in the decoding stage and the occupancy of the key-value cache will affect the maximum number of token batches supported by the decoding process, and the number of computing devices used in the prefill stage affects the number of token batches output by the prefill calculation, which in turn affects the resource utilization of the decoding stage. When the types of computing devices used in the prefill stage and the decoding stage are clear, the total memory of a single computing device can be known. Then, the double of the memory occupancy of the language task processing model, the memory occupancy of the intermediate activation calculation, and the total occupancy of the key-value cache corresponding to the maximum generation length involved in various language task processing scenarios can be calculated first, and the total memory of a single node can be calculated. After obtaining these information, compare the sum of the double memory occupancy demand, activation memory occupancy, and the memory occupancy demand of the key-value cache corresponding to the maximum generation length of the language task processing model with the total memory resources of a single host node, and accordingly determine whether to deploy on a single node or across nodes.
[0033] After determining the deployment architecture of the language task processing model, it is also necessary to determine the computing devices used in the pre-filling stage and the decoding stage respectively: obtain the pre-filling throughput in the pre-filling stage under different pre-filling resource configuration information, and the decoding throughput in the decoding stage under different decoding resource configuration information; when the pre-filling throughput and the decoding throughput meet the preset same or similar conditions, the corresponding target pre-filling resource configuration information and target decoding resource configuration information are used as the optimal resource configuration information of the language task processing model in the pre-filling stage and the decoding stage; when the pre-filling throughput and the decoding throughput cannot meet the preset same or similar conditions, and the pre-filling throughput is less than the decoding throughput, the total amount of pre-filling resources configured in the pre-filling resource configuration information should be greater than the total amount of decoding resources configured in the decoding resource configuration information, but it is necessary to meet the condition that the sum of the total amount of pre-filling resources and the total amount of decoding resources is less than or equal to the maximum value of the total resources.
[0034] In this embodiment, in order to determine the optimal resource configuration for the pre-filling stage and the decoding stage, different numbers of computing devices can be configured for the pre-filling stage in advance according to the total number of computing devices that can be deployed in parallel in the language task processing model, and the pre-filling maximum throughput corresponding to the different computing devices deployed in the pre-filling stage is analyzed, that is, the maximum number of word units output per second. Similarly, different numbers of computing devices can be configured for the decoding stage according to the total number of computing devices that can be deployed in parallel in the language task processing model, and the decoding maximum throughput corresponding to the different computing devices deployed in the decoding stage is analyzed, that is, the maximum number of word units output per second. Taking the computing device as a GPU, and the resource configuration information of the pre-filling stage and the decoding stage as the number of GPUs as an example, according to the number of GPUs that can be deployed in parallel in the model, different numbers of GPUs are configured for the pre-filling stage, and the pre-filling maximum throughput corresponding to the different numbers of GPUs deployed in the pre-filling stage is analyzed; according to the number of GPUs that can be deployed in parallel in the model, different numbers of GPUs are configured for the decoding stage, and the decoding maximum throughput corresponding to the different numbers of GPUs deployed in the decoding stage is analyzed. Through analysis, it can be determined that when the throughput of the pre-filling stage, that is, the pre-filling throughput, is close to the throughput of the decoding stage, that is, the decoding throughput, it is optimal. At this time, the resource utilization rate of the decoding stage is the highest, and this configuration is the optimal configuration, that is, the target pre-filling resource configuration information and the target decoding resource configuration information are obtained. For example, when the pre-filling throughput is close to the decoding throughput, the number of configured pre-filling GPUs and decoding GPUs is optimal, and the optimal configuration can improve the resource utilization rate of the GPU in the decoding stage. Among them, the preset same similarity condition of this embodiment is used to indicate that the pre-filling throughput is close to the decoding throughput. The condition can be flexibly set according to the actual situation. If the difference between the two is less than or equal to 10, it is considered to be close. Due to the limitations of the total number of computing devices and the performance configuration of the computing devices in the actual deployment, the number of computing devices deployed in the two stages is difficult to meet the optimal conditions under which the pre-filling throughput is close to the decoding throughput. Further, it can also be determined in combination with the total number of computing devices and the throughput of the two stages: when the throughput of the pre-filling stage is less than the throughput of the decoding stage, it will cause unsaturated resource utilization in the decoding stage, and the number of computing devices such as GPU deployment used in the pre-filling stage should be increased. Due to the memory limitation of the key-value cache in the decoding stage, the throughput of the pre-filling stage is greater than the throughput of the decoding stage, which will also lead to unsaturated resource utilization in the decoding stage. The word batches in the decoding stage can be scheduled through S203, and the word batches can be appropriately increased to increase the batch size processed in the decoding stage.
[0035] In the process of determining the resource allocation for the pre-filling stage and the decoding stage, not only the maximum and minimum numbers of deployable computing devices such as GPUs in actual deployment need to be considered, but also the minimum number of computing devices used in each of the pre-filling stage and the decoding stage needs to be determined to ensure the smooth execution of the task. In this embodiment, the computing devices used in the pre-filling stage are defined as pre-filling processors, and the computing devices used in the decoding stage are defined as decoding processors. According to the memory occupancy requirement of the language task processing model and the memory occupancy requirement of the key-value cache corresponding to the maximum generation length, the minimum memory occupancy is determined; according to the minimum memory occupancy and the total memory of the pre-filling processors, the minimum number of pre-filling processors used by the language task processing model in the pre-filling stage is determined; according to the minimum memory occupancy and the total memory of the decoding processors, the minimum number of decoding processors used by the language task processing model in the decoding stage is determined. In this embodiment, according to the total memory of a single computing device and the memory occupancy of a single model, the minimum number of model parallel splits is determined, that is, the minimum number of computing devices used. The minimum number of processors required for each of the pre-filling and decoding stages is: the minimum number of model parallel splits should be at least greater than the total memory occupancy of a single model and the memory occupancy of the key-value cache corresponding to the maximum generation length divided by the total memory of a single computing device such as a GPU.
[0036] As can be seen from the above, before deploying the inference calculation of the language task processing model in the pre-filling-decoding split parallel and model parallel modes, by determining the number of splits for model parallel deployment, the number of computing devices deployed in the pre-filling stage, the number of computing devices deployed in the decoding stage, and the selection of single-node or cross-node deployment methods for each of the pre-filling and decoding stages as resource configuration information, the resource utilization rate in the inference calculation process of the language task processing model can be effectively improved, and the problems of GPU resource utilization and inference latency can be effectively improved.
[0037] The above embodiments do not make any limitations on how to form a prompt batch. The present invention also provides a schematic implementation manner, which may include: obtaining the number of tokens of each task processing request, and selecting the smallest number of tokens therefrom as the truncation length; respectively obtaining the first n target tokens from each task processing request as the target request segment of each task processing request; the lengths of the target tokens composed of the target tokens of the same task processing request are all the same as the truncation length. The task processing request is a request containing a prompt, the prompt is composed of multiple tokens, and multiple task processing requests form a batch, that is, a request batch. Based on the truncation length, the token lengths of each task processing request in the request batch are truncated, and the truncation length is the minimum value of the number of tokens of each task processing request in the request batch. Taking Figure 5 as an example, the request batch includes 4 task processing requests, and each task processing request can be respectively expressed as: , , , , where a, b, c, and d respectively represent the tokens of each task processing request. The number of tokens of the first task processing request is the smallest, and the truncation length can be taken as 4. The first 4 tokens are obtained from each task processing request to form a prompt batch. That is, the prompt batch can be expressed as: .
[0038] Based on the above embodiments, the present invention also provides a more convenient batch request scheduling method. The batch request scheduling includes constructing a prompt batch in the request batch and scheduling it to the computing device corresponding to the pre-filling stage for pre-filling calculation, and also includes merging the token batches and scheduling them to the computing device corresponding to the decoding stage for decoding calculation, which may include the following: A batch processing queue is pre-constructed. The batch processing queue includes at least an undecoded batch queue and a decoded batch queue. In an exemplary implementation, the batch processing queue may also include a decoded batch queue. When the current request batch is received, each task processing request of the current request batch is moved into the undecoded batch queue. The first batch of target request segments with the number matching the pre-filling resource configuration information is obtained from the undecoded batch queue and moved into the decoded batch queue to perform pre-filling processing on each first batch of target request segments in the decoded batch queue in parallel. When the first batch of token batches corresponding to the first batch of target request segments is generated, if the first batch of token batches includes the first type of target tokens with end identification information, the target task processing requests corresponding to the first type of target tokens are completed. For the second type of target tokens without end identification information, when the second type of target tokens do not exist in the undecoded batch queue and the decoded batch queue, they are added to the corresponding request batch in the undecoded batch queue. If there are task processing requests of the current request batch in the undecoded batch queue, the second batch of target request segments with the number matching the pre-filling resource configuration information is obtained from the undecoded batch queue again and moved into the decoded batch queue. If there are no task processing requests of the current request batch in the undecoded batch queue, the token batches of the current request batch are merged based on the token batch length. When the batch processing queue also includes a decoded batch queue, correspondingly, when the first batch of token batches corresponding to the first batch of target request segments is generated, the first batch of target request segments is moved into the decoded batch queue. If the first batch of token batches includes the first type of target tokens with end identification information, the request segments of the task processing requests corresponding to the first type of target tokens are deleted from the decoded batch queue.
[0039] In this embodiment, based on the undecoded batch queue, the decoding batch queue, and the decoded batch queue, the task processing request for the user side, the word batch request output in the pre-filling stage and the decoding stage are managed and scheduled to the corresponding computing devices for calculation. Among them, the undecoded batch queue is used to store the batch data information that has not been scheduled, the decoding queue is used to store the batch data information that has been scheduled but has not completed the reasoning calculation, and the decoded batch queue is used to store the batch data information that has completed the reasoning calculation. In order to manage various types of batch data through batch queues, the data contained in the request batch, word batch, and prompt batch is defined as request information, and each request information can use a queue with a triple data structure. Description, each request information is represented as , Represents each word in the request information, The subscript of each word in the request information, that is, the word position. Each batch of request batches, word batches, and prompt batches is queued Description, batches are represented as , For each request message in the batch, Indicates the index of each request message in the batch. Use To describe, the batch information queue can be expressed as , Represents the index of each batch request. The undecoded batch queue, decoding batch queue, and decoded batch queue can be represented as , , .
[0040] When a task processing request is received, the task processing request and the current request batch Move to undecoded batch queue ; Calculate request batch The minimum number of tokens in each request , will request a batch Before each task processes the request words are not in the decoded batch queue Move to the decoding batch queue . Decoding batch queue All requests are waiting to be scheduled for execution. Each batch of requests is scheduled and executed in a static batch processing mode, that is, it is required to end after all requests in the batch are calculated. This is applicable to batch requests with the same length. When the first prompt batch is entered, the batch queue in the decoding The prompt batch information is sent to the pre-filling stage and the pre-filling calculation is performed on the corresponding computing device, such as Figure 5As shown, after the pre-filling calculation is completed, the current token batch is obtained . After the pre-filling calculation of the prompt batch or the decoding calculation of the token batch is completed, token batch information will be generated. The tokens in the input batch request corresponding to the token batch are moved from the decoding middle batch queue to the decoded batch queue ; then, it is judged whether there is an EOS end symbol in the tokens of the token batch. If so, the input batch request corresponding to the token batch is removed from the decoded batch queue , and a feedback message indicating that the task is completed is sent to the client. If the token is not an EOS end symbol, it is searched in the undecoded batch queue and the decoding middle batch queue . If it is not found in both, it is added to the request of the corresponding batch in the undecoded batch queue . Each batch in the undecoded batch queue is traversed to judge whether the current batch exists in the decoding middle batch queue . If not, the tokens in the undecoded batch queue are scheduled. If it exists, the next round of request batch scheduling starts in the above manner
[0041] The above embodiments do not make any limitations on the merging of the current token batch. The present invention also provides an exemplary implementation method: if the next token of the target prompt token of the current token batch is a known token in the current request batch, a first token batch is generated according to the known token corresponding to the target prompt token; if the combined length of the first token batch and the current token batch is less than the token batch length, the known token corresponding to the first token batch in the current request batch is obtained, and a second token batch is generated; if the combined length of the first token batch, the current token batch and the second token batch is the token batch length, the first token batch, the current token batch and the second token batch are decoded in parallel by using multiple pipelines. Taking Figure 5 as an example, the current token batch is , the token batch length is 3, that is, three current token batches need to be merged. The next token of the first target prompt token of the current token batch exists as a known token in the request batch, the next token of the second target prompt token exists as a known token in the request batch, and the next token of the third target prompt token exists as a known token in the request batch, and the first token batch is generated. The combined length of the first token batch and the current token batch is 2, so the first token batch is used as the current token batch to repeat the above process, and the second token batch , the combined length of the first word batch, the current word batch and the second word batch is 3, then the whole process ends. If the combined length of the first word batch, the current word batch and the second word batch is not equal to the word batch length, then the second word batch is used. Repeat the above process for the current word batch.
[0042] Based on the batch queue constructed above, this embodiment also provides a scheduling implementation method for implementing word batches based on the batch queue, which may include the following contents: If the current request batch is in the undecoded batch queue and decoding batch queue , for example, can traverse the undecoded batch queue Each batch of the current request batch Whether to batch queue during decoding If it does not exist, the current batch will be scheduled. The word unit scheduling process includes: extracting the undecoded batch queue The current batch corresponds to the first token of all requests in the request queue, forming a token batch, which can be expressed as , and the corresponding word is not decoded from the batch queue Move into the decoding batch queue Since the number of tokens in different requests is different, the next token of some requests is known. The above token batch merging method is used to continue to schedule the known next token in each request of the current batch to form a token batch. For example, the next token can be continuously scheduled in the undecoded batch queue. Extract the first token of all requests in the current batch to form a merged token batch, which can be expressed as , and the corresponding word is sent to the undecoded batch queue Move into the decoding batch queue Repeat this process to generate merged word batches until the number of merged word batches reaches the word batch length or all requested words in the current batch are empty. If the number of merged word batches in the previous step does not meet the set value, continue to traverse the undecoded batch queue Repeat the scheduling process of the previous step until the number of merged word batches meets the set value or all requested word units are empty, then exit the word unit batch request scheduling; finally, after completing the batch request scheduling, the decoding batch queue The word batches in are sent to the decoding GPU for decoding calculation. Figure 5 Examples include Figure 6 and Figure 7For the scheduling of each token batch request, the token batch scheduling of batch1 (batch 1) generates three token batches, and schedules the three token batches to three pipelines in the decoding stage, namely Stream 1, Stream 2, and Stream 3. Each token batch is executed on one stream, and the decoding operations of the token batches on the three streams are executed concurrently through the pipelines, indirectly increasing the number of token batches. Moreover, the token batches of different prompt batches, such as Figure 6 the remaining part of batch 1 and the first token of batch 2 can also be combined into one token batch. Figure 7 Batch 1, batch 2, and batch 3 of Figure 7 are combined into one token batch, thereby improving the resource utilization in the decoding stage.
[0043] As can be seen from the above, in this embodiment, the scheduling of request batches and token batches is implemented through a batch queue, which can simplify the processing complexity of batch requests, improve the scheduling efficiency of request batches and token batches, and is beneficial to improving the processing efficiency of language tasks. Further, the resource utilization rate can be further improved by combining the token batches of different batches.
[0044] The above embodiment does not make any limitation on how to perform decoding calculations in parallel through multiple pipelines. Based on the above embodiment, the present invention also gives an exemplary implementation manner, which may include the following content: According to the number of model layers of the language task processing model, each sub-model in the decoding stage where the language task processing model is deployed is respectively divided into multiple sub-sub-models, and the number of sub-sub-models of each sub-model is the same as the value of the token batch length; multiple pipelines are constructed, and each pipeline serially executes each sub-sub-model; the number of pipelines is the same as the value of the token batch length; through the first pipeline, each sub-sub-model is sequentially used to perform decoding processing on the first current token batch; when the first pipeline completes the decoding processing of the first current token batch using the first sub-sub-model, during the process of the first pipeline using the second sub-sub-model to decode the first current token batch, the second pipeline uses the target sub-sub-model corresponding to the first sub-sub-model to perform decoding processing on the second current token batch; when the second pipeline completes the decoding processing of the second current token batch using the target sub-sub-model, during the process of the second pipeline using the next sub-sub-model to decode the second current token batch, the third pipeline uses the sub-sub-model corresponding to the target sub-sub-model to perform decoding processing on the third current token batch. Of course, if the constructed pipelines also include the fourth pipeline, the fifth pipeline, the sixth pipeline, etc., the fourth pipeline, the fifth pipeline, the sixth pipeline, etc. perform the corresponding decoding processing in the same way that the execution of each sub-sub-model on the same stream is serial, the execution of sub-sub-models on different streams is asynchronous but synchronization is required between the same sub-sub-models.
[0045] It can be understood that the lengths of the original request batches are different. After the current token batch generated by the pre-padding calculation, when merging the tokens, the next tokens of some tokens in the token batch are known and do not need to be calculated and generated by the language task processing model. Multiple merged batches can be generated through the token batch merging scheduling method, and then concurrent pipelined execution of multiple token batches can be initiated in the decoding stage. However, the key-value cache data of the next token still needs to be obtained through the calculation of the previous token batch. Since the key-value cache of the tokens is obtained through layer-by-layer calculation of the model, considering the key-value data dependency relationship between the tokens of the two consecutive merged token batches, the model is divided into several parts based on the model layer. According to the key-value cache dependency relationship between the previous and subsequent token batches, the decoding calculations of each token batch are sent sequentially. Different token batches execute different parts of the model at the same time, enabling each token batch to execute concurrently in a pipeline. To ensure the correct dependencies of the previous and subsequent token batches in each part of the calculation model, synchronization is required when the previous and subsequent token batches calculate each part of the model. The number of model partitions and the number of merged token batches, that is, the length of the token batch, are the same. Since the number of model partitions is the same as the number of token batches, and the number of token batches is related to the token length in the batch request and the model partition, considering the synchronization overhead required for each token batch to execute each part of the model, the length of the token batch, that is, the number of pipeline stages, cannot be too large.
[0046] As Figure 8 shown, taking the computing device as the GPU and the token batch length as 3 as an example, after splitting the language task processing model into multiple sub-models using model parallelism, in order to achieve pipeline concurrency on the decoding GPU, the sub-models on the GPU are split into 3 sub-sub-models based on the number of model layers. For ease of description, the current token batch generated in the pre-padding stage is defined as the original token batch, and the token batch generated by merging the known tokens is defined as the merged token batch. The original token batch is serially executed on the first stream for three sub-sub-models, the first merged token batch is serially executed on the second stream for three sub-sub-models, and the second merged token batch is serially executed on the third stream for three sub-sub-models. However, the execution of each sub-sub-model on each stream needs to be synchronized with the execution of the corresponding sub-sub-model on the previous stream. That is, the execution of the first, second, and third sub-sub-models on the second stream needs to wait for the execution of the first, second, and third sub-sub-models on the first stream to complete respectively, and the execution of the first, second, and third sub-sub-models on the third stream needs to wait for the execution of the first, second, and third sub-sub-models on the second stream to complete respectively. That is to say, the execution of each sub-sub-model on the same stream is serial, and the execution of the sub-sub-models on different streams is asynchronous but synchronization is required between the same sub-sub-models.
[0047] As can be seen from the above, in this embodiment, by keeping the number of pipelines consistent with the token batch length and further dividing the sub-model in the decoding stage into multiple sub-sub-models with the same number as the token batch length, the correct dependencies of the front and back token batches on various parts of the computing model are ensured, which is beneficial to improving the resource utilization rate in the decoding stage.
[0048] Through the description of the above embodiments, those skilled in the art can clearly understand that the method according to the above embodiments can be implemented by means of software plus a necessary general hardware platform. Of course, it can also be implemented by hardware, but in many cases, the former is a better implementation method. The present invention also provides a corresponding device for the language task processing method, further making the method more practical. Among them, the device can be described from the perspective of functional modules and the perspective of hardware respectively. The language task processing device provided by the present invention will be introduced below, and this device is used to implement the language task processing method provided by the present invention. The description of the features in the corresponding embodiments of the language task processing device can refer to the relevant descriptions of the corresponding embodiments of the language task processing method, and the device embodiments will not be elaborated one by one.
[0049] First, from the perspective of functional modules, please refer to Figure 9 , Figure 9 which is a structural diagram of the language task processing device provided in this embodiment in a specific implementation manner. The device may include: A resource configuration module 901, configured to determine the resource configuration information of the language task processing model in the pre-filling stage and the decoding stage respectively according to the resource requirement information of the language task processing model during the execution of the language task, and determine the token batch length according to the pre-filling resource configuration information and the decoding resource configuration information.
[0050] A pre-filling stage processing module 902, configured to obtain, from each task processing request in the current request batch, a target request segment with a number matching the pre-filling resource configuration information, and perform pre-filling processing on each target request segment in parallel to generate the current token batch.
[0051] A decoding stage processing module 903, configured to, in the current request batch, generate multiple new token batches by obtaining the next token of each token in the latest generated token batch at the current moment and forming a new token batch with each next token, so as to meet the condition that the combined length of the current token batch and each new token batch reaches the token batch length; the lengths of each target request segment are the same and are partial contents of the corresponding task processing request; multiple pipelines are constructed by using the decoding resource configuration information, and each token batch is decoded in parallel through the multiple pipelines, and the corresponding language task processing result is obtained according to the decoding results of all request segments of each task processing request.
[0052] Exemplarily, in some embodiments of the present embodiment, the above resource configuration module 901 may further be configured to: when the sum of the double memory occupancy requirement of the language task processing model, the activation memory occupancy, and the memory occupancy requirement of the key-value cache corresponding to the maximum generation length is greater than or equal to the total memory resources of a single host node, use different host nodes to process the prefill stage and the decoding stage; when the sum of the double memory occupancy requirement of the language task processing model, the activation memory occupancy, and the memory occupancy requirement of the key-value cache corresponding to the maximum generation length is less than the total memory resources of a single host node, the language task processing model runs on a single host node.
[0053] Exemplarily, in some other embodiments of the present embodiment, the above resource configuration module 901 may further be configured to: obtain the prefill throughput of the prefill stage under different prefill resource configuration information, and the decoding throughput of the decoding stage under different decoding resource configuration information; when the prefill throughput and the decoding throughput meet the preset same or similar conditions, respectively corresponding target prefill resource configuration information and target decoding resource configuration information are used as the optimal resource configuration information of the language task processing model in the prefill stage and the decoding stage; when the prefill throughput and the decoding throughput cannot meet the preset same or similar conditions, and the prefill throughput is less than the decoding throughput, the total amount of prefill resources configured for the prefill resource configuration information is greater than the total amount of decoding resources configured for the decoding resource configuration information; the sum of the total amount of prefill resources and the total amount of decoding resources is less than or equal to the maximum value of the total resources.
[0054] Exemplarily, in some other embodiments of the present embodiment, the above resource configuration module 901 may further be configured to: determine the minimum memory occupancy according to the memory occupancy requirement of the language task processing model and the memory occupancy requirement of the key-value cache corresponding to the maximum generation length; determine the minimum number of prefill processors used by the language task processing model in the prefill stage according to the minimum memory occupancy and the total memory owned by the prefill processors; determine the minimum number of decoding processors used by the language task processing model in the decoding stage according to the minimum memory occupancy and the total memory owned by the decoding processors.
[0055] Exemplarily, in some other embodiments of the present embodiment, the above pre-filling stage processing module 902 may also be used to: obtain the number of tokens of each task processing request, and select the smallest number of tokens therefrom as the truncation length; respectively obtain the first n target tokens from each task processing request as the target request segment of each task processing request; the lengths of the target tokens of the same task processing request are all the same as the truncation length. Exemplarily, in some other embodiments of the present embodiment, the above pre-filling stage processing module 902 may also be used to: pre-construct a batch processing queue, the batch processing queue includes an undecoded batch queue and a decoded batch queue; when receiving the current request batch, move the task processing requests of the current request batch into the undecoded batch queue, obtain the first batch of target request segments with the number matching the pre-filling resource configuration information from the undecoded batch queue, and move the first batch of target request segments into the decoded batch queue to perform pre-filling processing on each of the first batch of target request segments in the decoded batch queue in parallel; when generating the first batch of token batches corresponding to the first batch of target request segments, if the first batch of token batches includes the first type of target tokens with end identification information, the target task processing request corresponding to the first type of target tokens is completed. For the second type of target tokens without end identification information, when the second type of target tokens do not exist in the undecoded batch queue and the decoded batch queue, add them to the corresponding request batch in the undecoded batch queue; if there are task processing requests of the current request batch in the undecoded batch queue, obtain the second batch of target request segments with the number matching the pre-filling resource configuration information from the undecoded batch queue again, and move the second batch of target request segments into the decoded batch queue; if there are no task processing requests of the current request batch in the undecoded batch queue, merge the token batches of the current request batch based on the token batch length.
[0056] As an exemplary implementation manner of the above embodiment, the above pre-filling stage processing module 902 may also be used to: the batch processing queue further includes a decoded batch queue; when generating the first batch of token batches corresponding to the first batch of target request segments, move the first batch of target request segments into the decoded batch queue, and if the first batch of token batches includes the first type of target tokens with end identification information, delete the request segment of the task processing request corresponding to the first type of target tokens from the decoded batch queue.
[0057] Exemplarily, in some other embodiments of this embodiment, the above decoding stage processing module 903 may also be used to: if the next token of the target prompt token in the current token batch is a known token in the current request batch, generate a first token batch according to the known token corresponding to the target prompt token; if the combined length of the first token batch and the current token batch is less than the token batch length, obtain the known tokens corresponding to the first token batch in the current request batch, and generate a second token batch; if the combined length of the first token batch, the current token batch, and the second token batch is the token batch length, use multiple pipelines to perform decoding processing on the first token batch, the current token batch, and the second token batch in parallel.
[0058] Exemplarily, in some other embodiments of this embodiment, the above decoding stage processing module 903 may also be used to: according to the number of model layers of the language task processing model, respectively divide each sub-model in the decoding stage where the language task processing model is deployed into multiple sub-sub-models, and the number of sub-sub-models of each sub-model is the same as the value of the token batch length; construct multiple pipelines, and each pipeline serially executes each sub-sub-model; the number of pipelines is the same as the value of the token batch length; through the first pipeline, use each sub-sub-model to perform decoding processing on the first current token batch in sequence; when the first pipeline completes the decoding processing of the first current token batch using the first sub-sub-model, during the process of the first pipeline using the second sub-sub-model to decode the first current token batch, use the second pipeline to use the target sub-sub-model corresponding to the first sub-sub-model to perform decoding processing on the second current token batch; when the second pipeline completes the decoding processing of the second current token batch using the target sub-sub-model, during the process of the second pipeline using the next sub-sub-model to decode the second current token batch, use the third pipeline to use the sub-sub-model corresponding to the target sub-sub-model to perform decoding processing on the third current token batch.
[0059] The language task processing device mentioned above is described from the perspective of functional modules. Further, the present invention also provides an electronic device, which is described from the perspective of hardware. The electronic device includes a memory and a processor. A computer program is stored in the memory, and the processor is configured to run the computer program to execute the steps in any of the above embodiments of the language task processing method.
[0060] An embodiment of the present application also provides a computer-readable storage medium. A computer program is stored in the computer-readable storage medium, wherein the computer program is configured to execute the steps in any of the above embodiments of the language task processing method when running. In an exemplary embodiment, the above computer-readable storage medium may include, but is not limited to: USB flash drives, ROM (Read-Only Memory), RAM (Random Access Memory), mobile hard disks, magnetic disks, or optical discs and other various media that can store computer programs.
[0061] An embodiment of the present application further provides a computer program product. The computer program product includes a computer program, and when the computer program is executed by a processor, the steps in any of the above language task processing method embodiments are implemented.
[0062] An embodiment of the present application further provides another computer program product, including a non-volatile computer-readable storage medium. The non-volatile computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the steps in any of the above language task processing method embodiments are implemented.
[0063] Finally, the present invention further provides a language task processing system. Refer to Figure 10 , the language task processing system may at least include a first processor 101 and a second processor 102. The first processor 101 is connected to the second processor 102. The resource configuration information of the first processor 101 and the second processor 102 is determined according to the resource configuration information of the language task processing model in the pre-filling stage and the decoding stage. The so-called resource configuration information is used to determine whether the second processor 102 is deployed on the same node or different nodes. The node is the processor included in the first processor 101, and to determine the number of processors included in the second processor 102. The first processor 101 obtains a target request segment with a number matching the pre-filling resource configuration information from each task processing request in the current request batch, and sends the target request segment as the current request batch to the second processor 102. The second processor 102 performs pre-filling processing on each target request segment in parallel to generate the current token batch, and sends it to the first processor 101. The lengths of the target request segments are the same and are part of the corresponding task processing requests. The first processor 101 obtains the next token of each token in the current token batch from the current request batch, and generates a new current token batch in this way to obtain multiple current token batches until the length of the combined current token batch reaches the token batch length, and combines each current token batch into a token batch and sends it to the second processor 102. The second processor 102 constructs multiple pipelines using the decoding resource configuration information, and performs decoding processing on each token batch in parallel through the multiple pipelines, and sends the decoding processing results to the first processor 101. The first processor 101 generates corresponding language task processing results according to the decoding results of all request segments of each task processing request.
[0064] Exemplarily, if the total memory of the second processor 102 is greater than or equal to the sum of twice the memory occupancy requirement of the language task processing model, the activation memory occupancy, and the memory occupancy requirement of the key-value cache corresponding to the maximum generation length, then the second processor 102 is deployed within the same node. That is, there is only one first processor 101. For ease of description, the processor that performs the pre-filling calculation is defined as the pre-filling processor group, and the processor that performs the decoding calculation is defined as the decoding processor group. The second processor 102 includes a pre-filling processor group and a decoding processor group; the pre-filling processor group and the decoding processor group are respectively connected to the first processor 101; the pre-filling processor group includes multiple pre-filling processors, and each pre-filling processor is connected in sequence; the decoding processor group includes multiple decoding processors, and each decoding processor is connected in sequence; among them, the minimum value of the total number of pre-filling processors and decoding processors is determined according to the total memory of a single processor and the memory occupancy requirement of the language task processing model; the total memory of a single processor is the maximum value of the total memory of the pre-filling processors and the total memory of the decoding processors; the optimal values of the number of pre-filling processors and the number of decoding processors respectively are the number of pre-filling processors included in the pre-filling processor group and the number of decoding processors included in the decoding processor group when the throughput of the pre-filling stage and the throughput of the decoding stage meet the preset same or similar conditions; when the throughput of the pre-filling stage is less than the throughput of the decoding stage, the number of pre-filling processors is greater than the number of decoding processors, and the sum of the number of pre-filling processors and the number of decoder processors is greater than or equal to the minimum value of the total number.
[0065] Exemplarily, if the total memory of the second processor 102 is less than the sum of twice the memory occupancy requirement of the language task processing model, the activation memory occupancy, and the memory occupancy requirement of the key-value cache corresponding to the maximum generation length, then the second processor 102 is deployed in different nodes, that is, the first processor 101 includes at least 2 nodes. For ease of description, the first processor 101 may include a first host node and a second host node. The first host node is connected to the second processor 102, the second host node is connected to a third processor, and the third processor is connected to the second processor 102. The second processor 102 includes a pre-filling processor group, and the third processor includes a decoding processor group. The total memory of the pre-filling processor group is greater than twice the memory occupancy requirement of the language task processing model, and the total memory of the decoding processor group is greater than twice the memory occupancy requirement of the language task processing model. The pre-filling processor group includes multiple pre-filling processors, and each pre-filling processor is connected in sequence. The decoding processor group includes multiple decoding processors, and each decoding processor is connected in sequence. Each pre-filling processor in the pre-filling processor group communicates through a point-to-point connection method, such as through the NVlink (a point-to-point connection method), and each decoding processor in the decoding processor group communicates through a point-to-point connection method. Each pre-filling processor in the pre-filling processor group is interconnected and communicates with each decoding processor in the decoding processor group through a remote direct access method.
[0066] To make the technical solution of the present invention clearer and more understandable to those skilled in the art, the present invention also describes the language task processing system by taking the first processor as a CPU and the second processor as a GPU as an example, such as Figure 11As shown in the figure, the language task processing system includes multiple CPUs and multiple GPUs. If the memory capacity of a single GPU cannot meet the memory requirements of the language task processing model, the model parallelism method can be used to split the model inference operation and calculate it on multiple GPUs. Correspondingly, the pre-filling calculation and the decoding calculation are each deployed on multiple GPUs. The CPU and the GPU can be connected through the PCIE method. The GPUs within the same CPU node can be connected and communicate through the NVlink method. The GPUs across CPU nodes are interconnected and communicate through the remote direct memory access method. Since the bandwidth of NVlink is higher than that of the direct memory access method, the single-node internal deployment is preferentially selected when the deployment conditions are met. Among them, the functions of receiving user prompt requests, scheduling and forwarding the prompt requests and tokens in batches, and text-sequence conversion on the CPU are encapsulated as functional modules by computer programs that implement these functions and solidified on the CPU. Correspondingly, the CPU includes a request scheduling module and a text-sequence conversion module. The text-sequence conversion module completes the conversion between the character text and the numerical sequence of each token in the generation request. Initially, the input prompt text is converted into numerical values through the encode function of the tokenizer module, and the numerical values are used as the model input for pre-filling and decoding inference calculations. After all decoding is completed, the predicted numerical values are input into the decode function of the tokenizer module to convert the numbers into text characters and complete text generation. The request scheduling module schedules the prompt batch requests from the user side and the token batch requests output from the pre-filling or decoding stage. The scheduling module schedules the generation inference calculation of the large language model to each GPU by managing three queues. The GPUs are divided into pre-filling GPUs and decoding GPUs according to the inference stage of the language task processing model. Both the pre-filling GPUs and the decoding GPUs can include multiple GPUs.
[0067] Based on Figure 11For the language task processing system, when the prompt requests of several users at the application layer are input into the request scheduling module, the request scheduling module truncates all the prompts in a batch to the minimum length uniformly and adopts a static batch processing mechanism to schedule and execute the prompt batch. The size of the prompt batch is determined according to the computing power of the pre-filled GPU. The prompt batch is input into the text-sequence conversion module to convert all the text characters of the prompts into digital sequences. According to the distributed parallel mode in the pre-filled computing stage, the prompt batch sequence is input into the pre-filled GPU for pre-filled computing. After the pre-filled stage, each prompt predicts and outputs a token and the key-value cache of all tokens. The parallel mode in the pre-filled stage depends on the model memory and the number of GPUs, and model parallelism and tensor parallelism can be adopted. When the token batch generated by the pre-filled computing is input into the request scheduling module, considering the problem of under-utilized GPU resources in the decoding stage, a batch processing scheduling method of token batch merging and pipeline concurrent execution is adopted to increase the batch size processed by the decoding GPU, thereby improving the resource utilization rate of the decoding GPU. Before the token batch executes the decoding calculation, the key-value cache of the prompt tokens calculated by the previous pre-filled GPU is transmitted to the decoding GPU, and the scheduled merged token batch is input into the decoding GPU to execute the decoding calculation in a pipeline concurrent manner. Each token in the batch outputs the next token and the key-value cache after the decoding calculation; repeat the token batch scheduling and decoding calculation in the previous two steps until all the tokens in the batch complete the decoding calculation. Finally, all the tokens in the batch are input into the text-sequence conversion module to convert the digital sequence into text characters, obtaining the processing result of the task processing request and completing the inference calculation of the language task.
[0068] As can be seen from the above, this embodiment uses a distributed computing platform composed of CPU+GPU to run the language task processing model, and splits the pre-filled stage and the decoding stage to different GPUs for execution, reducing resource coupling and improving the stability of the pre-filled latency. The pre-filled stage adopts sequence truncation and static batch processing technologies to reduce pipeline bubbles. The decoding stage adopts a batch processing scheduling method based on token batch merging and pipeline concurrent execution to increase the batch size processed by the GPU, thereby improving the resource utilization rate of the decoding GPU without significantly increasing the memory.
[0069] The above has introduced in detail a language task processing method, system, electronic device, computer-readable storage medium and computer program product provided by the present invention. Each embodiment in this specification is described in a progressive manner. The key point of each embodiment is the difference from other embodiments. For the same or similar parts between the embodiments, reference can be made to each other. Whether the units and algorithm steps of each example described in the disclosed embodiments are executed in the form of electronic hardware or computer software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods for each specific application to implement the described functions, and such implementation should not be considered to exceed the scope of the present invention. Without departing from the principle of the present invention, several improvements and modifications can also be made to the present invention, and these improvements and modifications also fall within the protection scope of the present invention.
Claims
1. A language task processing method, characterized in that: include: Determining resource configuration information of the language task processing model in a pre-filling stage and a decoding stage respectively according to resource demand information of the language task processing model in the process of executing the language task, and determining the word unit batch length according to the pre-filling resource configuration information and the decoding resource configuration information; From each task processing request of the current request batch, obtain target request segments of a number matching the pre-filled resource configuration information, perform pre-filling processing on each target request segment in parallel, and generate the current word unit batch; In the current request batch, multiple new word-unit batches are generated by obtaining the next word-unit of each word-unit in the word-unit batch that is most recently generated at the current moment, and each next word-unit constitutes a new word-unit batch, so as to meet the condition that the length of the current word-unit batch and each new word-unit batch combined reaches the length of the word-unit batch; each target request segment has the same length and is a partial content of the corresponding task processing request; Multiple pipelines are constructed using the decoding resource configuration information, and each word batch is decoded in parallel through the multiple pipelines, and the corresponding language task processing results are obtained according to the decoding results of all request segments of each task processing request.
2. The language task processing method according to claim 1, characterized in that: Determining resource configuration information of the language task processing model in a pre-filling phase and a decoding phase, respectively, according to resource requirement information of the language task processing model in a process of executing the language task, including: When the sum of the double memory usage requirement, the activation memory usage, and the memory usage requirement of the key-value cache corresponding to the maximum generation length of the language task processing model is greater than or equal to the total memory resources of a single host node, different host nodes are used to process the pre-filling stage and the decoding stage; When the sum of the double memory usage requirement, the activation memory usage and the memory usage requirement of the key-value cache corresponding to the maximum generation length of the language task processing model is less than the total memory resources of a single host node, the language task processing model runs on a single host node.
3. The language task processing method according to claim 1, characterized in that: Determining resource configuration information of the language task processing model in a pre-filling phase and a decoding phase, respectively, according to resource requirement information of the language task processing model in a process of executing the language task, including: Acquire the prefilling throughput of the prefilling stage under different prefilling resource configuration information, and the decoding throughput of the decoding stage under different decoding resource configuration information; When the pre-filling throughput and the decoding throughput meet the preset same or similar conditions, the corresponding target pre-filling resource configuration information and the target decoding resource configuration information are used as the optimal resource configuration information of the language task processing model in the pre-filling stage and the decoding stage; When the pre-filling throughput and the decoding throughput cannot satisfy preset identical or similar conditions, and the pre-filling throughput is less than the decoding throughput, the total amount of pre-filling resources configured for the pre-filling resource configuration information is greater than the total amount of decoding resources configured for the decoding resource configuration information; and the sum of the total amount of pre-filling resources and the total amount of decoding resources is less than or equal to the maximum total amount of resources.
4. The language task processing method according to claim 1, characterized in that: Determining resource configuration information of the language task processing model in a pre-filling phase and a decoding phase, respectively, according to resource requirement information of the language task processing model in a process of executing the language task, including: Determine a minimum memory usage value according to the memory usage requirement of the language task processing model and the memory usage requirement of the key-value cache corresponding to the maximum generation length; Determining the minimum number of pre-filled processors used by the language task processing model in the pre-filling phase according to the minimum memory usage and the total memory owned by the pre-filled processors; The minimum number of decoding processors used by the language task processing model in the decoding stage is determined according to the minimum memory occupancy value and the total memory owned by the decoding processor.
5. The language task processing method according to claim 1, characterized in that: From each task processing request in the current request batch, obtain the target request segments that match the pre-filled resource configuration information, including: Obtain the number of tokens in each task processing request, and select the minimum number of tokens as the truncation length; The first n target word units are obtained from each task processing request respectively as the target request segment of each task processing request to form a prompt batch; the length of each target word unit of the same task processing request is the same as the truncation length.
6. The language task processing method according to claim 1, characterized in that: From each task processing request in the current request batch, obtain the target request segments that match the pre-filled resource configuration information, including: Pre-building a batch processing queue, wherein the batch processing queue includes an undecoded batch queue and a decoded batch queue; When a current request batch is received, each task processing request of the current request batch is moved into the undecoded batch queue, a first batch of target request segments of a number matching the pre-filled resource configuration information is obtained from the undecoded batch queue, and the first batch of target request segments is moved into the decoded batch queue, so as to perform pre-filling processing on each first batch of target request segments in the decoded batch queue in parallel; When the first batch of word units corresponding to the first batch of target request segments is generated, if the first batch of word units includes the first type of target word units with end identification information, the target task processing request corresponding to the first type of target word units has been completed, and for the second type of target word units without end identification information, when the second type of target word units do not exist in the undecoded batch queue and the decoded batch queue, they are added to the corresponding request batch of the undecoded batch queue; If there are task processing requests for the current request batch in the undecoded batch queue, a second batch of target request segments whose number matches the pre-filled resource configuration information is obtained from the undecoded batch queue again, and the second batch of target request segments is moved into the decoded batch queue; if there are no task processing requests for the current request batch in the undecoded batch queue, the word-meta batches of the current request batch are merged based on the word-meta batch length.
7. The language task processing method according to claim 6, characterized in that: The batch processing queue also includes a decoded batch queue; When the first batch of word units corresponding to the first batch of target request segments are generated, the corresponding first batch of target request segments are moved into the decoded batch queue. If the first batch of word units includes a first type of target word units with end identification information, the request segment of the task processing request corresponding to the first type of target word units is deleted from the decoded batch queue.
8. The language task processing method according to claim 1, characterized in that: In the current request batch, multiple new word unit batches are generated by obtaining the next word unit of each word unit in the word unit batch that is most recently generated at the current moment, and forming each next word unit into a new word unit batch, including: If the next word of the target prompt word of the current word batch is a known word in the current request batch, generating a first word batch according to the known word corresponding to the target prompt word; If the combined length of the first word unit batch and the current word unit batch is less than the word unit batch length, obtaining the known word units corresponding to the first word unit batch in the current request batch, and generating a second word unit batch; If the combined length of the first word unit batch, the current word unit batch and the second word unit batch is the word unit batch length, multiple pipelines are used to decode the first word unit batch, the current word unit batch and the second word unit batch in parallel.
9. The language task processing method according to any one of claims 1 to 8, characterized in that: Multiple pipelines are constructed using the decoding resource configuration information, and each word batch is decoded in parallel through multiple pipelines, including: According to the number of model layers of the language task processing model, each sub-model deployed in the decoding stage of the language task processing model is divided into a plurality of sub-sub-models, and the number of sub-sub-models of each sub-model is the same as the value of the word batch length; Constructing multiple pipelines, each pipeline serially executing each sub-model; the number of pipelines is the same as the value of the word batch length; Through the first pipeline, each sub-model is used in turn to perform decoding processing on the first current word batch; After the first pipeline completes the decoding process of the first current word unit batch using the first sub-model, while the first pipeline is decoding the first current word unit batch using the second sub-model, the second pipeline is decoding the second current word unit batch using the target sub-sub-model corresponding to the first sub-model; After the second pipeline completes the decoding processing of the second current word batch using the target sub-sub-model, while the second pipeline is decoding the second current word batch using the next sub-sub-model, the third pipeline decodes the third current word batch using the sub-sub-model corresponding to the target sub-sub-model.
10. An electronic device, characterized in that: include: Memory for storing computer programs; A processor, configured to implement the steps of the language task processing method according to any one of claims 1 to 9 when executing the computer program.
11. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the steps of the language task processing method according to any one of claims 1 to 9 are implemented.
12. A computer program product comprising a computer program / instructions, characterized in that When the computer program / instructions are executed by a processor, the steps of the language task processing method according to any one of claims 1 to 9 are implemented.
13. A language task processing system, characterized in that: At least comprising a first processor and a second processor, wherein the first processor is connected to the second processor; resource configuration information of the first processor and the second processor is determined according to resource configuration information of the language task processing model in the pre-filling stage and the decoding stage; The first processor obtains target request segments of a number matching the pre-filled resource configuration information from each task processing request of the current request batch, and sends the target request segments as the current request batch to the second processor, and the second processor performs pre-filling processing on each target request segment in parallel to generate a current word unit batch, and sends it to the first processor; each target request segment has the same length and is a part of the corresponding task processing request; The first processor generates a plurality of new word unit batches in the current request batch by acquiring the next word unit of each word unit in the word unit batch most recently generated at the current moment and forming a new word unit batch with each next word unit, so as to meet the condition that the length of the current word unit batch and each new word unit batch is equal to the length of the word unit batch, and merges each current word unit batch into a word unit batch and sends it to the second processor; The second processor constructs a plurality of pipelines using the decoding resource configuration information, performs decoding processing on each word unit batch in parallel through the plurality of pipelines, and sends the decoding results to the first processor; The first processor generates a corresponding language task processing result according to the decoding results of all request segments of each task processing request.
14. The language task processing system according to claim 13, characterized in that: The total memory of the second processor is greater than or equal to twice the memory requirement of the language task processing model, the activation memory requirement and the sum of the memory requirement of the key-value cache corresponding to the maximum generation length, and the second processor includes a pre-filling processor group and a decoding processor group; The pre-filling processor group and the decoding processor group are respectively connected to the first processor; the pre-filling processor group includes a plurality of pre-filling processors, each of which is connected in sequence; the decoding processor group includes a plurality of decoding processors, each of which is connected in sequence; The minimum value of the total number of pre-filled processors and decoding processors is determined according to the total memory of a single processor and the memory occupation requirement of the language task processing model; the total memory of a single processor is the maximum value of the total memory of the pre-filled processor and the total memory of the decoding processor; The optimal values of the number of pre-filling processors and the number of decoding processors are the number of pre-filling processors included in the pre-filling processor group and the number of decoding processors included in the decoding processor group when the throughput of the pre-filling stage and the throughput of the decoding stage meet the same preset similar conditions; when the throughput of the pre-filling stage is less than the throughput of the decoding stage, the number of pre-filling processors is greater than the number of decoding processors, and the sum of the number of pre-filling processors and the number of decoder processors is greater than or equal to the minimum value of the total number.
15. The language task processing system according to claim 13, characterized in that: The total memory of the second processor is less than twice the memory requirement of the language task processing model, the activation memory requirement, and the sum of the memory requirement of the key-value cache corresponding to the maximum generation length, the first processor includes a first host node and a second host node, the first host node is connected to the second processor, the second host node is connected to the third processor, and the third processor is connected to the second processor; The second processor includes a pre-filling processor group, the third processor includes a decoding processor group, the total memory of the pre-filling processor group is greater than twice the memory occupation requirement of the language task processing model, and the total memory of the decoding processor group is greater than twice the memory occupation requirement of the language task processing model; the pre-filling processor group includes a plurality of pre-filling processors, each of which is connected in sequence; the decoding processor group includes a plurality of decoding processors, each of which is connected in sequence; The pre-filling processors of the pre-filling processor group communicate through a point-to-point connection, and the decoding processors of the decoding processor group communicate through a point-to-point connection; the pre-filling processors of the pre-filling processor group are interconnected and communicated with the decoding processors of the decoding processor group respectively through a remote direct access method.
Citation Information
Patent Citations
Large language model reasoning optimization method and device, computer equipment and storage medium
CN117194056A
Large language model reasoning optimization method and device, electronic equipment and storage medium
CN119150994A
Processing method for improving batch reasoning efficiency of large language model
CN119558398A
Large language model system and request response method thereof
CN119829282A
System and computer-executable program code for accelerated rescoring with recurrent neural net language models on hybrid CPU / GPU machines using a frame-wise, delayed dispatch of RNNLM score computation tasks to the GPU(s)
US10878806B1
Cited By
Large model service-oriented task parallel processing intelligent scheduling method and system
CN120256068A
Data processing method and device, equipment and storage medium
CN120336510A
Parameter optimization method and device, electronic equipment, storage medium and computer program product
CN120875051A
Model reasoning scheduling method and system, electronic equipment and storage medium
CN120909737A
Model inference scheduling method and system, electronic device, and storage medium
CN120909737B