A method, device, medium and computer program product for generating response information
By dynamically scheduling the number of word budgets for pre-filling and decoding tasks in parallel inference calculation of pre-trained language models, the contradiction between device pressure and generation performance is solved, and the execution performance and resource utilization efficiency of artificial intelligence question-and-answer tasks are improved.
Patent Information
- Application Number
- CN202510387245.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-31
- Publication Date
- 2025-07-22
- Estimated Expiration
- 2045-03-31
AI Technical Summary
In the prior art, parallel inference computing of pre-trained language models has the problem of device pressure and generation performance, resulting in insufficient resource utilization and low generation efficiency.
By dynamically scheduling the number of word budgets of pre-filling tasks and decoding tasks based on the computing power utilization rate of the equipment in a batch of inference calculations, the proportion of word budgets of the decoding tasks is negatively correlated with the computing power utilization rate of the pre-filling tasks, and the throughput and delay equalization of parallel inference calculations is achieved.
It improves the parallel inference performance of the pre-trained language model, improves the execution performance of artificial intelligence question-and-answer tasks, solves the contradiction between device pressure and generation performance, and achieves more efficient resource utilization.
Smart Images

Figure CN119884332B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of artificial intelligence technology, and in particular, to a method, device, medium, and computer program product for generating response information. Background Art
[0002] Artificial intelligence question-and-answer technology is a booming branch of artificial intelligence technology, and its applications cover multiple technical fields such as translation, article generation, abstract generation, information search, image generation, image parsing, code generation, etc. Pre-trained language models are commonly used models in artificial intelligence question-and-answer technology. Using pre-trained language models to perform parallel inference calculations on multiple user tasks can better utilize device resources and improve inference efficiency compared to single-task execution, but still brings a huge pressure on device performance, and the efficiency of generating response information also needs to be improved.
[0003] Solving the problems of device pressure and generation efficiency in generating response information for artificial intelligence question-and-answer technology is a technical problem that needs to be solved by those skilled in the art. Summary of the Invention
[0004] The present invention provides a method, device, medium, and computer program product for generating response information, so as to at least solve the problem of the contradiction between device pressure and generation performance in related technologies.
[0005] The present invention provides a method for generating response information, including:
[0006] Receiving a to-be-processed sequence of a response task;
[0007] Inputting the to-be-processed sequence into a pre-trained language model for inference calculation;
[0008] In a batch of inference calculations, determining the token budget quantity of the pre-filling task and the token budget quantity of the decoding task according to the computing power utilization rate of the device executing the pre-filling task, and inputting the corresponding to-be-processed sequence into the pre-trained language model for parallel inference calculation according to the token budget quantity;
[0009] Concatenating the output results of the decoding task to obtain the response information of the to-be-processed sequence;
[0010] Outputting the response information;
[0011] Wherein, the proportion of the token budget quantity of the decoding task is negatively correlated with the computing power utilization rate of the pre-filling task.
[0012] The present invention also provides an electronic device, including: a memory for storing a computer program; a processor for implementing the steps of the above response information generation method when executing the computer program.
[0013] The present invention also provides a computer-readable storage medium, in which a computer program is stored. When the computer program is executed by a processor, the steps of the above-mentioned response information generation method are implemented.
[0014] The present invention also provides a computer program product, including a computer program. When the computer program is executed by a processor, the steps of the above-mentioned response information generation method are implemented.
[0015] Through the present invention, for the inference calculation of the pre-trained language model, there are two tasks of pre-filling and decoding, which are respectively calculation-intensive and memory-intensive tasks. A dynamic scheduling scheme is provided. In a batch of inference calculations, according to the computing power utilization rate of the device executing the pre-filling task, the token budget quantity of the pre-filling task and the token budget quantity of the decoding task are determined. The proportion of the token budget quantity of the decoding task is negatively correlated with the computing power utilization rate of the pre-filling task. According to the token budget quantity, the corresponding sequence to be processed is input into the pre-trained language model for parallel inference calculation, so that the parallel inference calculation obtains a balance between throughput and latency. Thus, the problem of the contradiction between device pressure and generation performance in parallel inference scheduling in the related art can be solved, and the technical effect of improving the parallel inference performance of the pre-trained language model is achieved. The output results of the decoding task are spliced to obtain the response information of the sequence to be processed and output, improving the execution performance of the artificial intelligence question-and-answer task. BRIEF DESCRIPTION OF THE DRAWINGS
[0016] In order to more clearly illustrate the embodiments of the present invention, the drawings required for the embodiments will be briefly introduced below. Obviously, the drawings in the following description are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained based on these drawings.
[0017] Figure 1 It is a schematic diagram of the parallel inference process of a pre-trained language model;
[0018] Figure 2 It is a flowchart of a response information generation method provided by an embodiment of the present invention;
[0019] Figure 3 It is a flowchart of another response information generation method provided by an embodiment of the present invention;
[0020] Figure 4 It is a schematic diagram of the structure of a batch mixing manager provided by an embodiment of the present invention;
[0021] Figure 5 It is a schematic diagram of the structure of a response information generation device provided by an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0022] The technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.
[0023] It should be noted that in the description of the present invention, the terms "comprising", "including" or any other variant thereof are intended to cover a non-exclusive inclusion, so that a process, method, article or device including a series of elements not only includes those elements but also includes other elements not expressly listed, or further includes elements inherent to such process, method, article or device. The terms "first", "second", etc. in the present invention are used to distinguish similar objects and are not used to describe a specific order or sequence.
[0024] To enable those skilled in the art of the present technology to better understand the solution of the present invention, the present invention will be further described in detail below with reference to the accompanying drawings and specific embodiments.
[0025] Some key terms used in the embodiments of the present invention will be explained here first.
[0026] A pre-trained language model (PLM) generally refers to a large-scale neural network algorithm structure and parameters obtained by designing a language model training task based on a large-scale corpus (including, for example, language training materials such as sentences and paragraphs), training a large-scale neural network algorithm structure to learn and implement. Subsequently, for other tasks, feature extraction or task fine-tuning can be performed on the basis of this model to achieve specific task purposes. The idea of pre-training is to first train a task to obtain a set of model parameters, then use this set of model parameters to initialize the network model parameters, and then use the initialized network model to train other tasks to obtain a model adapted to other tasks. By pre-training on a large-scale corpus, a neural language representation model can learn powerful language representation capabilities and extract rich syntactic and semantic information from the text. The pre-trained language model can provide word elements (tokens) containing rich semantic information and sentence-level features for downstream tasks, or directly perform fine-tuning for downstream tasks on the pre-trained model to conveniently and quickly obtain a downstream-specific model.
[0027] The neural network algorithm structure for pre-training a language model can be a Convolutional Neural Network (CNN), a Recurrent Neural Network (RNN), a Long Short-Term Memory (LSTM) network, etc., or it can be a model constructed with an attention network, such as a Transformer model, a Bidirectional Encoder Representations from Transformers (BERT) model, a Contrastive Language-Image Pre-training (CLIP) model, etc. The present invention does not limit this here. An attention network refers to a network model that uses an attention mechanism for training. This model assigns different weights to each part of the input sequence, thereby extracting more important feature information from the input sequence and enabling the model to finally obtain a more accurate output.
[0028] Compute Express Link (CXL) is a high-speed serial protocol that allows for fast and reliable data transfer between different components within a computer system. It aims to address bottleneck problems in high-performance computing, including memory capacity, memory bandwidth, and I / O latency, etc. Compute Express Link can also enable memory expansion and memory sharing, and communicate with peripherals such as computing accelerators (e.g., GPUs, FPGAs), providing a faster and more flexible way of data exchange and processing.
[0029] During the inference process of a pre-trained language model, one reason that restricts its more efficient use of resources is that there are two distinct processes in the inference process of a pre-trained language model, namely the prefill (also known as the prompt phase) and the decode (also known as the token generation phase).
[0030] When the computing system receives a user task (Request), during the pre-filling stage, the pre-trained language model processes the input or prompt of all user tasks, processes the input tokens in parallel, and calculates the corresponding key-value cache data (KV Cache). It can fully utilize the computing power of the computing system in parallel and belongs to a compute-intensive task. Next, it enters the decoding stage, which sequentially generates one token at a time, and only calculates one token per memory access. Its requirement for computing power is relatively low and is mainly limited by the memory bandwidth, belonging to a memory-intensive task. The pre-trained language model repeats these two stages until the end-of-sequence (EOS) is generated or the user-set stop condition is reached.
[0031] Figure 1 It is a schematic diagram of the parallel inference process of a pre-trained language model.
[0032] As Figure 1 shown, x1 and x2 are two sequences with different lengths, that is, they include different numbers of tokens. The tokens of sequence x1 include "I", "think", "this", "is", "great", and the tokens of sequence x2 include "I", "love", "you". " <eos>” represents the termination symbol.
[0033] Padding will be performed ( Figure 1 After using left padding), sequences of equal length are fed into the pre-trained language model to perform the pre-padding task, and the first token is produced. The entire process of the pre-padding task is called an iteration (iteration, which can also be understood as an inference stage), that is Figure 1 Iteration 1 as shown. Next, the decoding task is performed on these two sequences.
[0034] It can be found that after one iteration, the inference of sequence x2 has been completed, while the inference of sequence x1 is still ongoing. Since in the traditional batching method, the sequences in a batch act together, even though the inference of sequence x2 has been completed, it still cannot be "released".
[0035] Next, sequence x1 undergoes two more iterations (iteration 2, iteration 3, iteration 4 is used to output the termination symbol). Now the inference of sequence x1 is also completed. Then the data in the entire batch can be truly "released". After the inference of this batch is completed. The remaining requests can then continue to form a new batch for the next round of inference.
[0036] Since the decoding task stage is token by token, the second iteration can only be performed after the first iteration produces a token. This results in idle time (bubble) of computing resources. Although the inference calculation of sequence x2 has been completed, it still occupies resources, which not only causes resource waste but also slows down the overall computing process.
[0037] In related technologies, the scheduling of the pre-padding task and the decoding task often adopts a fixed resource ratio or a decoding-task-priority scheduling. Here, the ratio is usually in units of tokens. A token is the smallest unit after text is tokenized. In the inference of a pre-trained language model, the input text is split into individual tokens, and the model processes these tokens one by one. More tokens usually mean higher computing resource requirements and higher storage resource requirements. During the execution stage of the decoding task, a key-value cache is used to store the key-value pairs of each token to reduce repeated calculations, and the key-value cache data usually requires a large amount of storage space. Considering the limited memory resources of computing accelerators such as GPUs, usually the decoding task corresponding to when the key-value cache with a large data volume fills up the memory space is considered, and after determining the ratio of the number of token budgets allocated to the pre-padding task and the decoding task according to the memory resources, it often no longer adjusts.
[0038] In practical applications, the lengths of the sequences to be processed in user tasks vary, and the workloads are different at different times. If a fixed proportion of the token budget is adopted, it will lead to the following situations: if the proportion of the token budget allocated to the pre-allocated tasks is too high, although the pre-filling tasks have been completed long ago, a large number of decoding tasks will pile up in the queue waiting, resulting in a high delay for the entire inference task; if the proportion of the token budget allocated to the decoding tasks is too high, the throughput will be too small and the parallel effect will be poor, and the computing resources cannot be fully utilized.
[0039] In order to improve the performance of generating response information while fully utilizing device resources, in view of the existence of two tasks, namely pre-filling and decoding, which are respectively computation-intensive and memory-intensive, in the inference calculation of the pre-trained language model, the present invention provides a dynamic scheduling scheme. In a batch of inference calculations, the token budget for the pre-filling task and the token budget for the decoding task are determined according to the computing power utilization rate of the device where the pre-filling task is executed. The proportion of the token budget for the decoding task is negatively correlated with the computing power utilization rate of the pre-filling task. According to the token budget, the corresponding sequence to be processed is input into the pre-trained language model for parallel inference calculation, so as to balance the throughput and delay of the parallel inference calculation, thereby solving the problem of the contradiction between device pressure and generation performance in parallel inference scheduling in the related art, achieving the technical effect of improving the parallel inference performance of the pre-trained language model, splicing the output results of the decoding tasks to obtain the response information of the sequence to be processed and outputting it, and improving the execution performance of the artificial intelligence question-and-answer task.
[0040] To facilitate understanding of the environment on which the response information generation method provided in the embodiments of the present invention depends, the software and hardware architecture on which the execution of the response information generation method provided in the embodiments of the present invention can be based will be described here.
[0041] The response information generation method provided in the embodiments of the present invention can be applied to a single computing device or a cluster including multiple computing devices. If it is applied to a computing cluster, the computing cluster may include multiple computing nodes, and each computing node is deployed with a pre-trained language model. The computing node is used to receive the sequence to be processed in the response task; input the sequence to be processed into the pre-trained language model for inference calculation; in a batch of inference calculations, determine the token budget for the pre-filling task and the token budget for the decoding task according to the computing power utilization rate of the device where the pre-filling task is executed, and input the corresponding sequence to be processed into the pre-trained language model for parallel inference calculation according to the token budget; splice the output results of the decoding tasks to obtain the response information of the sequence to be processed; output the response information; wherein, the proportion of the token budget for the decoding task is negatively correlated with the computing power utilization rate of the pre-filling task.
[0042] In an embodiment of the present invention, a computing node is a computing device. The types of computing nodes may include, but are not limited to, a Graphics Processing Unit (GPU), a Field Programmable Gate Array (FPGA), an Application Specific Integrated Circuit (ASIC), and a Data Processing Unit (DPU), or a computing device that uses one or more of them as accelerators. Other types of computing devices may also be used.
[0043] When the computing cluster executes the response information generation method provided by the embodiment of the present invention, it may adopt the method of Data Parallel Inference, split the user task into multiple subsets, and allocate these subsets to multiple computing devices (such as GPUs) for parallel processing. It may also adopt Model Parallel Inference, allocate different parts of the pre-trained language model to different computing devices for parallel processing, that is, each computing device only calculates a part of the model, and completes the inference of the entire model through communication between devices. It may also adopt the method of Hybrid Parallel Inference, simultaneously using the techniques of data parallel and model parallel to further improve the inference efficiency.
[0044] An embodiment of the present invention provides a response information generation method. Combining with the execution process of the response information generation method, the method will be described in detail below.
[0045] Figure 2 It is a flowchart of a response information generation method provided by an embodiment of the present invention.
[0046] As Figure 2 shown, the response information generation method provided by the embodiment of the present invention includes:
[0047] S201: Receive the sequence to be processed of the response task.
[0048] S202: Input the sequence to be processed into the pre-trained language model for inference calculation.
[0049] S203: In a batch of inference calculations, determine the token budget quantity of the prefill task and the token budget quantity of the decoding task according to the computing power utilization rate of the prefill task executed by the device where it is located, and input the corresponding sequence to be processed into the pre-trained language model for parallel inference calculation according to the token budget quantity.
[0050] S204: Concatenate the output results of the decoding task to obtain the response information of the sequence to be processed.
[0051] S205: Output the response information.
[0052] Among them, the proportion of the token budget quantity of the decoding task is negatively correlated with the computing power utilization rate of the pre-filling task.
[0053] In a specific implementation, the response information generation method provided by the embodiments of the present invention can be executed based on a computing device or a computing cluster including multiple computing devices.
[0054] The response task can be tasks such as image generation, speech generation, content parsing, and translation.
[0055] For S201, the sequence to be processed can be the input (input) or prompt of the response task given by the user.
[0056] For S202 and S203, one batch in the embodiments of the present invention is used to process multiple sequences to be processed in parallel. After the inference calculation of all the sequences to be processed in the current batch is completed, the next batch is combined.
[0057] Parallel inference calculation is performed using a pre-trained language model. The parallel manner can include at least one of the following two manners: multiple short input sequences are used as multiple sequences to be processed in parallel; a long input sequence is split into multiple short sequences to be processed in parallel.
[0058] Then, in S203, inputting the corresponding sequence to be processed into the pre-trained language model for parallel inference calculation according to the token budget quantity can include: performing the pre-filling task on multiple sequences to be processed in parallel to obtain the first sequence corresponding to the sequence to be processed; if there is a first sequence that meets the long sequence quantity, the first sequence is block-processed; performing the decoding task on multiple first sequences in parallel. By performing parallel inference calculation on multiple short sequences and splitting long sequences into multiple short sequences for parallel calculation, resource idle caused by parallel processing of different sequences with large length differences in the same batch is avoided.
[0059] According to specific inference strategies and goals, in one batch (batch) of inference calculation, one new token can be generated for each sequence to be processed, or multiple iterations can be performed for each sequence to be processed and multiple tokens can be generated.
[0060] During the inference process of the pre-trained language model, batches composed of different sequences to be processed may vary greatly in the input length of the pre-training task and the generated length of the decoding task. Embodiments of the present invention provide a dynamic scheduling scheme for pre-filling tasks and decoding tasks to adaptively adjust the token budget quantity of the pre-filling task and the token budget quantity of the decoding task. For a sequence to be processed with a long input (such as the first sequence that meets the long sequence quantity mentioned above), the token budget quantity of the pre-training task can be further segmented to ensure that the computational load for each input is appropriate and avoid increased latency caused by an excessive number of tokens in a single sequence to be processed.
[0061] In embodiments of the present invention, the token budget quantity of the pre-filling task and the token budget quantity of the decoding task are determined according to the computing power utilization rate of the device executing the pre-filling task. The adjustment principle is that the proportion of the token budget quantity of the decoding task is negatively correlated with the computing power utilization rate of the pre-filling task. Since the token budget quantity of the pre-filling task and the token budget quantity of the decoding task are allocated from the total token budget for the parallel inference task, the proportion of the token budget quantity of the decoding task is negatively correlated with the computing power utilization rate of the pre-filling task, which is equivalent to the proportion of the token budget quantity of the pre-filling task being positively correlated with the computing power utilization rate of the pre-filling task.
[0062] That is to say, when the computing power utilization rate of the pre-filling task drops significantly, indicating that the pre-filling task will soon be completed and is waiting for the decoding task to execute, at this time, by increasing the proportion of the token budget quantity of the decoding task, that is, allocating more token budget quantity to the decoding task, more decoding tasks can be loaded, thereby reducing the latency of the current batch. When the computing power utilization rate of the pre-filling task is relatively high, it means that the token budget quantity allocated to the pre-filling task is insufficient. At this time, more token budget quantity is allocated to the pre-filling task to increase the throughput of the current batch.
[0063] For S204, after reaching the stop condition set by the user, for each sequence to be processed, the output results of each decoding task are concatenated to obtain the response information of the sequence to be processed.
[0064] For S205, the response information is output to respond to the response task.
[0065] The response information generation method provided by the embodiments of the present invention aims at the two tasks of pre-filling and decoding, which are respectively computation-intensive and memory-intensive, in the inference calculation of the pre-trained language model. It provides a dynamic scheduling scheme. In a batch of inference calculations, according to the computing power utilization rate of the device executing the pre-filling task, the token budget quantity of the pre-filling task and the token budget quantity of the decoding task are determined, so that the proportion of the token budget quantity of the decoding task is negatively correlated with the computing power utilization rate of the pre-filling task. According to the token budget quantity, the corresponding sequence to be processed is input into the pre-trained language model for parallel inference calculation, so that the parallel inference calculation obtains the balance of throughput and latency, thereby solving the problem of the contradiction between device pressure and generation performance in parallel inference scheduling in the related art, achieving the technical effect of improving the parallel inference performance of the pre-trained language model. The output results of the decoding task are spliced to obtain the response information of the sequence to be processed and output, improving the execution performance of the artificial intelligence question and answer task.
[0066] In some optional embodiments of the embodiments of the present invention, determining the token budget quantity of the pre-filling task and the token budget quantity of the decoding task according to the computing power utilization rate of the device executing the pre-filling task in S203 includes: when the computing power utilization rate of the pre-filling task is less than the computing power threshold, increasing the proportion of the token budget quantity of the decoding task.
[0067] In the embodiments of the present invention, a computing power threshold is used to trigger the adjustment of the proportion of the token budget quantity of the pre-filling task and the decoding task. Only when the computing power utilization rate of the pre-filling task is less than the computing power threshold, the proportion of the token budget quantity of the decoding task is triggered to increase (that is, the proportion of the token budget quantity of the pre-filling task is reduced). Correspondingly, when the computing power utilization rate of the pre-filling task is greater than or equal to the computing power threshold, the proportion of the token budget quantity of the pre-filling task and the decoding task can adopt the initial setting value of the device, or the proportion of the token budget quantity of the decoding task can be reduced (that is, the proportion of the token budget quantity of the pre-filling task is increased) on the basis of the proportion of the token budget quantity at the previous moment.
[0068] In the embodiments of the present invention, the computing power utilization rate is the ratio of the actual computing power of the device when executing the pre-filling task to the peak computing power of the device.
[0069] Specifically, the actual computing power of the device when executing the pre-filling task can be the difference between the computing power consumed by each token of the pre-filling task of the current batch executed by the device and the computing power consumed by generating the first token when the device executes the pre-filling task of the current batch. At this time, the memory utilization rate of the pre-filling task can be expressed by the following formula:
[0070] ;
[0071] Among them, R represents the memory utilization rate of the pre-filling task, Represents the computing power metric of a single token for the device to perform the pre-filling task of the current batch. Represents the computing power consumed by each token for the device to perform the pre-filling task of the current batch. Represents the attenuation coefficient (a coefficient that decays over time, an inherent parameter related to the physical quality of the device). Represents the Time To First Token, that is, the time taken for the device to generate the first token when performing the pre-filling task of the current batch. Represents the peak computing power of the device when performing the pre-filling task of the current batch. Represents the peak computing power of the device.
[0072] Alternatively, when the device also includes extended memory, it can solve the problem that computing devices capable of efficiently performing pre-filling tasks often lack memory resources. At this time, communication losses and other losses will occur in the computing system. Then the actual computing power can be the difference between the computing power consumed by each token for the device to perform the pre-filling task of the current batch plus the communication loss of the device when performing the pre-filling task of the current batch minus the computing power consumed by the device to generate the first token when performing the pre-filling task of the current batch. At this time, the memory utilization rate of the pre-filling task can be expressed by the following formula:
[0073] ;
[0074] Among them, R represents the memory utilization rate of the pre-filling task. Represents the computing power metric of a single token for the device to perform the pre-filling task of the current batch. Represents the computing power consumed by each token for the device to perform the pre-filling task of the current batch. Represents the communication loss and other losses (which can be set to 50% of the peak computing power of the device with a scalable memory pool). Represents the attenuation coefficient (a coefficient that decays over time, an inherent parameter related to the physical quality of the device). Represents the Time To First Token, that is, the time taken for the device to generate the first token when performing the pre-filling task of the current batch. Represents the peak computing power of the device when performing the pre-filling task of the current batch. Represents the peak computing power of the device.
[0075] In the embodiments of the present invention, the computing power threshold may include one or more thresholds, and different adjustment scales of the token budget quantity ratio may be corresponding when different computing power thresholds are triggered.
[0076] In some other alternative embodiments of the embodiments of the present invention, determining the token budget quantity of the pre-filling task and the token budget quantity of the decoding task according to the computing power utilization rate of the device where the pre-filling task is executed in S203 may further include: determining the token budget quantity of the pre-filling task and the token budget quantity of the decoding task according to the computing power utilization rate of the pre-filling task and the memory utilization rate of the device; wherein, the proportion of the token budget quantity of the decoding task is negatively correlated with the memory utilization rate of the device.
[0077] Wherein, the proportion of the token budget quantity of the decoding task is negatively correlated with the computing power utilization rate of the pre-filling task, and the specific analysis is as described in the above embodiments.
[0078] The proportion of the token budget quantity of the decoding task is negatively correlated with the memory utilization rate of the device, that is, the proportion of the token budget quantity of the pre-filling task is positively correlated with the memory utilization rate of the device. For the memory utilization rate, when the memory utilization rate of the device is small, it indicates that the data volume of the generated key-value cache data is small, indicating that the throughput of the current batch is small and the token budget quantity allocated to the pre-filling task is insufficient. At this time, reduce the proportion of the token budget quantity of the decoding task, that is, increase the proportion of the token budget quantity of the pre-filling task to improve the throughput. When the memory utilization rate of the device is large, it indicates that the data volume of the generated key-value cache data is large, indicating that the load of the current batch is large. At this time, give priority to ensuring the token budget quantity of the decoding task to reduce the latency.
[0079] In the embodiments of the present invention, determining the token budget quantity of the pre-filling task and the token budget quantity of the decoding task according to the computing power utilization rate of the pre-filling task and the memory utilization rate of the device may include: when at least one of the conditions that the computing power utilization rate of the pre-filling task is less than the computing power threshold and the memory utilization rate of the device is less than the first memory threshold is satisfied, increase the proportion of the token budget quantity of the decoding task.
[0080] In the embodiments of the present invention, the calculation method of the computing power utilization rate and the determination method of the computing power threshold may be seen in the description of the above embodiments.
[0081] The memory utilization rate may be the ratio of the memory occupancy of the current batch to the total memory allocated by the device to the response task. The memory occupancy of the current batch may include the memory occupancy of the key-value cache data, the tokens of the current batch, and the model parameters of the pre-trained language model.
[0082] In some further alternative embodiments of the embodiments of the present invention, determining the token budget quantity of the pre-filling task and the token budget quantity of the decoding task according to the computing power utilization rate of the device executing the pre-filling task in S203 may further include: determining the token budget quantity of the pre-filling task and the token budget quantity of the decoding task according to the computing power utilization rate of the pre-filling task and the operating state parameters of the device; wherein, the operating state parameters include at least one of temperature, power, and the rotation speed of the cooling fan; the proportion of the token budget quantity of the decoding task is positively correlated with the operating state parameters.
[0083] In practical applications, the power of the device when executing the pre-filling task is usually higher than the power of the device when executing the decoding task. Correspondingly, the temperature of the device when executing the pre-filling task is also higher than the temperature of the device when executing the decoding task, and the fan rotation speed of the device when executing the pre-filling task is also higher than the fan rotation speed of the device when executing the decoding task. Therefore, it is also possible to combine the operating state parameters such as temperature, power, and the rotation speed of the cooling fan with the computing power utilization rate of the pre-filling task to determine whether the token budget quantity ratio needs to be adjusted. For example, when the temperature of the device is too high, it means that the token budget quantity allocated to the pre-filling task is too high. At this time, reduce the proportion of the token budget quantity of the pre-filling task, that is, increase the proportion of the token budget quantity of the decoding task, to achieve temperature regulation of the device.
[0084] In the embodiments of the present invention, determining the token budget quantity of the pre-filling task and the token budget quantity of the decoding task according to the computing power utilization rate of the pre-filling task and the operating state parameters of the device may include: when at least one of the conditions that the computing power utilization rate of the pre-filling task is less than the computing power threshold and the operating state parameter is greater than the operating state threshold is satisfied, increase the proportion of the token budget quantity of the decoding task.
[0085] In the embodiments of the present invention, the calculation method of the computing power utilization rate and the determination method of the computing power threshold can be found in the description of the above embodiments. The calculation method of the memory utilization rate and the determination method of the memory threshold can be found in the description of the above embodiments.
[0086] One or more of the operating state parameters can be considered, and corresponding operating state thresholds can be set for temperature, power, the rotation speed of the cooling fan, etc.
[0087] Referring to the above embodiments, determining the token budget quantity of the pre-filling task and the token budget quantity of the decoding task according to the computing power utilization rate of the device executing the pre-filling task in S203 may further include: determining the token budget quantity of the pre-filling task and the token budget quantity of the decoding task according to the computing power utilization rate of the pre-filling task, the memory utilization rate of the device, and the operating state parameters of the device; wherein, the operating state parameters include at least one of temperature, power, and the rotation speed of the cooling fan; the proportion of the token budget quantity of the decoding task is negatively correlated with the memory utilization rate of the device; the proportion of the token budget quantity of the decoding task is positively correlated with the operating state parameters.
[0088] In the embodiments of the present invention, determining the token budget quantity of the pre-filling task and the token budget quantity of the decoding task according to the computing power utilization rate of the pre-filling task, the memory utilization rate of the device, and the operating state parameters of the device may include: when at least one of the conditions that the computing power utilization rate of the pre-filling task is less than the computing power threshold, the memory utilization rate of the device is less than the first memory threshold, and the operating state parameters are greater than the operating state threshold is satisfied, increasing the proportion of the token budget quantity of the decoding task.
[0089] In the embodiments of the present invention, the calculation method of the computing power utilization rate and the determination method of the computing power threshold may be found in the description of the above embodiments. The calculation method of the memory utilization rate and the determination method of the memory threshold may be found in the description of the above embodiments. The calculation method of the operating state parameters and the determination method of the operating state threshold may be found in the description of the above embodiments.
[0090] In addition to adjusting the proportion of the token budget quantity by using the memory utilization rate and the operating state parameters to assist the computing power utilization rate in the above embodiments, the memory utilization rate and the operating state parameters may also be used to adjust the operating state of the device. Then, the response information generation method provided by the embodiments of the present invention may further include: when at least one of the conditions that the memory utilization rate reaches the memory threshold and the operating state parameters reach the operating state threshold is satisfied, performing a down-frequency operation on the device. That is to say, the system stability of the device can be monitored through the memory utilization rate and the operating state parameters to prevent the overall system from crashing.
[0091] Figure 3 It is a flowchart of another response information generation method provided by the embodiments of the present invention; Figure 4 It is a structural schematic diagram of a batch mixing manager provided by the embodiments of the present invention.
[0092] Since the pre-filling task and the decoding task are two types of tasks with diametrically opposite resource requirements, resulting in contradictions when using a single type of resource, in the embodiments of the present invention, a solution for separating the pre-filling task and the decoding task is provided. When using an accelerator as the computing device, an architecture for separating the pre-filling task and the decoding task can be implemented by combining extended memory.
[0093] Then, in the embodiments of the present invention, in S203, inputting the corresponding sequence to be processed into the pre-trained language model for parallel inference calculation according to the token budget quantity may include: performing a pre-fill task on the sequence to be processed to obtain a first sequence; reading the model weights of the pre-trained language model from the system memory to calculate the query vector and key-value matrix corresponding to the first sequence; storing the key-value matrix in the extended memory; and performing a decoding task according to the query vector, the key-value matrix read from the extended memory, and the tokens of the first sequence read from the system memory.
[0094] As Figure 3 shown, the embodiments of the present invention provide an architecture that separates the pre-fill task and the decoding task and consists of three parts: a computing module, a system memory, and an extended memory. Among them, the system memory is used to store the original data (sequence to be processed) that needs to be calculated for the pre-training task and the result vectors output by the model. The computing module is usually used for large-scale parallel computing to obtain the intermediate results and model outputs of the pre-trained language model. The extended memory is used to store the key-value cache data in the middle of the pre-trained language model to accelerate the inference speed of the model.
[0095] In the embodiments of the present invention, determining the token budget quantity of the pre-fill task and the token budget quantity of the decoding task according to the computing power utilization rate of the device executing the pre-fill task may include: determining the total token budget quantity allocated to the inference calculation according to the overall computing power index of the device and the memory utilization rate of the extended memory; after initializing the token budget quantity of the decoding task, determining the token budget quantity of the pre-fill task according to the remaining memory resources of the device; when the computing power utilization rate of the pre-fill task is less than the first computing power threshold, increasing the proportion of the token budget quantity of the decoding task; and when the computing power utilization rate of the pre-fill task is greater than or equal to the second computing power threshold, reducing the proportion of the token budget quantity of the decoding task.
[0096] In the embodiments of the present invention, the extended memory may adopt a computer express link extended memory, and specifically, a CXL Type 3 device (supporting the computer express link memory (CXL.mem) protocol) may be used.
[0097] Figure 3 In [reference], the calculation process of the query vector and the key-value matrix (QKV Linear) only needs to read the model weights from the memory of the computing module, and the calculation is independent of the length of the input sequence; when performing attention (Attn) calculation, it is not necessary to read the weights from the system memory; during the execution stage of the decoding task, it is necessary to read the key-value cache data (KV-cache) from the extended memory to accelerate the calculation.
[0098] As Figure 3 shown, in the current batch, assuming , , , Four pending sequences are processed for parallel reasoning calculations, each of which includes 5 tokens. At a certain moment in the parallel reasoning calculation, the batch mixing manager is designed and combined with the extended memory for scheduling. The steps of data flow include:
[0099] ①: Assume that at the current moment, 13 tokens ( , , , , , , , , , , , , ) performs pre-filled parallel computation, and the computation module reads the input from the system memory (of size [13, H]).
[0100] ②: In the linear transformation calculation module (QKV Linear), three matrices (of size [13, 3H]) are calculated: query vector (Q), key (K), and value (V).
[0101] ③: The batch mixing manager determines and controls the status of the extended memory through the status register read-write module and the control register read-write module.
[0102] ④: The host of the device stores the K and V matrices into the extended memory pool composed of CXL type-3 devices.
[0103] ⑤⑥: The Q matrix continues to flow to the subsequent computing module; read the key-value cache data in the extended memory; calculate and Attention vector and the attention vector ); The computing module initiates reading of the original input from the system memory again and starts decoding calculation, 7 token model input and In ( , , , )and( , , ) are decoded and calculated respectively to obtain the decoding vector.
[0104] ⑦⑧: The subsequent sequence calculation is the same as step ⑤⑥, including 4 identical operations, generating and Attention vector and the attention vector ).
[0105] ⑨: When the sequences have all completed the decoding calculation, the result sequences are concatenated and merged.
[0106] ⑩: The merged vector (with a size of [13, H]) is linearly transformed to obtain the output, which is the response information of the sequence to be processed.
[0107] As Figure 4 shown, the deployment batch mixing manager is used to interact with the extended memory and schedule the proportion of the number of tokens in the prefill task and the decoding task.
[0108] In a specific implementation, the device (the entire server) is roughly divided into three parts: system memory, computing module, and extended memory. The key-value cache data is stored in the extended memory. The input data of the model is stored in the system memory. The computing module is responsible for batch mixing during the inference process.
[0109] The batch mixing manager controls the access of the extended memory by reading and writing registers in the extended memory, and realizes the mixing strategy through calculating the dynamic regulation index. The dynamic regulation index considers the time decay coefficient of the hardware device and monitors the computing power index of a single token and the time consumption of the first token ( ). The batch mixing manager includes a communication transceiver module for coordinating multiple computing modules within the device.
[0110] Applying the architecture that separates the prefill task and the decoding task provided by the embodiments of the present invention, the steps executed in the batch mixing manager include:
[0111] After the input data is sent to the manager, the Q matrix is stored in the manager's memory, and the KV matrix (KV-cache) is stored in the extended memory through the control register reading and writing module.
[0112] The scheduling module of the batch mixing manager determines the proportion of the number of tokens in the prefill task and the decoding task by calculating the threshold, divides the tokens of the input sequence at the next moment, and passes the tokens in the execution stage of the decoding task to the decoding calculation module for calculation.
[0113] The computing power index data and the memory occupancy rate are sent to the manager. Memory occupancy will affect the running stability of the entire server. If the memory occupancy exceeds the set threshold, the overall system will be downclocked to prevent the overall system from crashing. The computing power index is used for the dynamic scheduling of the number of tokens in the prefill task and the decoding task in the computing module.
[0114] Utilize the cache coherence of CXL Type 3 devices to connect the manager with the managers of other computing modules to form an interconnection of multiple computing devices. The communication transceiver module transfers the status in the local computing module to other heterogeneous computing devices by modifying the field information carried in the CXL flit (CXL flight data unit) in the CXL protocol, realizing resource coordination among multiple devices.
[0115] Through the description of the above embodiments, those skilled in the art can clearly understand that the method according to the above embodiments can be implemented by means of software plus a necessary general hardware platform. Of course, it can also be implemented by hardware, but in many cases, the former is a better implementation method.
[0116] Figure 5 It is a schematic structural diagram of a response information generation device provided by an embodiment of the present invention.
[0117] Such as Figure 5 As shown, an embodiment of the present invention also provides a response information generation device, including:
[0118] A receiving module 501, configured to receive a to-be-processed sequence of a response task;
[0119] A computing module 502, configured to input the to-be-processed sequence into a pre-trained language model for inference calculation; in a batch of inference calculations, determine the token budget quantity of the pre-filling task and the token budget quantity of the decoding task according to the computing power utilization rate of the device where the pre-filling task is executed, and input the corresponding to-be-processed sequence into the pre-trained language model for parallel inference calculation according to the token budget quantity; splice the output results of the decoding task to obtain the response information of the to-be-processed sequence;
[0120] An output module 503, configured to output the response information;
[0121] Wherein, the proportion of the token budget quantity of the decoding task is negatively correlated with the computing power utilization rate of the pre-filling task.
[0122] In the embodiment of the present invention, the computing module 502 determines the token budget quantity of the pre-filling task and the token budget quantity of the decoding task according to the computing power utilization rate of the device where the pre-filling task is executed, which may include: when the computing power utilization rate of the pre-filling task is less than the computing power threshold, increasing the proportion of the token budget quantity of the decoding task.
[0123] In an embodiment of the present invention, the computing module 502 determines the token budget quantity of the pre-filling task and the token budget quantity of the decoding task according to the computing power utilization rate of the device where the pre-filling task is executed, and may further include: determining the token budget quantity of the pre-filling task and the token budget quantity of the decoding task according to the computing power utilization rate of the pre-filling task and the memory utilization rate of the device; wherein, the proportion of the token budget quantity of the decoding task is negatively correlated with the memory utilization rate of the device.
[0124] Wherein, determining the token budget quantity of the pre-filling task and the token budget quantity of the decoding task according to the computing power utilization rate of the pre-filling task and the memory utilization rate of the device may include: when at least one of the conditions that the computing power utilization rate of the pre-filling task is less than the computing power threshold and the memory utilization rate of the device is less than the first memory threshold is satisfied, increasing the proportion of the token budget quantity of the decoding task.
[0125] In an embodiment of the present invention, the computing module 502 determines the token budget quantity of the pre-filling task and the token budget quantity of the decoding task according to the computing power utilization rate of the device where the pre-filling task is executed, and may further include: determining the token budget quantity of the pre-filling task and the token budget quantity of the decoding task according to the computing power utilization rate of the pre-filling task and the operating state parameters of the device; wherein, the operating state parameters include at least one of temperature, power, and the rotation speed of the cooling fan; the proportion of the token budget quantity of the decoding task is positively correlated with the operating state parameters.
[0126] Wherein, determining the token budget quantity of the pre-filling task and the token budget quantity of the decoding task according to the computing power utilization rate of the pre-filling task and the operating state parameters of the device may include: when at least one of the conditions that the computing power utilization rate of the pre-filling task is less than the computing power threshold and the operating state parameters are greater than the operating state threshold is satisfied, increasing the proportion of the token budget quantity of the decoding task.
[0127] In an embodiment of the present invention, the computing power utilization rate may be the ratio of the actual computing power of the device when executing the pre-filling task to the peak computing power of the device.
[0128] In an embodiment of the present invention, the actual computing power may be the difference between the computing power consumed by each token of the device when executing the current batch of pre-filling tasks and the computing power consumed by the device when generating the first token when executing the current batch of pre-filling tasks.
[0129] In an embodiment of the present invention, the device includes an extended memory; then the actual computing power may be the difference between the computing power consumed by each token of the device when executing the current batch of pre-filling tasks plus the communication loss of the device when executing the current batch of pre-filling tasks and the computing power consumed by the device when generating the first token when executing the current batch of pre-filling tasks.
[0130] In an embodiment of the present invention, the computing module 502 inputs a corresponding sequence to be processed into a pre-trained language model for parallel inference calculation according to the token budget quantity, which may include: performing a pre-fill task on the sequence to be processed to obtain a first sequence; reading model weights of the pre-trained language model from the system memory to calculate a query vector and a key-value matrix corresponding to the first sequence; storing the key-value matrix in the extended memory; and performing a decoding task according to the query vector, the key-value matrix read from the extended memory, and the tokens of the first sequence read from the system memory.
[0131] Among them, the computing module 502 determines the token budget quantity of the pre-fill task and the token budget quantity of the decoding task according to the computing power utilization rate of the device where the pre-fill task is executed, which may include: determining the total token budget quantity allocated for inference calculation according to the overall computing power index of the device and the memory utilization rate of the extended memory; after initializing the token budget quantity of the decoding task, determining the token budget quantity of the pre-fill task according to the remaining memory resources of the device; when the computing power utilization rate of the pre-fill task is less than the first computing power threshold, increasing the proportion of the token budget quantity of the decoding task; and when the computing power utilization rate of the pre-fill task is greater than or equal to the second computing power threshold, reducing the proportion of the token budget quantity of the decoding task.
[0132] In an embodiment of the present invention, the computing module 502 inputs a corresponding sequence to be processed into a pre-trained language model for parallel inference calculation according to the token budget quantity, which may include: performing a pre-fill task on multiple sequences to be processed in parallel to obtain first sequences corresponding to the sequences to be processed; if there are first sequences meeting the long sequence quantity, performing block processing on the first sequences; and performing a decoding task on multiple first sequences in parallel.
[0133] For the description of the features corresponding to the embodiments of the response information generation device, reference may be made to the relevant descriptions of the embodiments corresponding to the response information generation method, which will not be elaborated here one by one.
[0134] An embodiment of the present invention further provides an electronic device, including a memory and a processor. A computer program is stored in the memory, and the processor is configured to run the computer program to execute the steps in any of the above embodiments of the response information generation method.
[0135] An embodiment of the present invention further provides a computer-readable storage medium, in which a computer program is stored. The computer program is configured to execute the steps in any of the above embodiments of the response information generation method when running.
[0136] In an exemplary embodiment, the above computer-readable storage medium may include, but is not limited to: various media such as USB flash drives, read-only memory (ROM), random access memory (RAM), external hard drives, magnetic disks, or optical discs that can store computer programs.
[0137] An embodiment of the present invention also provides a computer program product. The above computer program product includes a computer program, and when the computer program is executed by a processor, it implements the steps in any of the above embodiments of the response information generation method.
[0138] An embodiment of the present invention also provides another computer program product, including a non-volatile computer-readable storage medium. The non-volatile computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, it implements the steps in any of the above embodiments of the response information generation method.
[0139] Those skilled in the art can further realize that the units and algorithm steps of each example described in combination with the embodiments disclosed herein can be implemented by electronic hardware, computer software, or a combination of the two. To clearly illustrate the interchangeability of hardware and software, the components and steps of each example have been generally described according to their functions in the above description. Whether these functions are executed in a hardware or software manner depends on the specific application and design constraints of the technical solution. Skilled professionals can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of the present invention.
[0140] The above has introduced in detail a response information generation method, device, medium, and computer program product provided by the present invention. Specific examples are used herein to elaborate on the principles and implementation manners of the present invention. The description of the above embodiments is only used to help understand the method and its core idea of the present invention. It should be noted that for those of ordinary skill in the art in this technical field, without departing from the principle of the present invention, several improvements and modifications can be made to the present invention, and these improvements and modifications also fall within the protection scope of the present invention.< / eos>
Claims
1. A method for generating response information, characterized in that, including: a to-be-processed sequence for receiving a response task; inputting the to-be-processed sequence into a pre-trained language model for inference calculation; in a batch of inference calculations, determining the token budget quantity of the pre-filling task and the token budget quantity of the decoding task according to the computing power utilization rate of the device executing the pre-filling task, and inputting the corresponding to-be-processed sequence into the pre-trained language model for parallel inference calculation according to the token budget quantity; concatenating the output results of the decoding task to obtain the response information of the to-be-processed sequence; outputting the response information; wherein, the proportion of the token budget quantity of the decoding task is negatively correlated with the computing power utilization rate of the pre-filling task; determining the token budget quantity of the pre-filling task and the token budget quantity of the decoding task according to the computing power utilization rate of the device executing the pre-filling task includes: when the computing power utilization rate of the pre-filling task is less than the computing power threshold, increasing the proportion of the token budget quantity of the decoding task.
2. The method for generating response information according to claim 1, wherein determining the token budget quantity of the pre-filling task and the token budget quantity of the decoding task according to the computing power utilization rate of the device executing the pre-filling task includes: determining the token budget quantity of the pre-filling task and the token budget quantity of the decoding task according to the computing power utilization rate of the pre-filling task and the memory utilization rate of the device; wherein, the proportion of the token budget quantity of the decoding task is negatively correlated with the memory utilization rate of the device.
3. The response information generation method according to claim 2, wherein determining the token budget quantity of the pre-filling task and the token budget quantity of the decoding task according to the computing power utilization rate of the pre-filling task and the memory utilization rate of the device includes: when at least one of the conditions that the computing power utilization rate of the pre-filling task is less than the computing power threshold and the memory utilization rate of the device is less than the first memory threshold is satisfied, increasing the proportion of the token budget quantity of the decoding task.
4. The response information generation method according to claim 1, wherein determining the token budget quantity of the pre-filling task and the token budget quantity of the decoding task according to the computing power utilization rate of the device executing the pre-filling task includes: determining the token budget quantity of the pre-filling task and the token budget quantity of the decoding task according to the computing power utilization rate of the pre-filling task and the operating state parameters of the device; wherein, the operating state parameters include at least one of temperature, power, and the rotation speed of the cooling fan; the proportion of the token budget quantity of the decoding task is positively correlated with the operating state parameters.
5. The method for generating response information according to claim 4, wherein determining the token budget quantity of the pre-filling task and the token budget quantity of the decoding task according to the computing power utilization rate of the pre-filling task and the operating state parameters of the device includes: when at least one of the conditions that the computing power utilization rate of the pre-filling task is less than the computing power threshold and the operating state parameters are greater than the operating state threshold is satisfied, increasing the proportion of the token budget quantity of the decoding task.
6. The method for generating response information according to any one of claims 1 to 5, characterized in that The computing power utilization rate is the ratio of the actual computing power of the device when executing the pre-filling task to the peak computing power of the device.
7. The response information generation method according to claim 6, wherein The actual computing power is the difference between the computing power consumed by each token of the device executing the current batch of the pre-filling task and the computing power consumed by the device when generating the first token of the current batch of the pre-filling task.
8. The response information generation method according to claim 6, wherein The device includes an extended memory; The actual computing power is the difference between the computing power consumed by each token of the device to execute the pre-padding task of the current batch, plus the communication loss of the device when executing the pre-padding task of the current batch, and the computing power consumed by the device to generate the first token when executing the pre-padding task of the current batch.
9. The response information generation method according to claim 1, wherein Inputting the corresponding sequence to be processed into the pre-trained language model for parallel inference calculation according to the token budget quantity, including: Performing the pre-padding task on the sequence to be processed to obtain a first sequence; Reading the model weights of the pre-trained language model from the system memory to calculate the query vector and key-value matrix corresponding to the first sequence; Storing the key-value matrix in the extended memory; Performing the decoding task according to the query vector, the key-value matrix read from the extended memory, and the tokens of the first sequence read from the system memory.
10. The method for generating response information according to claim 9, wherein, Determining the token budget quantity of the pre-padding task and the token budget quantity of the decoding task according to the computing power utilization rate of the device to execute the pre-padding task, including: Determining the total token budget quantity allocated to the inference calculation according to the overall computing power index of the device and the memory utilization rate of the extended memory; After initializing the token budget quantity of the decoding task, determining the token budget quantity of the pre-padding task according to the remaining memory resources of the device; When the computing power utilization rate of the pre-padding task is less than the first computing power threshold, increasing the proportion of the token budget quantity of the decoding task; When the computing power utilization rate of the pre-padding task is greater than or equal to the second computing power threshold, reducing the proportion of the token budget quantity of the decoding task.
11. The response information generation method according to claim 1, characterized in that Inputting the corresponding sequence to be processed into the pre-trained language model for parallel inference calculation according to the token budget quantity, including: Performing the pre-padding task on multiple sequences to be processed in parallel to obtain the first sequences corresponding to the sequences to be processed; If there are first sequences that meet the long sequence quantity, performing block processing on the first sequences; Performing the decoding task on multiple first sequences in parallel.
12. An electronic device, characterized in that, Including: A memory for storing a computer program; A processor for implementing the steps of the response information generation method according to any one of claims 1 to 11 when executing the computer program.
13. A computer-readable storage medium, characterized in that, A computer program is stored in the computer-readable storage medium, wherein the computer program implements the steps of the response information generation method according to any one of claims 1 to 11 when executed by a processor.
14. A computer program product, comprising a computer program, characterized in that, The computer program implements the steps of the response information generation method according to any one of claims 1 to 11 when executed by a processor.
Citation Information
Patent Citations
Model reasoning scheduling method and device and server cluster
CN118897736A
Video memory management method and device for large language model reasoning, medium and product
CN119443173A