Multi-queue scheduling method based on total and incremental reasoning overhead merging

By adopting a multi-queue scheduling method based on the merger of full- and incremental inference overhead in multi-tenant, high-concurrency online inference scenarios, the problems of excessive latency and redundant computing caused by FCFS policies are solved, and higher system throughput and lower latency are achieved.

CN120216140APending Publication Date: 2025-06-27HANGZHOU DIANZI UNIV
View PDF 0 Cites 5 Cited by

Patent Information

Application Number
CN202510339900.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-21
Publication Date
2025-06-27

AI Technical Summary

Technical Problem

In multi-tenant, high concurrency online inference scenarios, FCFS strategy leads to problems of excessive delay and redundant calculation when waiting for queue requests, and existing optimization methods fail to effectively solve the problems of excessive resource occupancy and redundant calculation when long-sequence requests.

Method used

A multi-queue scheduling method based on the merger of full- and incremental inference overheads is adopted to predict incremental inference overheads through small models, and mixed evaluation of the overheads of full-scale inference and incremental inference is carried out to achieve more accurate inference task priority division and resource optimization.

Benefits of technology

It improves the overall system throughput, reduces the request delay, avoids the redundant calculation problem caused by insufficient KV Cache resources, and solves the problem of blocking subsequent requests for too long execution time of long-sequence requests.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120216140A_ABST
    Figure CN120216140A_ABST
Patent Text Reader

Abstract

The invention discloses a multi-queue scheduling method based on total and incremental reasoning overhead merging, which comprises the following steps of: firstly, calculating the total resource demand and memory overhead of a task according to the input length and the size of a predictor model in the process of performing total reasoning on a reasoning task; distributing tasks to queues with different priorities according to the total resource demand in combination with response time requirements and task importance; dynamically scheduling tasks according to a priority queue sequence, adjusting an execution sequence in combination with a system load, and preferentially executing tasks in a high-priority queue; distributing a time slice for each task, monitoring task execution time, reducing the priority of overtime tasks, moving the overtime tasks into a secondary priority queue, and triggering new task scheduling; and after preemptive scheduling, the key value cache of the latest scheduling task in the waiting queue is swapped from the accelerator memory to the CPU memory, and is swapped back to the accelerator memory before the task is re-executed.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of optimizing system resource utilization, and particularly to a multi-queue scheduling method based on the merging of full-scale and incremental inference overheads. Background Art

[0002] In recent years, with the rapid development of deep learning technology, generative models (such as GPT) have achieved remarkable performance breakthroughs in the fields of natural language processing, computer vision, etc. Through the training of massive amounts of data, these models can capture complex semantic information and long-distance dependency relationships, thus performing excellently in various tasks. However, the high performance of generative models also brings huge computational and storage requirements. Especially in the inference stage, problems such as overloading of resource usage or insufficient resource utilization often occur. How to make full use of existing system resources to efficiently deploy and run these generative models is crucial for improving the overall system throughput and reducing request latency.

[0003] During the inference task process of generative models, it can be divided into two forms: full-scale inference and incremental inference. Full-scale inference means that after all input data is loaded, the system processes all contents at once and outputs the result; incremental inference means that when the initial input is received, the model starts to gradually infer, generates intermediate results and outputs them. The way incremental inference processes input data is step-by-step, and the output results are also gradually returned along with the inference process. However, there are significant differences in resource requirements between full-scale inference and incremental inference, which pose great challenges to inference. In the full-scale inference mode, the system loads and processes all input data at once at the beginning, which consumes relatively high computational resources. In contrast, the step-by-step update mechanism of incremental inference requires the model to recalculate and update the state information at each step, resulting in the need for context information when generating each token. As the output sequence continues to expand, the lack of memory resources has become the main obstacle to improving the performance of incremental inference. At the same time, because the existing scheduling algorithms adopt the FCFS (First-Come-First-Served) strategy, a series of problems will occur in multi-tenant and high-concurrency online inference scenarios. First, long-sequence requests that arrive first will continuously occupy resources and positions in the running queue, significantly increasing the request latency in the waiting queue. Second, when the batch size of the system is too large, long-sequence requests occupy too much KV Cache resources, resulting in insufficient remaining KV Cache to support the inference of the next iteration. In this case, the system has to release the KV Cache of requests that have already been calculated, thus triggering redundant calculations and further affecting the system efficiency.

[0004] To solve a series of problems caused by the FCFS (First-Come, First-Served) policy in multi-tenant and high-concurrency online inference scenarios, existing optimization methods aim to provide a reasonable scheduling method so that the execution of overly long requests does not continuously block the requests in the waiting queue and occupy excessive resources. FastServe assigns appropriate priorities to requests according to the length of the prompts of the requests, then adds the requests to the corresponding priority queues according to the priorities. At the same time, a smaller time slice is assigned to the queue with a higher priority, and a larger time slice is assigned to the queue with a lower priority, so that requests with short prompts can obtain a higher priority and be executed earlier, without being blocked by long requests, and the response speed of short requests is improved. In addition, when the execution of a long request consumes the corresponding time slice, the priority of this request will be reduced, and requests in other priority queues will be scheduled for execution, avoiding the long execution time of long requests from blocking subsequent requests. FastServe reduces the processing delay of requests to a certain extent, but does not solve the problem that long-sequence requests occupy excessive resources and the redundant calculation problem caused by insufficient resources. At the same time, FastServe does not consider the resource consumption in the incremental inference process when assigning priorities, which may lead to unreasonable division of request priorities, resulting in a decrease in system throughput and an increase in request processing delay. Summary of the Invention

[0005] The object of the present invention is to provide a multi-queue scheduling method based on the combination of full-scale and incremental inference overheads for the problems of excessive waiting queue request delay and redundant calculation caused by the FCFS (First-Come, First-Served) scheduling policy and long-sequence requests in multi-tenant and high-concurrency online inference scenarios. Drawing on the speculative sampling function of small models, a small model is used to predict the incremental inference overhead, and the characteristics of inference tasks are deeply analyzed, and the overheads of full-scale inference and incremental inference are mixedly evaluated, so as to achieve more accurate inference task priority division and resource optimization.

[0006] To achieve the above object, the specific technical solutions adopted by the present invention are as follows:

[0007] A multi-queue scheduling method based on the combination of full-scale and incremental inference overheads, comprising the following steps:

[0008] S1, full-scale and incremental overhead prediction: During the full-scale inference process of an inference task, calculate the resource consumption C i and the total memory occupancy M i , as well as the incremental calculation resource consumption C incremental,i and the memory resource consumption M incremental of incremental inference, and comprehensively calculate the total resource requirement C total and the memory overhead M total of the task;

[0009] S2. Task priority division based on resource prediction: According to the total resource requirement C total , combined with the response time requirement and task importance, tasks are assigned to queues with different priorities;

[0010] S3. Task scheduling algorithm based on task priority: Dynamically schedule tasks according to the priority queue order, adjust the execution order in combination with the system load, and give priority to executing tasks in the high-priority queue;

[0011] S4: Preemptive scheduling based on time slice supervision: Assign time slices to each task, monitor the task execution time, tasks that time out are demoted in priority and moved to the next-priority queue, triggering new task scheduling;

[0012] S5: Memory optimization based on request mounting: After preemptive scheduling, swap out the key-value cache of the task with the latest scheduling in the waiting queue from the accelerator memory to the CPU memory, and swap it back to the accelerator memory before the task is re-executed.

[0013] Preferably, in the above S1, the full-inference computing resource consumption C i is estimated by the following formula:

[0014]

[0015] where α is a correlation coefficient, and L i is the task output length;

[0016] The total memory occupancy M i is estimated as:

[0017] M i = β * L i + γ * M model

[0018] where β and γ are correlation coefficients, and M model is the size of the inference model parameters.

[0019] Preferably, in the above S1, the incremental computing resource consumption C of incremental inference incremental,i is estimated by the following formula:

[0020] C incremental,i = α * |i| 2 + β * |Y i | 2

[0021] where |i| represents the task length of the input of the i-th round of inference, and Y i represents the output length generated in this round, and α and β are incremental inference consumption coefficients;

[0022] For the memory resource consumption during the incremental process of the inference task, it can be represented by the following formula:

[0023] M incremental = M weights + M intermediate + M ouput,i

[0024] Where, M weights represents the parameter weights for storing the model, M intermediate represents the storage size for generating intermediate results during the inference process, and M ouput,i represents the memory requirement for the output tokens generated in each round.

[0025] Preferably, in S2, the rule for priority division is: compare the total resource requirement C full with the preset thresholds M1, M2,..., M N-1 and divide them into the corresponding priority queues Q1, Q2,..., Q N , where Q1 is the highest priority queue and Q N is the lowest priority queue.

[0026] Preferably, the definition method of the priority queue is:

[0027]

[0028] Preferably, in S3, the task scheduling algorithm specifically includes: according to the order of the priority queues of the tasks, when the high-priority queue is empty, schedule the tasks in the next-priority queue; delay the low-priority tasks when the system load is high, and schedule the low-priority tasks in advance when the resources are idle.

[0029] Preferably, in S4, the time slice time slice is set through experimental tuning to balance the computational overhead and the response speed, and the priority of the timeout tasks is reduced and then moved to the next-priority queue.

[0030] Preferably, in S5, the memory optimization specifically includes: by calculating the next scheduling time t next of the task, select the request T latest with the longest scheduling time, move its KV Cache from the accelerator memory to the CPU memory, and reload it to the accelerator memory when the condition: t current + Δt ≈ t next (T latest ) is satisfied.

[0031] Preferably, the calculation method of the next scheduling time t next is:

[0032]

[0033] where: t next (T i ) is the next scheduling time of task T i ; t j (T j ) is the time required for each higher-priority task T j to execute; t hunger (T i ) is the starvation waiting time of task T i , which is used to limit the longest waiting time of low-priority tasks.

[0034] Preferably, the calculation method of the request T latest with the longest scheduling time is as follows:

[0035]

[0036] where T is the set of requests in all waiting queues.

[0037] The present invention has the following features and beneficial effects:

[0038] By predicting and inferring the computing resource requirements of tasks, this method optimizes the inference efficiency, reduces resource waste, and at the same time allocates priorities according to the total resource consumption of requests, making the priority division of requests more reasonable, thereby improving the overall throughput of the system and reducing the latency of requests. Further, when the KV Cache resources are insufficient to support the inference of the next iteration, the present invention will swap out the KV Cache resources occupied by the request from the GPU video memory to the CPU memory, and then swap it back into the GPU video memory when the request is executed again, thus avoiding the redundant calculation problem caused by insufficient KV Cache resources. When the execution time of a request is too long, the preemptive scheduling mechanism will lower its priority and schedule other requests to execute, thereby solving the problem that a long sequence of requests blocks subsequent requests due to too long execution time. BRIEF DESCRIPTION OF THE DRAWINGS

[0039] Figure 1 It is a schematic diagram of the model structure of the predictor in an embodiment of the present invention.

[0040] Figure 2 It is a flowchart of the operation of an embodiment of the present invention. DETAILED DESCRIPTION OF THE INVENTION

[0041] The present invention will be described in detail below with reference to specific embodiments. The following embodiments will help those skilled in the art to further understand the present invention, but do not limit the present invention in any form. It should be noted that, without conflict, the embodiments in the present invention and the features in the embodiments can be combined with each other.

[0042] A multi-queue scheduling method based on the combination of full-scale and incremental inference overheads, and its specific steps are as follows: Figure 2 as shown below, where:

[0043] S1 Full-scale and incremental overhead prediction: In the priority division and scheduling of inference tasks, accurately predicting the computational and memory consumption of full-scale inference and incremental inference is the greatest challenge in dividing the priorities of inference tasks. Different inference methods have significantly different requirements for resources (such as memory, computing time, etc.), so reasonable overhead prediction mechanisms need to be designed for full-scale inference and incremental inference respectively. During the full-scale inference process of an inference task, the computational resource C i and the memory resource M i requirements are mainly based on the input length and the size of the predictor model. In this embodiment, the predictor model uses the Bert model:

[0044]

[0045] where α is the correlation coefficient, and L i is the task output length.

[0046] And the memory resource M i is:

[0047] M i =β*L i +γ*M model

[0048] where β and γ are correlation coefficients, and M model is the size of the inference model parameters. To predict the incremental computational overhead of an inference task, the present invention predicts the incremental computational overhead of an inference task through the method of speculative sampling. In this embodiment, a small model based on the BERT-base pre-trained model is introduced as the predictor, and the model structure of the predictor is as follows: Figure 1As shown. It mainly includes the following processes: First, the input sequence is converted into an embedded vector representation for processing in the model. Each input is transformed into a vector of a fixed dimension through the embedding layer. Second, the input embedding undergoes three linear transformations to generate query (Q), key (K), and value (V) vectors. These three vectors are used for self-attention calculation to help the model capture the relationships between different positions in the input sequence. After that, through the calculation between the query (Q), key (K), and value (V), the self-attention module can identify the dependencies between different words in the input sequence. Specifically, the self-attention mechanism calculates the similarity scores between each pair of words in the input sequence and aggregates information in a weighted manner, enabling each word to focus on the relevant parts in the sequence. The result of the self-attention output is connected with the input through a residual connection, and then through a normalization operation to stabilize the model training. Finally, after a linear transformation, the result is passed to the next layer (or as the final output). After fine-tuning, this predictor can be used to estimate the resources required for incremental calculation. During the incremental inference process of the inference task, the model gradually generates the output, and the system needs to dynamically adjust the resource allocation according to the generation situation. Assume that during the incremental inference process, the full-scale inference of the first k tokens has been completed, and the output of |Y|-k tokens has been generated. Define the computational consumption of its i-th round as:

[0049] C incremental,i = α * |i| 2 + β * |Y i | 2

[0050] C incremental,i = α * |i| 2 + β * |Y i | 2

[0051] where, |i| represents the task length of the input for the i-th round of inference, and Y i represents the output length generated in this round. The parameters α and β are the consumption coefficients specific to incremental inference, reflecting the overhead during the gradual generation of the output. For the memory resource consumption during the incremental process of the inference task, it can be represented by the following formula:

[0052] M incremental = M weights + M intermediate + M ouput,i

[0053] where, M weights represents the parameter weights for storing the model, M intermediate represents the storage size for generating intermediate results during the inference process, and M ouput,i represents the memory requirement for the output tokens generated in each round.

[0054] Combined with the analysis of the overall and incremental overheads, we can comprehensively obtain the overall computational overhead C i required for an inference task T total as follows:

[0055]

[0056] The memory overhead M total is as follows:

[0057] M total = β * L i + γ * M model + M weights + M intermediate + M ouput,i

[0058] By conducting a detailed assessment of the computational resources and memory resources consumed by full - scale and incremental inferences, we can gain a more comprehensive understanding of the resource requirement characteristics of the inference task at different stages. The full - scale inference stage mainly focuses on the impact of input data and model scale on computational and memory requirements, while the incremental inference stage reflects the gradually increasing dynamic resource overhead during the generation process of the task. Based on the above analysis, the present invention realizes real - time prediction and dynamic scheduling of the resource requirements of the inference task through a multi - queue request scheduling method based on the prediction of full - scale and incremental inference overheads of the generative model, thereby optimizing system resource allocation and improving inference efficiency.

[0059] S2 Task priority division based on resource prediction: In the process of designing and analyzing the priority scheduling of inference tasks, the priority division strategy is a key factor, which determines which tasks can be served preferentially under the condition of limited computational resources, so as to optimize the system response time and resource utilization efficiency. By predicting the computational resources of full - scale and incremental inferences, a dynamic and flexible priority division mechanism can be designed. This mechanism not only considers the resource consumption of tasks but also can be adjusted in real - time during operation to meet the resource usage requirements of different tasks.

[0060] The present invention considers the priority division strategy from the following dimensions based on the combined prediction means of full - scale and incremental overheads: computational resource consumption, response time requirements, task importance or user level.

[0061] First, computational resource consumption is an important reference factor for priority division. The incremental inference resource consumption C incremental of the inference task is estimated by a small - model predictor, full and the full - scale inference resource consumption C

[0062] C total = C full + Cincremental

[0063] Based on this, the total computing resource requirement C of the task is defined total , and this is used as an important basis for task priority.

[0064] Generally speaking, tasks that consume less computing resources are given higher priorities because these tasks can complete inference faster and release resources for other tasks to use. Therefore, tasks with low computing resource consumption will be assigned to the high-priority queue, while tasks with higher resource consumption will be assigned to the low-priority queue.

[0065] Secondly, the response time is a key metric in the inference service of long-sequence tasks, and the system needs to ensure that user requests are responded to within a certain time. To achieve this goal, the system sets a time slice for each inference task to ensure that each task can obtain corresponding system resources within the specified time.

[0066] Finally, the importance of the task or the user level is also a dimension considered in priority division. In some application scenarios, the importance of different users or tasks may vary. For example, requests from VIP users should be processed first, or certain critical service tasks (such as system monitoring tasks) need to obtain higher priorities.

[0067] Assume that the total computing resource requirement for each task is C total , the system compares the value of C total with the set priority thresholds M1, M2,..., M N-1 to finally determine the priority queue to which the task belongs. There are N priority queues Q1, Q2,..., Q N , where Q1 represents the highest-priority queue and Q N is the lowest-priority queue. Then the relevant rules are shown in the following formula:

[0068]

[0069] Based on the above formula and the hybrid inference resource requirement C of the task total to reasonably divide the task priorities, thereby improving the overall efficiency of the system.

[0070] S3 Task scheduling algorithm based on task priority: After completing the priority division, based on the resource requirement characteristics of full-scale and incremental computing, the present invention proposes a priority scheduling method that integrates the characteristics of full-scale inference and incremental inference tasks. Algorithm 1 shows the detailed process of this method. First, this method will obtain the incremental inference resource consumption C incremental based on the small model predictor and the consumption C full of the full-scale inference resources obtained based on the perception of inference task characteristics, the total computing resource requirement C of the inference task will be obtained based on the resource consumption of the two total , which is used for priority division. The tasks are assigned to different priority queues, and the queue priority order is based on the total resource requirements, so as to dynamically adjust the priority processing order of the inference tasks.

[0071]

[0072]

[0073] After completing the task priority division, the dynamic scheduling of the inference tasks will be realized according to the management strategy of the priority queue and the current system load conditions. The scheduler not only executes tasks according to the priority queue order, but also makes dynamic adjustments in combination with the usage of system resources. The system schedules tasks from high to low priority. When the Q1 queue is empty, the system schedules Q2, and so on. When the system load is high, the scheduler will delay the execution of low-priority tasks to ensure that high-priority tasks can be completed as soon as possible; if the system resources are idle, the scheduler will schedule tasks in the low-priority queue in advance to improve resource utilization.

[0074] S4 Preemptive scheduling based on time slice supervision: Based on the task priority division described in S2 of the present invention, time slice supervision is performed on each inference task in the inference process. A time slice refers to a fixed execution time allocated by the system for each task during the inference process. The time slice is usually determined by the task priority. The higher the priority of the task, the shorter the time slice usually is, and the lower the priority of the task, the longer the time slice usually is.

[0075] Set a time slice time slice , whose unit is usually seconds or milliseconds, indicating the maximum duration that the task can continuously occupy computing resources. time exec (T i ) represents the execution time that task T i has used, and time rest (T i ) represents the remaining inference time that task T i still needs. Before a time slice expires, the task will continue to execute in the execution queue; when the execution time time i of task T exec (T i ) exceeds the time slice time slice , and the task has not completed the inference yet, then preemptive scheduling is triggered.

[0076] The time slice time sliceThe length setting needs to comprehensively consider the system throughput, task response time, and the resource utilization rate of the GPU accelerator. A too-long time slice will increase the response time of other tasks, while a too-short time slice will frequently cause task switching and increase the scheduling overhead. Therefore, when designing the method of the present invention, the value of time will be first tuned through experiments and analysis to find a balance between the computational overhead and the response speed. Through reasonable time slice setting and preemption scheduling mechanism, the system can avoid long-sequence requests from blocking the execution of other requests, thereby improving the throughput of the overall inference system and the utilization efficiency of computing resources. slice As shown in Algorithm 2. First, the system performs a time slice timeout detection on all tasks in the execution queue. When the time slice time of a task has expired and the inference task has not been completed, the system marks the task as "preemptible". Secondly, the timeout task T is removed from the current execution queue and its priority is reassigned. The task T is transferred to the sub-priority queue of the waiting queue. For example, if T originally belongs to the priority queue Q, it will be moved to the sub-priority queue Q after preemption. If Q is already the lowest priority queue, the task will continue to wait in the current queue until it is rescheduled. While removing the timeout task, the scheduler selects a new task from the waiting queue. Usually, a short-sequence inference request will be selected, and a candidate task T will be preferentially selected from the lower priority queue to enter the execution queue to maximize the overall throughput. Finally, the scheduler will execute the new task, that is, put the new task T into the execution queue and allocate the corresponding time slice time for inference.

[0077] As shown in Algorithm 2. First, the system performs a time slice timeout detection on all tasks in the execution queue. When the time slice time of a task has expired and the inference task has not been completed, the system marks the task as "preemptible". Secondly, the timeout task T slice is removed from the current execution queue and its priority is reassigned. The task T i is transferred to the sub-priority queue of the waiting queue. For example, if T i originally belongs to the priority queue Q i , then after preemption, it will be moved into the sub-priority queue Q p . If Q p+1 is already the lowest priority queue, the task will continue to wait in the current queue until it is rescheduled. While removing the timeout task, the scheduler selects a new task from the waiting queue. Usually, a short-sequence inference request will be selected, and a candidate task T p will be preferentially selected from the lower priority queue to enter the execution queue to maximize the overall throughput. Finally, the scheduler will execute the new task, that is, put the new task T j into the execution queue and allocate the corresponding time slice time j for inference. slice

[0078]

[0079]

[0080] S5 Memory Optimization Based on Request Mounting: In the inference task management of multi-queue scheduling, in order to achieve a fast response to users, the key-value cache (KV Cache) of the inference task usually needs to be pre-loaded into the accelerator memory. Due to the dynamic adjustment of task priorities and the preemption scheduling mechanism, the inference tasks during execution will be removed from the execution queue, which will cause a large number of key-value caches to be placed in the GPU accelerator. As a large number of waiting requests accumulate, the KV Cache stored in the accelerator memory may occupy a large amount of resources, resulting in new incoming requests being unable to obtain the required memory resources and affecting the system performance.

[0081] In view of the problem that excessive KV Cache resource occupation in the preemptive scheduling based on time slice supervision leads to the inability to execute new requests, the present invention proposes a memory optimization method based on request mounting. After each preemptive scheduling, the system will detect the waiting requests in all priority queues and dynamically adjust the allocation of memory resources. The specific approach is to select the request that was scheduled last, temporarily unload its KV Cache to the CPU memory, and reload its KV Cache back to the accelerator memory when the request is about to be executed, so as to optimize the utilization of memory. The main content of the method is divided into the following steps:

[0082] First, the system estimates the next scheduling time of the task according to the priority queue where the inference task is currently located and the starvation waiting time. This time is used to determine the moment when the task is rescheduled in the waiting queue. For each task T i , determine the priority queue Q i where it is located and its starvation waiting time t hunger . Then, calculate the total scheduling time of all task sets Q higher in the higher priority queues above it. That is, calculate the computing time required for each higher priority task T i . The calculation formula is as follows:

[0083]

[0084] where: t next (T i ) is the next scheduling time of task T i ; t j (T j ) is the execution time required for each higher priority task T j ; t hunger (T i ) is the starvation waiting time of task T i , which is used to limit the longest waiting time of low-priority tasks.

[0085] Secondly, select the request with the longest scheduling time. After calculating the next scheduling time of all waiting requests, the system will select the request T latest with the longest scheduling time. This request is selected because it will not be scheduled in the future for some time, so its KV Cache can be temporarily moved from the accelerator memory to the CPU memory to save accelerator memory resources. The selection condition for this request is:

[0086]

[0087] where T is the set of requests in all waiting queues.

[0088] For the latest scheduled request T selected latest , the system moves its kv cache from the accelerator memory to the CPU memory, releasing the memory space of the accelerator to provide resources for other requests with higher priorities. When T latest is about to be rescheduled, the system reloads its KV Cache from the CPU memory back to the accelerator memory. This requires ensuring that the KV Cache has returned to the accelerator memory before the task is about to execute to avoid affecting the inference response time. The conditions for reloading the KV Cache are as follows:

[0089] t current +Δt≈t next (T latest )

[0090] where Δt is the time for reloading the kv cache, and t current is the current time.

[0091] To ensure the efficient utilization of memory resources, the system repeats the above process after each preemption scheduling and dynamically adjusts the scheduling policy and memory resource allocation according to the tasks in the current waiting queue. When a new request arrives, the system promptly detects whether there is a task KV Cache that can be unloaded from the accelerator memory and minimizes the memory occupation as much as possible. When the inference task has its KV Cache removed after consuming all time slices, the system ensures that it will not have an obvious impact on the response time of the user because the KV Cache can be reloaded at any time when needed. By reasonably managing the KVCache of low-priority tasks, the system ensures that the memory resources are not occupied for a long time, thereby improving the overall memory utilization rate and ensuring that new requests obtain sufficient computing resources. Specifically, as shown in Algorithm 3:

[0092]

[0093]

[0094] Comparative example:

[0095] The following is the specific application of the method of the present invention, the experiments and analyses of the multi-queue request scheduling method based on the prediction of the full-scale and incremental inference overheads of the generative model:

[0096] The hardware environment settings adopted in the application are based on the Atlas 800 AI server, and a complete software environment is built based on the Ubuntu 20.04 operating system to support the inference and experiments of AI models, as shown in Table 1 below. The experimental toolchain includes the MindSpore

[70] deep learning framework that supports the Ascend inference chip, which is mainly used for the training and inference tasks of large-scale models to accelerate the inference process; the Ascend parallel computing architecture library CANN, which improves the parallel computing ability of the Ascend processor and makes matrix operations and data transmission in the experiment more efficient; the present invention integrates the functional code into the high-performance inference framework tool library MindSpore Serving. MindSpore Serving is a lightweight and high-performance service module designed to help MindSpore developers efficiently deploy online inference services in the production environment. When users complete model training using MindSpore and export the MindSpore model, they can use MindSpore Serving to create an inference service for this model. In addition, Python 3.9 is also used as the main programming language for the experiment, responsible for script writing and task scheduling.

[0097]

[0098] Table 1

[0099] At the same time, the present invention uses AI models of the Llama2 series, mainly the Llama2 model, with a model scale of 7B. In terms of the dataset, the experiments use Alpaca and SharedGP as the datasets for inference scheduling. This experiment combines the Llama2-7B model and the corresponding datasets, conducts performance tests for large model inference tasks, and evaluates the end-to-end performance improvement of the model and its throughput performance under different tasks.

[0100]

[0101]

[0102] Alpaca latency data

[0103]

[0104] ShareGPT latency data

[0105]

[0106] Alpaca throughput data

[0107]

[0108]

[0109] ShareGPT throughput data

[0110] Functions and effects of the embodiments

[0111] According to a multi-queue scheduling method based on the combination of full-scale and incremental inference overheads provided by this example, first, the resource overhead of the incremental inference task is estimated by fine-tuning the prediction model, and the resource overhead of the full-scale inference process is estimated by combining the feature perception technology. Based on the total resource overhead of the incremental and full-scale, the priority division of the inference task is completed; secondly, based on the priority and queue resource requirements, a dynamic scheduling strategy is adopted to batch allocate tasks to the large model inference system, optimizing the overall throughput of the system while ensuring the response speed of high-priority tasks; finally, the scheduling flexibility is improved through the preemptive scheduling strategy of time slice monitoring. When the task execution times out, the preemptive mechanism is triggered, and combined with the memory optimization technology based on request mounting, the dynamic scheduling of GPU and CPU memory is realized by using the Swap method, effectively reducing the memory loss. The present invention realizes the efficient scheduling of inference tasks and resource optimization, significantly improving the overall performance of the generative model inference system.

[0112] The above examples are only used to illustrate the specific implementation manners of the present invention, and the present invention is not limited to the descriptions of the above examples.

[0113] The above shows and describes the basic principles, main features and advantages of the present invention. Those skilled in the art of this industry should understand that the present invention is not limited by the above embodiments. The above embodiments and the descriptions in the specification are only preferred examples of the present invention and are not used to limit the present invention. Without departing from the spirit and scope of the present invention, the present invention will have various changes and improvements, and these changes and improvements all fall within the scope of the present invention claimed. The scope of protection claimed by the present invention is defined by the appended claims and their equivalents.

Claims

1. A multi-queue scheduling method based on merging full and incremental inference overhead, characterized in that: The steps include: S1. Full and incremental cost prediction: During the full inference process of the inference task, the resource consumption C is calculated based on the input length and the predictor model size. i And the total memory usage M i , and the incremental computing resource consumption C of incremental reasoning incremental,i and memory resource consumption M incremental , and comprehensively calculate the total resource requirements of the task C total and memory overhead M total ; S2. Task priority division based on resource forecast: According to the total resource demand C total ,Tasks are assigned to queues of different priorities based on the response time requirements and task importance; S3, Task scheduling algorithm based on task priority: dynamically schedule tasks according to the priority queue order, adjust the execution order according to the system load, and give priority to tasks in the high priority queue; S4: Preemptive scheduling based on time slice supervision: Allocate time slices to each task, monitor task execution time, and timed-out tasks are lowered in priority and moved into the secondary priority queue, triggering new task scheduling; S5: Memory optimization based on request mounting: After preemptive scheduling, the key-value cache of the latest scheduled task in the waiting queue is swapped out from the accelerator memory to the CPU memory, and then swapped back to the accelerator memory before the task is re-executed.

2. The method according to claim 1, characterized in that: In S1, the total inference computing resource consumption C i Estimated by the following formula: Among them, α is the correlation coefficient, L i Output the length for the task; Total memory usage M i The estimate is: M i =β*L i +γ*M model Among them, β, γ are correlation coefficients, M model is the inference model parameter size.

3. The method according to claim 1, characterized in that: In S1, the incremental computing resource consumption C of the incremental reasoning incremental,i Estimated by the following formula: C incremental,i =α*|i| 2 +β*|Y i | 2 Among them, |i| represents the task length of the i-th round of reasoning input, Y i represents the output length generated in this round, α and β are the incremental reasoning consumption coefficients; The memory resource consumption of the reasoning task during the incremental process can be expressed by the following formula: M incremental =M weights +M intermediate +M ouput,i Among them, M weights Represents the parameter weight of the storage model, M intermediate Indicates the storage size of the intermediate results generated during the inference process, M ouput,i Indicates the memory requirement for the output token generated in each round.

4. The method according to claim 1, characterized in that: In S2, the total resource requirement C of the calculation task is total , the expression is as follows: The total memory cost of computing tasks M total , the expression is as follows: M total =β*L i +γ*M model +M weights +M intermediate +M ouput,i 。 5. The method according to claim 1, characterized in that In S2, the priority division rule is: the total resource demand C full and the preset thresholds M1, M2, ..., M N-1 Compare and divide into corresponding priority queues Q1, Q2, ..., Q N , where Q1 is the highest priority queue, Q N The lowest priority queue.

6. The method according to claim 5, characterized in that The priority queue is defined as follows:

7. The method according to claim 1, characterized in that In S3, the task scheduling algorithm specifically includes: scheduling tasks in the secondary priority queue when the high priority queue is empty according to the priority queue order of the tasks; delaying low priority tasks when the system load is high, and scheduling low priority tasks in advance when resources are idle.

8. The method according to claim 1, characterized in that In S4, the time slice time slice The setting is optimized through experiments to balance computing overhead and response speed. The priority of timed tasks is reduced and moved to the secondary priority queue.

9. The method according to claim 1, characterized in that: In S5, memory optimization specifically includes: calculating the next scheduling time t of the task next , select the request T with the longest scheduling time latest , swap out its KV Cache from the accelerator memory to the CPU memory, and when the following conditions are met: current +Δt≈t next (T latest ) to the accelerator memory.

10. The method according to claim 9, characterized in that The next scheduling time t next The calculation method is: Where: t next (T i ) is task T i The next scheduling time; t j (T j ) is each higher priority task T j The time required for execution; t hunger (T i ) is task T i The starvation waiting time is used to limit the maximum waiting time of low-priority tasks.

11. The method according to claim 8, characterized in that The request with the longest scheduling time T latest The calculation method is: Where T is the set of all requests in the waiting queue.

Citation Information

Cited By

  • Task adjustment method and device, storage medium, electronic equipment and computer program product

    CN120448078A

  • Data merging scheduling method and device and storage medium

    CN120780681A

  • Data merging scheduling method and device, and storage medium

    CN120780681B

  • Mixed request queue scheduling method oriented to SLO perception of intelligent agent system

    CN121151342A

  • EDA simulation task scheduling method, device and equipment, medium and program product

    CN121387501A