A method, apparatus, and request management system for model inference request management
By predicting and analyzing the observation indicators of large-model inference services, determining scheduling strategies and suggestions, the problem that high-priority requests in the existing technology cannot be responded in a timely manner, and efficient utilization and reasonable allocation of resources are achieved.
Patent Information
- Application Number
- CN202411485019.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-10-23
- Publication Date
- 2025-06-20
- Estimated Expiration
- 2044-10-23
AI Technical Summary
In big model inference services, the existing technology has shortcomings in user priority distinction and resource scheduling, resulting in high-priority requests being unable to respond in a timely manner and resulting in waste of computing power.
By obtaining the observation indicators of the model inference service, metric prediction is carried out to determine the target memory utilization rate of the central processor and the target memory utilization rate of the graphics processor, scheduling strategies are determined based on the prediction indicators, and scheduling suggestions are determined in combination with the request queue, and request scheduling is carried out to optimize resource utilization.
Reasonable scheduling of requests of different priority levels is achieved, computing power is avoided, and efficient utilization and reasonable allocation of resources are improved.
Smart Images

Figure CN119336471B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of computer technology, and particularly to a method, device, and request management system for managing requests for model inference. Background Art
[0002] In the field of large model inference services, with the progress of technology and the growth of market demand, users are increasingly frequent in using large model inference services, and the load pressure on the enterprise's inference cluster is getting higher and higher. At the same time, since the identities of users using large model inference services are different, and the tasks of initiating large model inference requests are also diverse, this leads to different large model inference requests having different service priorities and urgencies. When hardware resources are limited, high-priority and high-urgency requests should be served prior to low-priority requests.
[0003] To meet the diverse user needs, the inference service framework not only needs to efficiently process large-scale data, but also needs to be able to reasonably allocate resources in a multi-user environment. However, although the current mainstream large model inference frameworks have excellent performance in processing large-scale data and complex computing tasks, there are many deficiencies in user priority differentiation and resource scheduling. In the field of cloud computing, the method of dividing request resource pools can be used to distinguish the resource quotas allocated to high-priority requests and low-priority requests, and resource pools with different resource amounts are specified for high-priority requests and low-priority requests respectively, so as to provide differentiated services for requests with different priorities and urgencies.
[0004] However, the method of dividing resource pools to meet the service requirements of requests with different priorities and urgencies is relatively rigid in request processing. When a request is routed to the corresponding resource pool, the request will only be served by the service instances in that resource pool. When there are vacancies in other resource pools, the request cannot be reallocated to the resource pool with vacancies. Even if the request can be rescheduled, the KVCache of the request Prompt needs to be recalculated, resulting in waste of computing power. Therefore, the method of using resource pools to serve requests with different priorities will inevitably cause waste of computing power, and this phenomenon is particularly obvious when the overall hardware resources are insufficient. Therefore, how to reasonably schedule requests during model inference becomes a problem to be solved. Summary of the Invention
[0005] The present invention provides a method, device, and request management system for managing requests for model inference to solve the problem of unreasonable request scheduling during model inference.
[0006] According to one aspect of the present invention, there is provided a method for managing requests for model inference, which is applied to a request processing engine in a request management system. The request management system further includes a model inference service. The method includes:
[0007] Obtain the observation metrics of the model inference service;
[0008] Perform metric prediction based on the observation metrics to obtain predicted metrics, where the predicted metrics include the target memory utilization rate of the central processing unit and the target video memory utilization rate of the graphics processing unit;
[0009] Determine a scheduling strategy according to the predicted metrics, determine a scheduling recommendation according to the scheduling strategy in combination with the request queue, and add the scheduling recommendation to the recommendation buffer queue;
[0010] When the model inference service executes request scheduling, read the scheduling recommendation from the recommendation buffer queue, determine a target scheduling decision according to the scheduling recommendation in combination with the scheduling decision of the model inference service, and control the model inference service to schedule corresponding requests to execute model inference according to the target scheduling decision.
[0011] According to another aspect of the present invention, there is provided a device for managing requests for model inference, which is applied to a request processing engine in a request management system. The request management system further includes a model inference service. The device includes:
[0012] An observation module for obtaining the observation metrics of the model inference service;
[0013] A state prediction module for performing metric prediction based on the observation metrics to obtain predicted metrics, where the predicted metrics include the target memory utilization rate of the central processing unit and the target video memory utilization rate of the graphics processing unit;
[0014] A request scheduler for determining a scheduling strategy according to the predicted metrics, determining a scheduling recommendation according to the scheduling strategy in combination with the request queue, and adding the scheduling recommendation to the recommendation buffer queue;
[0015] A side-loading scheduling module for, when the model inference service executes request scheduling, reading the scheduling recommendation from the recommendation buffer queue, determining a target scheduling decision according to the scheduling recommendation in combination with the scheduling decision of the model inference service, and controlling the model inference service to schedule corresponding requests to execute model inference according to the target scheduling decision.
[0016] According to another aspect of the present invention, there is provided a request management system, including: a request processing engine and a model inference service;
[0017] The request processing engine is configured to: obtain the observation metrics of the model inference service; perform metric prediction based on the observation metrics to obtain predicted metrics, where the predicted metrics include the target memory utilization rate of the central processing unit and the target video memory utilization rate of the graphics processing unit; determine a scheduling policy according to the predicted metrics, determine a scheduling recommendation in combination with the request queue according to the scheduling policy, and add the scheduling recommendation to the recommendation buffer queue; when the model inference service executes request scheduling, read the scheduling recommendation from the recommendation buffer queue, determine a target scheduling decision in combination with the scheduling decision of the model inference service according to the scheduling recommendation, and control the model inference service to schedule corresponding requests to execute model inference according to the target scheduling decision.
[0018] According to another aspect of the present invention, there is provided an electronic device, which includes:
[0019] at least one processor, and a memory communicatively connected to the at least one processor;
[0020] wherein, the memory stores a computer program executable by the at least one processor, and when the computer program is executed by the at least one processor, the at least one processor is enabled to execute the request management method for model inference according to any embodiment of the present invention.
[0021] According to another aspect of the present invention, there is provided a computer-readable storage medium storing computer instructions for causing a processor to implement the request management method for model inference according to any embodiment of the present invention when executed.
[0022] According to another aspect of the present invention, there is provided a computer program product including a computer program, where the computer program implements the request management method for model inference according to any embodiment of the present invention when executed by a processor.
[0023] The technical solution of the embodiment of the present invention is as follows: obtain the observation indexes of the model inference service; perform index prediction according to the observation indexes to obtain predicted indexes, where the predicted indexes include the target memory utilization rate of the central processing unit and the target video memory utilization rate of the graphics processing unit; determine a scheduling strategy according to the predicted indexes, determine a scheduling recommendation according to the scheduling strategy in combination with the request queue, and add the scheduling recommendation to the recommendation buffer queue; when the model inference service executes request scheduling, read the scheduling recommendation from the recommendation buffer queue, determine the target scheduling decision according to the scheduling recommendation in combination with the scheduling decision of the model inference service, and control the model inference service to schedule corresponding requests to execute model inference according to the target scheduling decision, which solves the problem of unreasonable request scheduling in the model inference process. By predicting through the observation indexes of the model inference service, predicted indexes are obtained, and the predicted indexes include the target memory utilization rate of the central processing unit and the target video memory utilization rate of the graphics processing unit. In the embodiment of the present application, the predicted indexes can be used to represent different indexes of the model inference service in the next time, that is, predict the indexes of the model inference service in the subsequent time; determine the scheduling strategy according to the predicted indexes, further determine the scheduling recommendation in combination with the request queue, and add it to the recommendation buffer queue; when the model inference service executes request scheduling, read the scheduling recommendation from the recommendation buffer queue, determine the target scheduling decision according to the scheduling recommendation in combination with the scheduling decision of the model inference service, intervene in the request scheduling of the model inference service, and finally the target scheduling decision controls the model inference service to schedule corresponding requests to execute model inference; interfere with the scheduling decision of the model inference service through index prediction, provide the target scheduling decision for the model inference service, and control the model inference service to execute model inference through reasonable request scheduling, so as to achieve efficient utilization and reasonable allocation of resources.
[0024] It should be understood that the content described in this part is not intended to identify the key or important features of the embodiments of the present invention, nor is it used to limit the scope of the present invention. Other features of the present invention will become easily understood through the following description. BRIEF DESCRIPTION OF THE DRAWINGS
[0025] In order to more clearly illustrate the technical solutions in the embodiments of the present invention, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the drawings in the following description are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained according to these drawings.
[0026] Figure 1 is a flowchart of a method for managing requests for model inference according to Embodiment 1 of the present invention;
[0027] Figure 2 is a flowchart of a method for managing requests for model inference according to Embodiment 2 of the present invention;
[0028] Figure 3 It is a schematic diagram of an implementation example of request scheduling provided in Embodiment 2 of the present invention;
[0029] Figure 4 It is a schematic structural diagram of a request management device for model inference provided in Embodiment 3 of the present invention;
[0030] Figure 5 It is a schematic structural diagram of a request management system provided in Embodiment 4 of the present invention;
[0031] Figure 6 It is a schematic diagram of an implementation example of request management provided in Embodiment 4 of the present invention;
[0032] Figure 7 It is a schematic structural diagram of an electronic device for implementing the request management method for model inference in an embodiment of the present invention. Detailed implementation manners
[0033] In order to enable those skilled in the art to better understand the solution of the present invention, the technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.
[0034] It should be noted that the terms "first", "second", etc. in the specification and claims of the present invention and the above drawings are used to distinguish similar objects, and do not necessarily need to describe a specific order or sequence. It should be understood that such data can be interchanged under appropriate circumstances so that the embodiments of the present invention described herein can be implemented in an order different from those illustrated or described herein. In addition, the terms "comprising" and "having" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product or device that includes a series of steps or units does not necessarily have to be limited to those steps or units clearly listed, but may include other steps or units not clearly listed or inherent to these processes, methods, products or devices.
[0035] Embodiment 1
[0036] Figure 1 It is a flowchart of a request management method for model inference provided in Embodiment 1 of the present invention. This embodiment is applicable to the situation of managing requests executed for model inference services. This method is applied to a request processing engine in a request management system, and the request management system further includes a model inference service. As Figure 1 shown, this method includes:
[0037] S101. Obtain the observation metrics of the model inference service.
[0038] In this embodiment, the model inference service can be the inference service of any type of model, such as the inference service of a large language model. The observation metrics can be understood as the metrics used to describe the relevant information during the model inference service process. The observation metrics can be the memory / video memory utilization rate of the CPU / GPU, the number of relevant requests during the model inference process, the key-value pair cache size of the tokens obtained by inference, and so on.
[0039] When the model inference service receives a request related to inference from the user, it performs the corresponding model inference. During the model inference process, the CPU (central processing unit) and GPU (graphics processing unit) will be used. For example, the GPU executes the model inference task. During the inference process, metrics such as the video memory utilization rate and core utilization rate of the GPU will change. Some data will be stored in the CPU memory during the model inference process, and the CPU memory utilization rate will change. At the same time, as the model inference progresses, some parameters involved in the model inference itself will also change. For example, the number of requests to be left for Prefilling, the request length, etc.
[0040] The observation metrics of the model inference service can be obtained through an interface. The request processing engine obtains the observation metrics of the model inference service by calling the corresponding interface. For example, the model inference service can be exposed to the request processing engine in the form of an API. The observation metrics can be obtained at a certain period or after meeting certain trigger conditions. For example, the trigger condition is that the model inference service needs to schedule a new request, that is, after the model inference service completes the scheduling of one request, it schedules other requests waiting to be processed.
[0041] S102. Perform metric prediction based on the observation metrics to obtain predicted metrics, where the predicted metrics include the target memory utilization rate of the central processing unit and the target video memory utilization rate of the graphics processing unit.
[0042] In this embodiment, the predicted metrics can be understood as the metric information of the model inference service obtained through prediction. The predicted metrics are the metric information for a period of time after the collection time of the observation metrics. For example, the collection time of the observation metrics is t1, which represents the observation metrics of the model inference service at time t1, and the predicted metrics are the predicted observation metrics of the model inference service at time (t1 + 10s). The target memory utilization rate can be understood as the predicted memory utilization rate of the central processing unit; the target video memory utilization rate can be understood as the predicted video memory utilization rate of the graphics processing unit.
[0043] Analyze the observation metrics. The observation metrics can reflect information about the model inference service from different perspectives. By analyzing the observation metrics, predict the future load status of the model inference service to obtain prediction metrics, which include the target memory utilization rate of the central processing unit and the target video memory utilization rate of the graphics processing unit. Models, functions, etc. can be pre-constructed to predict the observation metrics through the models and functions.
[0044] S103. Determine the scheduling policy according to the prediction metrics, determine the scheduling recommendation in combination with the request queue according to the scheduling policy, and add the scheduling recommendation to the recommendation buffer queue.
[0045] In this embodiment, the scheduling policy can be understood as a policy indication for assisting the model inference service in scheduling. For example, pausing requests, resuming requests, etc. The request queue can be understood as a queue formed by requests from different users. The request queue can be a first-in-first-out (FIFO) queue or other types of queues. The recommendation buffer queue can be understood as a queue formed by scheduling recommendations, which can be a first-in-first-out queue or other types of queues.
[0046] Analyze the processing capabilities of the CPU and GPU according to the prediction metrics, and determine the scheduling policy according to the processing capabilities of the CPU and GPU. For example, when the processing capability of the GPU is low, pause some requests to ensure the request processing speed. When the processing capability of the GPU is high, resume some requests to ensure the fast processing of requests. Select appropriate requests from the request queue according to the scheduling policy to generate scheduling recommendations. For example, the scheduling recommendation is to pause request 1 and resume request 2, etc.; add the scheduling recommendation to the recommendation buffer queue.
[0047] Since the request processing engine is asynchronous with the internal operation of the model inference service in request processing and scheduling decisions, when making scheduling decisions based on the collected metrics, this metric information lags behind the state of the model inference service itself at the time of decision-making. Therefore, when making scheduling decisions, the scheduling decisions should be made based on the predicted future state of the model inference service.
[0048] S104. When the model inference service executes request scheduling, read the scheduling recommendation from the recommendation buffer queue, determine the target scheduling decision in combination with the scheduling decision of the model inference service according to the scheduling recommendation, and control the model inference service to schedule the corresponding requests to execute model inference according to the target scheduling decision.
[0049] In this embodiment, the target scheduling decision can be understood as the policy finally adopted by the model inference service when scheduling requests.
[0050] The model inference service can execute requests sent by multiple users simultaneously. For example, the model inference service provides model inference-related services for multiple users. After receiving a request sent by a user, if the execution conditions are met, the model inference service can directly execute this request for model inference. If the execution conditions are not met, the request can also be placed in a waiting queue and executed after the execution conditions are met. The execution conditions can be that the number of requests being executed by the model inference service is less than a certain threshold, etc. After the model inference service finishes executing a request, it continues to obtain requests from the waiting queue and execute them. When the model inference service performs request scheduling, that is, when the model inference service obtains a new request and executes it, it reads the scheduling suggestion from the suggestion buffer queue. The scheduling suggestion and the scheduling decision of the model inference service are fused to determine the target scheduling decision. For example, determine the requests suggested to be paused or resumed in the scheduling suggestion, and judge whether the scheduling decision of the model inference service schedules this request in the same way. If so, determine that this request is to be paused or resumed; if not, the corresponding processing for this request can be determined according to the pre-set priority and other information. The model inference service is controlled according to the target scheduling decision to schedule the corresponding request to execute model inference.
[0051] After reading the scheduling suggestion from the suggestion buffer queue, this scheduling suggestion can be deleted from the suggestion buffer queue. In an embodiment of the present application, the trigger condition can also be set to that the scheduling suggestion in the suggestion buffer queue is 0, that is, after the scheduling suggestion in the suggestion buffer queue is read and there is no scheduling suggestion at this time, it is determined that the trigger condition is met, and the observation index is obtained to generate a new scheduling suggestion.
[0052] A method for managing requests for model inference according to an embodiment of the present invention solves the problem of unreasonable request scheduling in the model inference process. Through prediction using the observation index of the model inference service, a prediction index is obtained. The prediction index includes the target memory utilization rate of the central processing unit and the target video memory utilization rate of the graphics processing unit. In an embodiment of the present application, the prediction index can be used to represent different indexes of the model inference service in the next period of time, that is, predict the indexes of the model inference service in the subsequent time; determine the scheduling strategy through the prediction index, further combine the request queue to determine the scheduling suggestion, and add it to the suggestion buffer queue; when the model inference service performs request scheduling, read the scheduling suggestion from the suggestion buffer queue, and determine the target scheduling decision according to the scheduling suggestion combined with the scheduling decision of the model inference service, intervene in the request scheduling of the model inference service, and finally the target scheduling decision controls the model inference service to schedule the corresponding request to execute model inference; interfere with the scheduling decision of the model inference service through index prediction, provide the target scheduling decision for the model inference service, and control the model inference service to execute model inference through reasonable request scheduling, realizing the efficient utilization and reasonable allocation of resources.
[0053] Embodiment 2
[0054] Figure 2 This is a flowchart of a request management method for model inference provided in the second embodiment of the present invention. This embodiment is refined on the basis of the above embodiment. As Figure 2 shown, the method includes:
[0055] S201. Obtain the observation metrics of the model inference service.
[0056] The observation metrics include at least one of the processor parameter metrics and the request metrics.
[0057] In this embodiment, the processor parameter metrics can be understood as the parameters of the processor, which can be the same type of parameters of different types of processors or different types of parameters of the same processor. The request metrics can be understood as the relevant information of the requests.
[0058] Among them, the processor parameter metrics include at least one of the following:
[0059] The memory utilization rate of the central processing unit;
[0060] The video memory utilization rate of the graphics processing unit;
[0061] The core utilization rate of the graphics processing unit;
[0062] Among them, the request metrics include at least one of the following:
[0063] The number of requests to be left for Prefilling;
[0064] The length of the requests to be left for Prefilling;
[0065] The number of requests in the Decoding stage;
[0066] The key-value pair cache size of the requests to be swapped into the memory of the central processing unit;
[0067] The key-value pair cache size of the requests to be swapped into the video memory of the graphics processing unit.
[0068] Among them, the CPU memory usage rate and the GPU video memory usage rate are used to measure the usage of the memory and video memory by the current model inference service, so as to judge whether the current hardware storage resources are sufficient accordingly;
[0069] The GPU core utilization rate is used to measure the usage of the GPU core by the current model inference service, so as to judge whether the current hardware computing resources are sufficient accordingly;
[0070] The number of requests to be left for Prefilling is used to predict the future load status of the model inference service;
[0071] The request length to be left for Prefilling, which is used to predict the future load status of the model inference service;
[0072] The number of requests in the Decoding stage, which is used to predict the future load status of the model inference service;
[0073] The KV Cache size of the requests to be swapped into CPU memory, which is used to predict the future load status of the model inference service;
[0074] The KV Cache size of the requests to be swapped into GPU memory, which is used to predict the future load status of the model inference service.
[0075] The model based on the Transformer architecture, whose core is the Attention mechanism. The inference is divided into two stages: prefill (input understanding and initialization) and decoding (recursive inference and decoding output). After each request is initiated, the inference process will first go through a Prefill process. The Prefill process will calculate all the user inputs and generate the corresponding KV cache, and then go through several decoding processes. In each decoding process, the server will generate a character and put it into the KV cache. The predicted result of the inference is put into the input again, and so on, until the final result is inferred. When a new request comes in, after the prefill is completed, it will continuously iterate through the decoding process. After each decoding stage ends, the result will be returned to the user on the spot. Such a generation process is called streaming.
[0076] In the decoding phase, it is necessary to calculate the attention of the current token and all the previously generated tokens, so it is necessary to calculate the k and v vectors of all tokens. However, the KV values of the previous tokens are repeatedly calculated in each round of decoding. Therefore, they can be saved as two Tensors of [seq_len - 1, inner_dim], and only the kv values of the current token need to be calculated in each round of calculation.
[0077] The number of requests to be left for Prefilling refers to the number of requests to be processed in the Prefilling stage; the request length to be left for Prefilling refers to the length of the requests to be processed in the Prefilling stage. The length can be the number of tokens obtained after tokenizing the input sequence corresponding to all the user inputs calculated in the prefill process. The number of requests in the Decoding stage refers to the number of requests in the Decoding stage.
[0078] The cache size of the key-value pair of the request to be swapped into the memory of the central processing unit, that is, the KVCache size, is the size of the actually used memory. The cache size of the key-value pair of the request to be swapped into the video memory of the graphics processing unit, that is, the KVCache size, is the size of the actually used video memory. During the model inference process, the CPU memory and GPU video memory are pre-applied to store the key-value pairs generated during the inference process. The KVCache size in the embodiments of the present application is the size of the actually used video memory.
[0079] Optionally, the processor parameter metrics are updated when the first metric update condition is met;
[0080] Among them, the first metric update condition includes at least one of the following:
[0081] The system time meets the periodic update condition;
[0082] The model inference service completes the scheduling of a request.
[0083] In this embodiment, the first metric update condition can be understood as the condition for updating the processor parameter metrics. The periodic update condition can be understood as the condition for periodically updating the processor parameter metrics. For example, if the update period is set to T, when the time difference between the system time and the time of the last update is equal to T (or not less than T), the periodic update condition is met, and the processor parameter metrics need to be periodically updated. When the model inference service completes the scheduling of a request, at this time, the processor parameter metrics will also change, and the processor parameter metrics need to be updated. The processor parameter metrics in the embodiments of the present application can be updated when any of the above conditions is met.
[0084] Optionally, the request metrics are updated when the second metric update condition is met.
[0085] Among them, the second metric update condition includes: the model inference service completes the scheduling of a request.
[0086] In this embodiment, the second metric update condition can be understood as the condition for updating the processor parameter metrics. When the model inference service completes the scheduling of a request, the request metrics at this time will change, and the request metrics need to be updated. The request metrics in the embodiments of the present application can be updated after the model inference service completes the scheduling of a request.
[0087] In the embodiments of the present application, when collecting observation indicators, it is not necessary to collect various indicators in the model inference service in a synchronous process. Instead, at a certain time interval, the indicators are collected asynchronously and placed in an internal cache. At the same time, part of the indicator information is updated according to the internal request scheduling decision of each model inference. When the request processing engine accesses the indicator information through the API, it directly retrieves the indicator information from the cache and returns it. Compared with the synchronous collection method, the caching strategy can effectively ensure the performance of the indicator collection interface and the overall performance of the inference service framework.
[0088] S202. Use the memory size of the central processing unit, the video memory size of the graphics processing unit, and the observation indicators as inputs for prediction to determine the target memory utilization rate of the central processing unit and the target video memory utilization rate of the graphics processing unit.
[0089] Among them, the memory size of the central processing unit and the video memory size of the graphics processing unit are applied for when the program starts.
[0090] Apply for the memory size of the central processing unit and the video memory size of the graphics processing unit when the program starts. Use the memory size of the central processing unit, the video memory size of the graphics processing unit, and the observation indicators as inputs and input them into a pre-trained model, function, etc. for prediction to obtain the target memory utilization rate of the central processing unit and the target video memory utilization rate of the graphics processing unit.
[0091] Exemplarily, the indicator prediction of the present application can be implemented through algorithms such as linear regression. For example, through linear regression algorithm fitting, the linear relationship between the memory utilization rate of the central processing unit, the video memory utilization rate of the graphics processing unit, and the memory size of the central processing unit, the video memory size of the graphics processing unit, and the observation indicators is fitted to obtain a function expression. After determining the memory size of the central processing unit, the video memory size of the graphics processing unit, and the observation indicators, substitute them into the function expression to obtain the target memory utilization rate and the target video memory utilization rate; among them, the target memory utilization rate and the target video memory utilization rate can be calculated through one function expression or two function expressions. When performing indicator prediction, the memory utilization rate of the central processing unit, the video memory utilization rate of the graphics processing unit, the core utilization rate of the graphics processing unit, the number of requests to be Prefilled, the length of the requests to be Prefilled, the number of requests in the Decoding stage, the key-value pair cache size of the requests to be swapped into the memory of the central processing unit, and the key-value pair cache size of the requests to be swapped into the video memory of the graphics processing unit can be used as inputs for prediction.
[0092] S203. If the target video memory utilization rate is higher than the first preset threshold and the target memory utilization rate is lower than the second preset threshold, determine the scheduling strategy as suspending requests.
[0093] In this embodiment, the first preset threshold and the second preset threshold can be understood as preset thresholds, which can be set according to requirements such as the performance, speed, and service type of the processor. By comparing the target video memory utilization rate with the first preset threshold and the target memory utilization rate with the second preset threshold, if the target video memory utilization rate is higher than the first preset threshold and the target memory utilization rate is lower than the second preset threshold, it indicates that the memory available for the GPU to process requests is limited at this time. The scheduling policy is determined to be suspending requests to ensure the speed and performance of request processing. When suspending requests, it is recommended to suspend some requests with lower priorities and urgencies, so that the remaining requests can preempt the hardware resources used by these suspended requests. The KVCache of the preempted requests will be swapped to the CPU KVCache storage space.
[0094] S204. If the target video memory utilization rate is lower than the third preset threshold, determine the scheduling policy as resuming requests.
[0095] In this embodiment, the third preset threshold can be understood as a preset threshold, which can be set according to requirements such as the performance, speed, and service type of the processor. If the target video memory utilization rate is lower than the third preset threshold, it means that the resources available for the GPU to process requests are relatively abundant at this time. The scheduling policy is determined to be resuming requests to resume the suspended requests. When resuming requests, the preempted requests with higher priorities and urgencies can be resumed preferentially.
[0096] If both the target memory utilization rate of the GPU and the target video memory utilization rate of the CPU are higher than the corresponding thresholds, no suggestion is made because the exchange between the GPU KVCache and the CPU KVCache cannot be performed at this time.
[0097] S205. Determine a scheduling suggestion according to the scheduling policy in combination with the request queue, and add the scheduling suggestion to the suggestion buffer queue.
[0098] Optionally, the request queue includes a running queue and a pending queue. Determining a scheduling suggestion according to the scheduling policy in combination with the request queue includes:
[0099] A1. When the scheduling policy is suspending requests, traverse the running queue, sort the requests in the running queue according to the priority levels and urgencies of the requests in the running queue, screen out at least one request as a suspended request according to the sorting result, add the suspended request to the suspended request queue, and generate a scheduling suggestion according to the suspended request queue.
[0100] In this embodiment, the running queue is the Running queue; the suspended request can be understood as the request of the user that needs to be suspended; the suspended request queue can be understood as the queue for caching the requests that need to be suspended for processing.
[0101] When the scheduling policy is a suspension request, traverse the run queue, determine the priority level and urgency of each request in the run queue, sort the requests in the run queue according to the priority level and urgency. For example, arrange the requests with higher priority levels and higher urgencies in the front. Filter out at least one request with a lower priority level and a lower urgency as a suspension request according to the sorting result, add all the filtered suspension requests to the suspension request queue, and generate a scheduling recommendation based on the suspension request queue. For example, the scheduling recommendation is to suspend the suspension requests in the suspension request queue.
[0102] A2. When the scheduling policy is a resume request, traverse the suspension queue, sort the requests in the suspension queue according to the priority level and urgency of the requests in the suspension queue, filter out at least one request as a resume request according to the sorting result, add the resume request to the resume request queue, and generate a scheduling recommendation based on the resume request queue.
[0103] In this embodiment, the suspension queue is the Hanging queue; the resume request can be understood as the request of the user to be suspended; the resume request queue can be understood as the queue for caching the requests to be resumed.
[0104] When the scheduling policy is a resume request, traverse the suspension queue, determine the priority level and urgency of each request in the suspension queue, sort the requests in the run queue according to the priority level and urgency. For example, arrange the requests with higher priority levels and higher urgencies in the front. Filter out at least one request with a higher priority level and a higher urgency as a resume request according to the sorting result, add all the filtered resume requests to the resume request queue, and generate a scheduling recommendation based on the resume request queue. For example, the scheduling recommendation is to resume the resume requests in the resume request queue.
[0105] Optionally, the prediction metric further includes: the number of requests;
[0106] Correspondingly, the number of suspension requests or resume requests is the number of requests.
[0107] In this embodiment, the number of requests is the number of requests that should be suspended / resumed (pause / resume) in this round of scheduling. The number of requests can be predicted simultaneously with the target memory utilization rate and the target video memory utilization rate when performing metric prediction. When the prediction metric further includes the number of requests, when filtering the resume requests and suspension requests, filter out a preset number of suspension requests or resume requests.
[0108] In model inference services, requests with different priorities and urgencies have differences in hardware usage priorities. Requests with high priorities and high urgencies should have priority in occupying hardware resources for service. Therefore, it is necessary to schedule these tasks. When hardware resources are insufficient or the system load is high, it is necessary to determine which tasks should preempt other tasks for priority processing. In the embodiments of this application, by screening out some requests with the lowest priorities and urgencies and adding them to the pause request queue (Pause list), and then sending a pause scheduling recommendation to the model inference service; or screening out some requests with the highest priorities and urgencies from the hanging queue and adding them to the resume request queue (Resume list), and then sending a resume scheduling recommendation to the model inference service.
[0109] S206. When the model inference service executes request scheduling, read the scheduling recommendation from the recommendation buffer queue.
[0110] S207. Intercept and obtain the scheduling decision of the model inference service, and record it as the first scheduling decision.
[0111] In this embodiment, the first scheduling decision is a type of scheduling decision. For example, run request 1, pause request 1, etc. After a request is completed, the model inference service will schedule the next request. The model inference service itself will determine the scheduling decision. In the embodiments of this application, the scheduling decision of the model inference service can be intercepted and obtained through interfaces or other means, and the scheduling decision of the model inference service is recorded as the first scheduling decision.
[0112] S208. Determine the requests to be scheduled and the scheduling type according to the scheduling recommendation.
[0113] In this embodiment, the requests to be scheduled can be understood as the requests that need to be scheduled; the scheduling type can be pause, resume, etc., which is used to represent the type of operation performed on the requests in this round of scheduling. Analyze the scheduling recommendation to determine the requests to be scheduled and the scheduling type. Among them, the number of requests to be scheduled can be one or more.
[0114] S209. If the scheduling type of the requests to be scheduled is pause, and according to the first scheduling decision, it is determined that the scheduling method of the requests to be scheduled is swap out or discard, set the scheduling decision of the requests to be scheduled to swap out, and add the requests to be scheduled to the swap queue.
[0115] In this embodiment, the scheduling method refers to the method of scheduling requests, such as handling operations like swapping out, discarding, swapping back, etc.; the swap queue is the Swapped queue. If the scheduling type of the request to be scheduled is pause, it means that this round of scheduling needs to pause the request to be scheduled and analyze the first scheduling decision. The first scheduling decision includes which request to schedule and the scheduling method. If it is determined in the first scheduling decision that the scheduling method of the request to be scheduled is swapping out or discarding, set the scheduling decision of the request to be scheduled to swapping out and add the request to be scheduled to the swap queue to make it "transparent" to the model inference service, and the model inference service will not generate scheduling decisions related to this request subsequently.
[0116] S210. If the scheduling type of the request to be scheduled is resume, set the scheduling decision of the request to be scheduled to resume.
[0117] If the scheduling type of the request to be scheduled is resume, since the request to be scheduled is in the Swapped queue of the request processing engine and the model inference service will not generate relevant scheduling decisions, the scheduling decision of this request to be scheduled can be directly set to resume.
[0118] S211. Generate a target scheduling decision based on the scheduling decision of the request to be scheduled, and control the model inference service to schedule the corresponding request to execute model inference according to the target scheduling decision.
[0119] When there are multiple requests to be scheduled, perform the steps of S209 - S210 for each request to be scheduled to set its scheduling decision; generate the final target scheduling decision based on the scheduling decisions of all requests to be scheduled to achieve the scheduling of requests.
[0120] The existing solutions of the model inference service have the following drawbacks: 1. Simple task queue and basic scheduling algorithm: Some frameworks use simple task queues and scheduling algorithms to manage user requests. These algorithms usually follow the principle of "FIFO" and ignore the differences in user priorities and request urgencies. In the case of an increasing number of users or high load, this method may cause high-priority users not to be responded to in a timely manner, affecting the user experience and service satisfaction. 2. Static resource allocation strategy: Some model inference services adopt a predefined static resource allocation strategy to allocate computing resources and memory according to fixed rules. This method is simple and direct but lacks flexibility and is difficult to cope with the dynamic changes of user requests. 3. Strategies mainly focused on resource optimization and distributed processing: Some services pay more attention to the optimization of computing resources and distributed processing capabilities, aiming to improve the overall throughput and model inference efficiency of the system. However, these optimizations mainly focus on hardware resources and computing tasks rather than the priorities of user requests. Therefore, in applications facing multi-user scenarios, these frameworks are difficult to effectively distinguish different-priority users and provide differentiated services.
[0121] Embodiments of the present application generate scheduling suggestions based on observation indicators of the system, priority of requests, urgency, etc., intervene in the scheduling strategy of the model inference service itself through the scheduling suggestions, generate a target scheduling strategy to control the model inference service to execute corresponding scheduling requests, and enable requests with high priority and high urgency to be served prior to requests with low priority and urgency while taking into account the utilization efficiency of hardware resources.
[0122] Optionally, before determining the target scheduling decision based on the scheduling suggestion and the scheduling decision of the model inference service, the method further includes:
[0123] B1. If the type corresponding to the scheduling suggestion is a suspension request, determine a first candidate request for suspension according to the scheduling suggestion, take out the first candidate request from the requests being run, update the status of the first candidate request from running to suspended, and put the first candidate request into the swap queue.
[0124] In this embodiment, the first candidate request can be understood as a request that needs to be scheduled. If the type corresponding to the scheduling suggestion is a suspension request, that is, the scheduling suggestion indicates that the request is to be suspended, analyze the scheduling suggestion, determine the request that needs to be suspended indicated by it, use this part of the request as the first candidate request, take out the first candidate request from the requests being run, record the status of the first candidate request, update the status of the first candidate request from running to suspended, and put the first candidate request into the swap queue.
[0125] B2. If the type corresponding to the scheduling suggestion is a resume request, determine a second candidate request for resume according to the scheduling suggestion, take out the second candidate request from the swap queue, update the status of the second candidate request from suspended to running, and put the second candidate request into the requests being run.
[0126] In this embodiment, the second candidate request can be understood as a request that needs to be scheduled. If the type corresponding to the scheduling suggestion is a resume request, that is, the scheduling suggestion indicates that the request is to be resumed, analyze the scheduling suggestion, determine the request that needs to be resumed indicated by it, use this part of the request as the second candidate request, take out the second candidate request from the requests being run, record the status of the second candidate request, update the status of the second candidate request from suspended to running, and put the second candidate request into the queue of requests being run.
[0127] In the embodiments of the present application, the methods or functions related to scheduling in the model inference service are sliced to affect the internal scheduling decisions of the model inference service in a non-invasive manner, avoiding direct modification of the code of the scheduling methods or functions. When the scheduling suggestions from the external request processing engine arrive, they will first enter the Suggestions queue. Each time the request scheduling method or function inside the model inference service is executed, the request processing engine will take out the scheduling suggestions from the Suggestions queue and generate scheduling decisions to take out the request from the running requests or resume the running of some requests according to the scheduling suggestions:
[0128] For the scheduling suggestion to pause a specified request (i.e., the type of the scheduling suggestion corresponds to pausing the request), take it out from the running requests, record the current state of the request and put it into the Swapped queue (i.e., take out the first candidate request from the running requests, update the state of the first candidate request from running to paused, and put the first candidate request into the Swapped queue);
[0129] For the scheduling suggestion to resume a specified request (i.e., the type of the scheduling suggestion corresponds to resuming the request), take it out from the Swapped queue, resume its state and put it back into the running requests (take out the second candidate request from the Swapped queue, update the state of the second candidate request from paused to running, and put the second candidate request into the running requests);
[0130] Finally, fuse these decisions with the scheduling decisions generated by the scheduling methods or functions inside the model inference service to obtain the final target scheduling decision, and update the corresponding metrics according to the target scheduling decision.
[0131] Exemplarily, Figure 3Provides an example diagram of the implementation of request scheduling; taking the model inference service vLLM as an example, the internal scheduling module of vLLM (hereinafter referred to as Inner Scheduler) sets up three queues: Running, Swapped, and Waiting. The Waiting queue records the requests waiting to be executed, and all newly arrived requests will first be added to this queue for queuing; the Running queue records the requests currently being processed, and the Swapped queue records the requests that have been paused due to internal scheduling decisions. Each time vLLM executes the internal scheduling method, it will decide which requests need to be processed (taken out from the Waiting queue and put into the Running queue, i.e., compute), which requests need to be discarded (taken out from the Running queue and put back into the Waiting queue, i.e., drop out), which requests need to be swapped out (taken out from the Running queue and put into the Swapped queue, i.e., the swap out part in the dotted line in the figure), and which requests need to be swapped back (from the Swapped queue to the Running queue, i.e., the swap in part in the dotted line in the figure). Taking the request processing engine including the sidecar scheduling module Sidecar Scheduler as an example, Sidecar Scheduler adds an additional Swapped queue internally to record the requests that have been paused due to external scheduling decisions (outer decision). By slicing the internal scheduling method of vLLM, each time the internal scheduling method of vLLM is executed and returns a scheduling decision, Sidecar Scheduler will intercept these decisions and decide on the modification of these scheduling decisions based on the external scheduling decisions it receives (placed in the Suggestions queue). The specific strategy is as follows:
[0132] For the requests indicated in Suggestions that need to be paused, if Inner Scheduler makes a decision to swap out or drop out the request, modify it to a swap out scheduling decision and do not put the request into Inner Scheduler's Swapped queue, but into Sidecar Scheduler's Swapped queue, making it "transparent" to Inner Scheduler. Inner Scheduler will not generate subsequent scheduling decisions related to this request;
[0133] For requests that need to be resumed as indicated in Suggestions, since the requests are in the Swapped queue of the Sidar Scheduler and the Inner Scheduler will not generate relevant scheduling decisions, the release scheduling decision of the request can be directly put into the final target scheduling decision, or directly put into the Running queue of the Inner Scheduler. Figure 3 Taking the case of directly putting it into the Running queue of the Inner Scheduler as an example.
[0134] For requests not involved in Suggestions, no processing is done and they are directly retained as they are.
[0135] Through the above method, the fusion of the decisions of the internal scheduling method of the Sidecar Scheduler and the model inference service can be achieved.
[0136] As an optional embodiment of this embodiment, this optional embodiment is further optimized to include: when a new request is received, traverse the running queue to determine the user requests currently being executed; determine the execution time corresponding to each user request; if the execution time is greater than the service deadline of the new request, reject the new request; otherwise, accept the new request and send the new request to the model inference service.
[0137] The request processing engine can receive requests sent by different users. Usually, the user's requests need to be sent to the model inference service so that the model inference service can provide corresponding services for the users. For example, the user sends a request to the request processing engine, and the request is forwarded by the request processing to the model inference service, or the user sends a request to the model inference service, and the request processing engine intercepts or obtains the request sent by the user through an interface or other means. After the request processing engine receives a new request, it traverses the running queue to determine the requests in the running queue. These requests are the user requests currently being executed; based on the analysis of each user request, for example, according to the type of the request, the deadline requirement, etc., determine the execution time required to normally execute each user request. Determine the service deadline of the new request according to the relevant information of the new request. If the execution time is greater than the service deadline of the new request, it means that the requirement of the service deadline of the new request may not be met, and the new request is rejected; otherwise, the new request is accepted and the new request is sent to the model inference service.
[0138] In the embodiments of the present application, it is also possible to determine whether the system load is within an acceptable range by analyzing the currently executing user requests. If the load is too high, new requests may not meet the requirements of their service time limits, so the requests will be directly rejected. If the load is within the acceptable range, the requests are accepted and sent to the model inference service for specific processing through the API of the model inference service. The execution time can be used as a parameter for judging the load.
[0139] Optionally, after accepting a new request, the method further includes: adding the new request to a running queue.
[0140] After accepting a new request, the request will be added to a running queue and regulated by a request processing engine.
[0141] The embodiments of the present application provide a request management method for model inference. Through a request processing engine, based on the collection and prediction of the metric information of the model inference service, a scheduling suggestion for requests is issued. The request processing engine slices the scheduling functions and methods, introduces a suggestion buffer queue, and interferes with the request scheduling decision inside the inference framework according to the suggestion by a low-intrusive method, so as to achieve the effect of request preemption. The effect of request preemption is achieved by pausing and resuming the computing tasks of inference requests; by slicing the request scheduling methods or functions inside the model inference service, merging its own scheduling decision with the decision of the sliced scheduling function or method, and obtaining the final target scheduling decision, so as to interfere with the scheduling decision; on the basis of the existing model inference service, the scheduling is optimized through an external priority queue, and at the same time, the internal scheduling mechanism of the model inference service is extended by a low-intrusive method, thus avoiding major changes to the existing model inference service. This low-intrusive method simplifies the system integration and maintenance, making the scheduling mechanism of the embodiments of the present application can be easily integrated into the existing model inference service. By introducing state prediction, aligning the state of the model inference service seen by the request processing engine at the moment of making a scheduling decision with the actual state of the model inference service at that moment, and alleviating the scheduling decision errors caused by the asynchrony inside and outside the model inference service. The request processing engine schedules requests based on the predicted future state of the model inference service, making the scheduling more accurate. Through an external scheduling mechanism, the dynamic optimization of resource allocation is realized. According to the current system load, the priority and urgency requirements of requests, the requests occupying hardware resources are adjusted in real time. In the case of resource shortage, the execution of high-priority and high-urgency requests is preferentially guaranteed, and the effect of request preemption is achieved, so as to ensure the efficient utilization and reasonable allocation of hardware resources while taking into account the priority and urgency requirements of tasks.
[0142] Embodiment III
[0143] Figure 4The figure is a schematic structural diagram of a request management device for model inference provided in Embodiment 3 of the present invention, which is applied to a request processing engine in a request management system. The request management system further includes a model inference service. As Figure 4 shown, the device includes: an observation module 31, a state prediction module 32, a request scheduler 33, and a side-loading scheduler module 34.
[0144] The observation module 31 is configured to obtain observation metrics of the model inference service;
[0145] The state prediction module 32 is configured to perform metric prediction based on the observation metrics to obtain predicted metrics, where the predicted metrics include the target memory utilization rate of the central processing unit and the target video memory utilization rate of the graphics processing unit;
[0146] The request scheduler 33 is configured to determine a scheduling policy according to the predicted metrics, determine a scheduling recommendation according to the scheduling policy in combination with a request queue, and add the scheduling recommendation to a recommendation buffer queue;
[0147] The side-loading scheduler module 34 is configured to, when the model inference service executes request scheduling, read the scheduling recommendation from the recommendation buffer queue, determine a target scheduling decision according to the scheduling recommendation in combination with the scheduling decision of the model inference service, and control the model inference service to schedule corresponding requests to execute model inference according to the target scheduling decision.
[0148] A request management device for model inference according to an embodiment of the present invention solves the problem of unreasonable request scheduling in the process of model inference. By predicting through the observation metrics of the model inference service, predicted metrics are obtained. The predicted metrics include the target memory utilization rate of the central processing unit and the target video memory utilization rate of the graphics processing unit. In the embodiments of the present application, the predicted metrics can be used to represent different metrics of the model inference service in the next period of time, that is, predict the metrics of the model inference service in the subsequent time; determine a scheduling policy through the predicted metrics, further determine a scheduling recommendation in combination with the request queue, and add it to the recommendation buffer queue; when the model inference service executes request scheduling, read the scheduling recommendation from the recommendation buffer queue, and determine a target scheduling decision according to the scheduling recommendation in combination with the scheduling decision of the model inference service, intervene in the request scheduling of the model inference service, and finally the target scheduling decision controls the model inference service to schedule corresponding requests to execute model inference; interfere with the scheduling decision of the model inference service through metric prediction, provide a target scheduling decision for the model inference service, and control the model inference service to execute model inference through reasonable request scheduling, realizing efficient utilization and reasonable allocation of resources.
[0149] Among them, the side-loading scheduler module is deployed in the model inference service 32 to intervene in the scheduling policy of the model inference service 32.
[0150] Optionally, the observation metrics include at least one of processor parameter metrics and request metrics;
[0151] Among them, the processor parameter metrics include at least one of the following:
[0152] Memory utilization rate of the central processing unit;
[0153] Video memory utilization rate of the graphics processing unit;
[0154] Core utilization rate of the graphics processing unit;
[0155] Among them, the request metrics include at least one of the following:
[0156] Number of requests to be left for Prefilling;
[0157] Length of requests to be left for Prefilling;
[0158] Number of requests in the Decoding stage;
[0159] Key-value pair cache size of requests to be swapped into the memory of the central processing unit;
[0160] Key-value pair cache size of requests to be swapped into the video memory of the graphics processing unit.
[0161] Optionally, the processor parameter metrics are updated when the first metric update condition is met;
[0162] The request metrics are updated when the second metric update condition is met;
[0163] Among them, the first metric update condition includes at least one of the following:
[0164] The system time meets the periodic update condition;
[0165] The model inference service completes the scheduling of one request;
[0166] The second metric update condition includes:
[0167] The model inference service completes the scheduling of one request.
[0168] Optionally, the state prediction module is specifically configured to: use the memory size of the central processing unit, the video memory size of the graphics processing unit, and the observation metrics as inputs for prediction, and determine the target memory utilization rate of the central processing unit and the target video memory utilization rate of the graphics processing unit;
[0169] Among them, the memory size of the central processing unit and the video memory size of the graphics processing unit are applied for when the program starts.
[0170] Optionally, the request scheduler is specifically configured to: if the target video memory utilization rate is higher than a first preset threshold and the target memory utilization rate is lower than a second preset threshold, determine the scheduling policy as a pause request; if the target video memory utilization rate is lower than a third preset threshold, determine the scheduling policy as a resume request.
[0171] Optionally, the request queue includes a running queue and a pending queue;
[0172] The request scheduler is specifically configured to: when the scheduling policy is a pause request, traverse the running queue, sort the requests in the running queue according to the priority level and urgency of the requests in the running queue, filter out at least one request as a pause request according to the sorting result, add the pause request to the pause request queue, and generate a scheduling recommendation according to the pause request queue; when the scheduling policy is a resume request, traverse the pending queue, sort the requests in the pending queue according to the priority level and urgency of the requests in the pending queue, filter out at least one request as a resume request according to the sorting result, add the resume request to the resume request queue, and generate a scheduling recommendation according to the resume request queue.
[0173] Optionally, the prediction metric further includes: the number of requests;
[0174] Correspondingly, the number of the pause requests or resume requests is the number of requests.
[0175] Optionally, the sideload scheduling module is further configured to: if the type corresponding to the scheduling recommendation is a pause request, determine a first candidate request to be paused according to the scheduling recommendation, take out the first candidate request from the requests being run, update the status of the first candidate request from running to paused, and put the first candidate request into the swap queue; if the type corresponding to the scheduling recommendation is a resume request, determine a second candidate request to be resumed according to the scheduling recommendation, take out the second candidate request from the swap queue, update the status of the second candidate request from paused to running, and put the second candidate request into the requests being run.
[0176] Optionally, the sideload scheduling module is specifically configured to: intercept and obtain the scheduling decision of the model inference service, and record it as the first scheduling decision; determine the request to be scheduled and the scheduling type according to the scheduling recommendation; if the scheduling type of the request to be scheduled is a pause, and according to the first scheduling decision, the scheduling method of the request to be scheduled is swap out or discard, set the scheduling decision of the request to be scheduled to swap out, and add the request to be scheduled to the swap queue; if the scheduling type of the request to be scheduled is a resume, set the scheduling decision of the request to be scheduled to resume; generate a target scheduling decision based on the scheduling decision of the request to be scheduled.
[0177] Optionally, the request processing engine 31 further includes:
[0178] An early rejection module, configured to, when receiving a new request, traverse the running queue to determine the user requests currently being executed; determine the execution time corresponding to each of the user requests; if the execution time is greater than the service time limit of the new request, reject the new request; otherwise, accept the new request and send the new request to the model inference service.
[0179] Optionally, the early rejection module is further configured to: add the new request to the running queue.
[0180] The request management system provided by the embodiments of the present invention can execute the request management method for model inference provided by any embodiment of the present invention, and has the corresponding functional modules and beneficial effects of the execution method.
[0181] Embodiment 4
[0182] Figure 5 It is a schematic structural diagram of a request management system provided by Embodiment 4 of the present invention. As Figure 5 shown, the system includes: a request processing engine 41 and a model inference service 42.
[0183] The request processing engine 41 is configured to: obtain the observation indexes of the model inference service 42; perform index prediction according to the observation indexes to obtain prediction indexes, where the prediction indexes include the target memory utilization rate of the central processing unit and the target video memory utilization rate of the graphics processing unit; determine a scheduling strategy according to the prediction indexes, determine a scheduling recommendation according to the scheduling strategy in combination with the request queue, and add the scheduling recommendation to the recommendation buffer queue; when the model inference service 42 executes request scheduling, read the scheduling recommendation from the recommendation buffer queue, and determine a target scheduling decision according to the scheduling recommendation in combination with the scheduling decision of the model inference service, and control the model inference service 42 to schedule the corresponding request to execute model inference according to the target scheduling decision.
[0184] A request management system according to an embodiment of the present invention solves the problem of unreasonable request scheduling during the model inference process. By predicting through the observation metrics of the model inference service, predicted metrics are obtained. The predicted metrics include the target memory utilization rate of the central processing unit and the target video memory utilization rate of the graphics processing unit. In the embodiments of the present application, the predicted metrics can be used to represent different metrics of the model inference service in the next period of time, that is, to predict the metrics of the model inference service in the subsequent time. A scheduling strategy is determined based on the predicted metrics, and further combined with the request queue to determine a scheduling recommendation, which is added to the recommendation buffer queue. When the model inference service executes request scheduling, the scheduling recommendation is read from the recommendation buffer queue, and the target scheduling decision is determined based on the scheduling recommendation combined with the scheduling decision of the model inference service, so as to intervene in the request scheduling of the model inference service. Finally, the target scheduling decision controls the model inference service to schedule the corresponding request to execute model inference. By interfering with the scheduling decision of the model inference service through metric prediction, a target scheduling decision is provided for the model inference service, and the model inference service is controlled to execute model inference through reasonable request scheduling, realizing efficient utilization and reasonable allocation of resources.
[0185] Exemplarily, Figure 6 An implementation example diagram of request management is provided. The request processing engine 31 includes: an early rejection module 411, an observation module 412, a state prediction module 413, a request scheduler 414, a side load scheduler module 415, and a metric collection module 416. Among them, the side load scheduler module 415 and the metric collection module 416 are deployed on the model inference service 42.
[0186] When a new request arrives, the early rejection (Early Rejection) module 411 checks the current system load condition to determine whether the service time limit requirement of the new request is met. If not, the request is rejected; if so, the request is accepted and sent to the model inference service for processing through the API of the model inference service. Figure 6 Taking the request sent to the model inference service through the inference request processing interface as an example; at the same time, the request will be added to the Running queue and regulated by the request processing engine 41. After the request is processed by the model inference service 42, the corresponding result can be fed back to the user through the request processing engine 31 to complete the response of the request.
[0187] The observation (Observer) module 42 obtains the observation metrics of the model inference service through the API method. Figure 6Take the acquisition of observation metrics through the metric observation interface as an example. In each round of scheduling, the observation module 42 requests the metric API exposed by the model inference service to obtain the metrics of the model inference service, and hands them over to the State Predictor module 413 to predict the future load conditions of the model inference service, so that the request scheduling module can provide better scheduling suggestions.
[0188] The State Predictor module 413 is responsible for predicting the state of the model inference service when the request scheduler 414 makes a scheduling decision after a short period of time based on the observation metrics of the model inference service collected by the observation module 42. Since the external request processing engine 41 is asynchronous with the internal operation of the model inference service in request processing and scheduling decisions, when the request scheduler 414 makes a scheduling decision based on the observation metrics collected by the observation module 412, these metric information lags behind the state of the model inference service itself at the time of decision-making. Therefore, when making a scheduling decision, the request scheduler 414 should make a scheduling decision based on the future state (i.e., predicted metrics) of the model inference service predicted by the state prediction module 413.
[0189] In the large model inference service, requests with different priorities and urgencies have differences in priority in the use of hardware. Requests with high priority and high urgency should have priority in occupying hardware resources to obtain services. Therefore, the request scheduler 414 needs to schedule these tasks. When the hardware resources are insufficient or the system load is high, the request scheduler 414 needs to decide which tasks should preempt other tasks to be processed first. The request scheduler 414 determines the scheduling strategy to be executed based on the predicted GPU video memory utilization rate and CPU memory utilization rate.
[0190] The request scheduler 414 includes a Scanner and a Priority Selection. If a request needs to be paused, the Scanner will first traverse the Running queue and sort the requests in the Running queue according to the priority and urgency of the requests in the queue; if a request needs to be resumed, the Scanner will first traverse the Hanging queue and sort the requests in the Hanging queue according to the priority and urgency of the requests in the queue. The Priority Selection component will filter out some requests with the lowest priority and urgency from the Running queue that has been sorted by priority and add them to the Pause list, and then send a pause suggestion to the model inference service; or filter out some requests with the highest priority and urgency from the Hanging queue and add them to the Resume list, and then send a resume suggestion to the model inference service. Figure 6 Taking the scheduling suggestion sent to the sideload scheduling module 415 through the request suggestion interface as an example.
[0191] The metric Collector module 416 will collect various metrics of the model inference service at a certain period and expose them to the request processing engine through the API interface for scheduling decisions. Due to the existence of the state prediction module 412 in the request processing engine 41, the metric Collector module 416 does not need to collect various metrics in the model inference service in a synchronous process, but collects the metrics asynchronously at a certain time period and puts them into the internal cache, and at the same time updates some metric information according to each request scheduling decision inside the model inference service. When the request processing engine accesses the metric information through the API, the metric Collector module 416 will directly take out the metric information from the cache and return it as the observed metric. Compared with the synchronous collection method, the caching strategy of the metric Collector module 416 can effectively guarantee the performance of the metric collection interface and the overall performance of the inference service framework.
[0192] The Sidecar Scheduler module 415 affects the internal scheduling decisions of the model inference service in a non-invasive way by slicing the methods or functions related to scheduling in the model inference service, avoiding direct modification of the code of the scheduling methods or functions. When the scheduling suggestions from the external request processing engine arrive, they first enter the Suggestions queue of the Sidecar Scheduler module 415. Each time the request scheduling method or function inside the model inference service is executed, the Sidecar Scheduler module 415 takes out these suggestions from the Suggestions queue and generates scheduling decisions to take the request out of the currently running requests or resume the running of some requests according to the suggestions:
[0193] For the suggestion to pause a specified request, the Sidecar Scheduler module 415 takes it out of the currently running requests, records the current state of the request and puts it into the Swapped queue; for the suggestion to resume a specified request, the Sidecar Scheduler module 415 takes it out of the Swapped queue, resumes its state and puts it back into the currently running requests.
[0194] The Sidecar Scheduler module 415 fuses these decisions with the scheduling decisions generated by the internal scheduling methods or functions of the model inference service to obtain the final target scheduling decision, and notifies the Collector module 416 of the update of the corresponding metrics according to the target scheduling decision.
[0195] Embodiment Five
[0196] Figure 7 FIG. shows a schematic structural diagram of an electronic device 50 that can be used to implement the embodiments of the present invention. The electronic device is intended to represent various forms of digital computers, such as, laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as, personal digital assistants, cellular phones, smart phones, wearable devices (such as helmets, glasses, watches, etc.) and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely illustrative and are not intended to limit the implementation of the present invention described and / or claimed herein.
[0197] As Figure 7As shown, the electronic device 50 includes at least one processor 51 and a memory communicatively connected to the at least one processor 51, such as a read-only memory (ROM) 52, a random access memory (RAM) 53, etc. The memory stores a computer program executable by the at least one processor. The processor 51 can perform various appropriate actions and processes according to the computer program stored in the read-only memory (ROM) 52 or the computer program loaded from the storage unit 58 into the random access memory (RAM) 53. In the RAM 53, various programs and data required for the operation of the electronic device 50 can also be stored. The processor 51, the ROM 52, and the RAM 53 are connected to each other via a bus 54. An input / output (I / O) interface 55 is also connected to the bus 54.
[0198] Multiple components in the electronic device 50 are connected to the I / O interface 55, including: an input unit 56, such as a keyboard, a mouse, etc.; an output unit 57, such as various types of displays, speakers, etc.; a storage unit 58, such as a magnetic disk, an optical disc, etc.; and a communication unit 59, such as a network card, a modem, a wireless communication transceiver, etc. The communication unit 59 allows the electronic device 50 to exchange information / data with other devices via a computer network such as the Internet and / or various telecommunication networks.
[0199] The processor 51 can be various general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of the processor 51 include but are not limited to a central processing unit (CPU), a graphics processing unit (GPU), various dedicated artificial intelligence (AI) computing chips, various processors running machine learning model algorithms, a digital signal processor (DSP), and any appropriate processor, controller, microcontroller, etc. The processor 51 executes the various methods and processes described above, such as the method for managing requests for model inference.
[0200] In some embodiments, the method for managing requests for model inference can be implemented as a computer program tangibly embodied in a computer-readable storage medium, such as the storage unit 58. In some embodiments, part or all of the computer program can be loaded and / or installed onto the electronic device 50 via the ROM 52 and / or the communication unit 59. When the computer program is loaded into the RAM 53 and executed by the processor 51, one or more steps of the method for managing requests for model inference described above can be executed. Alternatively, in other embodiments, the processor 51 can be configured to execute the method for managing requests for model inference in any other appropriate manner (e.g., by means of firmware).
[0201] The various embodiments of the systems and techniques described above in this document can be implemented in digital electronic circuitry, integrated circuit systems, field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), application specific standard products (ASSPs), systems-on-chip (SOCs), complex programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments can include: being implemented in one or more computer programs that can be executed and / or interpreted on a programmable system including at least one programmable processor, which can be a special-purpose or general-purpose programmable processor that can receive data and instructions from a storage system, at least one input device, and at least one output device, and transmit the data and instructions to the storage system, the at least one input device, and the at least one output device.
[0202] The computer programs for implementing the methods of the present invention can be written in any combination of one or more programming languages. These computer programs can be provided to a processor of a general-purpose computer, a special-purpose computer, or other programmable data processing device, such that when the computer programs are executed by the processor, the functions / operations specified in the flowcharts and / or block diagrams are implemented. The computer programs can be executed entirely on the machine, partially on the machine, as a stand-alone software package partially on the machine and partially on a remote machine, or entirely on a remote machine or server.
[0203] Embodiments of the present invention provide a computer program product, which includes a computer program that, when executed by a processor, implements the method for managing requests for model inference according to any embodiment of the present invention.
[0204] In the context of the present invention, a computer-readable storage medium can be a tangible medium that can contain or store a computer program for use by or in connection with an instruction execution system, apparatus, or device. The computer-readable storage medium can include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. Alternatively, the computer-readable storage medium can be a machine-readable signal medium. More specific examples of the machine-readable storage medium would include electrical connections based on one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disc read-only memory (CD-ROM), optical storage device, magnetic storage device, or any suitable combination of the foregoing.
[0205] To provide interaction with a user, the systems and techniques described herein can be implemented on an electronic device having: a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user; and a keyboard and a pointing device (e.g., a mouse or a trackball) through which the user can provide input to the electronic device. Other kinds of devices can also be used to provide interaction with the user; for example, the feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including acoustic input, voice input, or tactile input).
[0206] The systems and techniques described herein can be implemented in a computing system including backend components (e.g., as a data server), or a computing system including middleware components (e.g., an application server), or a computing system including frontend components (e.g., a user computer having a graphical user interface or a web browser through which the user can interact with an implementation of the systems and techniques described herein), or a computing system including any combination of such backend components, middleware components, or frontend components. The components of the system can be interconnected to each other by digital data communication in any form or medium (e.g., a communication network). Examples of communication networks include: local area network (LAN), wide area network (WAN), blockchain network, and the Internet.
[0207] The computing system can include a client and a server. The client and the server are generally far from each other and usually interact through a communication network. The client-server relationship is created by computer programs running on the respective computers and having a client-server relationship with each other. The server can be a cloud server, also known as a cloud computing server or a cloud host, which is a host product in the cloud computing service system, solving the defects of difficult management and weak business scalability existing in traditional physical hosts and VPS services.
[0208] It should be understood that the various forms of processes shown above can be used, with steps reordered, added, or deleted. For example, the steps recited in the present invention can be executed in parallel, sequentially, or in a different order, as long as the desired results of the technical solution of the present invention can be achieved, and no limitation is made herein.
[0209] The above specific embodiments do not constitute a limitation on the protection scope of the present invention. Those skilled in the art should understand that various modifications, combinations, sub-combinations, and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of the present invention shall be included within the protection scope of the present invention.
Claims
1. A request management method for model reasoning, characterized in that: A request processing engine applied to a request management system, wherein the request management system further includes a model reasoning service, and the method includes: Obtaining observation indicators of the model inference service; An indicator prediction is performed according to the observed indicator to obtain a predicted indicator, wherein the predicted indicator includes a target memory utilization rate of a central processing unit and a target video memory utilization rate of a graphics processing unit; the predicted indicator is indicator information of a model inference service obtained through prediction, and is indicator information of a period of time later than the collection time of the observed indicator; Determine a scheduling strategy according to the prediction indicator, determine a scheduling suggestion according to the scheduling strategy combined with the request queue, and add the scheduling suggestion to the suggestion buffer queue; When the model reasoning service executes request scheduling, the scheduling suggestion is read from the suggestion buffer queue, and the target scheduling decision is determined based on the scheduling suggestion and the scheduling decision of the model reasoning service, and the model reasoning service is controlled to schedule the corresponding request to execute model reasoning according to the target scheduling decision.
2. The method according to claim 1, characterized in that The observation indicator includes at least one of a processor parameter indicator and a request indicator; The processor parameter indicator includes at least one of the following: CPU memory utilization; Graphics processor memory utilization; GPU core utilization; The request indicator includes at least one of the following: The number of requests left for prefilling; The length of the request left for prefilling; The number of requests in the decoding phase; The size of the requested key-value cache to be swapped into the CPU's memory; The requested key-value cache size to be swapped into the GPU's video memory.
3. The method according to claim 2, characterized in that The processor parameter indicator is updated when a first indicator update condition is met; The request indicator is updated when the second indicator update condition is met; The first indicator update condition includes at least one of the following: The system time meets the periodic update conditions; The model inference service completes the scheduling of a request; The second indicator update condition includes: The model inference service completes the scheduling of a request.
4. The method according to claim 1, characterized in that: The step of performing indicator prediction according to the observed indicator to obtain the predicted indicator includes: The memory size of the central processing unit, the video memory size of the graphics processing unit, and the observed index are used as inputs for prediction to determine a target memory utilization rate of the central processing unit and a target video memory utilization rate of the graphics processing unit; The memory size of the central processing unit and the video memory size of the graphics processing unit are applied for when the program is started.
5. The method according to claim 1, characterized in that The determining of the scheduling strategy according to the prediction index comprises: If the target video memory utilization is higher than a first preset threshold and the target memory utilization is lower than a second preset threshold, determining the scheduling strategy to be a pause request; If the target video memory utilization is lower than a third preset threshold, the scheduling strategy is determined to be a recovery request.
6. The method according to claim 1, characterized in that The request queue includes a running queue and a suspended queue, and the determining of the scheduling suggestion according to the scheduling policy in combination with the request queue includes: When the scheduling strategy is a pause request, traverse the run queue, sort the requests in the run queue according to the priority level and urgency of the requests in the run queue, select at least one request as a pause request according to the sorting result, add the pause request to a pause request queue, and generate a scheduling suggestion according to the pause request queue; When the scheduling strategy is a recovery request, the suspension queue is traversed, the requests in the suspension queue are sorted according to the priority level and urgency of the requests in the suspension queue, at least one request is screened out as a recovery request according to the sorting result, the recovery request is added to the recovery request queue, and a scheduling suggestion is generated according to the recovery request queue.
7. The method according to claim 6, characterized in that The prediction indicators also include: the number of requests; Correspondingly, the number of the pause requests or the resume requests is the request number.
8. The method according to claim 1, characterized in that Before determining the target scheduling decision according to the scheduling suggestion combined with the scheduling decision of the model reasoning service, the method further includes: If the type corresponding to the scheduling suggestion is a pause request, determine a first candidate request to be paused according to the scheduling suggestion, take the first candidate request out of the running requests, update the state of the first candidate request from running to paused, and put the first candidate request into the exchange queue; If the type corresponding to the scheduling suggestion is a recovery request, determine the second candidate request for recovery according to the scheduling suggestion, take the second candidate request out of the exchange queue, update the status of the second candidate request from paused to running, and put the second candidate request into the running request.
9. The method according to claim 1, characterized in that: Determine the target scheduling decision based on the scheduling suggestion combined with the scheduling decision of the model reasoning service, including: Intercept and obtain the scheduling decision of the model inference service and record it as the first scheduling decision; Determine the scheduling request and the scheduling type according to the scheduling suggestion; If the scheduling type of the request to be scheduled is pause, and it is determined according to the first scheduling decision that the scheduling mode of the request to be scheduled is swap out or discard, setting the scheduling decision of the request to be scheduled to swap out, and adding the request to be scheduled to the exchange queue; If the scheduling type of the request to be scheduled is recovery, setting the scheduling decision of the request to be scheduled to recovery; A target scheduling decision is generated based on the scheduling decision of the request to be scheduled.
10. The method according to any one of claims 1 to 9, characterized in that: Also includes: When a new request is received, the running queue is traversed to determine the currently executing user request; Determine the execution time corresponding to each of the user requests; If the execution time is greater than the service time limit of the new request, the new request is rejected; otherwise, the new request is accepted and sent to the model inference service.
11. The method according to claim 10, characterized in that After accepting the new request, the method further includes: The new request is added to the running queue.
12. A model reasoning request management device, characterized in that: A request processing engine applied to a request management system, wherein the request management system also includes a model reasoning service, wherein the device includes: An observation module, used to obtain observation indicators of the model reasoning service; A state prediction module is used to perform indicator prediction based on the observed indicator to obtain a predicted indicator, wherein the predicted indicator includes a target memory utilization rate of a central processing unit and a target video memory utilization rate of a graphics processing unit; the predicted indicator is the indicator information of the model inference service obtained by prediction, and is the indicator information of a period of time later than the collection time of the observed indicator; A request scheduler, configured to determine a scheduling strategy according to the prediction indicator, determine a scheduling suggestion according to the scheduling strategy in combination with the request queue, and add the scheduling suggestion to a suggestion buffer queue; The side-loaded scheduling module is used to read scheduling suggestions from the suggestion buffer queue when the model reasoning service executes request scheduling, and determine the target scheduling decision based on the scheduling suggestions and the scheduling decision of the model reasoning service, and control the model reasoning service to schedule the corresponding request to execute model reasoning according to the target scheduling decision.
13. A request management system, characterized in that: include: Request processing engine and model inference service; The request processing engine is used to obtain observation indicators of the model reasoning service; Indicator prediction is performed according to the observed indicator to obtain predicted indicators, wherein the predicted indicators include a target memory utilization of the central processing unit and a target video memory utilization of the graphics processing unit; the predicted indicator is indicator information of the model inference service obtained through prediction, and is indicator information of a period of time later than the collection time of the observed indicator; a scheduling strategy is determined according to the predicted indicator, a scheduling suggestion is determined according to the scheduling strategy in combination with the request queue, and the scheduling suggestion is added to the suggestion buffer queue; when the model inference service executes request scheduling, the scheduling suggestion is read from the suggestion buffer queue, and a target scheduling decision is determined according to the scheduling suggestion in combination with the scheduling decision of the model inference service, and the model inference service is controlled to schedule the corresponding request to execute model inference according to the target scheduling decision.
14. An electronic device, characterized in that: The electronic device comprises: at least one processor, and a memory communicatively coupled to the at least one processor; The memory stores a computer program executable by the at least one processor, and the computer program is executed by the at least one processor so that the at least one processor can execute the model reasoning request management method described in any one of claims 1-11.
15. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores computer instructions, and the computer instructions are used to enable a processor to implement the model reasoning request management method described in any one of claims 1-11 when executed.
16. A computer program product, characterized in that The computer program product comprises a computer program, which, when executed by a processor, implements the request management method for model reasoning according to any one of claims 1 to 11.
Citation Information
Patent Citations
Request scheduling method and request scheduling device
CN106775990A
Task scheduling optimization data processing system based on resource load prediction
CN113568722A