Multi-model reasoning service load balancer and method

By designing a load balancer for multi-model inference services, and using schedulers and detectors to monitor and schedule inference service instances in real time, the problems of unreasonable resource allocation and increased latency in the existing technology are solved, and load balancing and efficient inference service processing are achieved.

CN120163240APending Publication Date: 2025-06-17HUNAN UNIV
View PDF 0 Cites 4 Cited by

Patent Information

Application Number
CN202510190527.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-20
Publication Date
2025-06-17

AI Technical Summary

Technical Problem

The existing model scheduling algorithm cannot accurately predict changes in user request requirements, resulting in unreasonable resource configuration and inability to respond to changes in request traffic in a timely manner, resulting in increased latency of inference services.

Method used

Design a load balancer for multi-model inference services, including schedulers and detectors, to realize load balancing by monitoring the workload pressure in the cloud service cluster in real time, scheduling inference service instances are transferred from a lighter queue system to a heavier queue system to achieve load balancing.

Benefits of technology

It effectively realizes load balancing of queue systems in cloud service clusters, improves the processing efficiency of inference services, reduces response time delay, and improves the ability to respond to changes in user request requirements.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120163240A_ABST
    Figure CN120163240A_ABST
Patent Text Reader

Abstract

The load balancer of the multi-model reasoning service comprises a scheduler and a detector, and the detector detects the working load pressure of each queue system in a cloud service cluster. When the working load pressure of the queue system with the maximum working load pressure and the working load pressure of the queue system with the minimum working load pressure in the cloud service cluster meet scheduling conditions, the scheduler schedules one reasoning service instance from the queue system with the minimum working load pressure to the queue system with the maximum working load pressure; the scheduling condition comprises that the maximum working pressure in the cloud service cluster is greater than the product of the minimum working pressure and the threshold value. According to the load balancer and method for the multi-model reasoning service, the load balance of each queue system in a cloud service cluster can be effectively realized, so that the processing efficiency of the reasoning service of the cloud service cluster is improved; compared with a traditional greedy enumeration strategy, the load balancer and method for the multi-model reasoning service have more excellent performance in the aspects of working load, response time and response time distribution accuracy.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to load balancing in a cloud service cluster, and particularly to a load balancer and method for multi-model inference services. Background Art

[0002] Nowadays, as the scale and complexity of LLMs continue to increase, their increasingly demonstrated strong generalization ability meets the needs of a large number of users. Cloud service providers offer a variety of models based on the Transformer architecture to meet the needs of different users, such as chatbots, programming assistants, search engines, and personal assistants. The computations of these models are centrally deployed in a cloud service cluster, sharing the high-performance hardware resources in the cluster. These hardware resources are not only expensive but also limited in quantity. In fact, for cloud service providers offering multiple LLM services, the popularity of different models varies, and the user request traffic fluctuates greatly at different times. To meet the user's response latency requirements and the peak computing demands of the models, it is uneconomical to allocate too many computing resources to the models at any time because the extra allocated computing resources will be wasted most of the time. Therefore, to avoid the situation where some models are overloaded while others are underutilized, a flexible resource allocation scheme is the key. Considering the load conditions of different models, the load balance between multiple models can be achieved by redeploying the machines under the model with less access demand to the model with intensive requests.

[0003] Existing model scheduling algorithms use the request records of historical users and determine the model deployment scheme through a greedy enumeration strategy to achieve the optimal service level objective attainment rate (the proportion of requests meeting the SLO). This algorithm will redeploy the cluster regularly. In the emulator, the arrival time of the completed requests in the cluster and the required number of computational tokens are input to simulate the arrival of users. Then, the SLOs under different model configuration machines are calculated one by one to find the model deployment scheme that meets the optimal SLO. However, this machine reconfiguration scheme based on simulating historical request records has two obvious problems.

[0004] First, based on historical user requests, there is an obsolescence problem and it cannot accurately predict the user request demands in the next cycle. According to the simulated model deployment strategy, the optimal SLO can be achieved based on the user requests in the previous round. However, the model request demands in the next round may change. In the actual dataset, the gap between the current request rate distribution and the historical information can differ by several times. According to the experimental results, when the gap between the current data and the historical information is 2 times, compared with the ideal situation, the latency of the first token generation increases by 1.93 times.

[0005] Secondly, the periodic redeployment strategy cannot adapt to the incoming requests within the cluster in a timely manner. This is because an increase in request traffic requires an immediate increase in the number of inference service instances to provide computing power. However, the long reconfiguration cycle prevents the cluster from responding promptly, making it impossible to quickly provide sufficient inference service instances to handle the increased model workload. According to the experimental results, compared with the ideal situation, the latency of the first token generation increased by 1.33 times.

[0006] This raises a crucial question: for LLM services, is it possible to develop an algorithm that can accurately detect changes in the real-time workload of different models and adaptively schedule inference service instances to achieve low-latency responses and meet user SLOs; utilize the real-time nature of computing workloads to address the problem of outdated historical records, and adapt to changes in model request traffic through dynamic scheduling.

[0007] Precise scheduling of LLM service systems poses significant challenges. First, how to accurately calculate the workload pressure of different models for real-time scheduling. Most current LLM services are based on the Transformer architecture, and the completion time of each user service request is unpredictable. This is because the current inference service consists of two stages: the prefill stage and the decoding stage, both of which involve token generation. In the prefill stage, the request consists of multiple tokens, and the first token is generated after calculation. The decoding stage generates subsequent tokens based on the previous token until the context limit is reached or the token terminator is generated. Therefore, the autoregressive nature of LLM makes it impossible to determine the workload pressure of the inference service based on the user request queue.

[0008] Second, how to determine the timing of model scheduling in a timely manner; different models have different resource requirements, and the configuration and status of computing resources within the cluster are constantly changing. When the request traffic of multiple models changes, it may lead to resource adjustments among multiple models within the cluster. Especially among models with high request traffic, resources will be adjusted back and forth, resulting in ineffective scheduling. Summary of the Invention

[0009] The technical problem to be solved by the present invention is to overcome the deficiencies of the prior art and provide a load balancer and method for effectively implementing multi-model inference services.

[0010] To solve the above technical problem, the technical solution proposed by the present invention is: a load balancer for multi-model inference services, including a scheduler and a detector. The detector detects the workload pressure of each queue system in the cloud service cluster. When the workload pressures of the queue system with the maximum workload pressure and the queue system with the minimum workload pressure in the cloud service cluster meet the scheduling conditions, the scheduler schedules an inference service instance from the queue system with the minimum workload pressure to the queue system with the maximum workload pressure.

[0011] The scheduling condition includes that the maximum working pressure in the cloud service cluster is greater than the minimum working pressure multiplied by a threshold value.

[0012] For the load balancer of the multi-model inference service described above, preferably, multiple inference service instances are deployed on each node of the queue system.

[0013] For the load balancer of the multi-model inference service described above, preferably, the working load pressure is described by the following expression: ρ i is the working load pressure of queue system i, where the numerator λ i ×E(X i ) represents the average token arrival rate of queue system i, and the denominator c i ×μ i represents the average token processing rate of queue system i; E(X i ) is the expected number of request tokens of queue system i, λ i is the arrival rate of requests representing queue system i, c i is the number of inference service instances in queue system i, and μ i is the average processing time of each request in queue system i.

[0014] For the load balancer of the multi-model inference service described above, preferably, the number of tokens returned by the autoregressive model corresponding to each queue system is counted at the service entrance of the cloud service cluster.

[0015] For the load balancer of the multi-model inference service described above, preferably, different computational costs are attached to the returned tokens at the service entrance of the cloud service cluster. The computation of the first token generates a higher workload, and the weight value of each queue system is calculated; under the same workload, the number of inference service instances in the queue system with a higher weight value is greater than that in the queue system with a lower weight value.

[0016] For the load balancer of the multi-model inference service described above, preferably, the requests for inference services being performed by the scheduled inference service instances are rerouted to the inference service instance with the shortest queue under the same model, and the inference service is re-performed by other inference service instances.

[0017] A load balancing method for a multi-model inference service includes the following steps;

[0018] 1) Count the working load pressure of each queue system in the cloud service cluster;

[0019] 2) When the workload pressures of the queue system with the maximum workload pressure and the queue system with the minimum workload pressure meet the scheduling conditions, the scheduler schedules an inference service instance from the queue system with the minimum workload pressure to the queue system with the maximum workload pressure;

[0020] The scheduling conditions include that the maximum workload pressure in the cloud service cluster is greater than the minimum workload pressure × threshold;

[0021] 3) The inference service instance being scheduled is performing inference service, and the request is rerouted to the inference service instance with the shortest queue under the same model, and the inference service is re-performed by other inference service instances.

[0022] For the above load balancing method of multi-model inference service, preferably, the inference service instances are deployed on the nodes of the queue system; multiple inference service instances are deployed on each node of the queue system.

[0023] For the above load balancing method of multi-model inference service, preferably, the workload pressure is described by the following expression: ρ i is the workload pressure of queue system i, where the numerator λ i ×E(X i ) represents the average token arrival rate of queue system i, and the denominator c i ×μ i represents the average token processing rate of queue system i; E(X i ) is the expected number of request tokens of queue system i, λ i is the arrival rate of requests for queue system i, c i is the number of inference service instances in queue system i, and μ i is the average processing time of each request in queue system i.

[0024] For the above load balancing method of multi-model inference service, preferably, at the service entrance of the cloud service cluster, the number of tokens returned by the autoregressive model corresponding to each queue system is counted; different computational costs are attached to the returned tokens at the service entrance of the cloud service cluster, and the weight value of each queue system is calculated; the scheduler will balance the workload pressures of different queue systems to achieve load balancing of the queue systems within the cloud service cluster.

[0025] Compared with the prior art, the advantages of the present invention are as follows: The load balancer and method of the multi-model inference service of the present invention can effectively achieve load balancing of each queue system in the cloud service cluster, thereby improving the processing efficiency of the inference service in the cloud service cluster; compared with the traditional greedy enumeration strategy, the load balancer and method of the multi-model inference service of the present invention have better performance in terms of workload, response time, and accuracy of response time allocation. Description of the Drawings

[0026] Figure 1 It is the overall flowchart of TRE.

[0027] Figure 2 It is the framework diagram of the autoregressive model.

[0028] Figure 3 It is the framework diagram of the hybrid deployment of the LLM model.

[0029] Figure 4 It is the impact of redeployment on the original request.

[0030] Figure 5 It is the request access volume of LLM1 and the allocated inference service instances under a fixed request arrival rate.

[0031] Figure 6 It is the request access volume of LLM2 and the allocated inference service instances under a fixed request arrival rate.

[0032] Figure 7 It is the request access volume of LLM3 and the allocated inference service instances under a fixed request arrival rate.

[0033] Figure 8 It is the request access volume of LLM4 and the allocated inference service instances under a fixed request arrival rate.

[0034] Figure 9 It is the P90 first token generation latency under a fixed request arrival rate.

[0035] Figure 10 It is the first token generation latency under a fixed request arrival rate.

[0036] Figure 11 It is the request access volume of LLM1 and the allocated inference service instances under a loose request arrival rate.

[0037] Figure 12 It is the request access volume of LLM2 and the allocated inference service instances under a loose request arrival rate.

[0038] Figure 13 It is the request access volume of LLM3 and the allocated inference service instances under a loose request arrival rate.

[0039] Figure 14 It is the request access volume of LLM4 and the allocated inference service instances under a loose request arrival rate.

[0040] Figure 15 It is the P90 first token generation latency under a loose request arrival rate.

[0041] Figure 16 The latency of the first token generation under a loose request arrival rate.

[0042] Figure 17 The request access volume of LLM1 and the allocated inference service instances under a dense request arrival rate.

[0043] Figure 18 The graph of the request access volume of LLM2 and the allocated inference service instances under a dense request arrival rate.

[0044] Figure 19 The request access volume of LLM3 and the allocated inference service instances under a dense request arrival rate.

[0045] Figure 20 The request access volume of LLM4 and the allocated inference service instances under a dense request arrival rate.

[0046] Figure 21 The request access volume of LLM5 and the allocated inference service instances.

[0047] Figure 22 The request access volume of LLM6 and the allocated inference service instances.

[0048] Figure 23 The graph of the P90 first token generation latency under a dense request arrival rate.

[0049] Figure 24 The first token generation latency under a dense request arrival rate. Detailed implementation manners

[0050] To facilitate the understanding of the present invention, the present invention will be described more comprehensively and meticulously below in conjunction with preferred embodiments, but the protection scope of the present invention is not limited to the following specific embodiments.

[0051] It should be specifically noted that when a certain component is described as "fixed to, fixedly connected to, connected to, or communicated with" another component, it can be directly fixed, fixedly connected, connected, or communicated to the other component, or indirectly fixed, fixedly connected, connected, or communicated to the other component through other intermediate connectors.

[0052] Unless otherwise defined, all professional terms used hereinafter have the same meaning as commonly understood by those skilled in the art. The professional terms used herein are only for the purpose of describing specific embodiments and are not intended to limit the protection scope of the present invention.

[0053] Embodiment 1

[0054] In the present invention, the number of tokens generated periodically by the model is used as the evaluation index of the existing computational workload of the model (token generation rate). This index has two advantages: the token generation rate can reflect the computable ability of the model in real time, and it is easy to count without additional acquisition overhead. Although the time and required tokens of LLM service requests are unpredictable, all running requests will generate tokens regularly. Combining the token generation rate and the queue system status of the model, we can accurately understand the workload pressure of different models. In the cloud service cluster of LLM, each model corresponds to a queue system.

[0055] At the same time, accurate scheduling timing can reasonably allocate cluster resources, reduce waste of computing resources, and should preferentially allocate computing resources to models with higher computing requirements rather than to those models with a short-term surge in request traffic. The computing power allocation within the cluster can be gradually stabilized through a progressive scheduling method.

[0056] In the present invention, a Token Rate Equalizer (TRE) is introduced, which is a progressive algorithm for multi-model load balancing that takes into account both the model workload pressure and the computing resource distribution. The overall process of TRE is as Figure 1 shown. TRE can monitor the token generation rate at the entrance of the cloud service cluster in real time, thereby evaluating the current computing resource allocation and the workload pressure of the model. By evaluating the load pressure of the queue system, TRE can dynamically allocate computing resources and stabilize the request queue of the queue system. TRE will timely adjust the models with higher workload in the queue system and schedule them with the models in the queue system with lower workload pressure, adjusting the allocation of inference service instances of two different models to ensure cluster stability and maximize benefits. In addition, TRE further analyzes the tail request delay caused by model redeployment, selects the inference service instance that causes the least delay for switching, and uses a priority queue to reduce the latency of model re-response.

[0057] In the present invention, the inference service process of LLM and the inference service framework for providing multiple models in a cluster are presented, and then the drawbacks of the current model scheduling strategy and the challenges of implementing dynamic model scheduling are explored.

[0058] The LLM of the present invention is an autoregressive model, which enables it to capture complex contexts and generate coherent relevant texts. The inference service has two computational stages. (1) Prefill stage: The model processes the input sequence in parallel, converts it into a series of tokens, as Figure 1 shown, and predicts the output of the tokens, as Figure 2 shown. (2) Decoding stage: The model generates tokens one by one according to the processed input sequence, as Figure 3as shown until a complete output sequence is generated or the limit on the number of tokens is reached, such as Figure 3 as shown.

[0059] The two-stage characteristic of the LLM differentiates LLM services from traditional services. The two-stage nature of the LLM also gives rise to two special latency metrics: Time To First Token (TTFT), which represents the time required to generate the first token during the prefill stage; and Time Per Output Token (TPOT), which represents the average time interval between generating tokens for each request during the decoding stage. Together, they affect the latency of LLM services.

[0060] Currently, cloud providers have deployed multiple LLM models to provide users with more diverse and specialized services. Collaborative efforts are required for the hybrid deployment of multiple models. As Figure 3 shown, the LLM model hybrid deployment framework consists of three parts: a controller that schedules user requests and loads models, a repository that stores models, and host nodes that provide inference services. The request router is responsible for distributing and scheduling incoming user requests from clients and routing them to the designated nodes. When there are no routable nodes in the cluster, that is, when the required model has not been loaded, the model is loaded to the designated node through the model loading scheduler, and the user request is routed. At the same time, the model scheduler adjusts the model deployment within the cluster according to the specified planning requirements, loading the specified model from the cloud storage node into the GPU memory. The host nodes equipped with multiple GPUs provide inference services through LLM instances.

[0061] Embodiment 1

[0062] A load balancer for a multi-model inference service in this embodiment includes a scheduler and a detector. The detector detects the workload pressure of each queue system in the cloud service cluster. When the workload pressures of the queue system with the maximum workload pressure and the queue system with the minimum workload pressure in the cloud service cluster meet the scheduling conditions, the scheduler schedules an inference service instance from the queue system with the minimum workload pressure to the queue system with the maximum workload pressure.

[0063] The scheduling conditions include that the maximum working pressure in the cloud service cluster is greater than the minimum working pressure multiplied by a threshold.

[0064] In this embodiment, each model corresponds to a queue system, and multiple inference service instances are deployed for the model in each queue system.

[0065] In the present invention, the workload pressure is described by the following expression: ρ i is the workload pressure of queue system i, where the numerator λ i×E(X i ) represents the average token arrival rate of queue system i. The denominator c i ×μ i represents the average token processing rate of queue system i; E(X i ) is the expected number of request tokens of queue system i, and λ i is the arrival rate of requests representing queue system i, and c i is the number of inference service instances in queue system i, and μ i is the average processing time for each request in queue system i. At the service entry of the cloud service cluster, the number of tokens returned by the autoregressive model corresponding to each queue system is counted.

[0066] In the expression of the workload pressure in this embodiment, the only variable that can be adjusted through scheduling is c i , while other variables may change over time. When ρ i becomes too high for system i, the goal of this embodiment is to reduce the workload pressure of the system by increasing c i to ensure that the average processing rate of the system exceeds the average token arrival rate and ensure that the workload of the system is in a stable state.

[0067] In the present invention, since the total number of inference service instances in the cloud service cluster is fixed, any increase or decrease in the number of inference service instances of system i must be offset by adjusting the number of instances allocated to other queue systems. Therefore, there are inherent scheduling limitations when attempting to reduce the load of a certain system. Therefore, the overall goal of the scheduling process is to minimize the workload pressure of each system, thereby achieving workload pressure balance.

[0068] Based on this goal, the scheduling timing of this embodiment is designed as follows: When the ρ i of these systems are not equal, but in fact, since is proportional to c i ×μ i , we can reflect through the ratio. For example, consider two queue systems i and queue system j; by observing these systems at a given time, we can obtain λ×E(X) and R token of these two systems. If the ratio is significantly greater than indicating that the workload pressure of system i is greater, then the scheduling strategy is to gradually reallocate instances from queue system j to queue system i. is the token rate of queue system i, that is, when the time interval is small, the number of tokens output by queue system i within a given time interval.

[0069] In the present invention, computational costs are appended to the returned tokens at the service entrance of the cloud service cluster, and the weight values of each queue system are calculated. The calculation of the first token incurs a higher workload; the scheduler balances the workload pressure of different queue systems to achieve load balancing of the queue systems within the cloud service cluster.

[0070] In the present invention, different weights are set for models with different priorities, and the tokens returned by the queue system corresponding to the model are counted at the service entrance provided by the cloud service cluster, and computational costs are appended to each returned token. Models with higher weight values have higher computational requirements when calculating the workload, and the scheduler will allocate more inference service instances during scheduling, that is, allocate more computational resources.

[0071] During the requested inference service process, the computational cost required for the first token is different from that of subsequent tokens. We will calculate the first token separately, and the computational cost is the product of the model computation amount k and the prompt length l; Figure 10 To generate latency for the first token at a fixed request arrival rate.

[0072] In the present invention, the inference service instances being scheduled to perform the inference service for the request are left in the queue system with the minimum workload pressure, and the inference service instances left in the queue system with the minimum workload pressure re-perform the inference service for the request. In the present invention, when the scheduler mobilizes the inference service instances, the original request needs to suspend the service and respond again on the new queue system. There are two response methods. One is to transfer the request cache of the inference service instance being scheduled to perform the inference service to the new queue system and continue the inference service. The other is to let the inference service instance leaving the request in the queue system with the minimum workload pressure re-perform the inference service for the request. Since the data volume of the request of the inference service instance being scheduled to perform the inference service is often in the order of GB, and due to the limitation of the bandwidth between machines, the transmission overhead is often greater than the recomputation overhead. To reduce the time of re-response, the present invention selects to let the inference service instance leaving the request in the queue system with the minimum workload pressure re-perform the inference service for the request.

[0073] In the present invention, when the scheduler mobilizes the inference service instances, the redeployment overhead of preemptive scheduling is inevitable, and it is necessary to minimize the negative impact of each redeployment. The overhead of model redeployment comes from the lag time caused by re-providing the inference service for the original request and the inability to provide service during the model loading time. Knowing in advance the LLM with the least workload, the model weights can be downloaded from the remote model repository to the inference service instance in advance; each inference service instance needs to run the model, and the model file needs to be obtained from the remote model repository. This is because the request rate of the LLM with a small workload remains stable. The impact of redeployment on the original request is as Figure 4As shown. The elapsed computing time refers to the time consumed by the original request in the current inference service instance. The waiting time refers to the waiting time for the request to rejoin the instance. The recalculation time refers to the time to restore the original request state.

[0074] The present invention also provides a load balancing method for a multi-model inference service, including the following steps;

[0075] 1) Statistically calculate the workload pressure of each queue system in the cloud service cluster; the inference service instances are deployed on the host nodes within the cloud service cluster; there are multiple inference service instances for each model of each said queue system within the cloud service cluster.

[0076] 2) When the workload pressures of the queue system with the maximum workload pressure and the queue system with the minimum workload pressure satisfy the scheduling condition, the scheduler schedules an inference service instance from the queue system with the minimum workload pressure to the queue system with the maximum workload pressure;

[0077] The said scheduling condition includes that the maximum workload pressure in the cloud service cluster is greater than the minimum workload pressure × threshold value.

[0078] The workload pressure is described by the following expression: ρ i is the workload pressure of queue system i, where the numerator λ i ×E(X i ) represents the average token arrival rate of queue system i, and the denominator c i ×μ i represents the average token processing rate of queue system i; E(X i ) is the expected number of request tokens of queue system i, λ i is the arrival rate of requests for queue system i, c i is the number of inference service instances in queue system i, and μ i is the average processing time for each request in queue system i.

[0079] 3) The requests for which the scheduled inference service instance is performing the inference service remain in the queue system with the minimum workload pressure, and the inference service instances remaining in the queue system with the minimum workload pressure re-perform the inference service for the requests.

[0080] In the present invention, the number of tokens returned by the autoregressive model corresponding to each queue system is statistically calculated at the service entrance of the cloud service cluster; the calculation cost is attached to the returned tokens at the service entrance of the cloud service cluster, and the weight value of each queue system is calculated; the scheduler will balance the workload pressures of different queue systems to achieve load balancing of the queue systems within the cloud service cluster.

[0081] The load balancer and method for multi-model inference service of the present invention can effectively achieve load balancing for each queue system in the cloud service cluster, thereby improving the processing efficiency of the inference service in the cloud service cluster; compared with the traditional greedy enumeration strategy, the load balancer and method for multi-model inference service of the present invention have better performance in terms of workload, response time, and accuracy of response time allocation.

[0082] Embodiment 1

[0083] This embodiment introduces a Token Rate Equalizer (TRE), which is a progressive algorithm for multi-model load balancing that takes into account both model workload pressure and computing resource distribution. The overall process of TRE is as Figure 1 shown. TRE can monitor the token generation rate at the entrance of the cloud service cluster in real time to evaluate the current computing resource allocation and the workload pressure of the model. By evaluating the load pressure of the queue system, TRE can dynamically allocate computing resources and stabilize the request queue of the queue system. TRE will timely adjust the inference service instances with higher workload in the queue system and schedule them with the inference service instances in the queue system with lower workload pressure to ensure cluster stability and maximum benefit. In addition, TRE further analyzes the tail request latency caused by model redeployment, selects the machine that causes the least delay for switching, and uses a priority queue to reduce the latency of model re-response.

[0084] The production environment data referred to by the inference service instance in this embodiment is an inference server equipped with 8 DGX-H100 cards, and the deployed model is the Llama2-70B model (Touvron, Hugo, et al. "Llama 2: Open foundation and fine-tuned chat models." arXiv preprint arXiv:2307.09288 (2023)). In this embodiment, the smallest unit for executing the inference service is the inference service instance. In this environment, the inference service instance in this embodiment fits the characteristics of the first token latency varying with the number of prompt tokens and the inter-token latency varying with the batch size. Each inference service instance obtains the corresponding pre-fill processing latency according to the input tokens of the request by utilizing the characteristics of iterative batching, and determines the decoding stage latency according to the current batch size of the inference service instance. At each moment, the request status of each computing unit is updated according to the completion status of the request.

[0085] Initial settings of the load balancer. The number of inference service instances for initializing the load balancer is evenly distributed among each queue system in the cloud service cluster, that is, each model obtains the same number of inference service instances corresponding to the queue system. We evaluate the system with 1-hour user request traces and accordingly shorten the redeployment time. Each time the machine is redeployed, it will enter a 90s downtime and cannot provide services. Our request scheduling policy adopts a first-come, first-served policy, sets the batch size to 8, and merges requests into one batch and sends them to one of the inference service instances.

[0086] Comparative Example 1

[0087] In Comparative Example 1, the SOTA is set to a machine placement strategy that generates the optimal SLO solution by brute-force enumeration. We set the default value of SLO to 3 times the inference latency (SLO scale = 3, that is, the service request target for each request is three times the inference latency under non-interference conditions). By enumerating all placement schemes, SOTA can ensure obtaining the global optimal solution. According to the historical request traces, redeployment is performed every 30 minutes, and simulations are carried out based on the request arrival time and required computing time within these 30 minutes. The scheme with the largest SLO among all placement schemes is selected for redeployment. During redeployment, SOTA selects the inference service instance with the minimum switching cost.

[0088] The following analyzes the workload, response time, and accuracy of response time allocation of Example 1 and Comparative Example 1 under fixed request arrival rate, loose request arrival rate, and dense request arrival rate.

[0089] ① Fixed request arrival rate

[0090] Workload. We used four different models. As Figure 5 shown, the access rate of LLM 1 is 5req / s, and the access rates of LLM 2 and LLM 3 are 10req / s. Figure 6 The LLM 2 and LLM 3 described in Figure 11 are 10req / s, and the LLM 4 described in

[0091] is 20req / s. A total of five groups of experiments with different numbers of inference service instances were conducted to ensure the variability of computing resource supply. Figure 9 Example 1 has a faster response time.Describes the latency of the cloud service cluster in response to users. As the number of inference service instances in the cluster increases, the first token generation time of TRE shows a linear decrease, which is attributed to sufficient computing resources and reduced queue latency. The reduction rate of the first token generation time latency of SOTA gradually decreases, because under the condition of limited availability, the allocation efficiency of scarce computing resources is low. Compared with SOTA, the average first token generation latency of TRE is reduced by 10.5% - 24.2%, and the response time is shortened by 33.2ms - 113.8ms. The first token generation time at the P90 level is reduced by 14.4% - 33.6%, and the response time is shortened by 105ms - 400ms.

[0092] Example 1 has a more accurate allocation of inference service instances. To prove that our scheduling strategy can dynamically and accurately detect traffic changes, we plotted the request traffic of each model against the number of inference service instances generated by TRE and SOTA, as Figures 5 - 8 shown. In the first seven minutes, TRE initiated progressive scheduling by detecting the workload pressure of various LLMs and reached a relatively stable state after seven minutes. However, the machine allocation for each LLM is different: the number of machines allocated to LLM 1 is approximately 30, LLM 2 and LLM 3 are approximately 37, and LLM 4 is approximately 45. The reason for this fluctuation is that even under the assumptions of Poisson distribution and constant arrival rate, different LLMs may encounter significantly different request traffic at the same time, resulting in workload imbalance. Once a significant load gap is detected, TRE will intervene and adjust to balance the workload between each LLM.

[0093] ② Loose request arrival rate

[0094] Workload. We used a total of four models, and the request traffic is as Figures 11 - 14 shown. We used a total of four models, and the request traffic is shown in the figure. The basic access speed of each model is 5req / s. One hour of tracking is divided into four time periods. In each time period, we randomly select any two models to increase the access speed to 25req / s. To be consistent with the load in the real environment, the request arrival rate will gradually increase the model access speed to 25req / s. In addition, we also conducted five different experiments for different inference service instances to evaluate the changes in computing resource supply.

[0095] Example 1 has a faster response. Figure 15 and Figure 16Describes the latency of the cloud service cluster in response to users. As the number of inference service instances in the cluster increases, the latency of the first token generation of TRE and SOTA both show a linear downward trend. Compared with SOTA, the average latency of the first token generation of TRE is reduced by 6.3% - 11.3%, and the response time is shortened by 16ms - 27ms. The latency of the first token generation at the P90 level is reduced by 8.7% - 14.5%, and the response time is shortened by 53ms - 83ms.

[0096] Example 1 has a more precise allocation of inference service instances. As Figures 11 - 14 shown, during the one-hour user record, each LLM shows significant changes in request traffic. The number of inference service instances allocated by TRE is closely related to the trend of request traffic changes of the LLM. In Figure 11 , LLM 1 shows a change in the number of inference service instances in the last 30 minutes, with a fluctuation range of approximately 3 inference service instances. This is because the user request traffic of other LLMs fluctuates during this period, and the workload pressure of LLM 1 is the smallest during this period, so it becomes the best candidate for adjustment scheduling.

[0097] ③Dense request arrival rate

[0098] Workload. We used six models, and each model shows different request traffic trends. Among them, the access rate of LLM 1 drops from 15 req / s to 5 req / s, as Figure 16 shown. The access rate of LLM 2 is initially 5 req / s and then randomly increases to 13 req / s, as Figure 18 shown. LLM 3 starts with an access rate of 15 req / s, drops to 5 req / s at 25 minutes, and starts to increase the access rate at 45 minutes, as Figure 19 shown. LLM 4 rises from 5 req / s to 10 req / s at 30 minutes and returns to 5 req / s at 45 minutes, as Figure 20 shown. LLM 5 maintains a constant access rate of 10 req / s, as Figure 21 shown. The access rate of LLM6 increases from 5 req / s to 15 req / s, as Figure 22 shown.

[0099] Example 1 has a faster response time. Figure 23 and Figure 24Describes the latency of the cloud service cluster in responding to user requests. As the number of models increases, computing resources become more strained. Unreasonable resource allocation can lead to a significant increase in queue latency, which can be seen from the first token generation time. When the number of inference service instances reaches 150, due to the unreasonable computing resource allocation of LLM 4, the queuing time reaches 30 minutes, resulting in a decrease in the average first token generation time and the first token generation time reaching a peak of 2278 ms. In addition, the P90-level first token generation latency indicates that the average first token generation time is 2077 milliseconds. Compared with SOTA, the average first token generation latency of TRE is reduced by 15.4% - 76.9%, and the response time is shortened by 78.5 ms - 1752.5 ms. The P90-level first token generation latency is reduced by 18.8% - 43.1%, and the response time is shortened by 235 ms - 896 ms.

[0100] Example 1 has a more precise allocation of inference service instances. In the one-hour user record, the access traffic change trends of each LLM are different. The number of inference service instances allocated by TRE is closely related to the change trend of the request traffic of the LLM, and can timely sense the traffic change and perform scheduling, as Figures 17 - 22 shown, where the inference service instances allocated by SOTA for LLM 2, LLM 3, and LLM 4 are basically consistent with the past traffic when redeployed according to historical tracking, but do not match the currently modeled traffic.

Claims

1. A load balancer for a multi-model inference service, characterized in that: It includes a scheduler and a detector, wherein the detector detects the workload pressure of each queue system in the cloud service cluster, and when the workload pressure of the queue system with the maximum workload pressure and the queue system with the minimum workload pressure in the cloud service cluster meet the scheduling conditions, the scheduler schedules an inference service instance from the queue system with the minimum workload pressure to the queue system with the maximum workload pressure; The scheduling condition includes that the maximum working pressure in the cloud service cluster is greater than the minimum working pressure×threshold.

2. The load balancer for multi-model inference service according to claim 1, characterized in that: Each model of the queue system is deployed with multiple inference service instances.

3. The load balancer for multi-model inference service according to claim 1, characterized in that: The workload pressure is described by the following expression: ρ i is the workload pressure of queue system i, where the numerator λ i ×E(X i ) represents the average word arrival rate of queue system i, and the denominator c i ×μ i represents the average word processing rate of queue system i; E(X i ) is the expected number of request tokens for queue system i, λ i is the arrival rate of requests to queue system i, c i is the number of inference service instances in queue system i, μ i is the average processing time of each request in queue system i.

4. The load balancer for multi-model inference service according to claim 3, characterized in that: At the service entrance of the cloud service cluster, count the number of tokens returned by the autoregressive model corresponding to each queue system.

5. The load balancer for multi-model reasoning service according to claim 4, characterized in that: At the service entrance of the cloud service cluster, different computing costs are added to the returned words, and the weight value of each queue system is calculated. The calculation of the first word will generate a higher workload. Under the same workload, the number of inference service instances of the queue system with a high weight value is greater than the number of inference service instances of the queue system with a low weight value.

6. The load balancer for multi-model inference service according to claim 1, characterized in that: The request for the inference service being performed by the scheduled inference service instance will be rerouted to the inference service instance with the shortest queue under the same model, and the inference service will be re-performed by other inference service instances.

7. A load balancing method for multi-model reasoning services, characterized in that: The steps include: 1) Count the workload pressure of each queue system in the cloud service cluster; 2) When the workload pressures of the queue system with the maximum workload pressure and the queue system with the minimum workload pressure meet the scheduling conditions, the scheduler schedules an inference service instance from the queue system with the minimum workload pressure to the queue system with the maximum workload pressure; The scheduling condition includes that the maximum working pressure in the cloud service cluster is greater than the minimum working pressure × threshold; 3) The request for inference service being performed by the scheduled inference service instance will be rerouted to the inference service instance with the shortest queue under the same model, and the inference service will be performed again by other inference service instances.

8. The load balancing method for multi-model reasoning service according to claim 7, characterized in that: The inference service instance is deployed in the model of the queue system; each model of the queue system is deployed with multiple inference service instances.

9. The load balancing method for multi-model reasoning service according to claim 1, characterized in that: The workload pressure is described by the following expression: ρ i is the workload pressure of queue system i, where the numerator λ i ×E(X i ) represents the average word arrival rate of queue system i, and the denominator c i ×μ i represents the average word processing rate of queue system i; E(X i ) is the expected number of request tokens for queue system i, λ i is the arrival rate of requests to queue system i, c i is the number of inference service instances in queue system i, μ i is the average processing time of each request in queue system i.

10. The load balancing method for multi-model reasoning service according to claim 7, characterized in that: At the service entrance of the cloud service cluster, the number of words returned by the autoregressive model corresponding to each queue system is counted; at the service entrance of the cloud service cluster, different computing costs are added to the returned words, and the weight value of each queue system is calculated; the scheduler will balance the workload pressure of different queue systems to achieve load balancing of the queue systems within the cloud service cluster.

Citation Information

Cited By

  • Model reasoning scheduling method and system, electronic equipment and storage medium

    CN120909737A

  • Green ubiquitous distributed reasoning service method, device, equipment and medium

    CN121168655A

  • A green and ubiquitous distributed reasoning service method, apparatus, equipment and medium

    CN121168655B

  • Request processing method and device based on large model, electronic equipment and storage medium

    CN121328705A