Method, apparatus, device and medium for request rescheduling in large language model services
By obtaining the queuing situation of instances and key-value cache calculation scheduling measurements in the large language model service, determining the source instance and target instance, and performing migration optimization request allocation, the problem that traditional scheduling strategies cannot be effectively scheduled is solved, and resource utilization and response performance are improved.
Patent Information
- Application Number
- CN202510230430.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-28
- Publication Date
- 2025-07-08
- Estimated Expiration
- 2045-02-28
AI Technical Summary
Traditional scheduling strategies cannot effectively respond to the uncertainty and dynamic nature of large language model requests, resulting in low resource utilization and poor response performance.
By obtaining the queueing situation and key-value cache for each instance, calculating the scheduling quantity, determining the source and target instances to be scheduled, and optimizing request allocation based on the preset policy migration token and/or key-value cache when receiving the migration instruction.
It improves the system's resource utilization and response performance, and solves the problems of uncertainty and dynamics of traditional scheduling strategies when facing large language model requests.
Smart Images

Figure CN119718594B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of large language models, and particularly to a method, device, equipment and medium for request rescheduling in large language model services. Background Art
[0002] In recent years, large language models (LLMs) have been widely used in multiple domain tasks due to their excellent general performance and wide applicability. However, compared with traditional deep neural network (DNN) services, large language model requests exhibit high uncertainty in terms of input length, output length, and latency requirements, significantly increasing the complexity of scheduling and resource management.
[0003] Traditional DNN services are usually designed for specific tasks, with relatively fixed input and output sizes, and relatively uniform task complexity and latency requirements. Therefore, the request scheduling and resource management of traditional DNN services are relatively simple. Most requests are homogeneous, and the scheduling strategy is designed based on deterministic factors such as the processing duration and resource occupancy of the requests.
[0004] However, the inference process of large language models is usually generated step by step. Each request requires multiple iterations to generate output tokens, and the generation process of each token is dynamic. The total generation length is difficult to determine in advance, making the execution time and resource requirements of large language models highly uncertain. In addition, during the execution of large language models, due to the step-by-step generation characteristics, the generation of each token requires dynamic allocation of GPU memory, and the memory requirement increases with the number of generated tokens, resulting in significant memory fragmentation problems, which affect the resource scheduling among multiple instances.
[0005] To address the above problems, the scheduling strategies commonly adopted in related technologies include round-robin scheduling, resource-demand-based scheduling, and connection-number-based load balancing. However, these strategies often fail to effectively schedule when faced with the uncertainty of large language model requests. For example, the round-robin scheduling and connection-number-based load balancing methods assume that the resource requirements and processing times of requests are static, which may lead to the situation where some instances are overloaded while others are idle. The resource-demand-based scheduling relies on accurate prediction of request resource consumption before scheduling, which is difficult to achieve for large language model requests and urgently needs to be solved. Summary of the Invention
[0006] The present invention provides a method, device, equipment and medium for request rescheduling in large language model services to solve the problem that traditional scheduling strategies cannot effectively schedule in the face of the uncertainty and dynamics of large language model requests, thereby improving the resource utilization rate and response performance of the system.
[0007] To achieve the above object, an embodiment of the first aspect of the present invention provides a method for request rescheduling in large language model services, including the following steps:
[0008] Obtain the queuing situation and key-value cache of large language model service requests in each instance, and calculate the scheduling amount of each instance according to the queuing situation and key-value cache of large language model service requests in each instance;
[0009] Determine at least one source instance to be scheduled and the target instance corresponding to each source instance to be scheduled according to the scheduling amount of each instance;
[0010] When receiving a migration instruction, based on the type of large language model service request and a preset migration strategy, migrate the tokens and / or key-value cache in each source instance to be scheduled to the target instance corresponding to each source instance to be scheduled, and feedback the response result of the large language model service request to the user.
[0011] Through the above technical means, by realizing efficient request allocation in a dynamic environment, the problem that traditional scheduling strategies cannot effectively schedule in the face of the uncertainty and dynamics of large language model requests is solved, thereby improving the resource utilization rate and response performance of the system.
[0012] According to the method for request rescheduling in large language model services proposed by the embodiment of the present invention, by calculating the scheduling amount of each instance according to the queuing situation and key-value cache of large language model service requests, and determining each source instance to be scheduled and its corresponding target instance; when receiving a migration instruction, based on the type of large language model service request and a preset migration strategy, migrate the tokens and / or key-value cache in each source instance to be scheduled to the corresponding target instance, and feedback the response result of the large language model service request to the user. Thus, by realizing efficient request allocation in a dynamic environment, the problem that traditional scheduling strategies cannot effectively schedule in the face of the uncertainty and dynamics of large language model requests is solved, thereby improving the resource utilization rate and response performance of the system.
[0013] To achieve the above object, an embodiment of the second aspect of the present invention provides an apparatus for request rescheduling in large language model services, including:
[0014] An obtaining module, configured to obtain the queuing situation and key-value cache of large language model service requests in each instance, and calculate the scheduling amount of each instance according to the queuing situation and key-value cache of large language model service requests in each instance;
[0015] A determining module, configured to determine at least one source instance to be scheduled and the target instance corresponding to each source instance to be scheduled according to the scheduling amount of each instance;
[0016] A migration module, which, when receiving a migration instruction, migrates the tokens and / or key-value caches in each source instance to be scheduled to the corresponding target instance based on the type of the large language model service request and a preset migration strategy, and feeds back the response result of the large language model service request to the user.
[0017] The device for request rescheduling in the large language model service according to the embodiments of the present invention calculates the scheduling amount of each instance based on the queuing situation of the large language model service request and the key-value cache, and determines each source instance to be scheduled and its corresponding target instance; when receiving a migration instruction, it migrates the tokens and / or key-value caches in each source instance to be scheduled to the corresponding target instance based on the type of the large language model service request and a preset migration strategy, and feeds back the response result of the large language model service request to the user. Thus, by realizing efficient request allocation in a dynamic environment, the problem that traditional scheduling strategies cannot effectively schedule in the face of the uncertainty and dynamics of large language model requests is solved, thereby improving the resource utilization rate and response performance of the system.
[0018] To achieve the above object, an embodiment of the third aspect of the present invention proposes an electronic device, including: a memory, a processor, and a computer program stored on the memory and executable on the processor, where the processor executes the program to implement the method for request rescheduling in the large language model service as described in the above embodiments.
[0019] To achieve the above object, an embodiment of the fourth aspect of the present invention proposes a computer-readable storage medium, on which a computer program is stored, and the program is executed by a processor to implement the method for request rescheduling in the large language model service as described in the above embodiments.
[0020] Additional aspects and advantages of the present invention will be given in part in the following description, become apparent in part from the following description, or be learned through the practice of the present invention. Description of the Drawings
[0021] To more clearly illustrate the embodiments of the present invention, the following will briefly introduce the drawings required in the embodiments. Obviously, the drawings in the following description are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained based on these drawings.
[0022] Figure 1 is the traditional deep neural network service scheduling method;
[0023] Figure 2 is a schematic diagram of a request rescheduling system in the large language model service according to an embodiment of the present invention;
[0024] Figure 3 Flow chart of a method for request rescheduling in large language model services provided according to an embodiment of the present invention;
[0025] Figure 4 Schematic diagram of an instance migration process according to an embodiment of the present invention;
[0026] Figure 5 Block diagram of an apparatus for request rescheduling in large language model services provided according to an embodiment of the present invention;
[0027] Figure 6 Schematic diagram of the structure of an electronic device provided according to an embodiment of the present invention.
[0028] Reference numerals: 10 - apparatus for request rescheduling in large language model services, 100 - acquisition module, 200 - determination module, 300 - migration module; 601 - memory, 602 - processor, 603 - communication interface. Detailed implementation manners
[0029] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts belong to the protection scope of the present invention.
[0030] It should be noted that in the description of the present invention, the terms "include", "comprise" or any other variant thereof are intended to cover a non - exclusive inclusion, such that a process, method, article or device including a series of elements not only includes those elements but also includes other elements not explicitly listed, or further includes elements inherent to such process, method, article or device. The terms "first", "second", etc. in the present invention are used to distinguish similar objects and not to describe a specific order or sequence.
[0031] In order to enable those skilled in the art of the present technology to better understand the solution of the present invention, the present invention will be further described in detail below in conjunction with the accompanying drawings and specific implementation manners.
[0032] Next, a method, apparatus, device and medium for request rescheduling in large language model services proposed according to an embodiment of the present invention will be described with reference to the accompanying drawings.
[0033] Before introducing the method for request rescheduling in the large language model service of the embodiments of the present invention, first briefly introduce the scheduling strategies for requests in deep neural network services in the related art, as well as the request rescheduling system in the large language model service involved in the method for request rescheduling in the large language model service of the present invention.
[0034] Specifically, as Figure 1 shown, in the related art, the traditional deep neural network service scheduling method is that a user (Actor) sends a request to a traditional scheduling strategy module, and this module distributes the request to a traditional service for processing according to a preset strategy. The traditional service generally consists of multiple instances (or multiple replicas), and the functions of each instance are the same and can all process requests from users. The scheduling strategies for traditional deep neural network service requests mainly include round-robin scheduling, resource demand-based scheduling strategies, and connection number-based load balancing scheduling strategies.
[0035] Among them, round-robin scheduling is a simple scheduling strategy. When scheduling, requests are cyclically assigned to service instances in order, so each request is evenly assigned for execution. This scheduling strategy is applicable to the situation where the processing time differences of each request are relatively small.
[0036] The resource demand-based scheduling strategy schedules according to the resource requirements of each request (such as GPU video memory, CPU, etc.). Therefore, it is necessary to accurately model the resource requirements of the requests, otherwise it may lead to inefficient use of resources.
[0037] The connection number-based load balancing scheduling strategy means that in a multi-instance service, the scheduling system will select which instance to allocate a new request to based on the current connection numbers of each instance (that is, the number of requests being processed). Instances with fewer request connections will preferentially receive new requests. This method can avoid some nodes from being overloaded due to excessive requests and helps improve the overall performance of the system.
[0038] The above three scheduling strategies represent the design patterns of the current scheduling strategies for traditional deep neural network services, and their scheduling strategies based on predictability are feasible when solving the request scheduling problems of traditional deep neural network services.
[0039] Those skilled in the art can understand that during the large language model inference process, in order to improve efficiency, the large language model uses a key-value caching mechanism to help the model avoid recalculating previous input information during generation, thereby accelerating the inference process. The implementation method of the key-value caching mechanism is briefly described as follows:
[0040] In the self-attention mechanism of Transformer, each input token generates three vectors, namely Query, Key, and Value. During the standard inference process, all inputs need to recalculate these three vectors each time they are generated. When using key-value caching, the model caches the already calculated keys and values. In this way, when generating each token, the model can directly use these cached keys and values without having to recalculate all the inputs. This caching mechanism is particularly efficient during long text generation, reducing the computational load.
[0041] Prefill in the inference process of large language models refers to processing the input sequence in the user request through the model all at once to form a key-value cache for use in the Decode stage. Based on the context of the prefill stage, the model gradually generates new tokens until it reaches the target length or the end symbol. Decoding uses the key-value cache (KV Cache) to avoid repeated calculations and improve efficiency.
[0042] Inference of large language models usually involves multiple iterations, generating one token each time, and the total generation length is usually unpredictable. This makes the memory requirements for each large language model request change dynamically during the generation process. Also, large language model requests may have different processing durations for each request due to different tasks being processed (such as text summarization tasks that require a longer duration and real-time conversation tasks with a shorter processing duration).
[0043] Since the scheduling strategies for traditional deep neural network service requests are based on request determinism, they face difficulties when dealing with the unpredictable processing duration, resource occupancy, input, and output of large language model requests. For example, polling scheduling and load balancing scheduling strategies based on the number of connections usually assume that the resources and processing duration required for each connection are relatively static. This may lead to unreasonable request scheduling. For instance, when scheduling based on the number of connections, multiple requests with a high processing duration may be scheduled together, resulting in some instances being overloaded while other instances are idle. Adopting scheduling based on resource requirements requires accurate prediction of the resource consumption of requests during scheduling, which is difficult to achieve for large language model requests.
[0044] Based on the fact that the above traditional scheduling strategies cannot effectively schedule in the face of the unpredictability of large language model requests, the method of request rescheduling in large language model services proposed by the present invention reschedules the request according to the current state within the instance after the traditional scheduling strategy schedules the request to an instance, and schedules the request to other instances, solving the problem that the traditional scheduling strategy cannot effectively schedule in the face of the unpredictability of large language model requests.
[0045] Furthermore, as Figure 2 shown, Figure 2The request rescheduling system in the large language model service involved in the method for request rescheduling in the large language model service according to the embodiments of the present invention includes a scheduling module, a monitoring module, and a migration module. Among them, the scheduling module is located globally, and each instance includes a monitoring module and a migration module.
[0046] Among them, the scheduling module comprehensively analyzes and determines the instances that need to be scheduled based on the load data collected by the monitoring modules in all instances, and at the same time evaluates the instances with lower load and the ability to receive more requests. Then, the scheduling module establishes an association relationship between the source instance and the target instance, and sends a migration instruction to the migration module. It should be noted that the scheduling module does not directly determine the specific details of the request migration. Instead, after issuing the migration instruction, the migration module within the instance is responsible for controlling and executing the specific request migration operation. In addition, the monitoring module is responsible for real-time monitoring of the load situation of the instances, and regularly reports the load data to the scheduling module at set time intervals to support its decision-making process.
[0047] In addition, the scheduling module of the present invention collects monitoring data in real time and dynamically adjusts the scheduling scheme. Therefore, during the request processing process, as the key-value cache of the requests increases, it can quickly determine the requests that need to be scheduled in a timely manner, thereby timely fixing the irrationality of the traditional scheduling strategy.
[0048] The following details the method for request rescheduling in the large language model service using the above-mentioned request rescheduling system in the large language model service.
[0049] Figure 3 It is a flowchart of the method for request rescheduling in the large language model service according to an embodiment of the present invention.
[0050] Exemplarily, as Figure 3 shown, the method for request rescheduling in the large language model service includes the following steps:
[0051] In step S301, obtain the queuing situation and key-value cache of the large language model service requests in each instance, and calculate the scheduling volume of each instance according to the queuing situation and key-value cache of the large language model service requests in each instance.
[0052] Specifically, before obtaining the queuing situation and key-value cache of the large language model service requests in each instance, first determine whether a large language model service request is received, and when a large language model service request is received, schedule the large language model service request to multiple instances.
[0053] Through the above technical solution, it is ensured that the system can respond to the large language model service request in a timely manner, and achieve load balancing through multi-instance scheduling, providing a basis for subsequent resource management and optimized scheduling.
[0054] Further, in some embodiments, the scheduling volume of each instance is calculated according to the queuing situation of the large language model service requests and the key-value cache in each instance, including: obtaining the key-value cache occupied by the running requests, the virtual key-value cache expected to be occupied by the queuing requests, the total video memory capacity of the current instance, the video memory capacity occupied by the inherent number of parameters of the current large language model inference model, and the batch size of the current large language model inference model; calculating the scheduling volume of each instance according to the key-value cache occupied by the running requests, the virtual key-value cache expected to be occupied by the queuing requests, the total video memory capacity of the current instance, the video memory capacity occupied by the inherent number of parameters of the current large language model inference model, and the batch size of the current large language model inference model.
[0055] Among them, the key-value cache occupied by the running requests refers to the key-value cache occupied by the currently processed requests, the virtual key-value cache expected to be occupied by the queuing requests refers to the key-value cache expected to be occupied by the currently queuing requests waiting to be processed, the total video memory capacity of the current instance refers to the total GPU video memory capacity of the current large language model service instance, the video memory capacity occupied by the inherent number of parameters of the current large language model inference model refers to the video memory capacity occupied by the parameters (such as weights, biases, etc.) of the large language model inference model, and the batch size of the current large language model inference model refers to the number of requests processed at one time during the large language model inference process.
[0056] Specifically, in the embodiments of the present invention, the monitoring module in the request rescheduling system of the large language model service can be used to obtain the key-value cache occupied by the running requests and the virtual key-value cache expected to be occupied by the queuing requests. The total video memory capacity allocated to the current instance can be directly read from the hardware configuration or system resource management tool. The video memory capacity occupied by the inherent number of parameters of the current large language model inference model can be calculated according to the parameter scale at model loading. The batch size of the current large language model inference model can be obtained through the configuration file or API interface of the model operation framework, and no specific limitation is made here.
[0057] Further, for the set of running requests R, it is necessary to calculate the number of key-value caches occupied by these requests. The calculation formula for the key-value cache of the set of running requests R is:
[0058] ;
[0059] ;
[0060] Among them, is the key-value cache of the set of running requests R, is the sum of the cache in the pre-filling stage and the cache in the decoding stage, It is the cache for the prefill stage, and it is the cache for the decoding stage.
[0061] For the set of requests Q that are queuing, based on the priority P of the requests in the set of requests Q that are queuing, calculate the virtual key-value cache that is expected to be occupied:
[0062] ;
[0063] ;
[0064] Among them, is the virtual key-value cache that the set of requests Q that are queuing is expected to occupy, is the virtual key-value cache of each request in the set of queuing requests. At this time, the request is queuing and has not actually been executed on the physical machine and occupied the video memory, so it belongs to the theoretical data in the statistical sense and is therefore called "virtual" cache. It is calculated in two cases. The first case is when the request has not been executed at all. At this time, decoding has not started yet, so only the input sequence of the user request needs to be counted . The second case is when the request is suspended during execution due to tight remaining space in the video memory and then enters the queuing set. Since the request has entered the decoding stage, then at this time is equal to . is the request in the set of requests that are queuing, is the queuing time of the request, is the queuing time influence parameter, is the priority of the request, is the priority influence parameter, is the threshold, corresponds to the prefill stage in the LLM inference process. In this stage, according to the input sequence of the user request, the prefill key-value cache is calculated through the model at one time, Only the second case exists, that is, the suspended request. Since it has entered the decoding stage, when counting its cache, it is necessary to count the key-value cache that has been generated. Because when the request is scheduled to be executed again, the key-value caches of both the prefill and decoding stages will be swapped into the video memory, and the request will continue to execute from the suspended point.
[0065] Among them, the queuing time of the request is a number greater than 0. When the request starts to queue, the virtual key-value cache that the requests queuing are expected to occupy grows linearly. When the queuing time of the request is greater than the threshold T, the virtual key-value cache Is exponential growth. Queuing time impact parameter For balancing the impact of queuing time on the final degree of impact, the queuing time impact parameter Is a number greater than or equal to 1. When the large language model service is highly sensitive to the waiting time of users, it is necessary to increase the queuing time impact parameter . Priority impact parameter For balancing the impact of priority on the final degree of impact, the priority impact parameter Is a number greater than or equal to 1. When the large language model service focuses on quickly responding to high-priority user requests, then it is necessary to increase the priority impact parameter ,
[0066] It can be understood that because the large language model service needs to consider the impact of latency on users, so as the queuing time increases, when the queuing time is greater than the set threshold T, the set of requests Q that are queuing is expected to occupy the virtual key-value cache increases rapidly, thus triggering scheduling to avoid excessive latency of user requests. Among them, the threshold T can be a pre-set threshold, and the setting principle is the user's tolerance for latency. For high-priority requests, queuing should be avoided, so it is designed that when queuing, the virtual key-value cache occupied by the queuing requests is increased rapidly value, thus rapidly triggering scheduling. Among them, the priority P is an integer greater than or equal to 0, and the larger the priority P value, the higher the level.
[0067] Thus, by comprehensively considering the queuing duration and priority of the requests in the queue, it is possible to quickly discover which requests need to be scheduled in a timely manner, thereby increasing user friendliness.
[0068] Furthermore, in some embodiments, the scheduling amount of each instance is calculated according to the key-value cache occupied by the running requests, the virtual key-value cache expected to be occupied by the queuing requests, the total video memory capacity of the current instance, the video memory capacity occupied by the inherent number of parameters of the current large language model inference model, and the batch size of the current large language model inference model, including: obtaining the load value of the current instance according to the key-value cache occupied by the running requests, the virtual key-value cache expected to be occupied by the queuing requests, and the video memory capacity occupied by the inherent number of parameters of the current large language model inference model; obtaining the remaining space of the current instance according to the difference between the total video memory capacity of the current instance and the load value of the current instance; obtaining the scheduling amount of the current instance according to the ratio between the remaining space of the current instance and the batch size of the current large language model inference model.
[0069] Among them, the load value of the current instance is obtained from the sum of the key-value cache occupied by the running requests, the virtual key-value cache expected to be occupied by the queuing requests, and the video memory capacity occupied by the number of parameters inherent in the current large language model inference model.
[0070] Specifically, the load value of the current instance can be expressed as:
[0071] ;
[0072] Among them, is the load value of the current instance, is the key-value cache occupied by the running requests, is the virtual key-value cache expected to be occupied by the queuing requests, is the video memory capacity occupied by the number of parameters inherent in the current large language model inference model.
[0073] Furthermore, let C be the total video memory capacity of the instance. According to the difference between the total video memory capacity of the current instance and the load value of the current instance, the remaining space of the current instance can be obtained, that is, using the remaining space of the current instance can be calculated. Since tokens are also generated in batches, dividing the remaining space of the current instance by the batch size B of the current large language model inference model can calculate the consumption speed of these remaining spaces, that is, the scheduling amount. For ease of understanding, the scheduling amount can be expressed as:
[0074] ;
[0075] Among them, is the scheduling amount, is the total video memory capacity of the instance, is the key-value cache occupied by the running requests, is the virtual key-value cache expected to be occupied by the queuing requests, is the video memory capacity occupied by the number of parameters inherent in the current large language model inference model, and B is the batch size of the current large language model inference model.
[0076] It can be understood that the scheduling amount S may be negative. When the scheduling amount S is negative, it means that the video memory occupied by the queuing requests, the video memory capacity occupied by the number of parameters inherent in the current large language model inference model and the key-value cache occupied by the running requests have approached the total video memory capacity C of the instance. Therefore, new requests cannot obtain sufficient video memory and can only queue. Therefore, the purpose of designing the virtual cache in the present invention is to make S smaller or even negative, so that in the next scheduling module, the scheduling of these instances can be preferentially considered.
[0077] Thus, by comprehensively calculating the load value and remaining space of the current instance and determining the scheduling amount in combination with the model batch size, the resource usage and scheduling capabilities of the instance can be accurately evaluated, thereby achieving efficient utilization of resources.
[0078] In step S302, at least one source instance to be scheduled and a target instance corresponding to each source instance to be scheduled are determined according to the scheduling amount of each instance.
[0079] Further, in some embodiments, determining at least one source instance to be scheduled and a target instance corresponding to each source instance to be scheduled according to the scheduling amount of each instance includes: based on the scheduling amount of each instance, determining a set of source instances to be scheduled and a set of candidate target instances; determining whether the set of candidate target instances is an empty set; if the set of candidate target instances is not an empty set, then based on the set of candidate target instances, filtering out the instances that are on the same node as each source instance to be scheduled in the set of source instances to be scheduled and have the largest scheduling amount, to obtain the target instance corresponding to each source instance to be scheduled.
[0080] Among them, the set of source instances to be scheduled refers to the set of instances that need to migrate requests to other instances due to high load, and the set of candidate target instances refers to the set of instances with low load and sufficient resources to receive migration requests.
[0081] Further, in some embodiments, based on the scheduling amount of each instance, determining a set of source instances to be scheduled and a set of candidate target instances includes: based on the scheduling amount of each instance, filtering out the instance combinations with a scheduling amount less than a first preset threshold to obtain the set of source instances to be scheduled; based on the scheduling amount of each instance, determining whether there are instances with a scheduling amount greater than a second preset threshold; if there are instances with a scheduling amount greater than the second preset threshold, then combining the instances with a scheduling amount greater than the second preset threshold to obtain the set of candidate target instances.
[0082] Among them, the first preset threshold and the second preset threshold can be thresholds set by those skilled in the art according to actual situations, and are not specifically limited herein.
[0083] Specifically, the scheduling module collects the monitoring scheduling amount S from the set of instances of the large language model service in real time and sorts them in ascending order of the value of the scheduling amount S. Therefore, as shown in Table 1, , indicating that the instance at the k-th position after sorting is on the j-th node.
[0084] Table 1
[0085]
[0086] Further, in ascending order, the instance k is taken out in turn and judged Whether the value is less than the first preset threshold G. If the value is less than the first preset threshold G, start scheduling, use instance k as the source instance, add the source instance to the set of source instances to be scheduled, and then find a suitable target instance from largest to smallest:
[0087] .
[0088] It should be noted that when looking for a target instance , it is necessary to satisfy being greater than the second preset threshold H. When the scheduling volume S values of multiple instances are satisfied at the same time, add the instances that meet the conditions to the set of candidate target instances, and preferentially select from the set of candidate target instances and the candidate target instance that is on the same node and has the largest scheduling volume S value as the target instance corresponding to the current source instance.
[0089] Therefore, the embodiment of the present invention takes into account the problem that the communication bandwidth across nodes is less than the communication bandwidth between GPUs within a node. When scheduling, it preferentially selects instances within the same node, thus saving migration time, and can accurately identify instances with insufficient resources and instances with abundant resources, providing a basis for subsequent scheduling decisions, thereby optimizing resource allocation.
[0090] Furthermore, in some embodiments, after determining whether the set of candidate target instances is an empty set, it further includes: if the set of candidate target instances is an empty set, create a new instance for each source instance to be scheduled; use the new instance corresponding to each source instance to be scheduled as the target instance corresponding to each source instance to be scheduled.
[0091] Specifically, when the set of candidate target instances is an empty set, it means that the scheduling module fails to find a target instance. To ensure the smooth progress of scheduling and avoid the source instance affecting the service quality due to high load, it is necessary to create a new instance for each source instance to be scheduled, associate each source instance to be scheduled with its corresponding new instance, and set the target address to the new instance address, that is, designate the newly created instance as the target instance.
[0092] Therefore, through the above technical means, it is ensured that in the case of no suitable target instance, the system can still meet the scheduling requirements by dynamically creating new instances, ensuring the continuous operation of the system and the reasonable allocation of resources.
[0093] In step S303, when receiving the migration instruction, based on the type of the large language model service request and the preset migration strategy, migrate the tokens and / or key-value caches in each source instance to be scheduled to the target instance corresponding to each source instance to be scheduled, and feedback the response result of the large language model service request to the user.
[0094] Specifically, when the scheduling module determines the source instances to be scheduled and their corresponding target instances, the migration module starts migrating or rejects migration according to the preset migration strategy. First, the priorities of the requests are sorted, the requests are scheduled according to the sorting result of the priorities, and the requests with the same priority are scheduled according to the principle of first come, first served.
[0095] Through the above technical means, the migration process is simplified, the migration efficiency is improved, and at the same time, it is ensured that the target instance can quickly take over the unexecuted requests, guaranteeing the efficient operation of the system.
[0096] Furthermore, in some embodiments, the type of the large language model service request is a completely unexecuted large language model service request. Based on the type of the large language model service request and the preset migration strategy, the tokens and / or key-value caches in each source instance to be scheduled are migrated to the target instance corresponding to each source instance to be scheduled, including: based on the preset migration strategy, migrating the tokens in each source instance to be scheduled to the target instance corresponding to each source instance to be scheduled.
[0097] Specifically, for completely unexecuted large language model service requests, since the prefill and decoding phases have not started yet, there is no key-value cache at this time. At this time, the input tokens of these requests are extracted from the source instance, and the tokens of the requests are directly migrated to the target instance.
[0098] Through the above technical means, by focusing on meeting the video memory requirements, it is ensured that the target instance can seamlessly take over the requests in execution, guaranteeing the continuity of the task, and at the same time optimizing the resource allocation.
[0099] Furthermore, in some embodiments, the type of the large language model service request is a large language model service request waiting for video memory during execution. Based on the type of the large language model service request and the preset migration strategy, the tokens and / or key-value caches in each source instance to be scheduled are migrated to the target instance corresponding to each source instance to be scheduled, including: based on the preset migration strategy, migrating the key-value caches in each source instance to be scheduled to the target instance corresponding to each source instance to be scheduled.
[0100] Specifically, for requests waiting for video memory during execution, since there is already a key-value cache, the existing key-value cache is directly migrated. If this request waits for video memory during migration and starts to execute, then it is migrated according to the migration strategy for large language model service requests that are being executed.
[0101] Further, in some embodiments, the type of the large language model service request is an ongoing large language model service request. Based on the type of the large language model service request and a preset migration strategy, migrate the tokens and / or key-value caches in each source instance to be scheduled to the target instance corresponding to each source instance to be scheduled, including: based on the preset migration strategy, copy the existing key-value caches in each source instance to be scheduled to the target instance corresponding to each source instance to be scheduled, mark the end position of the current copy of each source instance to be scheduled, and append the newly generated key-value caches obtained by decoding calculation of each source instance to be scheduled to the existing key-value caches; repeatedly execute the above steps of copying the existing key-value caches in each source instance to be scheduled to the target instance corresponding to each source instance to be scheduled, marking the end position of the current copy of each source instance to be scheduled, and appending the newly generated key-value caches obtained by decoding calculation of each source instance to be scheduled to the existing key-value caches until there is only one token's key-value cache left in each source instance to be scheduled; copy the remaining one token's key-value cache to the target instance corresponding to each source instance to be scheduled, clear the key-value caches in each source instance to be scheduled, splice the key-value caches that have been migrated to the target instance corresponding to each source instance to be scheduled, and resume calculation after splicing is completed.
[0102] Specifically, for an ongoing request, a progressive migration strategy is adopted. Since the key-value cache of the large language model has an append feature, that is, after the pre-fill stage, the cache generated for each token in each iteration is appended to the existing key-value cache. Therefore, when migrating a request, a strategy that the calculation time for generating tokens iteratively can cover the existing key-value cache is used to execute.
[0103] That is to say, after the migration is started, start copying the existing key-value caches from the source instance to be scheduled to the target instance, and mark the end position W of the current copy. At the same time, the source instance to be scheduled continues to perform decoding calculation and append to the key-value cache. After the existing key-value caches are copied, then copy the newly calculated caches after the end position W of the previous copy to the target instance, and the source instance to be scheduled continues to perform decoding calculation.
[0104] Repeatedly execute the above processes of copying and calculation. Since the copying speed is much faster than the calculation speed, finally, there is only one token's key-value cache left in the source instance to be scheduled. Then copy the last key-value cache. The migration module stops copying after the last key-value cache is copied, and the source instance to be scheduled stops calculating, clears the key-value caches in the source instance to be scheduled, splices the copied key-value caches in the target instance, resumes calculation after splicing is completed, and notifies the scheduling module that the migration is successful. After the migration is successful, the scheduling module resets the routing of the user request to the target node.
[0105] As a possible implementation method, during the migration process, if the source instance to be scheduled has completed its calculation or an error occurs during the replication process, the migration is terminated, the key-value cache that has been migrated to the target instance is cleared, and the scheduling module is notified to reject the migration.
[0106] Through the above technical means, by gradually processing the key-value cache, it is ensured that the requests being executed can smoothly transition to the target instance, while minimizing the impact on the source instance, guaranteeing the continuity of tasks and the stability of the system.
[0107] Furthermore, as Figure 4 shown, although the migration duration of the entire request depends on the size of the key-value cache, thanks to the migration method of the present invention, the time loss during the migration process is very small, only the replication time of the key-value cache of one token, which can be almost ignored.
[0108] Thus, when migrating the key-value cache between the source instance and the target instance, the present invention adopts a progressive migration strategy for the requests being executed, so that the time loss during the migration can be ignored.
[0109] Furthermore, in some embodiments, the method for rescheduling requests in a large language model service further includes: determining whether there are multiple instances with a scheduling volume less than a third preset threshold within a preset duration; if there are multiple instances with a scheduling volume less than the third preset threshold within the preset duration, then based on a preset merging strategy, migrating and merging the multiple instances with a scheduling volume less than the third preset threshold within the preset duration, and deleting the idle instances.
[0110] Among them, the preset duration can be a duration preset by those skilled in the art, and the third preset threshold can be a threshold preset by those skilled in the art, which are not specifically limited herein.
[0111] Specifically, when the scheduling volume S value of a certain instance is lower than the third preset threshold L within a certain preset time D, that is <L, then a merging operation is triggered, the requests of multiple modules with too low loads are merged, and the idle instances are deleted, thereby releasing resources such as GPUs occupied by the instances and freeing up system resources.
[0112] In detail, when performing a merging operation on multiple instances with too low loads, the source instances to be scheduled are sequentially arranged in descending order of the scheduling volume S value, and the target instances are arranged in ascending order of the scheduling volume S value. When the scheduling volume S value of the target instance is greater than the third preset threshold H, the migration is performed.
[0113] Thus, through the above migration method, the number of migrations is reduced, and requests in instances with lower loads are completely migrated to instances with relatively higher loads, thereby releasing the resources of instances with lower loads.
[0114] To enable those skilled in the art to further understand the method for request rescheduling in the large language model service of the embodiments of the present invention, the following will be elaborated in detail with specific embodiments.
[0115] Specifically, the method for request rescheduling in the large language model service may include the following steps:
[0116] Step 1: Deploy a scheduling module and a large language model service in the cluster, and initialize each instance of the service. When initializing the instance, it is necessary to start the large language model inference model, as well as the monitoring module and migration module of the instance;
[0117] Step 2: The user starts using the large language model service, sends a request to the large language model service, and schedules the request to an instance using the traditional scheduling strategy;
[0118] Step 3: The monitoring module in the instance monitors each request entering the instance, including its queuing situation, the key-value cache during operation, and calculates the scheduling amount S value;
[0119] Step 4: The scheduling module obtains the scheduling amount S value from the monitoring module of each instance in real time, and sorts the scheduling amount S values of each instance. First, check whether the scheduling amount S value of each instance is less than the first preset threshold from smallest to largest. When the scheduling amount S value of an instance is less than the first preset threshold, it is determined as the instance to be scheduled (source instance). Secondly, start looking for the target instance from largest to smallest, and when the scheduling amount S value of an instance is greater than the second preset threshold, add it to the candidate target instance set. Finally, when the candidate target instance set has more than 1 instance, preferentially select the instance with the largest scheduling amount S value on the same node as the source instance as the target instance;
[0120] Step 5: The scheduling module sends a migration instruction to the migration modules of the source instance and the target instance;
[0121] Step 6: After receiving the request, the migration modules in the source instance and the target instance will start the migration process according to the principles of the highest priority and first-come, first-served;
[0122] Step 7: The migration module determines the state of the request. When the request has not been executed at all and is waiting for video memory during the execution process, directly migrate the token and key-value cache to the target instance; when the request is being executed, start gradually migrating the key-value cache, and finally resume the calculation on the target instance;
[0123] Step 8: After the request migration is successful, the scheduling module resets the route of the user request from the source instance to the target instance, so that the target instance returns the response result of the large language model request to the user.
[0124] According to the method for request rescheduling in large language model services proposed by the embodiments of the present invention, by calculating the scheduling amount of each instance based on the queuing situation of large language model service requests and the key-value cache, and determining each source instance to be scheduled and its corresponding target instance; when receiving a migration instruction, based on the type of large language model service request and a preset migration strategy, migrating the tokens and / or key-value cache in each source instance to be scheduled to the corresponding target instance, and feeding back the response result of the large language model service request to the user. Thus, by realizing efficient request allocation in a dynamic environment, the problem that traditional scheduling strategies cannot effectively schedule in the face of the uncertainty and dynamics of large language model requests is solved, thereby improving the resource utilization rate and response performance of the system.
[0125] Next, a device for request rescheduling in large language model services proposed by the embodiments of the present invention will be described with reference to the accompanying drawings.
[0126] Figure 5 It is a block diagram of a device for request rescheduling in large language model services according to an embodiment of the present invention.
[0127] As Figure 5 shown, the device 10 for request rescheduling in large language model services includes: an acquisition module 100, a determination module 200, and a migration module 300.
[0128] Among them, the acquisition module 100 is used to acquire the queuing situation of large language model service requests and the key-value cache in each instance, and calculate the scheduling amount of each instance according to the queuing situation of large language model service requests and the key-value cache in each instance; the determination module 200 is used to determine at least one source instance to be scheduled and the corresponding target instance of each source instance to be scheduled according to the scheduling amount of each instance; the migration module 300 is used to, when receiving a migration instruction, based on the type of large language model service request and a preset migration strategy, migrate the tokens and / or key-value cache in each source instance to be scheduled to the corresponding target instance of each source instance to be scheduled, and feed back the response result of the large language model service request to the user.
[0129] Further, in some embodiments, the obtaining module 100 is specifically configured to: obtain the key-value cache occupied by the running requests, the virtual key-value cache expected to be occupied by the queued requests, the total video memory capacity of the current instance, the video memory capacity occupied by the inherent number of parameters of the current large language model inference model, and the batch size of the current large language model inference model; calculate the scheduling amount of each instance according to the key-value cache occupied by the running requests, the virtual key-value cache expected to be occupied by the queued requests, the total video memory capacity of the current instance, the video memory capacity occupied by the inherent number of parameters of the current large language model inference model, and the batch size of the current large language model inference model.
[0130] Further, in some embodiments, the obtaining module 100 is specifically configured to: obtain the load value of the current instance according to the key-value cache occupied by the running requests, the virtual key-value cache expected to be occupied by the queued requests, and the video memory capacity occupied by the inherent number of parameters of the current large language model inference model; obtain the remaining space of the current instance according to the difference between the total video memory capacity of the current instance and the load value of the current instance; obtain the scheduling amount of the current instance according to the ratio between the remaining space of the current instance and the batch size of the current large language model inference model.
[0131] Further, in some embodiments, the load value of the current instance is obtained from the sum of the key-value cache occupied by the running requests, the virtual key-value cache expected to be occupied by the queued requests, and the video memory capacity occupied by the inherent number of parameters of the current large language model inference model.
[0132] Further, in some embodiments, the determining module 200 is configured to: determine a set of source instances to be scheduled and a set of candidate target instances based on the scheduling amount of each instance; determine whether the set of candidate target instances is an empty set; if the set of candidate target instances is not an empty set, then based on the set of candidate target instances, filter out the instance that is on the same node as each source instance to be scheduled in the set of source instances to be scheduled and has the largest scheduling amount, and obtain the target instance corresponding to each source instance to be scheduled.
[0133] Further, in some embodiments, the determining module 200 is configured to: filter out the instance combinations with a scheduling amount less than a first preset threshold based on the scheduling amount of each instance to obtain a set of source instances to be scheduled; based on the scheduling amount of each instance, determine whether there is an instance with a scheduling amount greater than a second preset threshold; if there is an instance with a scheduling amount greater than the second preset threshold, then combine the instances with a scheduling amount greater than the second preset threshold to obtain a set of candidate target instances.
[0134] Further, in some embodiments, after determining whether the candidate target instance set is an empty set, the determining module 200 is further configured to: if the candidate target instance set is an empty set, create new instances for each source instance to be scheduled; and use the new instances corresponding to each source instance to be scheduled as the target instances corresponding to each source instance to be scheduled.
[0135] Further, in some embodiments, the type of the large language model service request is a large language model service request that has not been executed at all. The migration module 300 is specifically configured to: based on a preset migration policy, migrate the tokens in each source instance to be scheduled to the target instance corresponding to each source instance to be scheduled.
[0136] Further, in some embodiments, the type of the large language model service request is a large language model service request waiting for video memory during execution. The migration module 300 is specifically configured to: based on a preset migration policy, migrate the key-value cache in each source instance to be scheduled to the target instance corresponding to each source instance to be scheduled.
[0137] Further, in some embodiments, the type of the large language model service request is a large language model service request that is being executed. The migration module 300 is specifically configured to: based on a preset migration policy, copy the existing key-value cache in each source instance to be scheduled to the target instance corresponding to each source instance to be scheduled, mark the end position of the current copy of each source instance to be scheduled, and append the newly generated key-value cache obtained by decoding calculation of each source instance to be scheduled to the existing key-value cache; repeat the above steps of copying the existing key-value cache in each source instance to be scheduled to the target instance corresponding to each source instance to be scheduled, marking the end position of the current copy of each source instance to be scheduled, and appending the newly generated key-value cache obtained by decoding calculation of each source instance to be scheduled to the existing key-value cache until there is a key-value cache with one remaining token in each source instance to be scheduled; copy the remaining key-value cache with one token to the target instance corresponding to each source instance to be scheduled, clear the key-value cache in each source instance to be scheduled, splice the key-value cache that has been migrated to the target instance corresponding to each source instance to be scheduled, and resume calculation after splicing is completed.
[0138] Further, in some embodiments, the apparatus 10 for request rescheduling in the large language model service is further configured to: determine whether there are multiple instances with a scheduling volume less than a third preset threshold within a preset duration; if there are multiple instances with a scheduling volume less than a third preset threshold within a preset duration, perform migration and merging on the multiple instances with a scheduling volume less than a third preset threshold within a preset duration based on a preset merging policy, and delete the idle instances.
[0139] Further, in some embodiments, before obtaining the queuing situation and key-value cache of the large language model service requests for each instance, the obtaining module 100 is further configured to: determine whether a large language model service request is received; when a large language model service request is received, schedule the large language model service request to multiple instances.
[0140] It should be noted that the foregoing explanation of the method embodiments for request rescheduling in large language model services also applies to the device for request rescheduling in large language model services in this embodiment, and will not be elaborated herein.
[0141] The device for request rescheduling in large language model services according to the embodiments of the present invention calculates the scheduling amount of each instance according to the queuing situation and key-value cache of the large language model service requests, and determines each source instance to be scheduled and its corresponding target instance; when a migration instruction is received, based on the type of the large language model service request and a preset migration policy, migrates the tokens and / or key-value cache in each source instance to be scheduled to the corresponding target instance, and feeds back the response result of the large language model service request to the user. Thus, by implementing efficient request allocation in a dynamic environment, the problem that traditional scheduling strategies cannot effectively schedule in the face of the uncertainty and dynamics of large language model requests is solved, thereby improving the resource utilization rate and response performance of the system.
[0142] Figure 6 The following is a schematic structural diagram of an electronic device provided by an embodiment of the present invention. The electronic device may include:
[0143] A memory 601, a processor 602, and a computer program stored on the memory 601 and executable on the processor 602.
[0144] When the processor 602 executes the program, it implements the method for request rescheduling in large language model services provided in the foregoing embodiments.
[0145] Further, the electronic device further includes:
[0146] A communication interface 603 for communication between the memory 601 and the processor 602.
[0147] The memory 601 is used to store a computer program executable on the processor 602.
[0148] The memory 601 may include a high-speed RAM (Random Access Memory) memory, and may also include a non-volatile memory, such as at least one disk memory.
[0149] If the memory 601, the processor 602, and the communication interface 603 are implemented independently, the communication interface 603, the memory 601, and the processor 602 can be interconnected via a bus and communicate with each other. The bus can be an ISA (Industry Standard Architecture) bus, a PCI (Peripheral Component Interconnect) bus, an EISA (Extended Industry Standard Architecture) bus, or the like. The bus can be divided into an address bus, a data bus, a control bus, etc. For ease of representation, Figure 6 only a thick line is used in Figure 6 , but this does not mean that there is only one bus or one type of bus.
[0150] Optionally, in a specific implementation, if the memory 601, the processor 602, and the communication interface 603 are integrated on a single chip, the memory 601, the processor 602, and the communication interface 603 can communicate with each other through an internal interface.
[0151] The processor 602 may be a CPU (Central Processing Unit), an ASIC (Application Specific Integrated Circuit), or one or more integrated circuits configured to implement the embodiments of the present invention.
[0152] The embodiments of the present invention also provide a computer-readable storage medium, on which a computer program is stored, and when the program is executed by a processor, the method for request rescheduling in the large language model service as described above is implemented.
[0153] Those skilled in the art can further realize that the units and algorithm steps of each example described in combination with the embodiments disclosed herein can be implemented by electronic hardware, computer software, or a combination of the two. To clearly illustrate the interchangeability of hardware and software, the composition and steps of each example have been generally described according to functions in the above description. Whether these functions are executed in a hardware or software manner depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of the present invention.
[0154] The above has introduced in detail a method for request rescheduling in a large language model service provided by the present invention. Specific examples are used in this article to elaborate on the principle and implementation manner of the present invention. The description of the above embodiments is only used to help understand the method and its core idea of the present invention. It should be noted that for those of ordinary skill in the art, without departing from the principle of the present invention, several improvements and modifications can still be made to the present invention, and these improvements and modifications also fall within the protection scope of the claims of the present invention.
Claims
1. A method for request rescheduling in large language model services, characterized in that, The steps include: Obtain the queuing situation and key-value cache of the large language model service requests in each instance, and calculate the scheduling volume of each instance according to the queuing situation and key-value cache of the large language model service requests in each instance; Determine at least one source instance to be scheduled and the target instance corresponding to each source instance to be scheduled according to the scheduling volume of each instance; When receiving a migration instruction, based on the type of the large language model service request and a preset migration policy, migrate the tokens and / or key-value cache in each source instance to be scheduled to the target instance corresponding to each source instance to be scheduled, and feedback the response result of the large language model service request to the user; Among them, calculating the scheduling volume of each instance according to the queuing situation and key-value cache of the large language model service requests in each instance includes: obtaining the key-value cache occupied by the running requests, the virtual key-value cache expected to be occupied by the queuing requests, the total video memory capacity of the current instance, the video memory capacity occupied by the inherent number of parameters of the current large language model inference model, and the batch size of the current large language model inference model; calculating the scheduling volume of each instance according to the key-value cache occupied by the running requests, the virtual key-value cache expected to be occupied by the queuing requests, the total video memory capacity of the current instance, the video memory capacity occupied by the inherent number of parameters of the current large language model inference model, and the batch size of the current large language model inference model.
2. The method for request rescheduling in the large language model service according to claim 1, wherein Calculating the scheduling volume of each instance according to the key-value cache occupied by the running requests, the virtual key-value cache expected to be occupied by the queuing requests, the total video memory capacity of the current instance, the video memory capacity occupied by the inherent number of parameters of the current large language model inference model, and the batch size of the current large language model inference model includes: Obtain the load value of the current instance according to the key-value cache occupied by the running requests, the virtual key-value cache expected to be occupied by the queuing requests, and the video memory capacity occupied by the inherent number of parameters of the current large language model inference model; Obtain the remaining space of the current instance according to the difference between the total video memory capacity of the current instance and the load value of the current instance; Obtain the scheduling volume of the current instance according to the ratio between the remaining space of the current instance and the batch size of the current large language model inference model.
3. The method for request rescheduling in the large language model service according to claim 2, wherein The load value of the current instance is obtained by the sum of the key-value cache occupied by the running requests, the virtual key-value cache expected to be occupied by the queuing requests, and the video memory capacity occupied by the inherent number of parameters of the current large language model inference model.
4. The method for request rescheduling in the large language model service according to claim 1, characterized in that, Determining at least one source instance to be scheduled and the target instance corresponding to each source instance to be scheduled according to the scheduling volume of each instance includes: Based on the scheduling volume of each instance, determine a set of source instances to be scheduled and a set of candidate target instances; Judge whether the set of candidate target instances is an empty set; If the candidate target instance set is not an empty set, based on the candidate target instance set, filter out the instance that is on the same node as each source instance to be scheduled in the source instance set to be scheduled and has the largest scheduling volume, and obtain the target instance corresponding to each source instance to be scheduled.
5. The method for request rescheduling in the large language model service according to claim 4, wherein, The determining the source instance set to be scheduled and the candidate target instance set based on the scheduling volume of each instance includes: Based on the scheduling volume of each instance, filter out the instance combinations with a scheduling volume less than the first preset threshold to obtain the source instance set to be scheduled; Based on the scheduling volume of each instance, determine whether there is an instance with a scheduling volume greater than the second preset threshold; If there is an instance with a scheduling volume greater than the second preset threshold, combine the instances with a scheduling volume greater than the second preset threshold to obtain the candidate target instance set.
6. The method for request rescheduling in the large language model service according to claim 4, characterized in that, After determining whether the candidate target instance set is an empty set, it further includes: If the candidate target instance set is an empty set, create a new instance for each source instance to be scheduled; Use the new instance corresponding to each source instance to be scheduled as the target instance corresponding to each source instance to be scheduled.
7. The method for request rescheduling in the large language model service according to claim 1, wherein The type of the large language model service request is a completely unexecuted large language model service request. The migrating the tokens and / or key-value caches in each source instance to be scheduled to the target instance corresponding to each source instance to be scheduled based on the type of the large language model service request and the preset migration strategy includes: Based on the preset migration strategy, migrate the tokens in each source instance to be scheduled to the target instance corresponding to each source instance to be scheduled.
8. The method for request rescheduling in the large language model service according to claim 1, wherein, The type of the large language model service request is a large language model service request waiting for video memory during execution. The migrating the tokens and / or key-value caches in each source instance to be scheduled to the target instance corresponding to each source instance to be scheduled based on the type of the large language model service request and the preset migration strategy includes: Based on the preset migration strategy, migrate the key-value caches in each source instance to be scheduled to the target instance corresponding to each source instance to be scheduled.
9. The method for request rescheduling in the large language model service according to claim 1, wherein The type of the large language model service request is an executing large language model service request. The migrating the tokens and / or key-value caches in each source instance to be scheduled to the target instance corresponding to each source instance to be scheduled based on the type of the large language model service request and the preset migration strategy includes: Based on the preset migration strategy, copy the existing key-value caches in each source instance to be scheduled to the target instance corresponding to each source instance to be scheduled, mark the end position of the current copy of each source instance to be scheduled, and append the newly generated key-value caches obtained by decoding and calculating each source instance to be scheduled to the existing key-value caches; Repeat the above steps of copying the existing key - value cache of each source instance to be scheduled to the corresponding target instance of each source instance to be scheduled, marking the end position of the current copy of each source instance to be scheduled, and appending the newly generated key - value cache obtained by decoding each source instance to be scheduled to the existing key - value cache until there is a key - value cache with one remaining token in each source instance to be scheduled; Copy the key - value cache with the remaining one token to the corresponding target instance of each source instance to be scheduled, clear the key - value cache in each source instance to be scheduled, splice the key - value caches that have been migrated to the corresponding target instances of each source instance to be scheduled, and resume calculation after splicing is completed.
10. The method for request rescheduling in the large language model service according to claim 1, wherein, It also includes: Judge whether there are multiple instances with a scheduling volume less than a third preset threshold within a preset time period; If there are such multiple instances with a scheduling volume less than a third preset threshold within a preset time period, based on a preset merging strategy, migrate and merge the multiple instances with a scheduling volume less than a third preset threshold within a preset time period, and delete the idle instances.
11. The method for request rescheduling in the large language model service according to claim 1, wherein Before obtaining the queuing situation of the large - language model service requests and the key - value caches in each instance, it also includes: Judge whether a large - language model service request is received; When a large - language model service request is received, schedule the large - language model service request to multiple instances.
12. An apparatus for request rescheduling in large language model services, characterized in that, It includes: An acquisition module, configured to acquire the queuing situation of the large - language model service requests and the key - value caches in each instance, and calculate the scheduling volume of each instance according to the queuing situation of the large - language model service requests and the key - value caches in each instance; A determination module, configured to determine at least one source instance to be scheduled and the corresponding target instance of each source instance to be scheduled according to the scheduling volume of each instance; A migration module, configured to, when receiving a migration instruction, based on the type of the large - language model service request and a preset migration strategy, migrate the tokens and / or key - value caches in each source instance to be scheduled to the corresponding target instance of each source instance to be scheduled, and feedback the response result of the large - language model service request to the user; Among them, the acquisition module is configured to: acquire the key - value cache occupied by the running requests, the virtual key - value cache expected to be occupied by the queuing requests, the total video memory capacity of the current instance, the video memory capacity occupied by the inherent parameter quantity of the current large - language model inference model, and the batch size of the current large - language model inference model; calculate the scheduling volume of each instance according to the key - value cache occupied by the running requests, the virtual key - value cache expected to be occupied by the queuing requests, the total video memory capacity of the current instance, the video memory capacity occupied by the inherent parameter quantity of the current large - language model inference model, and the batch size of the current large - language model inference model.
13. An electronic device, characterized in that, It includes: A memory, a processor, and a computer program stored on the memory and executable on the processor, wherein the processor executes the computer program to implement the method for request rescheduling in the large language model service according to any one of claims 1-11.
14. A computer-readable storage medium having a computer program stored thereon, characterized in that, The computer program is executed by a processor to implement the method for request rescheduling in the large language model service according to any one of claims 1-11.
Citation Information
Patent Citations
Data migration method and device, electronic equipment and storage medium
CN118394734A
Intelligent task planning method based on large model
CN119443695A