Request priority scheduling system and method for large model
By configuring the remaining waiting patience value and priority sorting scheduling for large model requests, the problem of satisfying user response speed without increasing computing power costs or adjusting model deployment is solved, realizing instant optimization of requests and improving user experience.
Patent Information
- Application Number
- CN202510284652.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-11
- Publication Date
- 2025-07-08
AI Technical Summary
The prior art is difficult to meet the request response speed requirements of different users without increasing the platform's computing power costs or large-scale adjustment of model deployment solutions, and the existing adjustment methods are complex and cannot be considered from the perspective of overall response speed.
By configuring the remaining waiting patience value for each request, determining whether the current request is forward-lined, using the priority sorting scheduling component to compare the patient consumption value and the remaining waiting patience value, setting the short request comparison termination condition and the long request becomes a short request after processing, realizing the immediate priority sorting of the request.
Improves the user experience, reduces the waiting or response time of all requests, ensures that the request response is within the user's patience, and optimizes the overall efficiency.
Smart Images

Figure CN120276837A_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the field of computer information processing, and more particularly, to a request priority scheduling system and method for large models. Background Art
[0002] Today, with the increasing popularity of large language models in the development of artificial intelligence, they are increasingly penetrating into people's lives. People are also becoming increasingly concerned about the request response requirements of large prediction model platforms. The speed of request response generally depends on the powerful computing power of the platform and the optimization of the model deployment architecture.
[0003] However, how to meet the request response speeds of different users without overly increasing the platform computing power cost or making large-scale adjustments to the model deployment plan when a large number of requests arrive has become a problem faced by the platform side or the model deployment side.
[0004] The inventors of the present disclosure have previously adjusted TPOT and TTFT from aspects of pre-fill and decoding step prediction. However, this adjustment method is very complex, and this adjustment method does not consider the overall response speed.
[0005] Therefore, people hope to provide a simple way to adjust the response time of different requests, so as to both generally reduce the waiting or response time of all requests and, without affecting the expected response waiting time of all users, achieve request response as fast as possible.
[0006] The above information disclosed in the background art section is only used to enhance the understanding of the background of the present disclosure, and thus it may include information that does not constitute the prior art known to those of ordinary skill in the art. Summary of the Invention
[0007] In view of this, the present disclosure provides an efficient large language model inference system for large models, especially large-scale hardware clusters, which improves the user experience by configuring a remaining waiting patience value for each request and determining whether to perform forward queue-jumping processing on the current request.
[0008] Other features and advantages of the present disclosure will become apparent from the following detailed description, or will be partially learned through the practice of the present disclosure.
[0009] According to one aspect of the present disclosure, a request priority scheduling system for large models is provided, including: a request information acquisition component that receives data processing requests in the order of request input time, obtains the number of tokens for each request through a tokenizer, and configures a remaining waiting patience value for each request; a priority sorting and scheduling component that, for each request, obtains the patience consumption cost value generated by the request to be processed preferentially, compares the patience consumption cost value of the current request with the remaining waiting patience value of the previous request arranged in chronological order in front of it, and when the patience consumption cost value of the current request is less than the remaining waiting patience value of the previous request, schedules the current request's arrangement priority to be advanced before the previous input request, and when the patience consumption cost value of the current request is greater than the remaining waiting patience value of the previous request, keeps the arrangement priority of the current request unchanged; and an information update component that, for each change in sorting by the priority sorting component, uses the result of subtracting the patience consumption cost value of the current request from the remaining waiting patience value of the previous request as the remaining waiting patience value of the previous request after the sorting change.
[0010] In the request priority scheduling system for large models according to the present disclosure, the remaining waiting patience value is determined according to the number of all tokens to be processed of the current request or according to the estimated execution time length of all tokens to be processed of the current request.
[0011] In the request priority scheduling system for large models according to the present disclosure, the priority sorting and scheduling component also, for any request, obtains the number of all tokens to be processed of it, and when the number of tokens to be processed of the previous request is less than that of the current request, does not perform priority sorting and scheduling processing.
[0012] In the request priority scheduling system for large models according to the present disclosure, the priority sorting and scheduling component does not perform priority sorting and scheduling processing when the number of tokens of the previous request is less than a predetermined number or the estimated execution time length is shorter than a predetermined time length.
[0013] In the request priority scheduling system for large models according to the present disclosure, the predetermined number is 4096 or 8192.
[0014] In the request priority scheduling system for large models according to the present disclosure, the remaining waiting patience value r is where f is the number of all tokens to be processed of the previous request, c is the patience value already consumed by any previous request, and the patience consumption cost value generated by the current request for the previous request assuming the arrangement priority is scheduled is e , wherein is the patience cost coefficient for each pre-filled token, is the patience cost coefficient for each pre-filled token, f is the remaining number of tokens that the current request needs to run, d is the number of tokens expected to be decoded for the current request.
[0015] According to the request priority scheduling system for large models of the present disclosure, wherein the patience cost coefficient of the pre-filled tokens is greater than or equal to 2, and the patience cost coefficient of the pre-filled tokens is the ratio of the average step efficiency in the decoding stage to the average step efficiency in the pre-filling stage of the platform where the large model is deployed.
[0016] According to another aspect of the present disclosure, there is also provided a request priority scheduling method for large models, including: a request information acquisition step of receiving data processing requests in the order of request input time, obtaining the number of tokens for each request through a tokenizer, and configuring a remaining waiting patience value for each request; a priority sorting and scheduling step of, for any request, obtaining the patience consumption cost value generated by its being preferentially processed, comparing the patience consumption cost value of the current request with the remaining waiting patience value of the previous request arranged in chronological order in front of it, and when the patience consumption cost value of the current request is less than the remaining waiting patience value of the previous request, moving the arranged priority order of the current request forward to before the previous input request, and when the patience consumption cost value of the current request is greater than the remaining waiting patience value of the previous request, keeping the arranged priority order of the current request unchanged; an information update step of, for each change in sorting by the priority sorting component, using the result of subtracting the patience consumption cost value of the current request from the remaining waiting patience value of the previous request as the remaining waiting patience value of the previous request after the sorting change; and repeating the priority sorting and scheduling step and the information update step for the current request that has completed the arranged priority order scheduling until the patience consumption cost value of the current request is greater than the remaining waiting patience value of the previous request.
[0017] According to the request priority scheduling method for large models of the present disclosure, wherein the priority sorting and scheduling step further includes: for any request, obtaining the number of all tokens to be processed, and when the number of tokens to be processed of the previous request is less than that of the current request, not performing priority sorting and scheduling processing.
[0018] The request priority scheduling method for large models according to the present disclosure, wherein the priority sorting and scheduling step includes: when the number of tokens of the previous request is less than a predetermined number or the estimated execution time length is shorter than a predetermined time length, no priority sorting and scheduling processing is performed.
[0019] According to the request priority scheduling system and method for large models of the present disclosure, by configuring a remaining waiting patience value for each request and comparing the patience consumption cost value of the current request with the remaining waiting patience value of the previous request arranged in front of it in chronological order, it can be determined whether to perform forward queue-jumping processing on the current request. On this basis, a short request comparison termination condition and a situation where a long request becomes substantially a short request after being processed are set, so that the priority sorting of the requests shows its immediate real-time state, thereby making the request response generally and immediately within the waiting patience value of the user, greatly improving the user experience.
[0020] It should be understood that the above general description and the following detailed description are only exemplary and do not limit the present disclosure. Brief Description of the Drawings
[0021] By referring to the accompanying drawings and describing its exemplary embodiments in detail, the above and other objects, features, and advantages of the present disclosure will become more apparent. The following described drawings are only some embodiments of the present disclosure. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.
[0022] Figure 1 is a block diagram of a first embodiment of a request priority scheduling system for large models shown according to an exemplary embodiment.
[0023] Figure 2 Shown is an example diagram of a request priority scheduling process executed by a request priority scheduling system for large models shown according to an exemplary embodiment.
[0024] Figure 3 is a flowchart of a request priority scheduling method for large models shown according to an exemplary embodiment. Detailed Description of Specific Embodiments
[0025] Example embodiments will now be described more fully with reference to the accompanying drawings. However, the example embodiments can be implemented in various forms and should not be construed as limited to the embodiments set forth herein; rather, these embodiments are provided so that this disclosure will be thorough and complete, and will fully convey the concept of the example embodiments to those skilled in the art. Identical reference numerals in the figures denote identical or similar parts, and thus their repeated description will be omitted.
[0026] In addition, the described features, structures, or characteristics may be combined in one or more embodiments in any suitable manner. In the following description, numerous specific details are provided to give a thorough understanding of the embodiments of the present disclosure. However, those skilled in the art will realize that the technical solutions of the present disclosure may be practiced without one or more of the specific details, or other methods, components, devices, steps, etc. may be employed. In other cases, well-known methods, devices, implementations, or operations are not shown or described in detail to avoid obscuring aspects of the present disclosure.
[0027] The block diagrams shown in the drawings are only functional entities and do not necessarily correspond to physically independent entities. That is, these functional entities may be implemented in software form, or implemented in one or more hardware modules or integrated circuits, or implemented in different networks and / or processor devices and / or microcontroller devices.
[0028] The flowcharts shown in the drawings are only illustrative and do not necessarily include all the contents and operations / steps, nor are they necessarily executed in the order described. For example, some operations / steps can be decomposed, while some operations / steps can be combined or partially combined, so the actual execution order may change according to the actual situation.
[0029] It should be understood that although terms such as first, second, and third may be used herein to describe various components, these components should not be limited by these terms. These terms are used to distinguish one component from another. Thus, the first computing device discussed below may be referred to as the second computing device without departing from the teachings of the concepts of the present disclosure. As used herein, the term "and / or" includes any one and all combinations of one or more of the associated listed items.
[0030] Those skilled in the art can understand that the drawings are only schematic diagrams of exemplary embodiments, and the modules or processes in the drawings are not necessarily essential for implementing the present disclosure, so they cannot be used to limit the protection scope of the present disclosure.
[0031] Figure 1 is a block diagram of a first embodiment of a request priority scheduling system for a large model shown according to an exemplary embodiment. As Figure 1 shown, the request priority scheduling system 100 for a large model includes: a request information acquisition component 110, a priority sorting and scheduling component 120, and an information update component 130. As Figure 1As shown, the request information acquisition component 110 receives data processing requests in the order of request input time, obtains the token count of each request through a tokenizer, and configures the remaining waiting patience value for each request. Today, as artificial intelligence is increasingly applied to people's lives, platforms handle a lot of long texts. However, in the scenario of processing long texts, some users may also input short texts. As the scheduler of a large model platform, it generally follows a strict queue order, that is, first in, first out. However, this strict sorting method sometimes results in unreasonable and overall low efficiency. This is the same as the situation in life. For example, when queuing up to check out in a supermarket, in front of the customer with 2 items in hand queuing up to check out, there is a customer in front pushing a cart full of goods. It takes the cashier 10 minutes to finish checking out the customer in front, while it only takes the cashier 20 seconds to finish checking out the customer with 2 items. Obviously, the customer with the cart has a longer psychological expectation for the time spent on checking out, while the customer with 2 items has a shorter psychological expectation for the time spent on checking out. If the customer with 2 items is checked out first, the overall waiting time of the two will be reduced by nearly 10 minutes.
[0032] Therefore, in the field of large language models, there is the same situation. For example, a server receives two requests of different lengths one after another. The first request has 100,000 tokens, and the second request has 100 tokens. Without considering decoding, assuming it takes 30s to execute the first request and 20ms (1 millisecond is 0.001 seconds) to execute the second request. If executed in the order of arrival, the first request needs to wait for 32s, and the second request needs to wait for 32.02s. If the short request is allowed to jump the queue and is executed before the long request. Then the second request needs to wait for 0.02s, and the first request needs to wait for 32.02s. Obviously, the second method of giving priority to short requests is better than first come, first served.
[0033] Another example of a scenario is that if the first request for 100,000 tokens has been executed halfway, with 50,000 tokens remaining and it still takes 24 seconds to complete. It should be noted that usually, since the execution time grows linearly, the first half of the tokens only takes 1 / 4 of the total time, and the second half takes 3 / 4 of the total time. Referring to the sum of the first three numbers 9 in the arithmetic sequence 1, 3, 5, 7, 9, 11, which is 1 / 4 of the total 36. At this time, a second request for 100 tokens comes in. It takes 0.02 seconds to complete the second request. If it is not allowed to prioritize and cut in the requests that are being executed, then the first request needs to wait for 32 seconds, and the second request needs to wait for 24.02 seconds. If it is allowed to cut in the requests that are being executed, then the first request needs to wait for 32.02 seconds, and the second request needs to wait for 0.02 seconds. Obviously, allowing the second request to cut in the requests that are being executed has a better effect, the overall waiting time will decrease significantly, and the overall effect is better.
[0034] Similarly, for a long request for which some requests have been executed, if the first request for 100,000 tokens has been executed for 90,000 tokens, with 10,000 tokens remaining and it still takes 6 seconds, and at this time a request for 50,000 tokens comes in and it takes 9 seconds to execute. Then obviously the effect of not allowing cutting in is better. This means that it is not determined whether a request is allowed to be cut in according to the total length of a request, but rather according to the remaining uncalculated tokens of the request.
[0035] Therefore, as described above, the present disclosure configures a remaining waiting patience value for each request, so as to be able to determine whether a request can be cut in by a request that is input later in time sequence based on the remaining waiting patience value.
[0036] For example, the first request (the previous request) has 100,000 tokens, and then 1000 short requests of 10,000 tokens each come in (the current requests). If each short request takes 0.6 s to execute. So every time a short request cuts in front of the first request, the short request that cuts in will wait 32 s less, while the first request will wait 0.6 s more, and the total waiting time will be reduced by 31.4 s. Then the natural way to minimize the total waiting time is for the first request to be cut in front of by all the short requests. Each short request will wait 32 s less, but the cost is that the first request will have to wait 600 s more. However, this can easily cause a problem. If short requests keep coming in continuously and cutting in line one after another, the long request will be delayed indefinitely. Moreover, a long request that is expected to return a result within half a minute still has no result after 8 minutes, which is a length of patience that users cannot accept. Just like vehicles merging from 2 lanes into 1 lane. According to traffic rules, vehicles enter the single lane at intervals from each lane. Just like a zipper, the teeth on both sides always engage at intervals. Regarding the request being cut in front of as one lane and the cutting-in requests as another lane, the waiting time of the request being cut in front of after it starts execution should preferably not exceed 2 times its own execution time. However, interval execution is not necessary. Although it is possible to first execute 10,000 tokens of the long request, then execute a 10,000-token short request, and then repeat this step. In this way, the long request will be cut in front of by 9 short requests. The waiting times (s) for this kind of cutting-in are as shown in Table 1 below: 1.2 3 5.4 8.4 12 16.2 21 26.4 32.4 0.6 0.6 1.2 0.6 1.8 0.6 2.4 0.6 3 0.6 3.6 0.6 4.2 0.6 4.8 0.6 5.4 0.6 6 From Table 1, the waiting time of the long request is 38.4 s, while the waiting times of the 9 short requests are 1.2 s, 3 s, 5.4 s, 8.4 s, 12 s, 16.2 s, 21 s, 26.4 s, 32.4 s respectively.
[0037] If, instead of interval execution, the first request allows 9 short requests to cut in line in advance, then the 9 short requests can be executed first. The waiting times (s) for this kind of cutting-in are as shown in Table 2 below: 0.6 1.2 1.8 2.4 3 3.6 4.2 4.8 5.4 0.6 0.6 0.6 0.6 0.6 0.6 0.6 0.6 0.6 0.6 33 In this way, the waiting time of the long request is still 38.4 s, while the waiting times of the 9 short requests are reduced to 0.6 s, 1.2 s, 1.8 s, 2.4 s, 3 s, 3.6 s, 4.2 s, 4.8 s, 5.4 s.
[0038] Therefore, as described above, set a patience value to estimate in advance the number of tokens allowed to be cut in front of. As long as the remaining patience value is sufficient, allow cutting in line.
[0039] For the patience value of the present disclosure, in order to simplify the program, it can be determined not by the execution time, but by the number of tokens. Optionally, it can also be determined by the execution time. Under the simplified program, even if the step predictor proposed by the present inventor before is not enabled, the function of short request queue jumping can still be used. Another advantage of being determined by the number of tokens is that the execution time of a long text is proportional to the square of its number of tokens. Because under the premise of a fixed batch token number, the execution time of each processing step is proportional to the centroid of the request token sequence, so the integral, or summation, of the execution time of the entire request is proportional to the square of the length of the pre-fill. Just like the sum of 1, 3, 5, 7, 9, 11 is 4 times the sum of 1, 3, 5; the sum of 1, 3, 5, 7, 9, 11, 13, 15, 17 is 9 times the sum of 1, 3, 5. Therefore, compared with determining the patience value according to the estimated execution time, the patience value determined according to the number of tokens can more restrict the queued time so that it will not be too long.
[0040] As can also be seen from the example in Table 2, when a long request with a length of 100,000 is queued by 9 short requests with a length of 10,000, the time only increases by 16.4%. If the patience value is calculated according to the estimated execution time, there will be 54 short requests queuing up, and the time can increase by 98.2%. Of course, using the latter can also achieve the purpose of priority scheduling of the present invention, but the effect will not be as good as that based on the number of tokens. It should be noted that setting the patience value according to the estimated execution time allows more short requests to queue up. The advantage is the decrease in the total waiting time. The disadvantage is that the waiting time of an overly long request exceeds the preset value of the platform (such as 10 minutes), resulting in the server directly giving a Time Out Error conclusion, and finally the long request no longer has a return value, so that the work of the pre-fill that has been executed is wasted, and patient users may do it again, which also leads to a waste of computing power.
[0041] In the case of determining the patience value by the number of tokens, considering that a long request can be queued during its execution, the queued request, that is, the first request in the above example, also becomes the previous request, and its remaining waiting patience value r is where f is the number of all tokens to be processed of the previous request, and c is the patience value consumed by any previous request. If the first request has not been queued yet, the consumed patience value c is 0, then its remaining waiting patience value is the initial patience value, that is, its initial number of tokens.
[0042] Other requests that are prepared to cut in line for the first request, that is, the current request. In the case where the permutation priority order is assumed to be scheduled (assuming it is prepared to cut in line), the cut-in result has a patience consumption cost value for the previous request as e , where is the patience cost coefficient for each pre-filled token, is the patience cost coefficient for each pre-filled token, f is the remaining number of tokens that the current request needs to run, and d is the number of tokens that the current request is expected to be decoded. Among them, according to the normal execution time relationship between pre-filling and decoding, can be set to 2, can be set to 8. It can also be adjusted according to the state of the specific deployment platform and the actual efficiency between different modules of the large model for processing data.
[0043] Although the difference between pre-filling and decoding is considered here for fine-grained use of the consumption cost value, for simplicity, the number of tokens of the current request can also be directly used as the patience consumption cost value brought to the request being cut in line after its cut-in e .
[0044] Next, as Figure 1 shown, the priority sorting and scheduling component 10 calculates the patience consumption cost value of the current request for the current request, and compares the patience consumption cost value of the current request with the remaining waiting patience value of the previous request arranged in chronological order in front of it. When the patience consumption cost value of the current request is less than the remaining waiting patience value of the previous request, the permutation priority order of the current request is scheduled to move forward to before the previous input request, and when the patience consumption cost value of the current request is greater than the remaining waiting patience value of the previous request, the permutation priority order of the current request remains unchanged. Obviously, the above scheduling method is based on the basic principle of setting the remaining waiting patience value to minimize the total waiting time for scheduling.
[0045] According to the priority sorting and scheduling component 10, when the current request enters the waiting queue or the ready queue, calculate the patience consumption cost value of the current request e , as long as e is less than the remaining waiting patience value of the previous request r , it means that the previous request still has patience margin for other requests to cut in line. Therefore, the waiting permutation order of these two requests can be exchanged, that is, the priority sorting and scheduling is executed. This means that the patience value already consumed by the previous request c will increase e, that is, the remaining waiting patience value of the previous request r reduce e The swapped current request can continue to be compared with its new previous request until the remaining waiting patience value of a new previous request reaches r Less than e , it means that the new preceding request has no patience margin for other requests to cut in line, so the waiting order of the two requests is stopped from being exchanged, that is, further priority sorting and scheduling are refused, that is, the queue cutting operation is stopped. Figure 2 FIG. 1 is a schematic diagram showing an example of a request priority scheduling process performed by a request priority scheduling system for a large model according to an exemplary embodiment. Figure 2 As shown, the number of tokens to be processed f When the current request (f2=128, d2=128, e2=1280) with (forward token num) of 128 enters the queue, it encounters the three requests before it, such as the previous request (f=6666, c=4096, r=2570), the previous request (f=44444, c=22222, r=22222), and the previous request (f=102400, c=20480, r=81920). All of them are sorted and scheduled based on the comparison, so they will slowly move forward from the end of the queue (from Figure 2 The reading direction is from right to left, or the display view direction is from top to bottom), until a new previous request (f1=5120, c1=4096, r1=1024) is encountered. After comparison, e2>r1 is obtained, which means that the previous request has no patience margin for other requests to interrupt the queue, so the waiting order of the two requests is stopped, that is, further priority sorting and scheduling are refused, that is, the queue interruption operation is stopped. Figure 2 It is shown that there is a previous request (f=8096, c=4096, r=4000) immediately before the new previous request (f1=5120, c1=4096, r1=1024). From a comparison perspective, the current request can be allowed to jump the queue. However, due to the existence of the new previous request (f1=5120, c1=4096, r1=1024), the priority sorting scheduling component 10 has terminated the queue-jumping request for the current request, and it is also impossible to compare with the previous request (f=8096, c=4096, r=4000) to perform priority sorting scheduling. Because if the current request can be queued due to the previous request (f=8096, c=4096, r=4000), it means that the previous request (f1=5120, c1=4096, r1=1024) needs to be queued first, which means that the remaining waiting patience value of the previous request (f1=5120, c1=4096, r1=1024) is rbe completely consumed, resulting in the time of the previous request (f1 = 5120, c1 = 4096, r1 = 1024) exceeding the time expectation set by the platform on which it is deployed, ultimately leading to the risk of the conclusion that the previous request (f1 = 5120, c1 = 4096, r1 = 1024) plus probability appears Time Out Error, and ultimately may result in no return value for the previous request (f1 = 5120, c1 = 4096, r1 = 1024).
[0046] After each priority sorting and scheduling, the information update component 130 will use the result of subtracting the patience consumption cost value of the current request from the remaining waiting patience value of the previous request as the remaining waiting patience value of the previous request after the sorting change for each sorting change of the priority sorting and scheduling component 120, and store it in the sorting information storage unit in the information update component 130. In this way, so that the requests input after the current request can all perform corresponding priority sorting and scheduling processing using the latest remaining waiting patience value. Figure 2 It shows the update change process of the information of the previous request cut in line after each priority sorting and scheduling. In addition, the stored sorting information in the sorting information storage unit will be aged and eliminated from the storage unit after any request is processed. At the same time, the information update component 130 will also target the number of pending tokens f (forward token num) of any previous request being processed. When it is not cut in line temporarily, as some of its tokens are executed, its number of pending tokens f (forwardtoken num) is continuously updated and reduced, and its remaining waiting patience value is also updated and reduced accordingly, but the consumed patience value c remains unchanged.
[0047] To sum up, according to the request priority scheduling system 100 for large models of the present disclosure, there are two conditions for judging whether priority sorting and scheduling can be performed for the previous request ( r 1 , f 1 , c 1 ) and the current request ( r 2 , f 2 , c 2 , e 2 ): (1) r 1 > 0 and (2) r 1> e 2 Although sometimes it is more effective for long requests to jump the queue ahead of short requests, and under the above two conditions, the probability of long requests jumping the queue ahead of short requests is also low, in order to prevent the phenomenon of long requests jumping the queue ahead of short requests, condition (3) can be added f 2 < f 1 to eliminate this situation.
[0048] The present disclosure can target the patience value for which the parameter has been consumed c、 as the patience cost coefficient for each pre-filled token and as the patience cost coefficient for each pre-filled token to adjust the range for simplifying the judgment process.
[0049] First of all, it is not allowed to jump the queue ahead of a short request that is a previous request, because the pre-filling times of short requests are roughly the same, and often multiple short requests are pre-filled in one batch. For short requests, they can be processed in the order of first come, first served. For this reason, the patience value that has been consumed for a short request c is initially set to the length of the short request. The length of this short request can be configured, generally 4096, or it can also be 2048 or other lengths. In the scenario of long texts, it can also be considered that a short request has a pre-filled token count less than 8192. Due to this setting for short requests, therefore, its remaining waiting patience value r, that is, the corresponding number of all tokens to be processed, will not exceed a predetermined number such as 4096 or 2048. As the priority sorting and scheduling is executed, the patience value that has been consumed for each previous request c will only increase after being jumped the queue on the basis of its initial value, and will not decrease. Therefore, when the initial value of the patience value that has been consumed for a short request c ≥4096 or c ≥8192, its remaining waiting patience value r, may be negative from the beginning, so there is no possibility of being jumped the queue.
[0050] As the pre-filling of requests is continuously executed, f it continuously decreases, c remains unchanged, and the remaining waiting patience value r (= f - c ) will definitely drop to a negative value. Therefore, this means that the remaining execution time of this previous request is very little and it cannot be jumped the queue. Taking supermarket checkout as an analogy, the checkout takes 10 minutes, but 9 minutes have been completed and there is 1 minute left to finish the checkout. At this time, it should not be jumped the queue anymore.
[0051] The request priority scheduling system 100 for large models of the present disclosure can also make condition (2) by defining the cost coefficient and the cost coefficient within a certain range so that condition (2) r 1 > e 2 includes condition (3) f 2 < f 1 . For example: ≥1 and ≥0. In this way, condition (3) can be included in condition (2), that is: f 1 > f 1 - c = r 1 > e 2 = C p * f 2 + C d *d2≥ f 2 The above inequality also implies condition [1], that is r 1 ≥ f 2 >0 However, by defining the range of the cost coefficient and the cost coefficient as ≥1 and ≥0, it will also lead to an unreasonable cut-in situation, although the probability of this situation is extremely low. For example, if the previous request takes 10 minutes, another request may be cut in line even if it only takes 9 minutes and 59 seconds. Therefore, considering that f and d are both integers, and the operation between integers and floating-point numbers in the computer involves type conversion, so can be set as an integer greater than or equal to 2. If so, for a 10-minute request, it is unreasonable for a 9-minute and 59-second request to cut in line, but it is reasonable for a request within two and a half minutes to cut in line. The reason for setting it to 2 is two and a half minutes. As mentioned before, the pre-fill time is proportional to the square of the number of tokens. When set to 2, without considering decoding, the pre-fill time will increase by at most 1 / 4. Usually, the ratio of the average step efficiency during decoding to the average step efficiency during pre-filling is used as Default value. (Measured with the qwen72b and glm9b models on the 4090 GUP, this value is approximately 16 to 22. On other platform models, it is basically within this range. Therefore, the impact of decoding on the preempted request is greater than that of prefill on the preempted request. For example, there may be a short request with 7 tokens, asking "What does A Dream of Red Mansions talk about", and a response of 256 tokens will be generated. The prefill is very short, but the decoding is long, and the decoding also limits the batch token number, reducing the efficiency. With such parameters, the preemption of short requests is also greatly restricted because the engine doesn't know when the decoding of a request ends. The engine only knows that the decoded tokens do not exceed a value (such as the default value of 512), but often the decoding of short requests ends around 250, some simple questions end around 100, and some professional questions may reach 2000 or even 4000. This makes it impossible to predict the accurate value of d. Therefore, the maximum value of d is used in this disclosure. At this time, C d needs to be adjusted accordingly. For example, C d is lowered to 8. Although this eliminates the above-mentioned problem, another problem is generated, that is, short requests are also calculated or configured for patience consumption according to the maximum value of decoding, which makes their patience consumption increase a lot. Therefore, in the actual test of this disclosure, it is observed that a short request with 7 tokens cannot preempt a long request with 8000 tokens, so that the effect of preemption is not obvious, and several short requests within 100 tokens pile up, with a waiting time of about 20 seconds, while the actual execution time of their prefill is within 1 second, which does not conform to the principle of short request priority in this disclosure. Considering user patience, this is also unreasonable. Therefore, in the case of adopting the above range limitation (if there is no such range limitation, there is no such problem, so it can be not considered), this disclosure adopts a short request priority strategy. For short requests with the number of prefill tokens less than the length of the short request (generally 4096, and in the scenario of long texts, it can also be considered that short requests with the number of prefill tokens less than 8192), its C d is directly set to 0. This greatly reduces the patience consumption of short requests c , making short request preemption more prioritized and avoiding excessive waiting time for short requests.
[0052] Figure 3 is a flowchart of a request priority scheduling method for a large model shown according to an exemplary embodiment. As Figure 3As shown, in the request priority scheduling method 300 for large models, first, in the request information acquisition step S310, data processing requests are received in the order of request input time. The token count of each request is obtained through a tokenizer, and a remaining waiting patience value is configured for each request. Subsequently, in the priority sorting and scheduling step S320. That is, in step 321 for any request, the patience consumption cost value generated by its being processed preferentially is obtained, and the patience consumption cost value of the current request is compared with the remaining waiting patience value of the previous request arranged in front of it in chronological order. If it is determined in step S321 that the patience consumption cost value of the current request is less than the remaining waiting patience value of the previous request, then in step S322, the arranged priority order of the current request is scheduled to be moved forward to before the previous input request; otherwise, when the patience consumption cost value of the current request is greater than the remaining waiting patience value of the previous request, the arranged priority order of the current request remains unchanged. Finally, in the information update step S330, for each change in the sorting of the priority sorting component, the result of subtracting the patience consumption cost value of the current request from the remaining waiting patience value of the previous request is used as the remaining waiting patience value of the previous request after the sorting change;. Subsequently, in step S340, it is determined whether there is still a previous request for the current request. If there is, the priority sorting and scheduling step and the information update step are repeated for the current request for which the arranged priority order has been completed, as Figure 3 shown, return to step S321; otherwise, the priority sorting and scheduling process ends.
[0053] In summary, compared with the traditional first-in-first-out sorting, the present disclosure can determine whether to cut in line for the current request by configuring a remaining waiting patience value for each request and comparing the patience consumption cost value of the current request with the remaining waiting patience value of the previous request arranged in front of it in chronological order. On this basis, a short request comparison termination condition and a situation where a long request becomes a short request after being processed are set, so that the priority sorting of the requests shows its immediate real-time state, thereby generally and immediately keeping the request response within the user's waiting patience value range and greatly improving the user experience.
[0054] Those skilled in the art can understand that the above-mentioned modules can be distributed in the device according to the description of the embodiments, or can be correspondingly changed and located in one or more devices different from the present embodiment. The modules of the above embodiments can be combined into one module, or further split into multiple sub-modules.
[0055] Through the description of the above embodiments, those skilled in the art can easily understand that the exemplary embodiments described herein can be implemented by software or by a combination of software and necessary hardware. Therefore, the technical solutions according to the embodiments of the present disclosure can be embodied in the form of a software product, which can be stored in a non-volatile storage medium (such as a CD-ROM, a USB flash drive, a mobile hard disk, etc.) or on a network, including several instructions to enable a computing device (such as a personal computer, a server, a mobile terminal, or a network device, etc.) to execute the method according to the embodiments of the present disclosure.
[0056] The exemplary embodiments of the present disclosure have been specifically shown and described above. It should be understood that the present disclosure is not limited to the detailed structures, settings, or implementation methods described herein; on the contrary, the present disclosure is intended to cover various modifications and equivalent settings included within the spirit and scope of the appended claims.
Claims
1. A request priority scheduling system for large models, comprising: A request information acquisition component, which receives data processing requests in the order of request input time, obtains the number of tokens for each request through a tokenizer, and configures a remaining waiting patience value for each request; A priority sorting and scheduling component, for each request, obtains the patience consumption cost value generated by the request to be processed preferentially, and compares the patience consumption cost value of the current request with the remaining waiting patience value of the previous request arranged in chronological order in front of it. When the patience consumption cost value of the current request is less than the remaining waiting patience value of the previous request, the arranged priority order of the current request is scheduled to be moved forward to before the previous input request, and when the patience consumption cost value of the current request is greater than the remaining waiting patience value of the previous request, the arranged priority order of the current request remains unchanged; And An information update component, for each change in sorting by the priority sorting component, uses the result of subtracting the patience consumption cost value of the current request from the remaining waiting patience value of the previous request as the remaining waiting patience value of the previous request after the sorting change.
2. The request priority scheduling system for large models according to claim 1, wherein, The remaining waiting patience value is determined according to the number of all tokens to be processed of the current request or according to the estimated execution time length of all tokens to be processed of the current request.
3. The request priority scheduling system for large models according to claim 2, wherein, The priority sorting and scheduling component obtains the number of all tokens to be processed for any request, and when the number of tokens to be processed of the previous request is less than that of the current request, no priority sorting and scheduling processing is performed.
4. The request priority scheduling system for large models according to claim 2, wherein, The priority sorting and scheduling component does not perform priority sorting and scheduling processing when the number of all tokens to be processed of the previous request is less than a predetermined number or the estimated execution time length is shorter than a predetermined time length.
5. The request priority scheduling system for large models according to claim 4, wherein, The predetermined number is 4096 or 8192.
6. The request priority scheduling system for large models according to any one of claims 1-5, wherein, the remaining waiting patience value r is where f is the number of all tokens to be processed in the previous requests, c is the patience value consumed by any previous request, and the patience consumption cost value of the current request for the previous requests assuming being scheduled in the permutation priority order is e , wherein is the patience cost coefficient for each pre-filled token, is the patience cost coefficient for each pre-filled token, f is the remaining number of tokens to be run for the current request, d is the number of tokens expected to be decoded for the current request.
7. The request priority scheduling system for large models as described in claim 6, wherein, The patience cost coefficient of the pre-filled token is greater than or equal to 2, and the patience cost coefficient of the pre-filled token is the ratio of the average step efficiency in the decoding stage to the average step efficiency in the pre-filling stage of the platform where the large model is deployed.
8. A request priority scheduling method for large models, comprising: A request information acquisition step, which receives data processing requests in the order of request input time, obtains the number of tokens for each request through a tokenizer, and configures a remaining waiting patience value for each request; A priority sorting and scheduling step, for the current request, obtains the patience consumption cost value generated by the request to be processed preferentially, and compares the patience consumption cost value of the current request with the remaining waiting patience value of the previous request arranged in chronological order in front of it. When the patience consumption cost value of the current request is less than the remaining waiting patience value of the previous request, the arranged priority order of the current request is scheduled to be moved forward to before the previous input request, and when the patience consumption cost value of the current request is greater than the remaining waiting patience value of the previous request, the arranged priority order of the current request remains unchanged; An information update step, for each change in sorting by the priority sorting component, uses the result of subtracting the patience consumption cost value of the current request from the remaining waiting patience value of the previous request as the remaining waiting patience value of the previous request after the sorting change; And Repeating the priority sorting and scheduling step and the information update step for the current request whose arranged priority order has been scheduled until the patience consumption cost value of the current request is greater than the remaining waiting patience value of the previous request.
9. The request priority scheduling method for large models according to claim 8, wherein, The priority sorting and scheduling steps include: for any request, obtaining the number of all tokens to be processed thereof, and when the number of tokens to be processed of the previous request is less than that of the current request, not performing priority sorting and scheduling processing.
10. The request priority scheduling method for large models according to claim 8, wherein, The priority sorting and scheduling steps include: when the number of all tokens to be processed of the previous request is less than a predetermined number or the estimated execution time length is shorter than a predetermined time length, not performing priority sorting and scheduling processing.