Large-model load balancing routing system and method

By designing a real-time cache-aware load balancing routing system, the problem of load imbalance in the inference engine of large language model is solved, efficient load balancing and parallel processing are achieved, and inference speed and resource utilization are significantly improved.

CN120144301AActive Publication Date: 2025-06-13BEIJING SILICONFLOW TECHNOLOGY CO LTD

Patent Information

Application Number
CN202510246765.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-04
Publication Date
2025-06-13
Estimated Expiration
2045-03-04

AI Technical Summary

Technical Problem

In the large language model inference engine, there is a problem of load imbalance, which leads to high load operation or waiting for some processing instances, resulting in waste of resources, and traditional request scheduling strategies are difficult to effectively optimize.

Method used

A load balancing routing system is designed to perceive the reusable results and vacant calculations of each parallel processing instance through real-time cache, and select the processing instance with the smallest comprehensive load as the forwarding destination to achieve the optimization of load balancing and parallel efficiency.

Benefits of technology

Through this system, the inference speed is significantly improved, resource consumption is reduced, load balancing is achieved, and waiting and resource waste in unbalanced situations are avoided.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120144301A_ABST
    Figure CN120144301A_ABST
Patent Text Reader

Abstract

The invention relates to a load balancing routing system and method for a large model. The system comprises a request analysis component for receiving a data processing request of a user to obtain token information of the data processing request; the subset segmentation component is used for dividing a token subset from the head of the token information of the received data processing request every time by taking a preset token number as a cardinal number, and generating a hash value of each token subset which is divided, so as to form one or more hash value sets; the load information acquisition component is used for acquiring a hash value set and the load amount of the data processing request being processed aiming at each destination processing instance; the scoring component scores each destination processing instance based on the matching degree between the current data processing request and the hash value set of each destination processing instance and the load capacity of each destination processing instance; and the forwarding component is used for forwarding the current data processing request route to the destination processing instance with the highest score or multiple highest scores.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the field of computer information processing, and more particularly, to a load balancing routing system and method for large models. Background Art

[0002] Large language model inference engines often receive a large number of user requests. Large language model inference usually requires a large amount of computing resources, but the hardware resources are limited and cannot handle all requests simultaneously. In actual application scenarios, in order to efficiently and stably process a large number of concurrent requests, while optimizing resource utilization and user experience, traditional large language model inference engines introduce a request scheduling layer to queue requests, wait for the inference engine to have idle computing resources, dynamically adjust the processing order according to priorities, and merge multiple requests to reuse the hardware parallel computing ability and improve throughput. For this reason, when traditional deep learning models are inferring, in order to make full use of computing power to accelerate the inference process, data parallel processing is usually adopted. In the parallel mode of DP (data parallel), requests are sent to different DP groups (or called processing instances). In the case of a large number of parallel processing instances, if the forwarding or task allocation method is not managed, it is easy to cause a load imbalance situation, resulting in the time of high load operation or processing waiting or queuing time of some cards being significantly higher than that of other cards, causing waiting of other cards and waste of computing power. Therefore, when designing a routing component or a component for distributing data, load balancing must be considered.

[0003] A traditional technique called continuous batching is to flexibly process input sequences of different lengths by combining multiple batches together and updating the batch in real time during the inference process to reduce resource idleness and inference latency. Traditional batching needs to wait for all requests to be prepared, which easily causes low hardware utilization and increased response time, while continuous batching allows the model to dynamically merge requests during the inference process, supports different sequence lengths, and immediately replaces new requests after one request is completed, thus significantly improving throughput and efficiency. This technique is particularly suitable for high-concurrency scenarios and can better balance the utilization of computing resources and response speed.

[0004] In addition, during the data inference process of traditional large models, the request scheduling strategies are usually the prefill-first strategy, the decode-first strategy, and the Chunked Prefill strategy. The prefill-first strategy gives priority to scheduling prefill requests into the engine when both prefill requests and decode requests exist. The decode-first strategy gives priority to scheduling decode requests into the engine when both prefill requests and decode requests exist. The Chunked Prefill strategy, when both prefill requests and decode requests exist, splits a prefill request into multiple parts and executes it in multiple rounds. In each round, it combines with the decode request to form a complete batch and sends it into the engine.

[0005] Obviously, considering the inference efficiency of large models and the stability of inference services, using the prefill-first strategy or the decode-first strategy will increase the latency of the inference service. The prefill-first strategy gives priority to scheduling prefill requests into the engine and delays the scheduling of decode requests, which has a greater impact on decode requests. The decode-first strategy, which gives priority to scheduling decode requests into the engine, will result in too few input tokens, reduced GPU utilization efficiency, and a significant increase in the latency of the first token, affecting the user experience. Similarly, the Chunked Prefill strategy can increase the number of tokens processed by the engine in each round and improve the inference throughput by combining prefill requests and decode requests into a complete batch for scheduling. Since the number of tokens processed by the inference engine in each round is relatively stable, the latency between tokens is relatively stable. However, its total latency is still affected by the insertion of prefill requests, delaying the processing speed of decode requests, and because prefill requests still occupy the video memory for a long time, the concurrency of decode requests is limited.

[0006] In addition, due to the existence of kv cache, in multi-round conversations, if they enter the same DP group, the same preamble part can directly reuse the previous kv cache without recalculation. Therefore, the router must also consider the situation of kv cache. More complex and likely to occur: an enterprise user adds the same long prompt to the request, then a large number of identical requests sent by this enterprise user will have the same prefix (kv cache). It is unreasonable to put all these requests in the same DP group and leave other groups empty.

[0007] Therefore, there is a desire to obtain a new routing system and method for data requests that does not need to specifically prioritize between prefill and decode, fully utilizes repeated calculations, and can fully consider the load conditions of specific parallel processing instances to optimize the routing of data requests and achieve load balancing.

[0008] The above information disclosed in the Background section is only for enhancing the understanding of the background of the present disclosure, and thus it may include information that does not constitute the prior art known to those of ordinary skill in the art. Summary of the Invention

[0009] In view of this, the present disclosure provides a load balancing routing system for large models, especially for large language models in large-scale hardware clusters. By enabling each parallel data processing instance in the real-time cache to reuse the number of already computed results and the vacant computing capacity, in the overall parallel processing environment, the processing instance with the smallest comprehensive load, which can complete the task in the shortest possible time and increase its own load as little as possible, is selected as the forwarding destination, thereby achieving the optimization of load balancing parallelism, improving the parallel efficiency of the entire model, and avoiding the situation of unbalanced load.

[0010] Other features and advantages of the present disclosure will become apparent from the following detailed description, or be learned in part through the practice of the present disclosure.

[0011] According to one aspect of the present disclosure, a load balancing routing system for a large model is provided, including: a request parsing component that receives a data processing request from a user to obtain the token information of the data processing request; a subset splitting component that, based on a predetermined number of tokens as a base, divides a token subset starting from the head of the token information of the received data processing request each time until the end of the token information is reached, and uses a hash function to generate the hash value of each divided token subset, thereby forming one or more hash value sets of the data processing request; a load information collection component that, for each destination processing instance, collects and caches the hash value set of each data processing request forwarded to it by the forwarding component and the load of the data processing requests being processed by each destination processing instance; a scoring component that scores each destination processing instance based on the matching degree between the hash value set corresponding to the current data processing request and the hash value sets corresponding to the data processing requests received by the load information collection component for each destination processing instance, and the load of each destination processing instance; and a forwarding component that routes and forwards the current data processing request to the destination processing instance with the highest score or one of the multiple highest scores.

[0012] In the load balancing routing system for the large model according to the present disclosure, when the subset splitting component divides the token subset, in two adjacent divisions, the number of token subsets divided in the latter division is twice the number of token subsets divided in the former division or the number of token subsets divided in the latter division increases by one base number compared to the number of token subsets divided in the former division.

[0013] The load balancing routing system of the large model according to the present disclosure, wherein the load information collection component allocates memory with a predetermined length for each destination processing instance, and when the memory is full, replaces the hash values cached in the memory in a first-in, first-out manner.

[0014] The load balancing routing system of the large model according to the present disclosure, wherein the scoring component scores the matching degree and the load amount respectively when scoring. The scoring of the matching degree calculates, for any destination processing instance, the ratio of the number of hash values of the current data processing request that match the hash values cached for the destination processing instance by the load information collection component to the number of the hash value set of the current data processing request as the matching hit rate, and assigns corresponding hit rate scores according to the multiples of the base number included in the divided page corresponding to the matched hash value, so as to accumulate the total hit rate scores of the current data processing request relative to the destination processing instance, and obtain the sum of the hit rate scores as the total hit rate score of the current data processing request for each destination processing instance; and the scoring of the load amount calculates, for any destination processing instance, the negative number of multiples of the preset length for each data processing request being processed by the destination data processing instance as the load amount score of each data processing request, thereby accumulating the load amount scores of all data processing requests being processed by the destination data processing instance as the total load amount score of the destination processing instance, and summing the total hit rate score and the total load amount score for each destination processing instance as the total score of the current data request for each destination processing instance.

[0015] The load balancing routing system of the large model according to the present disclosure, wherein when there are pre-filling components and decoding service components separately arranged in the destination processing instance, the forwarding component directly selects the destination processing instance with the highest total score in the pre-filling component for forwarding when the total hit rate score for any one destination processing instance is lower than the predetermined threshold of the total hit rate score.

[0016] The load balancing routing system of the large model according to the present disclosure, wherein when there are pre-filling components and decoding service components separately arranged in the destination processing instance, the forwarding component directly selects the destination processing instance with the highest total score in the decoding service component for forwarding when the total hit rate score for a certain destination processing instance is higher than the predetermined threshold of the total hit rate score.

[0017] The load balancing routing system of the large model according to the present disclosure, wherein when the length of the data processing request being processed by the destination data processing instance is less than the preset length, the scoring component directly assigns a score of -1.

[0018] According to another aspect of the present disclosure, there is provided a load balancing routing method for a large model, including: receiving a data processing request of a user through a request parsing component, and parsing to obtain token information of the data processing request; through a subset splitting component, taking a predetermined number of tokens as a base, each time dividing a token subset starting from the head of the token information of the received data processing request until reaching the end of the token information, and using a hash function to generate hash values of each divided token subset, thereby forming one or more hash value sets of the data processing request; through a load information collection component, for each destination processing instance, collecting the hash value set of each data processing request forwarded to it by the forwarding component and the load of the data processing requests being processed by each destination processing instance; through a scoring component, based on the matching degree between the hash value set corresponding to the current data processing request and the hash value sets corresponding to the data processing requests received by the load information collection component for each destination processing instance and the load of each destination processing instance, scoring each destination processing instance; and through a forwarding component, routing and forwarding the current data processing request to the destination processing instance with the highest score or one of the multiple highest scores.

[0019] In the load balancing routing method of the large model according to the present disclosure, when dividing the token subset through the subset splitting component, in two adjacent divisions, the number of token subsets divided in the latter division is twice the number of token subsets divided in the previous division or the number of token subsets divided in the latter division increases by one base number relative to the number of token subsets divided in the previous division.

[0020] In the load balancing routing method of the large model according to the present disclosure, the load information collection component allocates a memory with a predetermined length for each destination processing instance, and when the memory is full, replaces the hash values cached in the memory in a first-in, first-out manner.

[0021] The load balancing routing method of the large model according to the present disclosure, wherein when the scoring component performs scoring, it includes: for any destination processing instance, calculating the ratio of the number of hash values of the current data processing request that match the hash values cached for the destination processing instance collected by the load information collection component to the number of the hash value set of the current data processing request as the matching hit rate, and assigning corresponding hit rate scores according to multiples of the cardinality included in the partition page corresponding to the matched hash value, so as to accumulate the total hit rate scores of the current data processing request relative to the destination processing instance, and obtaining the sum of the hit rate scores as the total hit rate score of the current data processing request for each destination processing instance; for any destination processing instance, calculating the negative number of multiples of a preset length for each data processing request being processed by the destination data processing instance as the load score of each data processing request, and thus accumulating the load scores of all data processing requests being processed by the destination data processing instance as the total load score of the destination processing instance; and summing the total hit rate score and the total load score for each destination processing instance as the total score of the current data request for each destination processing instance.

[0022] The load balancing routing system of the large model according to the present disclosure, wherein the forwarding by the forwarding component further includes: in the case where the pre-filling component and the decoding service component are separately arranged in the destination processing instance, when the total hit rate score for any one destination processing instance is lower than the hit rate score predetermined threshold, directly selecting the destination processing instance with the highest total score in the pre-filling component for forwarding.

[0023] The load balancing routing system of the large model according to the present disclosure, wherein the forwarding by the forwarding component further includes: in the case where the pre-filling component and the decoding service component are separately arranged in the destination processing instance, when there is a total hit rate score for a certain destination processing instance that is higher than the hit rate score predetermined threshold, directly selecting the destination processing instance with the highest total score in the decoding service component for forwarding.

[0024] The load balancing routing system of the large model according to the present disclosure, wherein when the length of the data processing request being processed by the destination data processing instance is less than the preset length, the scoring component directly assigns a score of -1.

[0025] The load balancing routing and method of the large model according to the present disclosure optimize resources, perform parallel processing, and achieve load balancing by comprehensively considering the repeatable computing amount of each parallel processing instance and the data request load being processed, significantly improving the inference speed and reducing resource consumption to meet the requirements of different models and tasks. In other words, by introducing a cache-aware routing mechanism, the present disclosure polls each node (or processing instance), records the current load situation, and the routing system automatically distributes requests according to the load conditions of each node, making the load of each node relatively balanced and stable as much as possible to fully utilize all node resources, and enabling load balancing parallelism for the processing instances of padding calculation and decoding calculation respectively.

[0026] It should be understood that the above general description and the following detailed description are only exemplary and do not limit the present disclosure. Brief Description of the Drawings

[0027] By describing its exemplary embodiments in detail with reference to the accompanying drawings, the above and other objectives, features, and advantages of the present disclosure will become more apparent. The following described drawings are only some embodiments of the present disclosure. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.

[0028] Figure 1 It is a block diagram of the first embodiment of a load balancing routing system for a large model shown according to an exemplary embodiment.

[0029] Figure 2 It is a block diagram of the second embodiment of a load balancing routing system for a large model shown according to an exemplary embodiment.

[0030] Figure 3 It is a flowchart of the first embodiment of a load balancing routing method for a large model shown according to an exemplary embodiment.

[0031] Figure 4 It is a flowchart of the second embodiment of a load balancing routing method for a large model shown according to an exemplary embodiment. Detailed Description of the Specific Embodiments

[0032] Now, the exemplary embodiments will be described more comprehensively with reference to the accompanying drawings. However, the exemplary embodiments can be implemented in various forms and should not be construed as limited to the embodiments set forth herein; on the contrary, these embodiments are provided so that the present disclosure will be thorough and complete, and the concept of the exemplary embodiments will be fully conveyed to those skilled in the art. The same reference numerals in the figures denote the same or similar parts, and thus their repeated description will be omitted.

[0033] In addition, the described features, structures, or characteristics may be combined in any suitable manner in one or more embodiments. In the following description, numerous specific details are provided to give a thorough understanding of the embodiments of the present disclosure. However, those skilled in the art will realize that the technical solutions of the present disclosure may be practiced without one or more of the specific details, or other methods, components, devices, steps, etc. may be employed. In other cases, well-known methods, devices, implementations, or operations are not shown or described in detail to avoid obscuring aspects of the present disclosure.

[0034] The block diagrams shown in the drawings are only functional entities and do not necessarily correspond to physically independent entities. That is, these functional entities may be implemented in software form, or in one or more hardware modules or integrated circuits, or in different networks and / or processor devices and / or microcontroller devices.

[0035] The flowcharts shown in the drawings are only illustrative and do not necessarily include all the content and operations / steps, nor do they necessarily need to be executed in the described order. For example, some operations / steps may be decomposed, while some operations / steps may be combined or partially combined, so the actual execution order may change according to the actual situation.

[0036] It should be understood that although terms such as first, second, and third may be used herein to describe various components, these components should not be limited by these terms. These terms are used to distinguish one component from another. Thus, the first computing device discussed below may be referred to as the second computing device without departing from the teachings of the concepts of the present disclosure. As used herein, the term "and / or" includes any one and all combinations of one or more of the associated listed items.

[0037] Those skilled in the art can understand that the drawings are only schematic diagrams of exemplary embodiments, and the modules or processes in the drawings are not necessarily essential for implementing the present disclosure, so they cannot be used to limit the protection scope of the present disclosure.

[0038] Figure 1 is a block diagram of a first embodiment of a load balancing routing system for a large model shown according to an exemplary embodiment. As Figure 1 shown, the load balancing routing system 100 for a large model includes: a request parsing component 110, a subset splitting component 120, a load information collection component 130, a scoring component 140, and a forwarding component 150.

[0039] As Figure 1As shown, the request parsing component 110 receives a data processing request from the user to obtain the token information of the data processing request. In a large language model (LLM), a token is the basic unit of text, which can be a word, sub-word, or character. The model decomposes the input text into tokens for processing and understanding. The number of tokens determines the input length and computational complexity of the model. The text data input by the user is first converted into tokens and then input into the model. The request parsing component 110 parses the current data processing request into one or more tokens.

[0040] Next, the subset splitting component 120 divides one or more tokens of each data processing request into a token subset starting from the head of the token information of the received data processing request each time based on a predetermined number of tokens as the base, until the end of the token information is reached, and uses a hash function to generate the hash value of each divided token subset, thereby forming one or more hash value sets of the data processing request.

[0041] For example, in one splitting method, the subset splitting component 120 splits the tokens converted from the current data processing request according to the number of tokens per page of page attention, and draws the splitting line according to the rule of doubling each time. Assume that the number of tokens per page in the system is 32 (base), and the number of tokens in a request is 555. Then the splitting lines are [32, 64, 128, 256, 512]. Usually, the small number of tokens remaining when 555 is more than 512 will be ignored because it will not affect the accuracy of the final pre-filling or decoding. In this way, the first token subset has 32 tokens, the second token subset has 64 tokens, which includes the 32 tokens in the first token subset, and so on. In this way, the token subsets divided each time start from the 0th token of the request and are respectively tokens[0:32], tokens[0:64], tokens[0:128], tokens[0:256], tokens[0:512]. Tokens[0:i] represents from the 0th token to the ith token. For these token subsets (or token pages), 5 hash values can be converted through a fixed hash function. The advantage of this division is to maintain the overall linear complexity. For example, the complexity above is only 32 + 64 + 128 + 256 + 512 = 992 <= 2 * 555.

[0042] Optionally, it is also possible to perform a split every page, that is, based on a predetermined number of tokens, each time a subset of tokens is divided starting from the head of the current data processing request. In two adjacent splits, the number of the token subset divided in the latter split increases by one base number compared to the number of the token subset divided in the former split. For example, 32, 64, 96, 128, 160, …, 512, 544. However, its complexity will reach n^2: 32 + 64 + 96 + 128 + 160 + … + 512 + 544 = 32 * (1 + 17) * 17 / 2. Here, 17 or n will increase linearly as the total number of tokens increases, so the complexity is squared.

[0043] The load information collection component 130 collects and caches the hash value set of each data processing request forwarded by the forwarding component 150 to each destination processing instance and the load of the data processing requests being processed by each destination processing instance. That is to say, when the forwarding component 150 forwards each data processing request, the load information collection component 130 will obtain the ID of the processing instance as the forwarding destination and cache the hash value set of the forwarded data processing request. Correspondingly, it also knows how many data processing requests the forwarding component 150 forwards to each data processing request. That is, for each processing instance as the destination, a data structure is established to cache the relevant information of the received data processing requests. The data structure used for this information only needs to be able to record the hash values of the divided token subsets of the received data requests in the input order. For example, the load information collection component 130 uses a dict and a FIFO list to save the hash values of the existing requests from 0 to the dividing line for each processing instance used for data parallel inference. Therefore, the load information collection component 130 allocates memory of a predetermined length for each destination processing instance and, when the memory is full, replaces the hash values cached in the memory in a first-in, first-out manner.

[0044] The hash value in the dict here is a bit like the index of the tokens subset corresponding to the hash value. In other words, Dict is a mapping from the hash value to the number of instances of the hash value in the processing. FIFO list means First In First Out (in fact, the first in first out here means that the oldest one pops out first). Because the number of pages is limited, after the cache is used up, the new request will squeeze out the caches of the old request. So the FIFO list also has a size. Before the number of hashes reaches the maximum value of the FIFO list, the FIFO list will store the hash values ​​of all incoming requests according to the dividing line. Each request may store multiple hash values. The number of dividing lines, that is, the number of hash values, is determined by the length of tokens. When the number of hashes reaches the maximum value of the FIFO list, the FIFO list will pop out the oldest hash value each time, and then put a new hash value in, and repeat it multiple times. Multiple hash values ​​of a request will be entered in the hash value from small to large according to the number of pages. So multiple hash values ​​of the same request also have a sequence.

[0045] An example of this FIFO list is as follows: For the first processing instance, the load information collection component 130 records 3 requests.

[0046] “I ate apples and bananas” “I ate apples and pineapples” “I ate apples and bananas, they were so delicious I couldn’t stop eating them” For the convenience of description, it is assumed that the number of tokens in a page is 1, although it is impossible in reality, and it is generally 32 or 128. Here, it is simply assumed that each word and each punctuation mark is a token. At the same time, for the convenience of description, it is assumed that the corresponding hash value is a 4-digit number (the length of the hash value is very long, and the 4-digit number is used here just for the convenience of explanation). The hash value of the token subset after the token of the above request is divided (to simplify the description, it is assumed that it meets the predetermined division rule, and the division rule can be set manually, not necessarily the rule exemplified in this disclosure) is: "I" → 1234 "I eat" → 2222 "I ate an apple" → 3214 "I ate apples and bananas" → 4444 “I ate apples and pineapples” → 8888 “I ate apples and bananas, they were so delicious I couldn’t stop eating them” →5555 Assume the length of the FIFO list is 7 (for example, it can only store 7 hash values. In the actual scenario, any finite length can be set according to the data processing volume requirements). Before the first request comes in, both the FIFO list and the dict are empty.

[0047] FIFO_list = [], dict = {} After the first request "I ate apples and bananas" comes in, its data structure is as follows: FIFO_list = [1234, 2222, 3214, 4444] Dict = {1234: 1, 2222: 1, 3214: 1, 4444: 1}, where 1 means that there is only 1 occurrence of this hash value in the FIFI_list.

[0048] If the second request "I ate apples and pineapples" comes in, the data structure becomes: FIFO_list = [2222, 3214, 4444, 1234, 2222, 3214, 8888] Dict = {1234: 1, 2222: 2, 3214: 2, 4444: 1, 8888: 1}. 1234 appears only once; 2222 appears 2 times, at the 1st and 5th positions; 3214 appears 2 times, at the 2nd and 6th positions; 4444 appears only once; 8888 appears only once. It can be seen that in the FIFO_list, the oldest 1234 was pushed out before the last hash 8888 came in.

[0049] If the third request "I ate apples and bananas and couldn't stop eating" comes in, the data structure becomes: FIFO_list = [3214, 8888, 1234, 2222, 3214, 4444, 5555] Dict = {8888: 1, 1234: 1, 2222: 1, 3214: 2, 4444: 1, 5555: 1} Because the FIFO list was full before the third request came in, so every time a new hash value is inserted, an old hash value is pushed out.

[0050] With these two data structures, in the first partitioning, maintaining the mapping only requires a complexity of O(1). Not only can it confirm whether a certain hash value exists in the corresponding processing instance in O(1) complexity in the first place, but it can also squeeze out the oldest hash value in chronological order after the FIFO_list is full, perfectly simulating the process of the kv cache, so that the load information collection component 130 can restore the real load situation of each processing instance as much as possible. If the simulation is not good, for example, it is thought that there are two requests a and b in the first processing instance, but in fact, there are two requests c and d in the first processing instance, or there are 100 requests in the first processing instance while the second processing instance is empty. When encountering subsequent problems with request a, because the simulation is not good and the real load situation is not figured out, and it is still placed in the first processing instance, it will increase the load of the first processing instance.

[0051] In addition, the load information collection component 130 collects the load of the data processing requests being processed by each destination processing instance for each destination processing instance, which is obtained by detecting the current running state of each processing instance. Specifically, after each processing instance completes a data processing request or a subset of tokens, it can feedback information on task completion to the load information collection component 130. Thus, based on the number or the number and order of subsets of tokens of the data processing requests that the load information collection component has previously learned that the forwarding component 150 forwards to each data processing request, subtracting the completed number correspondingly, it can know the number of data requests that each processing instance is currently processing and the number of tokens of each data request.

[0052] Subsequently, the scoring component 140 scores each destination processing instance based on the matching degree between the set of hash values corresponding to the current data processing request and the set of hash values corresponding to the data processing requests received by the load information collection component for each destination processing instance, and the load of each destination processing instance. Specifically, the scoring of the matching degree is calculated by, for any destination processing instance, calculating the ratio of the number of hash values of the current data processing request that match the hash values cached by the load information collection component for the destination processing instance to the number of the set of hash values of the current data processing request as the matching hit rate, and assigning corresponding hit rate scores according to the multiple of the base number included in the partition page corresponding to the matched hash value, so as to accumulate the total hit rate scores of the current data processing request relative to the destination processing instance, and obtain the sum of the hit rate scores as the total hit rate score of the current data processing request for each destination processing instance.

[0053] For example, for a processing instance, after the first processing instance has calculated the request "I ate apples and bananas", the Dict for the first processing instance is as follows: Dict = {1234: 1, 2222: 1, 3214: 1, 4444: 1}, After the second processing instance has calculated a request: "I ate apples and pineapples", the Dict for the second processing instance is as follows: Dict = {1234: 1, 2222: 1, 3214: 1, 8888: 1} Then, there is a current data processing request: "I ate apples and bananas and couldn't stop because they were so delicious". The hash values for each page after segmentation are: Hash: 1234, 2222, 3214, 4444, 5555 As described above, for a request, multiple hash values are calculated according to the delimiter. Since scores are assigned to different hashes, and the score is the number of pages. That is, one point is obtained for each page hit. The scores for hitting different hash values will be different. The scores are assigned as multiples of a predetermined number of tokens (for example, 32 tokens) as the base. Therefore, the scores for hitting the first to fifth hash values are: 1, 2, 4, 8, 16 respectively. As mentioned before, assuming the delimiter is obtained by doubling the tokens of one page, then hitting the first hash value, which is one page of tokens, can reduce the calculation of one page of tokens during actual data inference; hitting the second hash value can reduce the calculation of two pages of tokens; hitting the third hash value can reduce the calculation of 4 pages of tokens; hitting the fourth hash value can reduce the calculation of 8 pages of tokens, and so on. Therefore, for the current data processing request: "I ate apples and bananas and couldn't stop because they were so delicious", the hash value set is: Hash: 1234, 2222, 3214, 4444, 5555, and it is matched against the hash values in the Dict cached by the first processing instance and the second instance respectively. In this way, for the hash value hit scoring of the current data processing request, for the first processing instance, hitting the first four hashes, the total hit rate is 15, while for the second processing instance, hitting the first three hashes, the total hit rate score is 7. Therefore, the subsequent forwarding component 150 will route the current data processing request to the data processing instance with the highest matching hit rate score in the pre-population or decoding service component 10, that is, the first processing instance, thereby reducing the calculation and / or access volume of the first processing instance for performing data population or decoding services.

[0054] Furthermore, the load amount is scored by calculating the negative value of the multiple of the preset length for each data processing request being processed by the destination data processing instance for any destination processing instance as the load amount score for each data processing request. Thus, the load amount scores of all data processing requests being processed by the destination data processing instance are accumulated as the total load amount score of the destination processing instance, and the sum of the total hit rate score and the total load amount score for each destination processing instance is used as the total score of the current data request for each destination processing instance. In the consideration of load balancing, it is necessary to consider both the computing time and the memory load simultaneously. Simply put, the memory occupation of a request with a length of 100,000 tokens is definitely different from that of a request with a length of 20 tokens. However, when it reaches the router, there is only the prefill part and no decode part. Maybe a request with 100,000 tokens only generates 1 token and then ends. Maybe a request with 20 tokens will generate 8,000 tokens (a simple example of such a request is: "Help me write an 8,000-word composition"). The generated length cannot be predicted, but fortunately, based on statistics, the average length of general requests can be known. The same is true for the computing time. The prefill length of request a is 5,000 times that of request b, but the actual computing time of a may only be 50 times that of request b. In the LLM model, for each decoded token, the attention part grows linearly with the increase in the number of tokens, but the time of the MLP is fixed. It should be noted that the time to generate one token for a request with a length of 100,000 tokens will definitely be longer than the time to generate one token for a request with 20 tokens, but it will definitely not be 5,000 times longer.

[0055] For example, for a specific model, an empirical token value of a request can be set based on experience to evaluate the score of the load. To show that the more the load, the lower the score, a negative score is assigned to the load increased on a specific processing instance. That is, for each additional load request of one basic unit (empirical token value), -1 point is assigned, and for any token subset, even if its length is less than one basic unit, -1 point is assigned. Therefore, the formula Score = -(token_len / / 2048 + 1) can be used to calculate the score, where / / represents the quotient (without remainder, no decimal part), token_len is the length of the token set, and 2048 is the empirical length. Of course, for simplicity, the score can also be directly assigned according to the number of load requests, without calculating the length of the load requests, that is, without using the empirical value. Because, from the perspective of normal distribution, when there are a large number of requests, the overall average length basically does not change much. Therefore, by directly using the counting method, for each processing instance, for each additional load request, -1 point is assigned, and the purpose of the present disclosure can also be achieved. Of course, as long as various methods that can satisfy the result that the assigned score decreases linearly with the increase of the load amount can achieve the technical effect of the present disclosure. Optionally, for a specific purpose, a non-linear decrease or change method can also be used to assign negative scores. This means that for each request received by a processing instance, the increased score will be added to the total load score of the processing instance. Note that this total load score is a negative score, that is, the more requests a processing instance receives, the less the total load score of the processing instance, and new requests will tend to enter other processing instances.

[0056] In addition, it should be noted that the total load score of the data processing requests being processed by the data processing instance gradually increases as the tasks arranged therein are completed. That is, for each request that ends from the processing instance, the corresponding score will be deducted from the total load score corresponding to the processing instance. (Because the score is a negative score, the score value of the processing instance increases when the request ends). At the same time, due to some special requests, it is possible that a request will not end normally from the processing instance (for example, it is cancelled), so the processing instance clears requests whose duration exceeds a certain time (temporarily set to 80 seconds. This is sufficient for short requests. For scenarios that require generating long requests, or models that run relatively slowly, such as Deepseek-R1, it can be set to 10 minutes or 30 minutes). That is to say, the load information acquisition component 130 will obtain the situation of the completion or cancellation of the data processing requests fed back by each processing instance, or in other words, the load information acquisition component 130 will obtain the number and respective lengths of the data processing requests being processed by each processing instance. The load information acquisition component 130 will also obtain the survival status of each processing instance.

[0057] Obviously, the present disclosure can score only considering the hit rate to obtain the total hit rate score as the final score for subsequent forwarding processing, or can score only considering the load balancing to obtain the total load amount score for forwarding processing, or can comprehensively score considering both the hit rate and the load balancing to obtain the sum of the total hit rate score and the total load amount score for forwarding processing, that is, sum the total hit rate score and the total load amount score for each destination processing instance as the total score of the current data request for each destination processing instance.

[0058] Finally, the forwarding component 150 routes and forwards the current data processing request to the destination processing instance with the highest score or one of the multiple highest scores. The specific forwarding process will not be elaborated here, and conventional forwarding technical means can be adopted. The destination processing instance is a processing instance for pre-filling processing or decoding processing, such as Figure 1 the processing instances 1-n in the pre-filling or decoding component 10 of the kind.

[0059] Figure 2 is a block diagram of a second embodiment of a load balancing routing system 200 of a large model shown according to an exemplary embodiment. Compared with Figure 1 the load balancing routing system 100 of the large model shown in the kind, the pre-filling or decoding component 10 is divided into two independent pre-filling components 20 or decoding service components 30, and the reference signs of the other identical parts are the same as Figure 1 those in it, and the difference is only that they start with the number "2". Therefore, the description of the same parts adopts the description for Figure 1 it, and will not be elaborated one by one here.

[0060] Since pre-filling calculation is computationally intensive, it can be executed on a machine with stronger computing power (metric tflops). Decoding is memory access (bandwidth) limited and can be placed on a machine with a larger video memory bandwidth. Therefore, in order to split the two stages of pre-filling calculation and decoding calculation to different machines for execution without interference, so as to achieve the purpose of optimizing the first token latency and the token-to-token latency, as Figure 2As shown, the prefill component 20 and the decoding service component 30 are separately deployed. For example, the prefill component 20 is deployed on the first computing device, and the decoding service component 30 is deployed on the second computing device. Optionally, they can also be deployed on different physical parts of the same computing device. Therefore, from the perspective of the large model, the traditionally monolithic large model is divided into a sub-model for performing prefill inference and a sub-model for performing decoding inference. Through this separate deployment, the situation where one has to neglect the other due to the priority consideration between prefill inference and decoding inference is eliminated, that is, the defect that the prefill priority strategy has a greater impact on decoding requests, while the decoding priority strategy results in too small an input token number, reduced GPU utilization efficiency, and a significant increase in the first token latency, affecting the user experience, is eliminated. Therefore, the defect that the prefill priority strategy or the decoding priority strategy will increase the inference service latency is also eliminated. Similarly, the defect that the total latency of the Chunked Prefill strategy is still affected by the insertion of prefill requests, delaying the processing speed of decoding requests, and restricting the concurrency of decoding requests because the prefill requests still occupy the video memory for a long time is also eliminated.

[0061] Therefore, in the case of separately deploying the prefill component 20 and the decoding service component 30, although it is still possible to route according to the Figure 1 shown routing system, after the forwarding component 250 receives the total hit rate score, the total load score, and the total score of the sum of the two given by the scoring component 240, it can also judge the length of the request based on the total hit rate score. Specifically, when the total hit rate score for any destination processing instance by the forwarding component 250 is lower than the predetermined threshold of the total hit rate score, it directly selects the destination processing instance with the highest total score in the prefill component 20 for forwarding, while when there is a total hit rate score for a certain destination processing instance that is higher than the predetermined threshold of the total hit rate score, the forwarding component 250 directly selects the destination processing instance with the highest total score in the decoding service component for forwarding.

[0062] Requests where the total hit rate score is judged to be lower than the predetermined threshold of the total hit rate score are called long requests, which usually belong to relatively new requests, and there are relatively few context identifiers related to long requests cached in the entire model. Since the decoding stage is a process where the model generates output tokens one by one, with the characteristic of serialization, frequent memory access (such as loading model parameters and KV caches) is required each time a token is generated. Due to the relatively small amount of computation and the underutilization of the parallel computing power of the GPU, the performance bottleneck lies mainly in the memory bandwidth. Therefore, the decoding stage is a typical "memory access limited" stage. For this reason, it is a better solution to first route long requests to the prefill component 20 for prefilling and then perform decoding, and it is also a better solution to directly route short requests to the decoding service component 30. Therefore, in order to determine the length of a request, in addition to the length of the request text itself, it is also necessary to consider whether the context representation cached by the trained model exists or most of it already exists for the request. If most of it already exists, it means that a request with a long text itself is a relatively short request for the large model because only the remaining part of the tokens needs to be decoded. Therefore, the relative length under the existing data of the large model can be determined based on the tokens cached by the model. The predetermined threshold of the total hit rate score here can be 50%, 60%, or 70%. Usually, setting it to 60% can achieve a compromise state of efficiency. If a higher predetermined threshold of the total hit rate score is needed for judgment, the second way of dividing the token information of the request can be adopted, so that when calculating the score, the matching degree of more than half of the request will exceed 50%.

[0063] Such as Figure 2As shown, the prefill component 20 performs prefill processing on the long requests forwarded from the forwarding component 250 by its processing instance, generates an initial context representation (KV cache), and its processing instance updates the information including its own cache, load, and survival status in real time to the load information collection component 230. At the same time, it transmits the original long request and the generated initial context to the decoding service component 30. The decoding service component 30, by its processing instance, performs decoding processing based on the short requests forwarded from the forwarding component 250 and the context associated with the short requests that has been cached, generates output tokens corresponding to the short requests, and performs decoding processing based on the original long requests transmitted from the prefill component 20 via a communication component (not shown) and the initial context generated for the long requests, generates output tokens corresponding to the long requests, and its processing instance updates the information including its own cache, load, and survival status in real time to the load information collection component 230. The communication component between the decoding service component 30 and the prefill component 20 is an RDMA communication component. KVCache is a cache mechanism for storing the intermediate states generated during the inference process of the Transformer model. In the Transformer model, for each inference step (whether in the prefill stage or the decoding stage), corresponding Keys and Values need to be generated based on the current input, and these intermediate states will be reused in subsequent inference steps. The role of KVCache is to cache these intermediate states to avoid repeated calculations during each inference, thereby significantly improving the inference efficiency. In the case where prefill and decoding are separated in the present disclosure, after the prefill instance completes the prefill stage, it is necessary to transmit the KVCache to the decoding instance. The core goal of the KV Cache transmission module is to achieve high-speed and low-overhead data transmission to ensure that the performance advantages of the inference system in the prefill and decoding separation mode can be fully demonstrated. Therefore, RDMA is adopted. RDMA is a high-performance network transmission technology that allows network devices to directly access the memory of remote hosts without the CPU participating in data copying, thereby significantly reducing transmission latency and CPU overhead. The implementation of this RDMA communication can be achieved by setting up the network hardware of the computing devices where the prefill component and the decoding service component are deployed with a conventional RDMA. Optionally, zero-copy technology can also be used to further reduce transmission latency by directly transmitting the KVCache from the memory of the Prefill node to the memory of the Decode node by avoiding multiple copies between the kernel space and the user space, or an efficient network protocol can be used to ensure low-latency and high-bandwidth data transmission in a high-speed network environment by using a high-performance network protocol (such as RoCEv2 or InfiniBand).The communication component can also be implemented by using the RDMA communication ports of the computing devices respectively deployed by the decoding service component 30 and the pre-filling component 20.

[0064] When the hit rate total score for a certain destination processing instance in the forwarding component 250 is higher than the hit rate total score predetermined threshold, the forwarding component 250 directly selects the destination processing instance with the highest total score in the decoding service component for forwarding. This means that, regardless of whether the new processing request is formally long or short, most of it has already been processed by the processing instance of a certain destination. Therefore, having the destination processing instance process it again will save a lot of computing and memory. For the entire data processing system, the load will be reduced and the processing speed of the new request will be accelerated.

[0065] Although the token of the request can be directly used for matching to evaluate the hit rate, using the hash value of the split token subset obtained by splitting the request as described above to evaluate the length of the request by matching the hit rate will make the evaluation of the hit rate faster and more efficient, and load balancing can also be performed inside the pre-filling component 20 or the decoding service component 30. Specifically, after the forwarding component 250 forwards each received data processing request to the pre-filling component 20 or the decoding service component 30, the load information collection component 230 collects and caches the hash values of the data processing requests routed and forwarded by the destination data processing instances from the forwarding component, and calculates the matching hit rate scores between the hash value corresponding to the current data processing request and the hash values of the data processing requests processed by each data processing instance. Then, based on the calculated matching hit rate scores, the forwarding component 250 routes and forwards the currently received data processing request to the data processing instance with the highest matching hit rate score in the pre-filling component 20 or the decoding service component 30. Generally speaking, it is to allocate the new request to the processing instance that has processed a request highly similar to the new request as much as possible, so as to make full use of the calculated results that have been carried out during the processing of the new processing request and reduce the amount of calculation.

[0066] Optionally, when both the pre-filling component 20 and the decoding service component 30 are separated, the request delivery can also be performed after scoring while only considering load balancing. After the forwarding component 250 forwards each received data processing request to the pre-filling component or the decoding service component, the load information collection component 230 collects and caches the hash values of the data processing requests routed and forwarded from the forwarding component to the destination data processing instance, so as to count the load quantity score of the data processing requests routed and forwarded from the routing component being processed by the destination data processing instance. And the forwarding component 250 routes and forwards the currently received data processing request to the data processing instance with the highest load quantity score in the pre-filling component 20 based on the counted load quantity score. The higher the number of data processing requests being processed by the data processing instance, the lower the load quantity score. At this time, the length of the data request is not considered.

[0067] Optionally, when both the pre-filling component 20 and the decoding service component 30 are separated, after the forwarding component 250 receives the total hit rate score, the total load quantity score, and the total score of the sum of the two given by the scoring component 240, in the case of performing request delivery after comprehensive scoring while considering both the matching hit rate and load balancing, after judging the length of the request based on the total hit rate score, the two scores can be directly added, or from the perspective of preferentially matching the hit rate score, the perspective of preferentially matching the load score, or the perspective of no preference, a weighted method can be used to obtain a weighted sum score to select the processing instance to be routed. That is, if the hit rate score of the first processing instance is 70 and the load quantity score is -30, and the hit rate score of the second processing instance is 60 and the load quantity score is -10, since the comprehensive score of the first processing instance is 40 and the comprehensive score of the second processing instance is 50, the currently received data processing request will be routed by the forwarding component 250 to the second processing instance for data processing.

[0068] Optionally, when both the pre-filling component 20 and the decoding service component 30 are separated, a SmoothQuant W8A8 smoothing quantization component (not shown) can be deployed in association with the pre-filling component 20. The activation values and weights accessed by the pre-filling component 20 are rescaled using a smoothing factor, and the smoothed weights and activation values are quantized and converted into the INT8 numerical format for the pre-filling component 20 to perform pre-filling processing. Correspondingly, when both the pre-filling component 20 and the decoding service component 30 are separated, a WeightOnly quantization component (not shown) can be deployed in association with the decoding service component 30. The weight values to be accessed by the decoding service component 30 are quantized and converted into the INT4 numerical format for the decoding service component 30 to perform decoding processing.

[0069] In traditional large language model inference, by quantifying the weights, the model's video memory overhead can be significantly reduced. Correspondingly, the overhead of the model reading weights is reduced, thereby accelerating the inference. Currently, the quantization schemes for large models are mainly divided into two categories: WeightOnly quantization only quantifies the weight part of the model without quantifying the activation values, reducing the model's storage overhead. During the model inference process, the weights are first dequantized and then normal-precision operations are performed with the activation values; this scheme can usually quantize the weights to 8 bits or 4 bits. SmoothQuant W8A8 quantization rescales the activations and weights through a smoothing factor, reduces the scale of outliers in the activations, and performs fixed-point operations during inference, while reducing both storage and computational overhead. However, in the quantization schemes adopted in the traditional single-instance deployment (i.e., the traditional unified deployment of prefill and decoding), it is difficult to coordinate the quantization of the prefill and decoding scenarios. Specifically, the prefill instance has a relatively large computational load due to the large number of tokens processed, and it is a computation-limited scenario. From an optimization perspective, a quantization scheme with higher computational efficiency should be selected. For example, the SmoothQuant W8A8 quantization strategy; while the decoding instance processes a smaller number of tokens, has a relatively low computational load, and needs to frequently read the KVCache from memory, which is a memory-access-limited scenario. A quantization scheme with higher memory-access efficiency should be selected, such as the WeightOnly quantization scheme. Moreover, the WeightOnly scheme needs to dequantize the weights back to the FP16 data type and perform FP16 matrix multiplication, and its computational efficiency is often lower than that of SmoothQuant when used in the prefill stage. Therefore, in the case of unified hybrid deployment of the prefill and decoding scenarios, it is difficult to reconcile and achieve a satisfactory quantization effect for both, resulting in more workload for later adjustment and optimization, and instead losing the meaning of later optimization.

[0070] To this end, the prefill and decoding two-scenario separation method of the present disclosure provides respective quantization strategies for the two scenarios and achieves the respective quantization optimization effects, thus achieving the purpose. Therefore, on the one hand, through the SmoothQuant W8A8 quantization component, the activation values and weights used to access the prefill component 20 are rescaled, and the smoothed weights and activation values are quantized and converted into the INT8 numerical format, which enables the prefill component 20 to make full use of the INT8 tensor core in the hardware to perform INT8 matrix operations. The computational intensity is often twice that of the unquantized FP16 matrix operation, accelerating the prefill calculation process. On the other hand, since the decoding service component 30 processes a small number of tokens, its computational load is not large, and it needs to frequently read the KVCache from memory. After the weight values to be accessed by the decoding service component 30 are quantized and converted into the INT4 numerical format through the WeightOnly quantization component, they are used for the decoding service component 30 to perform decoding processing. The weight reading amount is half of the SmoothQuant W8A8 quantization strategy, and it has better decoding performance.

[0071] Therefore, after adopting the prefill and decoding two-scenario separation method, it is possible to freely select appropriate quantization schemes for the prefill and decoding scenarios respectively. For the prefill component 20, by associatively deploying the more computationally efficient SmoothQuant W8A8 quantization component, better Prefill performance can be obtained. For the decoding service component 30, by associatively deploying the WeightOnly 4bit quantization component, higher memory access efficiency can be obtained, thereby making the performance of the overall inference system better than that of the inference system in the case of unified hybrid deployment.

[0072] Figure 3 It is a flowchart of the first embodiment of the load balancing routing method 300 of a large model shown according to an exemplary embodiment. As Figure 3 shown, in the load balancing routing method 300 of a large model, after the user inputs the current data processing request, at step S310, the request parsing component 110 or 210 receives the user's data processing request to obtain the token information of the data processing request. In a large language model (LLM), a token is the basic unit of text and can be a word, subword, or character. The model decomposes the input text into tokens for processing and understanding. The number of tokens determines the input length and computational complexity of the model. The text data input by the user will first be converted into tokens and input into the model. The request parsing component 110 parses the current data processing request into one or more tokens.

[0073] Next, at step S320, the subset splitting component 120 or 220 divides one or more tokens of each data processing request into a token subset starting from the head of the token information of the received data processing request each time based on a predetermined number of tokens until the end of the token information is reached, and generates a hash value for each divided token subset using a hash function, thereby forming one or more hash value sets of the data processing request. When the subset splitting component 120 divides the token subset, in two adjacent divisions, the number of the token subsets divided in the latter division is twice the number of the token subsets divided in the former division or the number of the token subsets divided in the latter division increases by one base number relative to the number of the token subsets divided in the former division.

[0074] After the division, at step S330, the scoring component 140 or 240 scores each destination processing instance based on the matching degree between the hash value set corresponding to the current data processing request and the hash value sets corresponding to the data processing requests received by the load information collection component for each destination processing instance and the load of each destination processing instance. At step S350, the load information collection component collects, for each destination processing instance, the hash value set of each data processing request forwarded to it by the forwarding component and the load of the data processing requests being processed by each destination processing instance. Although step S350 is shown here to be after step S330 in the described order, due to the cyclic nature of the entire data processing, for any new data processing request, step S350 is before step S330. Finally, at step S340, the forwarding component 150 or 250 routes and forwards the current data processing request to the destination processing instance with the highest score or one of the multiple highest scores. The specific processing procedures for each step can be referred to the content described for Figure 1 and 2 and will not be elaborated one by one.

[0075] Figure 4 is a flowchart of a second embodiment of the load balancing routing method 400 of a large model shown according to an exemplary embodiment. Figure 4 Same steps as Figure 3 are numbered with basically the same step numbers, and the difference in the numbers is that “S4” is used as the beginning instead of Figure 3"S3" in it. The only difference between the steps of the two is that step S431 is added between step S430 and step S440, step S340 is divided into step S441 and step S442, and step S360 is divided into step S461 and step S462. Other identical steps will not be elaborated here one by one.

[0076] As Figure 4 shown, after scoring each destination processing instance at step S430 by the scoring component 140 or 240 based on the matching degree between the hash value set corresponding to the current data processing request and the hash value set corresponding to the data processing requests received by the load information collection component for each destination processing instance, and the load of each destination processing instance, at step S431, when the total hit rate score for any one destination processing instance by the forwarding component 250 is lower than the predetermined threshold of the total hit rate score ("Yes"), the forwarding component 250 directly selects the destination processing instance with the highest total score in the pre-filling component 20 for forwarding, while when there is a destination processing instance for which the total hit rate score is higher than the predetermined threshold of the total hit rate score, the forwarding component 250 directly selects the destination processing instance with the highest total score in the decoding service component 30 for forwarding.

[0077] At step S431, after directly selecting the destination processing instance with the highest total score in the pre-filling component 20 for forwarding, at step S441, the processing instance with the highest routing score in the pre-filling component 20 is selected as the destination for routing and forwarding. It should be noted that in order to perform the subsequent pre-filling process, the maximum number of tokens for the long request is set to 1. Subsequently, at step S461, the pre-filling component 20 pre-fills the long request forwarded from the forwarding component 250 by its processing instance to generate an initial context representation (KV cache), and its processing instance updates the information including its own cache, load, and survival status to the load information collection component 230 in real time, and at the same time transmits the original long request and the generated initial context to the decoding service component 30. Then at step S462, the decoding service component 30 performs decoding processing based on the original long request transmitted from the pre-filling component 20 via the communication component (not shown) and the initial context generated for the long request to generate output tokens corresponding to the long request, and its processing instance updates the information including its own cache, load, and survival status to the load information collection component 230 in real time. The communication component between the decoding service component 30 and the pre-filling component 20 is an RDMA communication component, which will not be elaborated here.

[0078] Similarly, at step S431, when the hit rate total score for a certain destination processing instance in the forwarding component 250 is higher than the hit rate total score predetermined threshold, after directly selecting the destination processing instance with the highest total score in the decoding service component 30 for forwarding, then at step S462, the decoding service component 30 performs decoding processing by its processing instance based on the short request forwarded from the forwarding component 250 and the context already cached and associated with the short request, and generates an output token corresponding to the short request. When the hit rate total score for a certain destination processing instance in the forwarding component 250 is higher than the hit rate total score predetermined threshold and it directly selects the destination processing instance with the highest total score in the decoding service component for forwarding, it means that regardless of whether the new processing request is formally long or short, most of it has already been processed by the processing instance of a certain destination. Therefore, having the destination processing instance process it again will save a lot of computing and memory. For the entire data processing system, the load will be reduced and the processing speed of the new request will be accelerated.

[0079] In summary, compared with traditional routing mechanisms that usually adopt globally fixed routing, the systems and methods of the present disclosure select specific parallel processing instances for resource optimization, parallel processing, and load balancing by comprehensively considering the repeatable computation amount of each parallel processing instance and the data request load being processed, significantly improving the inference speed and reducing resource consumption to meet the requirements of different models and tasks. In other words, by introducing a cache-aware routing mechanism, the present disclosure polls each node (or processing instance), records the current load situation, and the routing system automatically distributes requests according to the load conditions of each node, making the loads of each node relatively balanced and stable as much as possible to fully utilize all node resources, enabling load-balanced parallelism for the processing instances of padding calculation and decoding calculation respectively. Moreover, the present disclosure also splits the two stages of pre-padding calculation and decoding calculation into different machines for execution without interference, achieving the purpose of optimizing the first token latency and the inter-token latency. Based on the separation of pre-padding calculation and decoding calculation, by introducing a cache-aware routing mechanism, it polls each node (or processing instance), records the current load situation, and the routing component automatically distributes requests according to the load conditions of each node, making the loads of each node relatively balanced and stable as much as possible to fully utilize all node resources, enabling load-balanced parallelism for the processing instances of padding calculation and decoding calculation respectively. In addition, based on the structure of separating pre-padding calculation and decoding calculation, the present disclosure introduces the RDMA high-performance network transmission technology. The network device can directly access the host memory for data transmission without the CPU participating in the data copy process, reducing the transmission overhead and significantly reducing the transmission latency. Finally, based on the structure of separating pre-padding calculation and decoding calculation, the present disclosure supports the pre-padding instance and the decoding calculation instance to select appropriate model quantization strategies according to the resource requirements of the inference scenario to improve the utilization rate of computing resources.

[0080] Those skilled in the art can understand that the above-mentioned modules can be distributed in the device according to the description of the embodiments, or can be correspondingly changed and arranged in one or more devices different from the present embodiment. The modules of the above embodiments can be combined into one module, or further split into multiple sub-modules.

[0081] From the description of the above embodiments, those skilled in the art can easily understand that the example embodiments described here can be implemented by software, or by software in combination with necessary hardware. Therefore, the technical solutions according to the embodiments of the present disclosure can be embodied in the form of a software product, which can be stored in a non-volatile storage medium (such as a CD-ROM, USB flash drive, mobile hard disk, etc.) or on the network, including several instructions to enable a computing device (such as a personal computer, server, mobile terminal, or network device, etc.) to execute the methods according to the embodiments of the present disclosure.

[0082] The exemplary embodiments of the present disclosure have been specifically shown and described above. It should be understood that the present disclosure is not limited to the detailed structures, arrangements or implementation methods described herein; on the contrary, the present disclosure is intended to cover various modifications and equivalent arrangements included within the spirit and scope of the appended claims.

Claims

1. A large-model load balancing routing system, comprising: Request parsing component, receives the user's data processing request to obtain the token information of the data processing request; A subset partitioning component, taking a predetermined number of tokens as a base, divides a token subset from the head of the token information of the received data processing request each time until the end of the token information is reached, and generates a hash value of each divided token subset by using a hash function, thereby forming one or more hash value sets of the data processing request; A load information collection component collects and caches, for each destination processing instance, a hash value set of each data processing request forwarded to it by the forwarding component and the load of the data processing request being processed by each destination processing instance; A scoring component scores each destination processing instance based on the matching degree between the hash value set corresponding to the current data processing request and the hash value set corresponding to the data processing request received by each destination processing instance collected by the load information collection component and the load of each destination processing instance; as well as The forwarding component routes the current data processing request to a destination processing instance with the highest score or multiple highest scores.

2. A large-model load-balancing routing system as described in claim 1, wherein when the subset splitting component divides the token subsets, in two adjacent divisions, the number of token subsets divided in the latter division is twice the number of token subsets divided in the former division or the number of token subsets divided in the latter division is increased by a cardinality relative to the number of token subsets divided in the former division.

3. A large-model load balancing routing system as described in claim 1 or 2, wherein the load information collection component allocates a memory of a predetermined length for each destination processing instance, and when the memory is full, a first-in-first-out method is used to replace the hash value cached in the memory.

4. The load balancing routing system of the large model as described in claim 1, wherein the scoring component scores the matching degree and the load volume respectively during the scoring, and the matching degree is scored by calculating the ratio of the number of matches between the hash value of the current data processing request and the hash value cached by the load information collection component for the destination processing instance and the number of hash value sets of the current data processing request as the matching hit rate for any destination processing instance, and assigning a corresponding hit rate score according to the multiple of the cardinality contained in the partition page corresponding to the matched hit hash value, thereby accumulating the current data processing request relative to the destination processing instance. All hit rate scores are obtained, and the sum of the hit rate scores is obtained as the total hit rate score of the current data processing request for each destination processing instance; and the load is scored by calculating, for any destination processing instance, the negative of a multiple of a preset length for each data processing request being processed by the destination data processing instance as the load score of each data processing request, thereby accumulating the load scores of all data processing requests being processed by the destination data processing instance as the total load score of the destination processing instance, and summing the total hit rate score and the total load score for each destination processing instance as the total score of the current data request for each destination processing instance.

5. A load balancing routing system for a large model as described in claim 4, wherein the forwarding component directly selects the destination processing instance with the highest total score in the pre-filled component for forwarding when the pre-filled component and the decoding service component are separated in the destination processing instance and when the total hit rate score for any destination processing instance is lower than a predetermined threshold of the total hit rate score.

6. A load balancing routing system for a large model as described in claim 4, wherein the forwarding component directly selects the destination processing instance with the highest total score in the decoding service component for forwarding when a pre-filling component and a decoding service component are separated in the destination processing instance, and when the total hit rate score for a certain destination processing instance is higher than a predetermined threshold of the total hit rate score.

7. A load balancing routing system for a large model as described in any one of claims 4-6, wherein the scoring component directly assigns a score of -1 when the length of the data processing request being processed by the destination data processing instance is less than a preset length.

8. A load balancing routing method for a large model, comprising: Receive the user's data processing request through the request parsing component, and parse and obtain the token information of the data processing request; By using a subset partitioning component, a token subset is divided from the head of the token information of the received data processing request each time based on a predetermined number of tokens until the end of the token information is reached, and a hash value of each divided token subset is generated by using a hash function, thereby forming one or more hash value sets of the data processing request; Through the load information collection component, for each destination processing instance, a hash value set of each data processing request forwarded to it by the forwarding component and the load of the data processing request being processed by each destination processing instance are collected; By means of the scoring component, each destination processing instance is scored based on the matching degree between the hash value set corresponding to the current data processing request and the hash value set corresponding to the data processing requests received by each destination processing instance collected by the load information collection component and the load of each destination processing instance; as well as The current data processing request is routed and forwarded to a destination processing instance with the highest score or multiple highest scores through the forwarding component.

9. The load balancing routing method for a large model as described in claim 8, wherein when dividing token subsets through the subset splitting component, in two adjacent divisions, the number of token subsets divided in the latter time is twice the number of token subsets divided in the former time or the number of token subsets divided in the latter time increases by a cardinality relative to the number of token subsets divided in the former time.

10. A load balancing routing method for a large model as described in claim 8 or 9, wherein the load information collection component allocates a memory of a predetermined length to each destination processing instance, and when the memory is full, a first-in-first-out method is used to replace the hash value cached in the memory.

11. The load balancing routing method for a large model as claimed in claim 8, wherein the scoring component comprises: By calculating, for any destination processing instance, the ratio of the number of matches between the hash value of the current data processing request and the hash value cached for the destination processing instance collected by the load information collection component to the number of hash value sets of the current data processing request as the matching hit rate, and assigning a corresponding hit rate score according to the multiple of the cardinality contained in the partition page corresponding to the matched hash value, thereby accumulating all hit rate scores of the current data processing request relative to the destination processing instance, and obtaining the sum of the hit rate scores as the total hit rate score of the current data processing request for each destination processing instance; By calculating, for any destination processing instance, the negative of a multiple of a preset length of each data processing request being processed by the destination data processing instance as the load score of each data processing request, thereby accumulating the load scores of all data processing requests being processed by the destination data processing instance as the total load score of the destination processing instance; as well as The total hit rate score and the total load score of each destination processing instance are summed up as the total score of the current data request for each destination processing instance.

12. The load balancing routing system of a large model as claimed in claim 11, wherein forwarding by the forwarding component further comprises: In the case where the destination processing instance has a pre-filling component and a decoding service component separated, when the total hit rate score for any destination processing instance is lower than the predetermined threshold of the total hit rate score, the destination processing instance with the highest total score in the pre-filling component is directly selected for forwarding.

13. The load balancing routing system of a large model as claimed in claim 11, wherein forwarding by the forwarding component further comprises: In the case where the destination processing instance has a pre-filled component and a decoding service component, when the total hit rate score for a certain destination processing instance is higher than the predetermined threshold of the total hit rate score, the destination processing instance with the highest total score in the decoding service component is directly selected for forwarding.

14. A load balancing routing system for a large model as described in any one of claims 11-13, wherein the scoring component directly assigns a score of -1 when the length of the data processing request being processed by the destination data processing instance is less than a preset length.

Citation Information

Patent Citations

  • High-concurrency workflow arrangement method and device, computer equipment and storage medium

    CN115796766A

  • Load balancing method, device and system, electronic equipment and storage medium

    CN117278562A

  • Model reasoning scheduling method and device and server cluster

    CN118897736A

  • Distribution task adaptive load prediction method and system oriented to edge computing, and medium

    CN119110350A

  • Data storage method based on large model

    CN119292524A

Cited By

  • Large language model reasoning effective throughput optimization method, system, equipment and medium

    CN121996437A

  • Inference service request scheduling system and method and electronic equipment

    CN122317166A