Large language model high concurrency reasoning method and system
By introducing executors and schedulers into the large language model inference system, dynamically allocating video memory and implementing continuous batch processing, the problem of insufficient concurrency and throughput of large language model inference is solved, and the resource utilization and service capabilities are significantly improved.
Patent Information
- Application Number
- CN202510660889.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-22
- Publication Date
- 2025-06-20
- Estimated Expiration
- 2045-05-22
AI Technical Summary
The concurrency and throughput of large language models during the inference stage are insufficient, resulting in low resource utilization and long inference, making it difficult to support multiple users to use model inference services at the same time.
A high concurrency inference method of large language model is adopted. Through the executor, the executor calculates the size of the video memory block and allocates the video memory space. The scheduler converts the request sequence and dynamically allocates the video memory, allocates the video memory blocks according to the type and number of the request sequence, and realizes continuous batch processing and task scheduling to increase the concurrency.
It effectively improves the concurrency and throughput of large-model inference, reduces the inference waiting time, and improves resource utilization and service capabilities.
Smart Images

Figure CN120181245A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of artificial intelligence technology, and in particular, to a high-concurrency inference method and system for large language models. Background Art
[0002] Artificial intelligence technology is a hot topic in the current field of computer science, and large language models are an important research direction in the field of artificial intelligence. As an emerging large language model, its special feature is that the number of parameters far exceeds that of traditional models, usually in the tens of billions, which also leads to higher training and inference costs for the model. For a model with tens of billions of parameters, such as Qwen-14B, its inference speed in the inference stage is about 25-35 tokens / s when the batch size = 1 and two GPUs of model 3090 are used. For a conventional processing scheme (batch size = 1, serial mode), it is difficult to support multiple users to use the model inference service simultaneously.
[0003] For most traditional models (such as image processing models), as long as the batch size is increased under the condition of sufficient video memory, the concurrency of model inference can be improved and the user experience can be optimized. Unfortunately, this method cannot be simply migrated to models with the transformer architecture. The main reason is that the generative model with the transformer architecture has an unfixed output for any input. Therefore, by simply increasing the batch, it is usually necessary to allocate memory space sufficient to accommodate the maximum length output for each input, and wait for the longest sample output in this batch to complete before returning uniformly. This will not only cause a large waste of video memory space, but sometimes it may take a long waiting time to wait for a simple answer. To solve the above problems, more effective space allocation methods and batch processing scheduling methods are needed.
[0004] Therefore, how to improve resource utilization and inference request throughput, reduce inference time consumption, and enhance the capabilities and effects of large model-related services and applications during the inference process of large language models is a technical problem that urgently needs to be solved. Summary of the Invention
[0005] The present invention provides a high-concurrency inference method and system for large language models, aiming to solve at least one of the above technical problems.
[0006] To achieve the above object, the present invention provides a high-concurrency inference method for large language models, including the following steps: S1: System initialization, the executor calculates the video memory block size by simulating peak data, allocates video memory space according to the calculation result, and synchronizes the video memory resource information to the scheduler; S2: The scheduler receives an inference HTTP request, converts the inference HTTP request into a request sequence through preprocessing, and places the obtained request sequence into the waiting queue of the scheduler; S3: When the waiting queue and the running queue are not both empty, the scheduler checks the video memory usage of the current running queue, determines whether each request sequence can perform the next inference. If not, it allocates corresponding video memory blocks for each request sequence until each request sequence can perform the next inference; S4: The scheduler checks the video memory usage of the current running queue, calculates the video memory requirements of the request sequences in the waiting queue in the order of priority, and transfers the request sequences in the waiting queue to the running queue according to the calculation results of the video memory requirements; S5: Define the request sequences transferred from the waiting queue to the running queue in step S4 as the pre-fill type, and define the request sequences in the running queue in step S3 as the decoding type. According to the number of pre-fill type request sequences and the number of decoding type request sequences, allocate the number of video memory blocks for performing pre-fill inference or for performing decoding inference; S6: Call the executor to perform pre-fill inference or decoding inference based on the allocated number of video memory blocks; S7: The scheduler obtains the result returned by the executor, updates the status and information of the request sequence, and changes the type of the request sequence of the pre-fill type; S8: Traverse all the request sequences in the running queue, determine whether each request sequence meets the end condition. If so, remove the request sequence from the running queue, release the video memory, and respond to the HTTP request.
[0007] Optionally, in step S1, the video memory block size is configured as the number of tokens stored in the video memory block.
[0008] Optionally, in step S1, allocating video memory space according to the calculation results further includes: if the video memory is insufficient, directly exit the program, replace with a smaller large language model or reduce the maximum length.
[0009] Optionally, in step S2, the request sequence includes the metadata information and the token sequence of the HTTP request.
[0010] Optionally, in step S3, determining whether each request sequence can perform the next inference. If not, allocating corresponding video memory blocks for each request sequence until each request sequence can perform the next inference further includes: Determine whether the video memory of the current running queue can support each request sequence to perform the next inference. If not, allocate new video memory blocks for the running queue; When the allocated new video memory blocks are insufficient, remove the request sequences in the running queue according to the priority and release the corresponding video memory blocks until each request sequence can perform the next inference.
[0011] Optionally, in step S4, calculate the video memory requirements of the request sequences in the waiting queue in the order of priority, and transfer the request sequences in the waiting queue to the running queue according to the calculation result of the video memory requirements, specifically including: Calculate the video memory requirements of the request sequences in the waiting queue in the order of priority; wherein, the priority order is configured as the arrival time order of each http request corresponding to the request sequence; If the video memory is sufficient, put the requests in the waiting queue into the running queue until the video memory is insufficient or the threshold conditions of the maximum batch character length and the maximum batch quantity are reached.
[0012] Optionally, in step S5, allocate the number of video memory blocks for performing pre-fill inference or for performing decoding inference according to the number of pre-fill types and the number of decoding types of the request sequences, specifically including: If the number of request sequences of the pre-fill type is greater than or equal to the number of request sequences of the decoding type, perform pre-fill inference and allocate corresponding numbers of video memory blocks for each request sequence of the pre-fill type; If the number of request sequences of the pre-fill type is less than the number of request sequences of the decoding type, put the request sequences of the pre-fill type back into the waiting queue, perform decoding inference, and allocate the required video memory blocks for each request sequence of the decoding type.
[0013] Optionally, in step S6, perform pre-fill inference or decoding inference, specifically including: If performing pre-fill inference, allocate video memory resources for all pre-filled request sequences and only send the pre-filled request sequences during inference; If performing decoding inference, put all pre-filled request sequences back into the waiting queue, allocate video memory for the next token for each decoded request sequence, and send all request sequences for inference.
[0014] Optionally, in step S7, update the status and information of the request sequences, and perform type change on the request sequences of the pre-fill type, specifically including: Update the status and information of the request sequences; If performing pre-fill inference, change all request sequences of the pre-fill type to the decoding type.
[0015] In addition, to achieve the above object, the present invention also provides a large language model high-concurrency inference system, including: an executor and a scheduler, and the executor and the scheduler are configured to execute the large language model high-concurrency inference method described in any one of the above.
[0016] The beneficial effects of the present invention are as follows: A high-concurrency inference method and system for large language models are proposed. The executor calculates the size of the video memory block to allocate video memory space; the scheduler converts the request sequence and puts it into the waiting queue of the scheduler; the scheduler allocates corresponding video memory blocks to each request sequence until each request sequence can perform the next inference; the scheduler calculates the video memory requirements of the request sequences in the waiting queue in the order of priority and transfers the request sequences in the waiting queue to the running queue; according to the number of pre-filled types and the number of decoding types of the request sequence, the number of video memory blocks used for performing pre-filled inference or for performing decoding inference is allocated; thus, the present invention adopts a continuous batch processing, dynamic space allocation mechanism and task scheduling framework, makes full use of the parallel inference ability of continuous batch processing, improves the concurrency and throughput of large model inference, and solves the limitation of traditional continuous batch processing that requires pre-allocation of space. BRIEF DESCRIPTION OF THE DRAWINGS
[0017] Figure 1 It is a schematic flowchart of the high-concurrency inference method for large language models in an embodiment of the present invention; Figure 2 It is a schematic diagram of the principle of converting traditional batch processing into continuous batch processing in an embodiment of the present invention; Figure 3 It is a schematic diagram of the situation of insufficient video memory in an embodiment of the present invention; Figure 4 It is a schematic structural diagram of the high-concurrency inference system for large language models in an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0018] In order to make the objectives, technical solutions and advantages of the present invention clearer, the present invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention.
[0019] An embodiment of the present invention provides a high-concurrency inference method for large language models, referring to Figure 1 , Figure 1 It is a schematic flowchart of the high-concurrency inference method for large language models in an embodiment of the present invention.
[0020] In this embodiment, the high-concurrency inference method for large language models includes the following steps: S1: System initialization. The executor calculates the size of the video memory block by simulating peak data, allocates video memory space according to the calculation result, and synchronizes the video memory resource information to the scheduler; S2: The scheduler receives the inference http request, converts the inference http request into a request sequence through preprocessing, and puts the obtained request sequence into the waiting queue of the scheduler; S3: When the waiting queue and the running queue are not all empty, the scheduler checks the video memory usage of the current running queue, determines whether each request sequence can perform the next inference. If not, it allocates corresponding video memory blocks for each request sequence until each request sequence can perform the next inference; S4: The scheduler checks the video memory usage of the current running queue, calculates the video memory requirements of the request sequences in the waiting queue according to the priority order, and transfers the request sequences in the waiting queue to the running queue according to the calculation results of the video memory requirements; S5: Define the request sequences transferred from the waiting queue to the running queue in step S4 as the pre-filling type, and define the request sequences in the running queue in step S3 as the decoding type. According to the number of pre-filling type request sequences and the number of decoding type request sequences, allocate the number of video memory blocks for performing pre-filling inference or for performing decoding inference; S6: Call the executor to perform pre-filling inference or decoding inference based on the allocated number of video memory blocks; S7: The scheduler obtains the return result of the executor, updates the status and information of the request sequence, and changes the type of the request sequence of the pre-filling type; S8: Traverse all request sequences in the running queue, determine whether each request sequence meets the end condition. If so, remove the request sequence from the running queue, release the video memory, and respond to the http request.
[0021] In this embodiment, the executor is used to calculate the video memory block size to allocate the video memory space; the scheduler is used to transfer the request sequence and put it into the waiting queue of the scheduler; the scheduler allocates corresponding video memory blocks for each request sequence until each request sequence can perform the next inference; the scheduler calculates the video memory requirements of the request sequences in the waiting queue according to the priority order, and transfers the request sequences in the waiting queue to the running queue; according to the number of pre-filling type request sequences and the number of decoding type request sequences, allocate the number of video memory blocks for performing pre-filling inference or for performing decoding inference. The present invention adopts a continuous batch processing, dynamic space allocation mechanism and task scheduling framework, makes full use of the parallel inference ability of continuous batch processing, improves the concurrency and throughput of large model inference, and solves the limitation that traditional continuous batch processing requires pre-allocation of space.
[0022] To more clearly explain the present invention, the following provides a specific example of the high-concurrency inference method for the large language model of the present invention.
[0023] In traditional deep learning models, to increase the concurrency of the model, batch processing techniques are usually used. Leveraging the parallel computing power of the GPU, multiple requests are merged and processed simultaneously for the calculation results. The batch processing technique generally follows the following steps: (1) Instead of performing calculations immediately when requests arrive, wait for a certain period of time or until the number of requests reaches a specified quantity; (2) Concatenate multiple request tensors (Tensors) into one Tensor; (3) Pre-allocate video memory for each request, sufficient to accommodate the model inference results; (4) Use the forward propagation of the model to calculate the inference results; (5) Split the results and distribute them back to each request. The above batch processing method is only applicable to traditional discriminative models and is not suitable for direct migration and use in existing transformer architecture generative models. The reasons are as follows: First, current generative models have variable-length result outputs for each model input. For example, the model may output dozens of characters, or it may output thousands of characters until the maximum length is reached, and there is no judgment basis. Therefore, it is necessary to allocate video memory space required for the maximum length for each request to prevent errors due to insufficient video memory, resulting in low video memory utilization. Second, the variable-length output results cause requests with short outputs in the same batch to wait until the request with the longest output in the batch is calculated and completed before being returned in a batch, wasting a lot of time on unnecessary waiting. Finally, in models with a transformer architecture, it is usually divided into a prefill stage and a decode stage, and the calculation modes of the two stages are different. In the prefill stage, it is a parallel calculation once, and the Key and Value values of all characters (tokens: the unit of the model) are cached for use in the decode stage. This step uses a space-for-time method to reduce the calculation overhead of inference. The decode stage then uses the cached Key and Value that have been calculated, continuously iterates a new token and appends and caches this token until an end condition appears. The two stages cannot be simply placed in one batch for calculation.
[0024] Many problems caused by these reasons are the main factors restricting the model inference speed and concurrency. To solve the above problems, the present invention implements an efficient parallel inference method, as Figure 2 shown, and the key steps are as follows: (1) Use the continuous batch method to replace the traditional batch method; (2) In continuous batch processing, generate only one token and return it for each inference; (3) Use the virtual block caching scheme to allocate video memory; (4) Dynamically allocate video memory space, and each request only allocates the video memory required for the current inference; (5) Separate the execution of the prefill stage and the decode stage, that is, requests in the prefill stage and requests in the decode stage cannot exist simultaneously in one execution batch.
[0025] An efficient parallel inference method proposed by the present invention adopts the following specific execution process: Step 1: Module initialization. The executor calculates the size of the video memory block through simulated peak data. If the video memory is insufficient, the program exits directly, and a smaller model needs to be replaced or the maximum length needs to be reduced. After successful calculation, allocate video memory space and synchronize relevant video memory resource information to the scheduler; Step 2: The scheduler receives an inference http request, converts it into a sequence through preprocessing, which includes metadata information of the request, token sequence, etc., and puts it into the waiting queue of the scheduler; Step 3: When the waiting queue and the running queue are not all empty, the scheduler checks the video memory usage of the current running queue. Confirm whether each sequence can perform the next inference. If a new video memory block needs to be allocated, whether there are enough video memory blocks. If the video memory blocks are insufficient, remove the sequence in the running queue according to the priority and release the corresponding video memory blocks until each sequence can perform the next inference; Step 4: The scheduler checks the video memory usage of the current running queue, calculates the video memory requirements of the sequences in the waiting queue in the order of priority. If the video memory is sufficient, put the requests in the waiting queue into the running queue until the video memory is insufficient or threshold conditions such as the maximum batch character length and the maximum batch quantity are reached; Step 5: The sequence types transferred from the waiting queue to the running queue in Step 4 are all set to prefill; the sequence types in the running queue in Step 3 are all set to decode. If the number of prefill type sequences is greater than or equal to the number of decode type sequences, perform prefill inference and allocate corresponding numbers of video memory blocks for each prefill type sequence; if the number of prefill type sequences is less than the number of decode type sequences, put the prefill type sequences back into the waiting queue, perform decode inference, and allocate the required video memory blocks for each decode type sequence; Step 6: Call the executor to perform prefill inference or decode inference; Step 7: The scheduler obtains the return result of the executor, updates the status and information of the sequence. If prefill inference is performed, all prefill type sequences also need to be changed to decode type; Step 8: Traverse all sequences in the running queue. If a sequence meets the end condition, remove it from the running queue, release the video memory, and respond to the HTTP request.
[0026] A method for improving the inference concurrency of large language models provided by the present invention can achieve the effects of shorter time consumption, higher resource utilization rate, and increased inference request throughput, effectively reducing the time overhead of inference waiting, and enhancing the capabilities and effects of large model-related services and applications.
[0027] The overall solution is divided into two parts, a scheduler and an executor. Parameters such as the video memory block size (block_size), the maximum number of characters per batch (max_token_for_batch), and the maximum batch size (max_batch_size) need to be specified. The video memory block size refers to how many tokens can be stored in a video memory block; the maximum number of characters per batch refers to the maximum number of tokens that all sequences in the prefill stage of a batch do not exceed; the maximum batch size refers to the upper limit of the number of sequences that can exist in a batch.
[0028] (1) The scheduler is responsible for receiving external HTTP requests, maintaining two queues (waiting, running), maintaining a video memory block mapping table, and allocating video memory resources. The specific process is as follows: 1. When an external request arrives, it is preprocessed into a request sequence (sequence) that can be processed in the module and placed in the waiting queue. One request is one sequence; 2. Whenever the running queue finishes an inference, ensure that all sequences in the running queue have enough space to accommodate the video memory required for the next inference generation. If the video memory is insufficient, move the sequences in the running queue to the waiting queue according to the priority and release the video memory blocks occupied by these sequences until the video memory is sufficient; After that, calculate the video memory requirements for the sequences in the waiting queue according to the priority. If the video memory is sufficient, move the sequences in the waiting queue to the running queue until the video memory is insufficient or reaches threshold conditions such as the maximum number of characters per batch length (max_token_for_batch) (note that all sequences in the waiting queue have not passed the prefill stage); 3. After the scheduling is completed, select whether to perform prefill inference or decode inference based on the number of prefill stages and the number of decode stages. If the number of prefill stages >= the number of decode stages, enter prefill; otherwise, enter decode. 4. If prefill inference is performed, allocate video memory resources for all sequences in the prefill stage, and only send the sequences in the prefill stage during inference. If decode inference is performed, put all sequences in the prefill stage back into the waiting queue, allocate video memory for the next token of each sequence in the decode stage, and send all sequences for inference. 5. Call the executor, return the result, and update the status and information of all sequences in the running queue. 6. For sequences that meet the end condition, remove them from the running queue, release the video memory block, and return the http request. Among them, the priority is sorted according to the request arrival time, and the earlier the sequence arrives, the higher the priority. The end condition is usually reaching the specified maximum length or the appearance of an end flag. The situation of insufficient video memory is shown in the appendix Figure 3 as shown.
[0029] It should be noted that according to the set block_size, when performing the next inference, it is not necessary to allocate a new video memory block every time, but only when the space in the video memory block is used up. For example, if block_size = 4, it means that a new video memory block needs to be allocated every 4 rounds of inference for a sequence.
[0030] (2) The executor is responsible for performing the tasks of model inference. The main process is as follows: 1. Initialize the number of computing video memory block resources and synchronize the video memory block information to the scheduler. 2. Determine whether it is the decode stage or the prefill stage, and call the model for forward inference. 3. If it is the decode stage, perform a decode iteration batch-wise to generate a new token, and cache the key and value of this token. If it is the prefill stage, construct a prefill matrix for each sequence, perform a prefill calculation, and cache the key and value of all tokens in the sequence. The initialization phase is an important phase, mainly including performing simulated inference, calculating the peak of model inference, calculating the number of VRAM blocks and allocating corresponding space, etc., to ensure the stable operation of the system.
[0031] The purpose of each VRAM block is to cache the values that a token needs to participate in the calculation in the model, that is, the key and value values of a token in each layer of the model. Its function is to exchange space for time and reduce the calculation of the key and value values of tokens (if the caching method is not used, the key and value values of all tokens in the sequence need to be recalculated every time inference is performed).
[0032] Among them, the calculation formula of the VRAM block is as follows: In the formula, represents the memory occupancy of the VRAM block, block_size represents the size of the VRAM block, num_layers represents the number of decoder stack layers of the transformer architecture, num_heads represents the number of attention heads, and head_size represents the head dimension. represents the number of bytes after the bit width conversion of the data type; For example, for a model with a 32-layer transformer architecture, 32 heads, a head dimension of 128, and using the fp16 format, when the block_size is set to 4, the memory occupancy of one block is: During initialization, simulated data is constructed through the given maximum batch character length (max_token_for_batch) and maximum batch size (max_batch_size), and the peak memory peak_memory required for inference is calculated, and then the remaining VRAM can be allocated to the Block. For example, when max_token_for_batch is 8192 and max_batch_size is 128, the length of each sentence is 64, and simulated data is constructed to calculate the inference peak, and during subsequent operation, it is ensured that the total number of tokens in each prefill batch does not exceed 8192.
[0033] In the formula, represents the number of VRAM blocks, represents the peak memory, represents the memory occupancy of the VRAM block, / / represents the integer division operator; Finally, the available number of VRAM blocks is obtained according to the above formula, and the block VRAM space is pre-allocated.
[0034] The parallel inference method implemented by the present invention can effectively increase the concurrency of model inference. When using the Qwen-14B model on two 3090 GPUs, the highest generation speed can reach 200 tokens / s, which is more than 5 times the token generation speed compared to the conventional solution.
[0035] Refer to Figure 4 , Figure 4 which is a schematic structural diagram of the high-concurrency inference system for large language models according to an embodiment of the present invention.
[0036] As Figure 4 shown, the high-concurrency inference system for large language models proposed in the embodiment of the present invention includes: an executor and a scheduler, and the executor and the scheduler are configured to execute the high-concurrency inference method for large language models described in any one of the above.
[0037] For other embodiments or specific implementation manners of the high-concurrency inference system for large language models of the present invention, reference may be made to the above method embodiments, and details are not described herein again.
[0038] It can be understood that in the description of this specification, the descriptions with reference to terms such as "one embodiment", "another embodiment", "other embodiments", or "the first embodiment to the Nth embodiment" mean that the specific features, structures, materials, or characteristics described in connection with the embodiment or example are included in at least one embodiment or example of the present invention. In this specification, the schematic representations of the above terms do not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials, or characteristics described can be combined in a suitable manner in any one or more embodiments or examples.
[0039] It should be noted that in this article, the term "comprising", "including" or any other variation thereof is intended to cover a non-exclusive inclusion, so that a process, method, article or system including a series of elements not only includes those elements but also includes other elements not expressly listed, or further includes elements inherent to such process, method, article or system. Without further limitation, the element defined by the statement "including one..." does not exclude the existence of additional identical elements in the process, method, article or system including the element.
[0040] The above are only the preferred embodiments of the present invention, and do not limit the patent scope of the present invention accordingly. Any equivalent structural or equivalent process transformation made by using the content of the specification and drawings of the present invention, or directly or indirectly applied in other related technical fields, shall be equally included in the patent protection scope of the present invention.
Claims
1. A method for high-concurrency inference of large language models, characterized in that, It includes the following steps: S1: System initialization. The executor calculates the size of the video memory block through simulated peak data, allocates video memory space according to the calculation result, and synchronizes the video memory resource information to the scheduler; S2: The scheduler receives the inference http request, converts the inference http request into a request sequence through preprocessing, and puts the obtained request sequence into the waiting queue of the scheduler; S3: When the waiting queue and the running queue are not all empty, the scheduler checks the video memory usage of the current running queue, determines whether each request sequence can perform the next inference. If not, allocate corresponding video memory blocks for each request sequence until each request sequence can perform the next inference; S4: The scheduler checks the video memory usage of the current running queue, calculates the video memory requirements of the request sequences in the waiting queue in the order of priority, and transfers the request sequences in the waiting queue to the running queue according to the calculation result of the video memory requirements; S5: Define the request sequences transferred from the waiting queue to the running queue in step S4 as the pre-fill type, and define the request sequences in the running queue in step S3 as the decoding type. Allocate the number of video memory blocks for performing pre-fill inference or decoding inference according to the number of pre-fill type request sequences and the number of decoding type request sequences; S6: Call the executor to perform pre-fill inference or decoding inference based on the allocated number of video memory blocks; S7: The scheduler obtains the result returned by the executor, updates the status and information of the request sequence, and changes the type of the request sequence of the pre-fill type; S8: Traverse all request sequences in the running queue, determine whether each request sequence meets the end condition. If so, remove the request sequence from the running queue, release the video memory, and respond to the http request.
2. The method for high-concurrency inference of large language models according to claim 1, characterized in that, In step S1, the size of the video memory block is configured as the number of tokens stored in the video memory block.
3. The method for high-concurrency inference of large language models according to claim 1, characterized in that, In step S1, when allocating video memory space according to the calculation result, it also includes: if the video memory is insufficient, directly exit the program, replace with a smaller large language model or reduce the maximum length.
4. The method for high-concurrency inference of large language models according to claim 1, characterized in that, In step S2, the request sequence includes the metadata information of the http request and the token sequence.
5. The method for high-concurrency inference of large language models according to claim 1, characterized in that, In step S3, when determining whether each request sequence can perform the next inference. If not, allocate corresponding video memory blocks for each request sequence until each request sequence can perform the next inference, it also includes: Judge whether the video memory of the current running queue can support each request sequence to perform the next inference. If not, allocate a new video memory block for the running queue; When the allocated new video memory block is insufficient, remove the request sequences in the running queue according to the priority and release the corresponding video memory blocks until each request sequence can perform the next inference.
6. The method for high-concurrency inference of large language models according to claim 1, characterized in that, In step S4, when calculating the video memory requirements of the request sequences in the waiting queue in the order of priority and transferring the request sequences in the waiting queue to the running queue according to the calculation result of the video memory requirements, it specifically includes: Calculate the video memory requirements of the request sequences in the waiting queue in the order of priority; where the priority order is configured as the arrival time order of the http requests corresponding to each request sequence; If the video memory is sufficient, the requests in the waiting queue are put into the running queue until the video memory is insufficient or the threshold conditions of the maximum batch character length and the maximum batch quantity are reached.
7. The method for high-concurrency inference of large language models according to claim 1, characterized in that, In step S5, according to the number of pre-fill types and the number of decoding types in the request sequence, the number of video memory blocks used to perform pre-fill inference or used to perform decoding inference is allocated, specifically including: If the number of request sequences of the pre-fill type is greater than or equal to the number of request sequences of the decoding type, pre-fill inference is performed, and the corresponding number of video memory blocks is allocated to each request sequence of the pre-fill type; If the number of request sequences of the pre-fill type is less than the number of request sequences of the decoding type, the request sequences of the pre-fill type are put back into the waiting queue, decoding inference is performed, and the required video memory blocks are allocated to each request sequence of the decoding type.
8. The method for high-concurrency inference of large language models according to claim 1, characterized in that, In step S6, pre-fill inference or decoding inference is performed, specifically including: If pre-fill inference is performed, video memory resources are allocated to all pre-filled request sequences, and only the pre-filled request sequences are sent during inference; If decoding inference is performed, all pre-filled request sequences are put back into the waiting queue, and video memory for the next token is allocated to each decoded request sequence, and all request sequences are sent for inference.
9. The method for high-concurrency inference of large language models according to claim 1, characterized in that, In step S7, the status and information of the request sequence are updated, and the type of the request sequence of the pre-fill type is changed, specifically including: Update the status and information of the request sequence; If pre-fill inference is performed, all request sequences of the pre-fill type are changed to the decoding type.
10. A high-concurrency inference system for large language models, characterized in that,Including: An executor and a scheduler, the executor and the scheduler are configured to execute the high-concurrency inference method for large language models according to any one of claims 1-9.
Citation Information
Patent Citations
Method and device for improving throughput of large language model
CN117349032A
Key value cache management method and device, model reasoning method and device and data processing method and device of large language model
CN118860573A
Request processing method, electronic equipment, storage medium and program product
CN118966362A
Large language model reasoning acceleration method and device, equipment and medium
CN119440817A
Cited By
Task allocation method, device and equipment, storage medium and computer program product
CN120762926A
Task allocation methods, apparatus and equipment, storage media and computer program products
CN120762926B
Distributed large language model reasoning system
CN121413773A