Inference method and device for large language model, equipment and medium
Through the CPU and GPU collaboratively asynchronous execution of multiple iterations of inference processes, the resource utilization of large language models is optimized, the problems of high computing resources and response delay are solved, and efficient inference performance is achieved.
Patent Information
- Application Number
- CN202510551114.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-29
- Publication Date
- 2025-08-12
AI Technical Summary
In the inference process of existing large language models, the computing resource requirements are high, the response latency and throughput capabilities are insufficient, making it difficult to meet the needs of real-time interaction and large-scale deployment.
The CPU and GPU are used to coordinate the inference process of multiple iterations. Each iteration includes tasks executed asynchronously, CPU processes unfinished sequences and generates word identity, GPU calculates probability distribution and sampling, and performs different tasks in parallel, hides the overhead of CPU tasks and optimizes resource utilization.
The inference delay of a single iteration is greatly compressed, the overall throughput capability is improved, and the performance is linearly improved when GPU is expanded.
Smart Images

Figure CN120469800A_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the field of artificial intelligence technology, and in particular to an inference method, apparatus, device, and medium for a large language model. Background Art
[0002] In recent years, the scale and number of parameters in AI models have grown exponentially. Large language models have become widely used in scenarios such as question answering, code generation, and content creation. As model size continues to expand, the demands on computing resources, response latency, and throughput for inference have also increased. The industry is generally concerned about how to further reduce inference costs and improve overall performance while maintaining generation quality to meet the needs of real-time interaction and large-scale deployment. Summary of the Invention
[0003] According to one aspect of the present disclosure, a method for reasoning a large language model is provided, comprising: utilizing a CPU and a GPU to collaboratively complete a multi-iteration reasoning process, wherein each iteration includes the following tasks executed sequentially: a first task, configured to determine, on the CPU, multiple unfinished sequences and last generated word-gram identifiers of the multiple sequences; a second task, configured to determine, on the CPU, a model input based on the last generated word-gram identifiers of the multiple sequences; a third task, configured to, on the GPU, calculate, based on the model input, a probability distribution of multiple next word-grams to be generated for the multiple sequences; a fourth task, configured to, on the GPU, perform sampling based on the probability distribution of the multiple next word-grams to obtain word-gram identifiers of the multiple next word-grams; and a fifth task, configured to, on the CPU, update the completion status of the multiple sequences based on the word-gram identifiers of the multiple next word-grams, wherein the (n+1)th iteration and the (n)th iteration are executed asynchronously, the first and second tasks of the (n+1)th iteration are executed in parallel with the third task of the (n)th iteration, and the third task of the (n+1)th iteration is executed in parallel with the fifth task of the (n)th iteration.
[0004] According to another aspect of the present disclosure, an inference device for a large language model is provided, which is used to use a CPU and a GPU to collaboratively complete a multi-iteration inference process, wherein each iteration includes a first task, a second task, a third task, a fourth task, and a fifth task that are executed in sequence, and the device includes: a first processing unit for executing the first task, the first task being configured to determine, on the CPU, multiple unfinished sequences and the last generated word element identifiers of the multiple sequences; a second processing unit for executing the second task, the second task being configured to determine, on the CPU, a model input based on the last generated word element identifiers of the multiple sequences; a third processing unit for executing the third task, the third task being configured to execute, on the GPU U calculates the probability distribution of multiple next words to be generated for multiple sequences based on the model input; a fourth processing unit is used to perform a fourth task, the fourth task is configured to perform sampling on the GPU based on the probability distribution of the multiple next words to obtain word-word identifiers of the multiple next words; and a fifth processing unit is used to perform a fifth task, the fifth task is configured to update the completion status of the multiple sequences based on the word-word identifiers of the multiple next words on the CPU, wherein the (n+1)th iteration and the (n)th iteration are executed asynchronously, the first task and the second task of the (n+1)th iteration are executed in parallel with the third task of the (n)th iteration, and the third task of the (n+1)th iteration and the fifth task of the (n)th iteration are executed in parallel.
[0005] According to yet another aspect of the present disclosure, a computer device is provided, comprising: at least one processor; and a memory on which a computer program is stored, wherein when the computer program is executed by the processor, the processor executes the above method.
[0006] According to yet another aspect of the present disclosure, a computer-readable storage medium is provided, on which a computer program is stored. When the computer program is executed by a processor, the processor executes the above method.
[0007] According to yet another aspect of the present disclosure, a computer program product is provided, comprising a computer program, wherein when the computer program is executed by a processor, the processor executes the method as claimed in claim 1 .
[0008] According to some embodiments, the present disclosure interleaves the reasoning processes of two adjacent iterations: while the GPU is still processing the forward propagation calculation and sampling of the nth iteration (i.e., the third and fourth tasks), the CPU has already completed the model input preparation work for the next round of iteration (i.e., the first and second tasks) in parallel. Subsequently, only one communication interaction between the CPU and GPU is required to start the GPU's forward propagation calculation for the n+1th iteration (the third task). In this way, all CPU-side overhead is completely hidden within the GPU computing window, eliminating the traditional mutual waiting between sequential processes of adjacent iterations, significantly reducing the latency of the reasoning process of a single iteration and improving the overall throughput.
[0009] These and other aspects of the disclosure will be apparent from and elucidated with reference to the embodiments described hereinafter. BRIEF DESCRIPTION OF THE DRAWINGS
[0010] Further details, features and advantages of the present disclosure are disclosed in the following description of exemplary embodiments in conjunction with the accompanying drawings, in which:
[0011] Figure 1 is a schematic diagram illustrating an example system in which the various methods described herein may be implemented, according to an exemplary embodiment;
[0012] Figure 2 is a flowchart illustrating an inference method for a large language model according to an exemplary embodiment;
[0013] Figure 3 FIG. 1 is a schematic diagram illustrating data and resource dependency within a single iteration and between adjacent iterations during an inference process according to an exemplary embodiment;
[0014] Figure 4 is a schematic diagram illustrating an inference process of asynchronously executing iterations according to an exemplary embodiment;
[0015] Figure 5 is a schematic diagram illustrating input processing according to an exemplary embodiment;
[0016] Figure 6 is a schematic diagram illustrating a parallel sampling process according to an exemplary embodiment;
[0017] Figure 7 is a schematic diagram illustrating an inference process according to an exemplary embodiment;
[0018] Figure 8 is a schematic diagram illustrating a performance comparison between the solution of the present disclosure according to an exemplary embodiment and the related art;
[0019] Figure 9 is a schematic diagram illustrating a performance comparison between the solution of the present disclosure according to an exemplary embodiment and the related art;
[0020] Figure 10 is a block diagram illustrating a structure of an inference apparatus for a large language model according to an exemplary embodiment; and
[0021] Figure 11 is a block diagram illustrating an exemplary computer device that can be used with the exemplary embodiments. DETAILED DESCRIPTION
[0022] In this disclosure, unless otherwise specified, the use of terms such as "first" and "second" to describe various elements is not intended to limit the positional relationship, temporal relationship, or importance relationship of these elements. Such terms are only used to distinguish one element from another. In some examples, the first element and the second element may refer to the same instance of the element, while in some cases, based on the context of the description, they may also refer to different instances.
[0023] The terms used in the description of the various examples described in this disclosure are for the purpose of describing specific examples only and are not intended to be limiting. Unless the context clearly indicates otherwise, if the number of elements is not specifically limited, the element can be one or more. As used herein, the term "plurality" means two or more, and the term "based on" should be interpreted as "based at least in part on". In addition, the terms "and / or" and "at least one of..." cover any one of the listed items and all possible combinations.
[0024] Before introducing exemplary embodiments of the present disclosure, several terms used herein are first explained.
[0025] In Large Language Model (LLM) reasoning, the reasoning engine (e.g., vLLM) calls the trained model to generate output based on the prompt text (e.g., user input). Before reasoning, the prompt text is divided into smaller units (e.g., subwords) and mapped to token IDs through a tokenizer (e.g., Byte Pair Encoding or SentencePiece). This disclosure uses sequence to refer to a set of ordered token IDs corresponding to the prompt text and its generated output. In addition, unless otherwise specified, the model, large model, etc. in this disclosure all refer to the large language model LLM. Before generating the complete output, the sequence needs to go through two stages: prefill and decode.
[0026] Prefill: In the prefill phase, the model processes all prompt token IDs in the sequence (i.e., prompt text) through one or more batch forward propagation processes and generates a probability distribution corresponding to all input token IDs.
[0027] Decode: During the decoding phase, the model iteratively generates a new token ID based on the prompt token ID and the previously generated output token ID. In some embodiments, the next token ID can be predicted using a probability distribution calculated using a softmax function. The inference engine can use strategies such as greedy sampling to select the next token ID (referred to as the sampled token ID) and append it to the sequence. This process repeats until a stopping condition is met.
[0028] Parallel Inference: A parallel inference system deploys the same model for inference on multiple GPUs by launching multiple worker processes (each of which can correspond to a specific GPU). A scheduler can run on the master process (e.g., the first of all worker processes) and is responsible for dispatching requests, collecting outputs, and distributing the computational load to other worker processes, thereby achieving parallel GPU execution.
[0029] Iteration: In LLM reasoning, iteration refers to the process of generating a word in the pre-filling or decoding phase.
[0030] Exemplary embodiments of the present disclosure will be described in detail below with reference to the accompanying drawings.
[0031] Figure 1 is a schematic diagram illustrating an example system 100 in which the various methods described herein may be implemented, according to an exemplary embodiment.
[0032] refer to Figure 1 , the system 100 includes a client device 110 , a server 120 , and a network 130 communicatively coupling the client device 110 and the server 120 .
[0033] The client device 110 includes a display 114 and a client application (APP) 112 that can be displayed via the display 114. The client application 112 can be an application that needs to be downloaded and installed before running or a small program (liteapp) that is a lightweight application. In the case where the client application 112 is an application that needs to be downloaded and installed before running, the client application 112 can be pre-installed on the client device 110 and activated. In the case where the client application 112 is a small program, the user 102 can directly run the client application 112 on the client device 110 by searching for the client application 112 in the host application (for example, by the name of the client application 112, etc.) or scanning a graphic code (for example, a barcode, a QR code, etc.) of the client application 112, without installing the client application 112. In some embodiments, the client device 110 can be any type of mobile computer device, including a mobile computer, a mobile phone, a wearable computer device (for example, a smart watch, a head-mounted device, including smart glasses, etc.) or other types of mobile devices. In some embodiments, client device 110 may alternatively be a stationary computer device, such as a desktop computer, a server computer, or other type of stationary computer device.
[0034] The server 120 is typically a server deployed by an Internet Service Provider (ISP) or an Internet Content Provider (ICP). The server 120 may represent a single server, a cluster of multiple servers, a distributed system, or a cloud server that provides basic cloud services (such as cloud databases, cloud computing, cloud storage, and cloud communications). It will be understood that although Figure 1 1. The server 120 is shown communicating with only one client device 110, but the server 120 may provide background services to multiple client devices simultaneously.
[0035] Examples of network 130 include a local area network (LAN), a wide area network (WAN), a personal area network (PAN), and / or a combination of communication networks such as the Internet. Network 130 can be a wired or wireless network. In some embodiments, data exchanged through network 130 is processed using technologies and / or formats including Hypertext Markup Language (HTML), Extensible Markup Language (XML), etc. In addition, encryption technologies such as Secure Sockets Layer (SSL), Transport Layer Security (TLS), Virtual Private Network (VPN), Internet Protocol Security (IPsec) can also be used to encrypt all or some links. In some embodiments, customized and / or dedicated data communication technologies can also be used to replace or supplement the above-mentioned data communication technologies.
[0036] For the purpose of the embodiments of this disclosure, Figure 1In the example, the client application 112 may be an electronic map application that can provide various electronic map-based functions, such as navigation, route query, and location search. Accordingly, the server 120 may be a server used in conjunction with the electronic map application. The server 120 may provide online map services, such as online navigation, online route query, and online location search, to the client application 112 running on the client device 110 based on the road network data. Alternatively, the server 120 may provide the road network data to the client device 110, and the client application 112 running on the client device 110 may provide local map services based on the road network data.
[0037] Figure 2 FIG2 is a flow chart illustrating an inference method 200 for a large language model according to an exemplary embodiment. The inference method 200 includes a multi-iteration inference process using a CPU and a GPU to collaboratively perform multiple inference iterations. Each iteration includes a first task, a second task, a third task, a fourth task, and a fifth task, which are executed sequentially.
[0038] The method 200 may be performed on a client device (e.g., Figure 1 , that is, the execution subject of each step of the method 200 may be a client device 110 shown in FIG. Figure 1 In some embodiments, the method 200 may be performed on a server (e.g., Figure 1 In some embodiments, method 200 may be performed by a combination of a client device (e.g., client device 110) and a server (e.g., server 120). Below, the various tasks performed in method 200 are described in detail, taking client device 110 as an example.
[0039] The GPU used in the inference method of the present disclosure may be a GPU cluster, which may include multiple working processes (or working nodes), and each working process corresponds to a GPU graphics card.
[0040] refer to Figure 2 The first task 210 is configured to determine, on the CPU, a plurality of unfinished sequences and a last generated word element identifier of the plurality of sequences.
[0041] The first task can also be called the scheduling task (Scheduling, T1). The current reasoning framework usually adopts a scheduling strategy called iteration batching. At the beginning of each iteration, a portion of unfinished sequences can be selected based on a certain priority strategy (such as LRU in SGLang or FCFS in vLLM), and scheduling outputs (Scheduling Outputs) are generated at the same time. The selected sequences (that is, multiple unfinished sequences) will complete the reasoning process in the current iteration in collaboration on the CPU and GPU. At the end of each iteration, the completed sequences will be removed and new sequences will be added, thereby minimizing the waiting time of the sequences to be processed. It is understandable that the unfinished sequences may refer to sequences that have not yet been generated, for example <eos>In some embodiments, the above-mentioned scheduling process of the first task can be implemented by a scheduler thread running on the CPU.
[0042] In some embodiments, the scheduling output (ie, the processing result of the first task) may include the token ID of the last word generated by the multiple sequences, and may also include the number of the multiple sequences, as well as information such as the prompt length and generated length of each sequence.
[0043] According to some embodiments, the multiple sequences may include multiple pre-filling sequences in the pre-filling stage and multiple decoding sequences in the decoding stage. Since the pre-filling stage does not need to generate word element identifiers, the last word element identifier generated by the multiple sequences may be the last word element identifier generated by the multiple decoding sequences. The pre-filling stage still needs to perform reasoning to calculate the probability distribution corresponding to the word element identifier in the prompt. The first task can be configured to: determine the word element identifier last generated by the multiple decoding sequences in the CPU; and determine the prompt word element identifier of the multiple pre-filling sequences input in the current iteration.
[0044] In some embodiments, the pre-fill sequence may input all token IDs corresponding to the prompt in a single iteration, or may input a portion of the token IDs, as will be described below.
[0045] The second task 220 is configured to determine the model input based on the last generated word element identifiers of the multiple sequences on the CPU.
[0046] The second task can also be referred to as the input processing task (Input Processing, T2). The second task can convert the scheduling output obtained by the first task into a model input. For example, the token ID list finally generated by each sequence that needs to be inferred in the current iteration can be converted into a tensor and transmitted to the GPU. In some embodiments, the second task can also calculate the metadata required for the reasoning process, such as the number of sequences inferred in the current iteration, the control parameters required for the sampling process (i.e., sampling metadata), etc. These metadata can also be transmitted to the GPU as part of the model input. When using tensor parallelism, the second task can also be configured to broadcast the calculated metadata and model input to all working processes. In some embodiments, the second task can be implemented by an input processor (Input Processor) thread running on the CPU.
[0047] The third task 230 is configured to calculate, on the GPU, the probability distribution of multiple next tokens to be generated in multiple sequences based on the model input.
[0048] The third task can also be called the decoding forward propagation task (Decoding Forward, T3). When the model input is ready, the trained model can calculate the probability distribution of the possible next word (i.e., output token) by means of a softmax function or the like. This probability distribution is then passed to the next task as log probabilities (logits). The third task can be executed in parallel on the GPU. In some embodiments, a single GPU can perform forward propagation calculations on multiple sequences in parallel. In some embodiments, the GPU includes multiple working processes (e.g., multiple GPUs), and the multiple working processes can each perform forward propagation calculations on a portion of the sequence; or using tensor parallelism, each working process performs forward propagation calculations for a portion of the vocabulary dimension.
[0049] In some embodiments, since the pre-filled sequence does not need to generate the next word, the third task can perform different processing on the decoding sequence and the pre-filled sequence. The third task can be configured to: calculate the probability distribution of multiple next words to be generated for multiple decoding sequences (that is, the probability distribution of multiple next words to be generated for multiple sequences) on the GPU; and perform forward propagation calculations on the prompt word identifiers input in the current iteration of multiple pre-filled sequences. The forward propagation calculations on the pre-filled sequences can be used to generate the probability distribution of the prompt word identifiers. After obtaining the probability distribution corresponding to all token IDs of prompt, the decoding stage can be entered to obtain the first generated token by sampling.
[0050] The fourth task 240 is configured to perform sampling on the GPU based on the probability distribution of the multiple next word-grams to obtain word-gram identifiers of the multiple next word-grams.
[0051] The fourth task can also be called the sampling task (Sampling, T4). The sampling process selects the next tokenID from the vocabulary (i.e., the set of all discrete tokens that the model can understand and generate) based on the probability distribution predicted by the third task. In some embodiments, the sampling process for the next token of a sequence can use only the probability distribution of the next token, or it can use the probability distribution of multiple tokens at the end of the sequence. The fourth task can use strategies such as greedy search, beam search, top-k sampling, nucleus sampling, etc., or other sampling strategies, which are not limited here. In some embodiments, the fourth task can be implemented by a parallel sampler. Multiple sampler instances of a parallel sampler can run on multiple GPUs.
[0052] After obtaining the word-unit identifier of the next word-unit in a sequence, it can be considered that the word-unit identifier generation process of the current iteration round is completed.
[0053] The fifth task 250 is configured to update, on the CPU, the completion status of the plurality of sequences based on the word-gram identifiers of the plurality of next word-grams.
[0054] The fifth task 250 can also be called the output processing task (T5). This task adds the sampled token ID to the sequence, converts it into a human-readable string through detokenization, and gradually adds the detokenized text to the output. At the same time, it checks the termination condition (such as generating an end tag). <eos>) and updates the state of the sequence. In some embodiments, the fifth task can be implemented by an output processor (Output Processor) thread running on the CPU.
[0055] In some embodiments, the inference process on the pre-filled sequence can skip sampling (the fourth task) and output processing (the fifth task).
[0056] In some embodiments of the related art, a sequential workflow is often used, whereby tasks T1 to T5 are executed sequentially within each inference iteration, and iterations of different rounds are also executed sequentially. This approach has poor scalability because, of all tasks, only T3 can be accelerated in parallel using multiple GPUs, while the remaining tasks cannot be parallelized and optimized by adding GPUs.
[0057] Figure 3 FIG2 is a schematic diagram showing the data and resource dependency relationship within a single iteration and between adjacent iterations in the inference process according to an exemplary embodiment of the present disclosure. i n Indicates task T in the nth iteration i Execution of P i Indicates T i The average proportion of time consumed in the entire reasoning process, t is the parallelism, and A is the acceleration ratio. Figure 3 As can be seen, the execution of T4 requires the sampling metadata generated by T2, and the execution of T3 requires the model input generated by T2. There is a resource dependency between the next iteration's T1 and the previous iteration's T5. The execution of the next iteration's T2 requires the previous iteration's T5 to insert the sampled token ID into the sequence. Therefore, using traditional methods, only T3 can be shortened by adding GPUs (scalable), while other tasks cannot be shortened (non-scalable).
[0058] In an experiment, the Llama-2-13B model was used and tested on 8 H100 GPUs based on the vLLM framework. The results showed that the GPUs were idle 47.5% of the time, waiting for other tasks to complete. Due to the non-scalability of T1, T2, T4, and T5, even if more GPUs were added to improve the performance of T3, the overall inference performance could not be significantly improved. However, due to the T (n+1) Dependency T n The output of , causes the reasoning scheme in related technologies to only adopt this sequential workflow.
[0059] Figure 4 FIG. 1 shows a schematic diagram of an iterative reasoning process according to an exemplary embodiment of the present disclosure. Figure 4 As shown, this disclosure proposes a new inference scheme: asynchronous execution between iterations n+1 and n, with the first and second tasks of iteration n+1 executing in parallel with the third task of iteration n, and the third task of iteration n+1 executing in parallel with the fifth task of iteration n. This inference scheme hides the overhead of T1, T2, and T5, with the performance of tasks primarily determined by T3 and T4, allowing overall performance to scale linearly with the number of GPUs.
[0060] Thus, by interleaving the inference processes of two adjacent iterations, while the GPU is still processing the forward propagation calculations and sampling of the nth iteration (i.e., the third and fourth tasks), the CPU has already completed the model input preparation work for the next round of iteration (i.e., the first and second tasks) in parallel. Subsequently, only one communication interaction between the CPU and GPU is required to start the GPU's forward propagation calculations for the n+1th iteration (the third task). In this way, all CPU-side overhead is completely hidden within the GPU computing window, eliminating the traditional waiting between sequential processes between adjacent iterations, significantly reducing the latency of the inference process for a single iteration and improving overall throughput. Because the critical path of the multi-iteration inference process is only the GPU's forward propagation calculations and sampling, its time consumption is shortened almost inversely with the number of newly added GPUs, resulting in an ideal linear relationship between system performance and parallelism t, ensuring that large model inference achieves an acceleration effect proportional to the number of GPU cards when GPU cards are expanded.
[0061] In some embodiments, the present disclosure formalizes the scheduling process to clearly demonstrate and When scheduling a sequence, the inference framework considers multiple scheduling budgets, including the upper limit B on the number of new Token IDs that can be generated in each iteration. t , the upper limit of the number of sequences that the model can process simultaneously B seq , and the available GPU memory. In addition, in some embodiments, PagedAttention or other methods can be used to allocate each B in the GPU memory. c The key-value cache corresponding to each Token ID is grouped into a block, and these blocks are managed as "pages" in the operating system. Therefore, the number of available blocks B b is regarded as the GPU memory budget. During the scheduling process, let the current sequence set be The length of the sequence seq (that is, the number of Token IDs contained) is L seq , the number of new Token IDs that need to be generated is N seq .
[0062] The reasoning framework needs to be based on the defined scheduling strategy. In an exemplary embodiment, Select a subset And ensure that the following conditions are met:
[0063]
[0064] It is understandable that, in addition to the above methods, other scheduling strategies can also be used to determine the multiple sequences that need to be inferred.
[0065] Asynchronous dispatch to obtain the correct collection There are challenges when processing the output of the previous iteration because the previous iteration may not have completed yet. Some sequences may terminate early, thus changing their status. Directly affects scheduling Choice, incorrect The selection may violate the resource constraints defined in the above inequality. In addition, in asynchronous scheduling, the L seq and N seq Inaccurate estimates of can also lead to violations of these resource dependencies.
[0066] Based on this, the present disclosure adopts an optimistic prediction strategy to break the resource dependency between adjacent tasks. In short, the strategy allocates memory under the assumption that "the request never encounters a stop condition" and calculates L based on this assumption. seq and N seq This approach ensures that the scheduling of the next iteration can start before the current iteration is completed. Under this assumption, the recurrence relation can be easily derived before the sequence seq is completed: Among them, the initial length and The value of can be inferred from the stage of the sequence (prefill or decoding), which is determined by the current length The next challenge is how to accurately predict the completion status of the sequence in asynchronous scheduling to ensure the correct identification of the sequence set
[0067] This disclosure proposes a sequence management mechanism for asynchronous scheduling optimization. In the stage, a sequence seq that participates in the n-1th iteration may have two uncertain lengths: Previously And in After that This dual-state condition introduces potential ambiguity for the scheduler. and Sequential execution will weaken the advantages of asynchronous execution. To solve this problem, this disclosure introduces Iteration-Dependent Sequence Management, which eliminates the problem of inconsistent sequence states by assigning a virtual state to each tokenID in each sequence.
[0068] According to some embodiments, the first task may be configured to determine the sequence status of each of the plurality of sequences. This process may be implemented by a scheduler. The sequence status may include the following three states:
[0069] The expected length of the corresponding sequence at the end of the current iteration (EL, Expected Length);
[0070] The current length (CL) of the corresponding sequence at the beginning of the current iteration; and
[0071] The number of tokens generated by the corresponding decoding sequence in the current iteration or the number of prompt token identifiers (NNT) input by the corresponding pre-filled sequence in the current iteration.
[0072] The current length of each of the multiple sequences at the start of the current iteration can be determined based on the expected length of the sequence at the end of the previous iteration in which the sequence was inferred. For example, at the n+1th iteration of a sequence, although the nth iteration has not yet ended, the first task of the nth iteration (i.e., the scheduling phase) determines that the EL of the sequence should be CL+NNT based on the optimistic prediction strategy. Therefore, at the n+1th iteration, the CL of the n+1th iteration can be determined based on the EL of the nth iteration (even though the end token of the sequence may not have been sampled at this time).
[0073] This elegant sequence management strategy allows the first task (e.g., by the scheduler) to query the future state of the sequence (e.g., EL and NNT) based on the number of iterations that the sequence has already been scheduled, leading to more efficient pre-scheduling. This approach ensures consistent state tracking for all sequences during asynchronous execution.
[0074] In addition, the sequence state can also be used to determine the position index of the next token to be generated in multiple decoding sequences. This position index can be used by the GPU to help the GPU locate the next token to be generated, so as to calculate the probability distribution at the correct position.
[0075] In the LLM reasoning process, the key (key, K) and value (value, V) of the inferred word are fixed and will be used in subsequent reasoning. Therefore, these KVs can be cached in the GPU memory to speed up the reasoning process. Based on this, the present disclosure proposes a KV cache management mechanism for LLM reasoning, which is used to efficiently manage GPU memory in multi-sequence parallel reasoning scenarios. As described above, each B in the GPU memory can be cached. c The KV group corresponding to each Token ID is a cache block.
[0076] According to some embodiments, the first task may be configured to determine at least one KV cache block allocated to each of the plurality of sequences to obtain a KV cache block index for each of the plurality of sequences. The model input may include the KV cache block index for each of the plurality of sequences.
[0077] Therefore, through this method, the key values corresponding to multiple words in each sequence only need to be calculated once and can be reused between different iterations. By passing these KV cache block indexes to the GPU as model input, the GPU can locate the corresponding KV cache block according to the index when reasoning about the sequence of the current iteration to obtain the cached key values of each sequence, thereby achieving fast reasoning.
[0078] It is understandable that, in different iterations, the content of the cached key value in at least one KV cache block of the same sequence remains unchanged, and a newly calculated new key value can be written to the unfilled KV cache block in each iteration. In some embodiments, the third task can be configured to, for a sequence among multiple sequences, determine at least one corresponding KV cache block based on the KV cache block index of the sequence; obtain the cached key value of the sequence from the corresponding at least one KV cache block; and calculate the probability distribution of the next token of the sequence using the cached key value of the sequence, and write the new key value generated during the calculation process to the unfilled KV cache block in the corresponding at least one KV cache block.
[0079] Since the third task of iteration n and the first task of iteration n+1 are executed in parallel, the new key values generated during the calculation of the third task of iteration n may not have been written into the corresponding KV cache blocks when the first task of iteration n+1 is executed. To ensure that the new key values generated during the calculation of the third task of each iteration are written into the correct location, the number, index, and vacancy information of the KV cache blocks of each sequence can be maintained in the first task.
[0080] In some embodiments, the first task may be configured to determine, in at least one KV cache block of each of the multiple sequences, a KV cache slot for the current iteration of each of the multiple sequences; the model input determined by the second task may include the KV cache slot for the current iteration of each sequence; and the third task may be configured to write the new key value generated during the calculation process into the KV cache slot. In this way, it is possible to ensure that the key values generated by the calculation process of the same sequence in different iterations are written to different locations in the KV cache block.
[0081] In some embodiments, the KV cache block index can be implemented as the sequence number or subscript of the corresponding at least one KV cache block in the global block list. The KV cache vacancy can be implemented as a KV cache offset. The KV cache offset represents the token number within the corresponding block, or the byte offset in the GPU address space.
[0082] The core of asynchronous scheduling can involve two key aspects: (A1) determining the number of KV cache blocks required for a sequence in each scheduling iteration, and (A2) predicting whether the sequence should continue to generate the next token ID (i.e., whether further processing is required).
[0083] For A1, the core is to calculate the number of KV cache blocks C required by the sequence in the nth iteration n .
[0084] According to some embodiments, a single KV cache block can accommodate a second preset number of word-unit corresponding keys and values. Determining at least one KV cache block allocated to each of a plurality of sequences to obtain a KV cache block index for each of the plurality of sequences may include: determining the number C of KV cache blocks allocated to the sequence based on the sequence length and the second preset number of the sequence. n .
[0085] In one exemplary embodiment, C n The calculation can be performed interactively in the following ways:
[0086]
[0087] Here, L0=0, N p Indicates the number of prompt token IDs in the sequence, N c Indicates the block size of token IDs processed in the pre-fill phase. n-1 <N p When), you can use the preset N c To update L n , to improve the efficiency of the pre-filling process. In contrast, in the decoding phase, only one token ID is generated per iteration, so it is only necessary to increase the current length by 1.
[0088] According to some embodiments, the number of word units generated in the current iteration determined for each of the plurality of decoding sequences is 1, and the number of prompt word units input in the current iteration determined for each of the plurality of pre-filled sequences does not exceed a first preset number N. c , the first preset number N c Not less than 2.
[0089] In some embodiments, N p Can be N c An integer multiple of N, so that each iteration of the pre-fill phase can process the same number of prompt token IDs. c It can also be set to make N p ÷N c The remainder is as close to N as possible c , so that the last iteration of the pre-filling phase can make better use of computing resources.
[0090] It is understandable that, in addition to the above methods, other methods can be used to calculate the number of KV cache blocks C required for each sequence in the nth iteration. n .
[0091] In some embodiments, after obtaining the number of KV cache blocks C required for each sequence in the nth iteration, n This can then be compared to the number of KV cache blocks already allocated for each sequence. If the number of KV cache blocks C required for the nth iteration of a sequence is n If the number of KV cache blocks exceeds the number already allocated for the sequence, additional KV cache blocks need to be allocated for the sequence, and the number, index, and vacancy information of the KV cache blocks of the sequence need to be updated accordingly. If the number is the same, no new KV cache blocks need to be allocated, but the vacancy information of the KV cache blocks of the sequence still needs to be updated.
[0092] In some embodiments, after a sequence completes, all KV cache blocks assigned to the sequence may be reclaimed.
[0093] When multiple sequences of varying lengths and progress are being inferred simultaneously, using traditional contiguous memory to manage the KV cache can lead to severe memory fragmentation and waste, and the frequent allocation and deallocation of memory also incurs additional overhead. This paper proposes a flexible KV cache management mechanism that better supports multi-sequence parallel inference scenarios, effectively alleviating GPU fragmentation and waste, and improving GPU utilization and overall inference efficiency.
[0094] For A2, an optimistic prediction strategy can be adopted, assuming that each sequence in each iteration needs to continue inference. In practice, this strategy is very effective because for a sequence of N tokens, the number of successful predictions is N-1, and only the last prediction fails, thus generating a total of N+1 token IDs.
[0095] Furthermore, by using single-iteration asynchronous scheduling, the n+1th iteration is scheduled during the nth iteration, which brings at least two benefits. First, referring to Figure 4 This approach completely overlaps the CPU computation (the first and second tasks) of iteration n+1 with the GPU computation (the third and fourth tasks, primarily the decoding forward propagation process) of iteration n to hide the CPU computation overhead, while only adding a slight time overhead between the fourth task of iteration n and the third task of iteration n+1. In some embodiments, the slight additional time overhead can include the GPU (e.g., the parallel sampler running on it) writing the sample output of iteration n to a queue and the GPU removing the model input of iteration n+1 from the queue (in preparation for the third task of iteration n+1). This process typically takes no more than 80 microseconds.
[0096] Secondly, in online services, when new requests arrive, the pre-populated task schedule must be recalculated to ensure consistent latency in generating the first token. Therefore, in asynchronous scheduling, the more iterations advanced, the higher the likelihood of schedule failure. Single-iteration asynchronous scheduling achieves an effective balance between computational overhead and performance improvement.
[0097] Therefore, by decoupling the scheduling process (i.e., the first task) from the sequential execution workflow and making it an asynchronous operation, the input processing (i.e., the second task) can be performed asynchronously in advance, thereby significantly shortening the total duration of the reasoning process of multiple iterations.
[0098] In some embodiments, when generating the n+1th token ID in the n+1th iteration, the model requires the nth token ID in the sequence as part of the input. Without this nth token ID, the model will generate an incorrect token ID, resulting in inference errors.
[0099] In order to further decouple input processing and output processing from the sequential execution workflow and eliminate the dependence of input processing on output processing, the present disclosure introduces an early feedback backfill mechanism. This mechanism establishes a fast channel from the sampling task (fourth task) to the output processing task (fifth task) of the same iteration and the input processing task (first task) of the next iteration. This channel can be established, for example, between the sampler running on the GPU and the input processor and output processor running on the CPU. During the sampling process, each newly generated token ID is immediately forwarded to these processors, so that the model input is aligned with the input in the sequential execution workflow and potential inference errors are corrected.
[0100] The input processing of the n+1th iteration depends on the scheduling output of the n+1th iteration and the sampling output of the nth iteration. These dependencies are the reason why existing reasoning frameworks cannot perform input processing asynchronously. In the solution proposed in the present disclosure, the input processing task (the second task) of the n+1th iteration can directly access the scheduling output of the n+1th iteration described in the asynchronous scheduling (the processing result of the first task), thereby reducing the data dependency to only the sampling output of the nth iteration (the intermediate result of the fourth task of the previous iteration).
[0101] According to some embodiments, determining the unfinished multiple sequences and the last generated word-gram identifiers of the multiple sequences in the first task may include: in response to determining that the multiple sequences include the target sequence determined in the first task of the previous iteration, using a placeholder as the last generated word-gram identifier of the target sequence. Determining the model input based on the last generated word-gram identifiers of the multiple sequences in the second task may include: determining a first tensor, the first tensor including the last generated word-gram identifiers of the multiple sequences; and after sampling the target word-gram identifier generated in the previous iteration of the target sequence in the fourth task of the previous iteration, backfilling the target word-gram identifier into the placeholder in the first tensor, wherein the model input includes the backfilled first tensor.
[0102] Although the token ID at the end position of the target sequence of the n+1th iteration depends on the sampling result of the fourth task of the nth iteration, the shape of the first tensor (the word element identifier of the last sample of multiple sequences) is determined only by the scheduling output. This allows the input processor to determine the first tensor immediately after receiving the processing result of the first task without waiting for the sampling process to complete. For the target sequence being inferred in the previous iteration, since the "last generated word element identifier" of the sequence has not actually been generated yet, a placeholder is temporarily used as the word element identifier. After the fourth task of the previous iteration completes sampling, the sampling result corresponding to the target sequence (that is, the token ID last generated by the target sequence) can be immediately backfilled into the corresponding placeholder in the first tensor to ensure that the model input including the first tensor can provide accurate tokenIDs to the GPU.
[0103] According to some embodiments, determining the model input based on the last generated word element identification of multiple sequences in the second task may also include: determining a second tensor, the second tensor including the position indexes of multiple next words, the model input including the first tensor; and sending the second tensor to the GPU before backfilling.
[0104] In addition to the first tensor including the token ID of the last sample of each sequence, a second tensor can also be determined in the second task. The second tensor includes the position indexes of the multiple next tokens to be generated for the multiple sequences in the corresponding sequences. These position indexes can be determined based on the sequence states of the multiple sequences determined in the first task mentioned above, and can be used on the GPU to locate the token to be generated, so as to calculate the probability distribution at the correct position. Since these position indexes do not depend on the sampling process of the previous iteration, the determination of the second tensor can be completed immediately after the start of the second task, and the second tensor can be sent to the GPU before backfilling without waiting for the sampling process to be completed.
[0105] In some embodiments, after backfilling the placeholders in the first tensor is completed, the backfilled first tensor can be sent to the GPU.
[0106] In some cases, the multiple sequences scheduled for the current iteration may not include any sequences scheduled for the previous iteration. In other words, each of the multiple sequences scheduled for the current iteration has already completed the previous token generation process. Therefore, the last generated token ID for each of the multiple sequences can be obtained in the first task.
[0107] According to some embodiments, determining the model input based on the last generated word element identifier of multiple sequences in the second task may also include: in response to determining that the multiple sequences do not include any sequence determined in the first task of the previous iteration, after determining the first tensor, directly sending the first tensor to the GPU.
[0108] In this case, after the first tensor is determined in the input processing stage, it can be sent directly to the GPU without waiting for backfill processing.
[0109] In some embodiments, the second task can also be configured to calculate metadata. Model input can also include metadata. Metadata can include the batch size required for the inference process of the third task, the current length of each sequence, the index and vacancy of the KV cache block of each sequence, etc. It can also include sampling control parameters required for the sampling phase of the fourth task, such as the sampling strategy, etc. The calculation of metadata does not require the last generated token ID, so it can also be started immediately after the first task ends.
[0110] Figure 5 FIG2 shows a schematic diagram of input processing according to an exemplary embodiment of the present disclosure. Figure 5 After the decoding forward propagation of the nth iteration, the sampling process of the nth iteration can be performed, and then the output processing of the nth iteration can be performed. In addition, after obtaining the scheduling output of the n+1th iteration (the upper left box), the input processing of the n+1th iteration can be performed to obtain the model input X. Where X M is metadata; X T is the first tensor, i.e. the last generated word element identifier of multiple sequences; XX M -X T For other tensors, such as the second tensor, that is, the position index of multiple next tokens to be generated in the current iteration of multiple sequences. It can be seen that for seq2 and seq3 scheduled in the nth iteration, the early feedback backfill mechanism can be used to backfill the sampling results to the first tensor X T For seq4 and seq5, which are not scheduled in iteration n, the last generated token IDs of these sequences have been sampled in earlier iterations, so they can be directly obtained from the scheduling results. After obtaining the model input, the decoding forward propagation of iteration n+1 can be performed.
[0111] According to some embodiments, the second task may be configured to write the model input to a first queue, and the third task may be configured to retrieve the model input from the first queue, where the first queue is used for communication from the CPU to the GPU.
[0112] By using the first queue, efficient connection from the second task to the third task in the same iteration can be achieved.
[0113] According to some embodiments, the fourth task may be configured to write the word-gram identifier of the next word-gram of the second sequence into the second queue, and the fifth task may be configured to retrieve the word-gram identifier of the next word-gram of the second sequence from the second queue, where the second queue is used for communication from the GPU to the CPU.
[0114] By using the second queue, efficient connection between the fourth task and the fifth task in the same iteration can be achieved.
[0115] In some embodiments, the early feedback backfill mechanism from the fourth task of the nth iteration (sampler on the GPU) to the second task of the n+1th iteration (input processor on the CPU) can be implemented through a queue or other means, which are not limited here.
[0116] According to some embodiments, the first task of the n+2th iteration is executed after the fifth task of the nth iteration. Before the fifth task of the nth iteration is completed, it is not known whether the sequence of the nth iteration has completed reasoning, and its completion status has not been updated. If the first task of the n+2th iteration is started at this time, the sequence that completed reasoning on the nth iteration may still be scheduled in the n+2th iteration. Figure 4 It can be seen that the total duration of the fifth task of the nth iteration, the first task of the n+2th iteration, and the second task is still less than the third task of the n+1th iteration. Therefore, the first task of the n+2th iteration can be postponed to after the fifth task of the nth iteration.
[0117] In some embodiments of the related art, the sampling process can be formalized as:
[0118]
[0119]
[0120] Y=Sample(Pr),
[0121] Where X is the model output, Pr is the probability of the output of task T3, g represents the gather() operation, and Y represents the sampling result. The probability matrix consists of s×v elements, representing the probability of v tokens in the vocabulary in s requests. represents partitioning the matrix along the vocabulary dimension.
[0122] There are two main reasons why sampling in existing frameworks is kept in single thread execution.
[0123] First, parallel sampling is incompatible with model parallelism. Sampling is performed along the vocabulary dimension, which is already split across tensor parallel groups. Furthermore, sampling relies on the output of the model's last layer, which means that only the final stage of pipeline parallelism can perform sampling directly. As a result, distributing sampling to other workers in model parallelism requires expensive collective communication, adding milliseconds of latency to each iteration.
[0124] Second, sampling requires additional metadata as input, which introduces significant communication overhead. Specifically, each worker node must receive the settings (such as temperature, penalty term) and sampling status for each sequence. Using a module like Pickle to directly serialize this data into a binary format incurs a significant overhead, for example, for a batch size of 256, the overhead per iteration can exceed ten milliseconds. Therefore, broadcasting this data between workers can severely impact performance (for example, the throughput loss for vLLM reaches 37%).
[0125] The inventors note that in non-parallel sampling, operations on s×v logits are independent in the sequence dimension. This independence allows the sampling workload to be divided among multiple GPUs (also referred to as multiple GPU worker processes in this disclosure), thereby improving scalability.
[0126] According to some embodiments, after the third task is completed, each of the multiple working processes of the GPU may have a probability distribution of a partial vocabulary of all or part of the decoding sequences in the multiple decoding sequences. The fourth task may be configured to: perform a many-to-many data interaction operation between the multiple working processes so that each of the multiple working processes has a probability distribution of the full vocabulary of a partial decoding sequence in the multiple decoding sequences; use each of the multiple working processes to sample the probability distribution of the full vocabulary of the next word of the partial decoding sequence possessed by the working process; and aggregate the sampling results of the multiple working processes to obtain the word identifier of the next word of each of the multiple decoding sequences.
[0127] Compared with the original sampling operation, the sequential parallel sampling process can be expressed as follows:
[0128]
[0129]
[0130]
[0131]
[0132] where g and Represents gather and all-to-all operations respectively. The sampling workflow can include four steps:
[0133] 1. The driver node uses the scatter operation to distribute the sampling metadata of different sequences to each worker node (worker process).
[0134] 2. Each worker node performs an all-to-all operation after the decoding forward calculation to exchange the probability of the vocabulary, and the data is split according to the sequence dimension.
[0135] 3. The core sampling calculation remains unchanged, but each worker node processes a smaller subset of requests independently.
[0136] 4. The driver node aggregates the token IDs sampled by all worker nodes into an array and sends it through the early feedback backfill mechanism.
[0137] The above process greatly improves processing efficiency and reduces performance bottlenecks by parallelizing the sampling process and processing each request in a distributed manner.
[0138] According to some embodiments, the present disclosure adopts a more efficient logit exchange method. In a traditional setting, the driver worker node collects the probabilities of the vocabulary dimension from all worker nodes for sampling. However, once the sampling workload is distributed to all worker nodes, each worker node must access these probabilities. In order to optimize this process, the present disclosure replaces the traditional gather with an all-to-all operation, which collects the vocabulary dimension probabilities while splitting them according to the batch size. This approach eliminates the need for the worker nodes to separately split the output of the allgather operation. Therefore, the amount of data transmitted by each worker node is the same as that of a single gather operation, ensuring that parallel sampling does not introduce additional communication overhead at this stage.
[0139] The driver node can collect results only after the sampling task is completed. Each result includes a single token ID and sequence ID for each sequence. Therefore, even with a large number of requests, the communication data remains very small. In one exemplary embodiment, only about 200 microseconds of latency was observed for 256 requests.
[0140] According to some embodiments, the second task may be configured to determine multiple sampling metadata for multiple sequences, and the sampling metadata for each of the multiple sequences may be sent to the GPU when executing the third task. The fourth task may be configured to perform sampling on the GPU based on the sampling metadata for each of the multiple sequences and the probability distribution of multiple next tokens to be generated.
[0141] In some embodiments, the forward pass before sampling is a GPU-bound operation, and a large amount of sampling metadata is only needed after the forward calculation is completed. Based on this, the present disclosure proposes a method of delaying the scatter operation until the forward pass begins, which effectively hides the communication overhead in the unavoidable GPU calculation time. The ratio of the sampling metadata scatter time to the forward pass time is defined as R s In most cases, even on an H100 GPU, R s It is also below 20% because the metadata size is small (about 1.5KB per request for a sequence of length 1000). This complete overlap effectively eliminates the communication overhead associated with scatter.
[0142] According to some embodiments, the model input may further include original indices of the plurality of sequences, and the indices of the plurality of sequences in the plurality of working processes may be determined by performing a modulo operation on the original indices of the plurality of sequences according to the number of the plurality of working processes.
[0143] In some embodiments, the original indices of each of the multiple sequences may constitute a third tensor (also referred to as an index tensor). The third tensor may be determined in the second task and sent to the GPU before backfilling (if any).
[0144] During sampling, the index tensor is used to select specific rows in the logits matrix for matrix multiplication and related calculations (e.g. Figure 6 [0,1,...,s] in [0,1,...,s]). When the sampling is distributed to t worker nodes, each worker node only holds s / t rows of logits, which makes the original index invalid. To address this problem, this paper introduces a remapping process: (1) realigning the index with its corresponding target, and (2) applying a modulo operation on the original index based on the number of worker nodes to ensure an accurate mapping from each worker node's index ([0,1,...s / t]) to the logits target row.
[0145] According to some embodiments, the fourth task may be configured to: before performing many-to-many data interaction, fill multiple sequences with a filling sequence so that the number of the filled multiple sequences is divisible by the number of multiple work processes; and after aggregating the sampling results of the multiple work processes, discard the sampling results of the filling sequence.
[0146] In some embodiments of the related art, existing collective communication frameworks typically require that the number of data items be divisible by the number of worker nodes, and problems arise when the batch size is not divisible by the number of worker nodes. In order to enable serial parallel sampling in all cases, the present disclosure proposes a batch padding technique so that s'mod t=0. This technique is implemented by adding synthetic metadata to placeholder (virtual) requests and padding the logit tensor with additional rows. Subsequently, the driver worker node discards the sampling results of these padding requests before performing the early feedback backfill operation. Although this introduces some additional communication padding overhead (metadata padding is less than 10KB), the advantages of parallel sampling far outweigh these costs, ultimately accelerating the inference process.
[0147] Figure 6 A schematic diagram of a parallel sampling process according to an exemplary embodiment of the present disclosure is shown. First, a padding operation can be performed on multiple sequences s to obtain multiple padded sequences s'. For a GPU cluster with a parallelism of t, each worker process includes the probability distribution of all sequences s' in v / t vocabulary dimensions, where v is the total vocabulary dimensions. The metadata and logits corresponding to s' can be subjected to a scatter operation ① and a many-to-many data interaction operation ②. After the scatter operation ①, each worker process has the metadata of the s' / t sequences that the process needs to process, and after the many-to-many data interaction operation ②, each worker process has the probability distribution of the s' / t sequences that the process needs to process in all vocabulary dimensions. Furthermore, sampling ③ can be performed on each worker process, and each worker process obtains sampling results for s' / t sequences. The sampling results of different worker processes are aggregated ④ to obtain sampling results for s' sequences. The above process implements a sequence-parallel sampling process. Then, the driver node can complete the de-padding operation, discarding the sampling results of the padded sequences, and obtain the original sampling results of the multiple sequences s. In addition, multiple sequences s can be remapped by different worker processes (e.g., t GPUs from GPU 0 to GPU t).
[0148] According to some embodiments, the same third preset number of random numbers may be pre-generated on multiple work processes, and each work process may use a subset of the third preset number of random numbers when executing the third task and / or the fourth task.
[0149] Random numbers are crucial in operations like multinomials, where a fixed random seed aims to produce deterministic results. However, different worker nodes do not share the state of the random number generator. Therefore, independently generating k / t random numbers on t worker nodes may produce different results than generating all k random numbers on a single worker node. To address this issue and maintain full compatibility with the original approach, all k random numbers can be pre-generated on all t GPUs so that each worker node only uses a subset of them. This approach trades off additional GPU memory (about 128KB per request) for lower communication overhead (saving about 1 millisecond of remote GPU read time).
[0150] Figure 7 A schematic diagram of an inference process according to an exemplary embodiment of the present disclosure is shown, wherein the nth decoding forward propagation starts on the GPU, and the scheduling of the n+1th iteration is simultaneously started on the CPU.
[0151] On the CPU side, the scheduler evaluates the state of the current sequence and generates scheduling outputs for n+1 iterations based on a service policy (such as FCFS). These outputs are forwarded to the predictor and input processor (①). The predictor predicts the number of KV cache blocks required for each scheduled sequence in the next iteration and actively updates their state to prepare for the scheduling of the next iteration (②). At the same time, the input processor converts the scheduling output into model input, generates corresponding sampling metadata, and caches it in a queue.
[0152] On the GPU side, the model input and sampling metadata for the nth iteration can be obtained from the Input Processor's queue and then transmitted to the model and the Parallel Sampler simultaneously (③). After the decoding forward pass is completed, the Parallel Sampler distributes the probability distribution logits (Pr) from the model to the sampler instances (④). Each instance independently performs parallel sampling using the sampling metadata.
[0153] Once the sampled token ID is determined, the parallel sampler immediately sends it to the output processor to backfill the model input for the n+1 iteration (⑤). This operation is crucial in asynchronous scheduling scenarios, because in this case, some sequences may appear in both the nth and n+1 iterations, but their last generated token ID is not yet determined at the time of scheduling. In this case, the scheduler inserts a placeholder at the corresponding position, and the backfill operation ensures the correctness of inference. Once sampling is completed, the parallel sampler passes the output to the output processor (⑥).
[0154] Meanwhile, the output processor continuously extracts sampling results from the queue, updates the status of unfinished sequences, and removes completed sequences (⑦).
[0155] In some embodiments, the inventors evaluated the solution proposed in this disclosure on three test platforms, each equipped with 8 GPUs of a specific type: NVIDIA H100 (NVLink), NVIDIA A100 (NVLink), and NVIDIA A100 (PCIe), all with 80GB of memory. Each platform was also equipped with 2TB of RAM. The systems were configured with a Platinum 8468 CPU with 192 logical cores and a 200Gbps network interface card (NIC) connected to GPFS. All systems ran Ubuntu 22.04, NVIDIA driver version 535.129.03, and CUDA version 12.3.
[0156] In the test, the maximum sequence output length is set to 1000, the maximum batch size is set to 256, and Figure 8 The average throughput of different inference frameworks is measured in part (a).
[0157] The inventors observed that: (1) the proposed solution achieved 49% to 92% improvement on different models. (2) the proposed solution maintained a consistent improvement ratio under different maximum sequence output lengths.
[0158] In order to verify whether the solution proposed in this disclosure can bring performance improvement on different GPUs, the inventors conducted the same experiment on a test platform equipped with A100 (PCIe) and A100 (NVLink), such as Figure 8 Part (b) and Figure 8 As shown in part (c) of the , on these two test platforms, the proposed solution achieved performance improvements of 27% to 66% and 56% to 2.5 times, respectively, compared to vLLM. This proves that the proposed solution is GPU-independent and can achieve performance improvements on a variety of GPUs.
[0159] Furthermore, the inventors observed that the proposed solution on A100 (NVLink) achieved better acceleration than A100 (PCIe). This is because forward propagation requires data exchange between GPUs, and the communication overhead via NVLink is lower than PCIe.
[0160] A key design goal of the solution proposed in this disclosure is to improve GPU utilization and achieve balanced workload distribution. Figure 9 (a), (b), and (c) show the average GPU utilization of the proposed scheme measured every 20 milliseconds compared to vLLM. The results show that: (1) through asynchronous execution, the proposed scheme improves the average GPU utilization on Llama-2-13B by 28% and Gemma-2-9B by 34% by reducing the GPU idle time in T1, T2, and T5; (2) GPU utilization is evenly distributed among all worker threads (see Figure 9 (c) of ), because the parallel sampler distributes the sampling tasks—which are usually concentrated in the threads of the driver node in other frameworks—evenly to all worker threads.
[0161] The solution proposed in this disclosure improves GPU utilization while achieving energy-efficient reasoning. Figure 9 Part (d) shows the average GPU power consumption measured every 20 milliseconds when using different frameworks during the inference process of 256 requests. The results show that the proposed scheme reduces the average GPU power consumption by 15% and shortens the inference time by 47%, thus achieving a 54% energy saving.
[0162] According to another aspect of the present disclosure, an inference device for a large language model is proposed, which is used to collaboratively complete a multi-iteration inference process using a CPU and a GPU, wherein each iteration includes a first task, a second task, a third task, a fourth task, and a fifth task that are executed in sequence. Figure 10 As shown, the apparatus 1000 includes: a first processing unit 1010 for performing a first task, the first task being configured to determine, on a CPU, multiple unfinished sequences and the last generated word-gram identifiers of the multiple sequences; a second processing unit 1020 for performing a second task, the second task being configured to determine, on a CPU, a model input based on the last generated word-gram identifiers of the multiple sequences; a third processing unit 1030 for performing a third task, the third task being configured to calculate, on a GPU, a probability distribution of multiple next word-grams to be generated for the multiple sequences based on the model input; a fourth processing unit 1040 for performing a fourth task, the fourth task being configured to perform sampling on a GPU based on the probability distribution of multiple next word-grams to obtain word-gram identifiers of multiple next word-grams; and a fifth processing unit 1050 for performing a fifth task, the fifth task being configured to update, on a CPU, the completion status of the multiple sequences based on the word-gram identifiers of the multiple next word-grams. The (n+1)th iteration and the (n)th iteration are executed asynchronously, the first and second tasks of the (n+1)th iteration are executed in parallel with the third task of the (n)th iteration, and the third task of the (n+1)th iteration is executed in parallel with the fifth task of the (n)th iteration.
[0163] It should be understood that Figure 10 The various modules of the apparatus 1000 shown in FIG. 1 can be used in conjunction with the reference Figure 2 The various tasks in the described method 200 correspond to each other. Therefore, the operations, features and advantages described above for the method 200 are also applicable to the apparatus 1000 and the modules included therein. For the sake of brevity, some operations, features and advantages are not repeated here.
[0164] According to one aspect of the present disclosure, a computer device is provided, comprising a memory, a processor, and a computer program stored in the memory. The processor is configured to execute the computer program to implement the steps of any one of the method embodiments described above.
[0165] According to one aspect of the present disclosure, a non-transitory computer-readable storage medium is provided, on which a computer program is stored. When the computer program is executed by a processor, the steps of any method embodiment described above are implemented.
[0166] According to one aspect of the present disclosure, a computer program product is provided, which includes a computer program. When the computer program is executed by a processor, the steps of any one of the method embodiments described above are implemented.
[0167] In the following, combined Figure 11 Illustrative examples of such a computer device, non-transitory computer-readable storage medium, and computer program product are described.
[0168] Figure 11 1 shows an example configuration of a computer device 1100 that can be used to implement the methods described herein. For example, Figure 1 The server 120 and / or the client device 110 shown in FIG may include an architecture similar to the computer device 1100. The above apparatus 1000 may also be implemented in whole or at least in part by the computer device 1100 or a similar device or system.
[0169] Computer device 1100 can be a variety of different types of devices. Examples of computer device 1100 include, but are not limited to, desktop computers, server computers, laptop or netbook computers, mobile devices (e.g., tablet computers, cellular or other wireless phones (e.g., smartphones), notepad computers, mobile stations), wearable devices (e.g., eyeglasses, watches), entertainment devices (e.g., entertainment appliances, set-top boxes communicatively coupled to a display device, game consoles), televisions or other display devices, automotive computers, and the like.
[0170] The computer device 1100 may include at least one processor 1102, memory 1104, communication interface(s) 1106, a display device 1108, other input / output (I / O) devices 1110, and one or more mass storage devices 1112, all capable of communicating with one another, such as via a system bus 1114 or other appropriate connections.
[0171] The processor 1102 may be a single processing unit or multiple processing units, all of which may include a single or multiple computing units or multiple cores. The processor 1102 may be implemented as one or more microprocessors, microcomputers, microcontrollers, digital signal processors, central processing units, state machines, logic circuits, and / or any device that manipulates signals based on operational instructions. Among other capabilities, the processor 1102 may be configured to retrieve and execute computer-readable instructions stored in the memory 1104, mass storage device 1112, or other computer-readable media, such as program code for an operating system 1116, program code for application programs 1118, program code for other programs 1120, and the like.
[0172] Memory 1104 and mass storage device 1112 are examples of computer-readable storage media for storing instructions that are executed by processor 1102 to implement the various functions described above. For example, memory 1104 may generally include both volatile memory and non-volatile memory (e.g., RAM, ROM, etc.). In addition, mass storage device 1112 may generally include a hard drive, a solid-state drive, removable media, including external and removable drives, memory cards, flash memory, floppy disks, optical disks (e.g., CDs, DVDs), storage arrays, network attached storage, storage area networks, etc. Memory 1104 and mass storage device 1112 may be collectively referred to herein as memory or computer-readable storage media, and may be non-transitory media capable of storing computer-readable, processor-executable program instructions as computer program code that may be executed by processor 1102 as a specific machine configured to implement the operations and functions described in the examples herein.
[0173] A plurality of programs may be stored on the mass storage device 1112. These programs include an operating system 1116, one or more application programs 1118, other programs 1120, and program data 1122, and they may be loaded into the memory 1104 for execution. Examples of such applications or program modules may include, for example, computer program logic (e.g., computer program code or instructions) for implementing the following components / functionality: the client application 112 (including a sending module 112a, a receiving module 112b, a pattern generation module 112c, and a presentation module 112d), the resource transfer application 128 (including a receiving module 128a, a resource transfer module 128b, a sending module 128c, a credit management module 128d, an access control module 128e, a transaction status setting module 128f, and a message generation module 128g), the method 200 and / or the method 300 (including any suitable steps of the methods 200, 300), and / or other embodiments described herein.
[0174] Although Figure 8 1100, but modules 1116, 1118, 1120, and 1122, or portions thereof, may be implemented using any form of computer-readable media accessible by the computer device 1100. As used herein, "computer-readable media" includes at least two types of computer-readable media, namely, computer-readable storage media and communication media.
[0175] Computer-readable storage media include volatile and non-volatile, removable and non-removable media implemented by any method or technology for storing information, such as computer-readable instructions, data structures, program modules or other data. Computer-readable storage media include, but are not limited to, RAM, ROM, EEPROM, flash memory or other memory technology, CD-ROM, digital versatile disk (DVD), or other optical storage devices, magnetic cassettes, magnetic tape, magnetic disk storage devices or other magnetic storage devices, or any other non-transmission media that can be used to store information for access by a computer device. In contrast, communication media can embody computer-readable instructions, data structures, program modules or other data in a modulated data signal such as a carrier wave or other transmission mechanism. Computer-readable storage media as defined herein does not include communication media.
[0176] One or more communication interfaces 1106 are used to exchange data with other devices, such as via a network, direct connection, and the like. Such communication interfaces can be one or more of the following: any type of network interface (e.g., a network interface card (NIC)), a wired or wireless (such as an IEEE 802.11 wireless LAN (WLAN)) wireless interface, a Worldwide Interoperability for Microwave Access (Wi-MAX) interface, an Ethernet interface, a Universal Serial Bus (USB) interface, a cellular network interface, a Bluetooth™ interface, a Near Field Communication (NFC) interface, and the like. The communication interface 1106 can facilitate communication within a variety of network and protocol types, including wired networks (e.g., LAN, cable, and the like) and wireless networks (e.g., WLAN, cellular, satellite, and the like), the Internet, and the like. The communication interface 1106 can also provide communication with external storage devices (not shown), such as storage arrays, network attached storage, storage area networks, and the like.
[0177] In some examples, a display device 1108 such as a monitor may be included for displaying information and images to the user. Other I / O devices 1110 may be devices that receive various inputs from the user and provide various outputs to the user, and may include a touch input device, a gesture input device, a camera, a keyboard, a remote control, a mouse, a printer, an audio input / output device, and the like.
[0178] The technology described herein can be supported by these various configurations of the computer device 1100 and is not limited to the specific examples of the technology described herein. For example, the functionality can also be implemented in whole or in part on a "cloud" by using a distributed system. The cloud includes and / or represents a platform for resources. The platform abstracts the underlying functionality of the hardware (e.g., servers) and software resources of the cloud. Resources can include applications and / or data that can be used when performing computing processing on a server away from the computer device 1100. Resources can also include services provided over the Internet and / or through a subscriber network such as a cellular or Wi-Fi network. The platform can abstract resources and functions to connect the computer device 1100 to other computer devices. Therefore, the implementation of the functionality described herein can be distributed throughout the cloud. For example, the functionality can be implemented partially on the computer device 1100 and partially through a platform that abstracts the functionality of the cloud.
[0179] Although the present disclosure has been illustrated and described in detail in the drawings and the foregoing description, such illustration and description are to be considered illustrative and exemplary and not restrictive; the disclosure is not limited to the disclosed embodiments. Variations to the disclosed embodiments will be understood and effected by those skilled in the art in practicing the claimed subject matter by studying the drawings, the disclosure and the appended claims. In the claims, the word "comprising" does not exclude other elements or steps that are not listed, the indefinite article "a" or "an" does not exclude a plurality, the term "plurality" means two or more, and the term "based on" is to be interpreted as "based at least in part on". The mere fact that certain measures are recited in mutually different dependent claims does not indicate that a combination of these measures cannot be used to advantage.< / eos> < / eos>
Claims
1. A method for reasoning with a large language model, comprising: The CPU and GPU collaborate to complete multiple iterations of the inference process, where each iteration includes the following tasks performed sequentially: A first task is configured to determine, on the CPU, a plurality of unfinished sequences and a last generated word element identifier of the plurality of sequences; The second task is configured to determine, on the CPU, a model input based on a word element identifier last generated by the plurality of sequences; A third task is configured to calculate, on the GPU, a probability distribution of multiple next tokens to be generated for the multiple sequences based on the model input; A fourth task is configured to perform sampling on the GPU based on the probability distribution of the multiple next word-grams to obtain word-gram identifiers of the multiple next word-grams; and The fifth task is configured to update the completion status of the multiple sequences based on the word-gram identifiers of the multiple next word-grams on the CPU, The (n+1)th iteration and the (n)th iteration are executed asynchronously, the first and second tasks of the (n+1)th iteration are executed in parallel with the third task of the (n)th iteration, and the third task of the (n+1)th iteration is executed in parallel with the fifth task of the (n)th iteration.
2. The method according to claim 1, wherein Determining the unfinished plurality of sequences and the last generated word element identifiers of the plurality of sequences includes: In response to determining that the plurality of sequences include a target sequence determined in the first task of the previous iteration, using a placeholder as a last generated word-gram identifier of the target sequence, The step of determining the model input based on the last generated word element identifier of the plurality of sequences includes: determining a first tensor comprising last generated word-gram identifiers of the plurality of sequences; and After the fourth task of the previous iteration samples the target word element identifier generated by the target sequence in the previous iteration, the target word element identifier is backfilled into the placeholder in the first tensor, wherein the model input includes the backfilled first tensor.
3. The method according to claim 2, wherein: Determining the model input based on the last generated word element identifier of the multiple sequences further includes: determining a second tensor, wherein the second tensor includes position indices of the plurality of next tokens, the model input including the first tensor; and Before the backfilling, the second tensor is sent to the GPU.
4. The method according to claim 3, wherein: Determining the model input based on the last generated word element identifier of the multiple sequences further includes: In response to determining that the plurality of sequences do not include any sequence determined in the first task of the previous iteration, after determining the first tensor, the first tensor is directly sent to the GPU.
5. The method according to any one of claims 1 to 4, wherein The first task of the (n+2)th iteration is executed after the fifth task of the (n)th iteration.
6. The method according to claim 3, wherein: The plurality of sequences include a plurality of decoding sequences and a plurality of pre-filling sequences, wherein the first task is configured to: Determining a word element identifier last generated by the plurality of decoding sequences; and Determine the prompt word element identifier of the plurality of pre-filled sequences input at the current iteration, The third task is configured as follows on the GPU: Calculating probability distributions of the multiple next tokens to be generated for the multiple decoding sequences; and Perform forward propagation calculation on the prompt word-unit identifiers inputted at the current iteration of the multiple pre-filled sequences.
7. The method according to claim 6, wherein: The first task is configured to determine a sequence status of each of the plurality of sequences, wherein the sequence status includes: The expected length of the corresponding sequence at the end of the current iteration; the current length of the corresponding sequence at the start of the current iteration; and The number of tokens generated by the corresponding decoding sequence in the current iteration or the number of prompt token identifiers input by the corresponding pre-filled sequence in the current iteration, The current length of each sequence in the plurality of sequences at the beginning of the current iteration is determined based on the expected length of the sequence at the end of the last iteration in which the sequence is inferred.
8. The method according to claim 7, wherein: The second task is configured as follows: The second tensor is determined based on a sequence state of each of the plurality of decoded sequences.
9. The method according to claim 7, wherein: The number of word units generated in the current iteration determined for each of the multiple decoding sequences is 1, and the number of prompt word units input in the current iteration determined for each of the multiple pre-filled sequences does not exceed a first preset number, and the first preset number is not less than 2.
10. The method according to any one of claims 1 to 4, wherein: The second task is configured to determine sampling metadata for each of the multiple sequences, and the sampling metadata for each of the multiple sequences is sent to the GPU when executing the third task. The fourth task is configured to perform sampling on the GPU based on the sampling metadata for each of the multiple sequences and a probability distribution of the next token to be generated.
11. The method according to any one of claims 1 to 4, wherein: After the third task is completed, each of the multiple working processes of the GPU has a probability distribution of a partial vocabulary of all or part of the multiple sequences, wherein the fourth task is configured as follows: Performing a many-to-many data interaction operation between the plurality of working processes so that each of the plurality of working processes has a probability distribution of the entire vocabulary of a portion of the plurality of sequences; Using each of the plurality of working processes to sample the probability distribution of the entire vocabulary of the next word of the partial sequence possessed by the working process; and The sampling results of the multiple working processes are aggregated to obtain the word-gram identifier of the next word-gram of each of the multiple sequences.
12. The method according to claim 11, wherein The model input further includes original indices of the plurality of sequences, where the indices of the plurality of sequences in the plurality of working processes are determined by performing a modulo operation on the original indices of the plurality of sequences according to the number of the plurality of working processes.
13. The method according to claim 11, wherein The fourth task is configured to: Before performing the many-to-many data interaction, the plurality of sequences are padded with a padding sequence so that the number of the padded plurality of sequences is divisible by the number of the plurality of working processes; as well as After aggregating the sampling results of the multiple working processes, the sampling results of the filling sequence are discarded.
14. The method according to any one of claims 1 to 4, wherein: The same third preset number of random numbers is pre-generated on the multiple work processes, and each work process uses a subset of the third preset number of random numbers when executing the third task and / or the fourth task.
15. The method according to any one of claims 1 to 4, wherein: The first task is configured to determine at least one KV cache block allocated to each of the multiple sequences to obtain KV cache block indexes of the multiple sequences, wherein the model input includes the KV cache block indexes of the multiple sequences.
16. The method according to claim 15, wherein A single KV cache block can accommodate keys and values corresponding to a second preset number of tokens, wherein determining at least one KV cache block allocated to each of the plurality of sequences to obtain a KV cache block index for each of the plurality of sequences includes: The number of KV cache blocks allocated to the sequence is determined based on the sequence length of the sequence and the second preset number.
17. An inference device for a large language model, configured to utilize a CPU and a GPU to collaboratively complete a multi-iteration inference process, wherein: Each iteration includes a first task, a second task, a third task, a fourth task, and a fifth task that are performed sequentially, and the apparatus includes: a first processing unit for performing the first task, wherein the first task is configured to determine, on the CPU, a plurality of unfinished sequences and a last generated word element identifier of the plurality of sequences; a second processing unit for performing the second task, wherein the second task is configured to determine a model input based on a last generated word element identifier of the plurality of sequences on the CPU; a third processing unit for performing the third task, wherein the third task is configured to calculate, on the GPU, a probability distribution of a plurality of next tokens to be generated for the plurality of sequences based on the model input; a fourth processing unit for performing the fourth task, wherein the fourth task is configured to perform sampling on the GPU based on the probability distribution of the multiple next word-grams to obtain word-gram identifiers of the multiple next word-grams; and a fifth processing unit for performing the fifth task, wherein the fifth task is configured to update the completion status of the plurality of sequences based on the word-gram identifiers of the plurality of next word-grams on the CPU; The (n+1)th iteration and the (n)th iteration are executed asynchronously, the first and second tasks of the (n+1)th iteration are executed in parallel with the third task of the (n)th iteration, and the third task of the (n+1)th iteration is executed in parallel with the fifth task of the (n)th iteration.
18. A computer device comprising: at least one processor; as well as a memory having a computer program stored thereon, When the computer program is executed by the processor, the processor is caused to perform the method according to any one of claims 1 to 16.
19. A computer-readable storage medium having a computer program stored thereon, wherein when the computer program is executed by a processor, the processor is caused to perform the method according to any one of claims 1 to 16.
20. A computer program product comprising a computer program, which, when executed by a processor, causes the processor to perform the method according to any one of claims 1 to 16.
Citation Information
Cited By
Data transmission method and device, storage medium and electronic equipment
CN120892163A
Data transmission method and apparatus, storage medium, and electronic device
CN120892163B