Request processing method and device for large language model
By setting multiple generation segments in the output text of the large language model and using speculative sampling method, the problem of high latency for LLM processing user requests in the prior art is solved, and more efficient request processing is achieved, reducing latency and improving throughput.
Patent Information
- Application Number
- CN202510081518.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-17
- Publication Date
- 2025-05-16
AI Technical Summary
The prior art is difficult to effectively reduce latency or improve throughput when calling large-scale language models (LLMs) to process user requests, mainly due to the high latency caused by the inability to fully utilize hardware computing performance and serial dependency characteristics in the decoding stage.
By presetting multiple generation segments in the output text of the large language model, each generation segment is configured with a starting word element, a termination word element and a corresponding query sampling corpus, and the speculative sampling method is used for processing. The specific steps include determining the corresponding generated segment based on the word sequence in the request, and speculating based on the query sampling corpus and large language model of the segment to generate a continuous word sequence.
By improving the prediction accuracy of the speculative sampling algorithm, the number of forward calculations in the LLM decoding phase is reduced, the delay in processing user requests is significantly reduced and the service throughput is improved.
Smart Images

Figure CN120012781A_ABST
Abstract
Description
Technical Field
[0001] One or more embodiments of the present specification relate to the field of machine learning technology, and more particularly, to a method and device for processing requests for a large language model. Background Art
[0002] Large Language Model (LLM) is a deep learning model trained on massive text data. It can understand the meaning of text and generate natural language text based on given text, and handle various natural language tasks such as text summarization, question answering, translation, etc.
[0003] Currently, when calling LLM to process user requests, we hope to effectively reduce latency or increase throughput. Summary of the invention
[0004] The embodiments of this specification describe a method and device for processing requests for a large language model, which can meet higher requirements in practical applications.
[0005] According to a first aspect, a method for processing a request of a large language model is provided, the implementation of which is based on a plurality of generation segments pre-set for the output text of the large language model, wherein each generation segment is configured with a corresponding start word and end word, and is configured with a corresponding query sampling corpus. The method comprises: for a first request, when it is determined that it is in a decoding stage, determining a corresponding first generation segment according to a first word sequence included in the request; based on the first query sampling corpus corresponding to the first generation segment and the large language model, performing speculative sampling processing on the first word sequence, obtaining a target word sequence following the first word sequence as a processing result of the first request.
[0006] In some embodiments, determining the corresponding first generation segment according to the first word-gram sequence included therein includes: determining the most recently generated starting word-gram in the first word-gram sequence, thereby determining the generation segment corresponding to the starting word-gram as the first generation segment.
[0007] In some embodiments, the large language model undergoes the following fine-tuning processing: obtaining a fine-tuning sample, which includes an original first label word-meta sequence; inserting the start word and the end word corresponding to each generated segment according to the division position of the label word-meta sequence corresponding to each generated segment, to obtain a modified second label word-meta sequence; using the fine-tuning sample in which the first label word-meta sequence is replaced with the second label word-meta sequence, the pre-trained large language model is fine-tuned.
[0008] In some embodiments, each generation segment is configured with a corresponding speculative sampling algorithm; wherein, based on the first query sampling corpus corresponding to the first generation segment, speculative sampling is performed on the first word sequence, including: using the first speculative sampling algorithm corresponding to the first generation segment to perform the speculative sampling, the first speculative sampling algorithm is one of the following: single branch prediction, depth-first search algorithm under multi-branch prediction, breadth-first search algorithm under multi-branch prediction.
[0009] In some embodiments, the speculative sampling process includes: performing query sampling on the first query sampling corpus according to the first word-gram sequence to obtain a plurality of predicted word-gram sequences following the first word-gram sequence; calling the large language model to process the plurality of inputs constructed based on the plurality of predicted word-gram sequences to obtain a plurality of inference results; and using the plurality of inference results to verify the plurality of predicted word-gram sequences to obtain the target word-gram sequence.
[0010] Further, in some specific embodiments, the first query sampling corpus includes a plurality of key-value pairs, and the key and value in each key-value pair are in the form of a word sequence, and have frequency information of the value being connected after the key. The query sampling includes: determining a first key based on the first word sequence; searching for a plurality of key-value pairs containing the first key in the first query sampling corpus, and determining the plurality of predicted word sequences based on the values in a predetermined number of key-value pairs whose frequencies are ranked in the front range.
[0011] In some other specific embodiments, the construction of the multiple inputs includes: for any first prediction sequence among the several predicted word-element sequences, splitting it into multiple sub-sequences, each of the multiple sub-sequences containing a different number of consecutive word-elements starting from the first word-element in the first prediction sequence; respectively splicing the multiple sub-sequences after the first word-element sequence to obtain multiple spliced sequences; and classifying the first word-element sequence and the multiple spliced sequences as the multiple inputs.
[0012] In some embodiments, after obtaining the target word-gram sequence, the method further includes: updating the first query sampling corpus based on the target word-gram sequence when the last word-gram of the target word-gram sequence is the terminal word-gram corresponding to the first generated segment.
[0013] In some embodiments, after obtaining the target word-gram sequence, the method further includes: when the last word of the target word-gram sequence is not a text-ending word-gram, or when the length of the generated sequence up to the target word-gram sequence does not reach a length threshold, constructing a new request to be processed based on the target word-gram sequence.
[0014] In some embodiments, the method further includes: scheduling a batch of requests participating in this batch calculation based on a current request queue, wherein any one of the requests is used as the first request.
[0015] According to the second aspect, a method for processing requests of a large language model is provided, wherein the implementation of the method is based on a plurality of generation segments pre-set for the output text of the large language model, wherein each generation segment is configured with a corresponding start word and end word, and is configured with a corresponding preliminary model. The method comprises: for a first request, when it is determined that it is in a decoding stage, determining a corresponding first generation segment according to a first word sequence included therein. Based on the first preliminary model corresponding to the first generation segment and the large language model, a speculative sampling process is performed on the first word sequence to obtain a target word sequence following the first word sequence as a processing result of the first request.
[0016] In some embodiments, the speculative sampling process includes: inputting the first word-gram sequence into the first preliminary model to obtain a predicted word-gram sequence following the first word-gram sequence; calling the large language model to process the multiple inputs constructed based on the predicted word-gram sequence to obtain multiple inference results; and using the multiple inference results to verify the predicted word-gram sequence to obtain the target word-gram sequence.
[0017] According to a third aspect, a request processing device for a large language model is provided, wherein the output text of the large language model is pre-set as a plurality of generated segments, wherein each generated segment is configured with a corresponding start word and end word, and is configured with a corresponding query sampling corpus. The device comprises: a generated segment determination module, configured to determine the corresponding first generated segment according to the first word sequence included in the first request when it is determined that the request is in the decoding stage; and a speculative sampling module, configured to perform speculative sampling processing on the first word sequence based on the first query sampling corpus corresponding to the first generated segment and the large language model, and obtain a target word sequence following the first word sequence as the processing result of the first request.
[0018] According to a fourth aspect, a request processing device for a large language model is provided, wherein the output text of the large language model is pre-set as a plurality of generated segments, wherein each generated segment is configured with a corresponding start word and end word, and is configured with a corresponding preliminary model. The device comprises: a generated segment determination module, configured to determine the corresponding first generated segment according to the first word sequence included in the first request when it is determined that the request is in the decoding stage; and a speculative sampling module, configured to perform speculative sampling processing on the first word sequence based on the first preliminary model corresponding to the first generated segment and the large language model, and obtain a target word sequence following the first word sequence as the processing result of the first request.
[0019] According to a fifth aspect, a computer-readable storage medium is provided, on which a computer program is stored. When the computer program is executed in a computer, the computer is caused to execute the method provided in the first aspect or the second aspect.
[0020] According to a sixth aspect, a computing device is provided, including a memory and a processor, wherein the memory stores executable code, and when the processor executes the executable code, the method provided in the first aspect or the second aspect is implemented.
[0021] In summary, by adopting the above method and device disclosed in the embodiments of this specification, when providing LLM model services, the delay in processing user requests can be effectively reduced and the throughput of service requests can be improved. BRIEF DESCRIPTION OF THE DRAWINGS
[0022] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the following briefly introduces the drawings required for use in the description of the embodiments. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without creative work.
[0023] Figure 1 A schematic diagram of the use of the segmented vocabulary disclosed in the embodiments of this specification;
[0024] Figure 2 This is one of the flowcharts of the request processing method of the large language model disclosed in the embodiments of this specification;
[0025] Figure 3 For the embodiments disclosed in this specification Figure 2 Schematic diagram of the sub-step flow of step S220;
[0026] Figure 4 The single branch prediction algorithm disclosed in the embodiment of this specification is used to execute Figure 3 Example diagram of the process;
[0027] Figure 5 A schematic diagram of multiple prediction sequences obtained by using a multi-branch prediction algorithm disclosed in an embodiment of this specification;
[0028] Figure 6 This is a second flow chart of the request processing method of a large language model disclosed in the embodiment of this specification;
[0029] Figure 7 This is one of the structural schematic diagrams of the request processing device of the large language model disclosed in the embodiments of this specification;
[0030] Figure 8 This is the second structural diagram of the request processing device for a large language model disclosed in the embodiments of this specification. DETAILED DESCRIPTION
[0031] The solution provided in this specification is described below in conjunction with the accompanying drawings.
[0032] As mentioned above, when calling LLM to process user requests, we hope to effectively reduce latency or improve throughput.
[0033] Specifically, since LLM usually uses an autoregressive approach to generate subsequent tokens one by one according to the prompts input by the user, for example, assuming that LLM generates the tth token at the tth time step, then at the t+1th time step, the tth token needs to be used as input to generate the t+1th token. This means that each execution of the forward calculation of LLM can only generate one token.
[0034] When the scale of LLM is relatively large, the time per output token (TPOT) for calculating a single word will increase significantly. A large amount of time is spent on memory access operations, and due to the serial dependency of model reasoning, the decoding stage cannot be calculated in parallel, making it difficult for the decoding stage of LLM reasoning to fully utilize the computing performance of the hardware.
[0035] In order to solve the problem that the decoding stage cannot fully utilize hardware features and serial dependency, which causes high TPOT and affects user experience, the industry has developed a variety of algorithms based on speculative sampling to improve the parallelism of the decoding stage. Such algorithms usually first quickly predict the word sequence that may be generated by LLM, such as using a lightweight preliminary model (draft model) for prediction, and then verify the correctness of the predicted word sequence in the actual reasoning stage of LLM, thereby accelerating the reasoning of the decoding stage.
[0036] Therefore, the accuracy of predicting generated words has a crucial impact on the acceleration effect of the speculative sampling algorithm on LLM reasoning. The higher the prediction accuracy of the speculative sampling algorithm, the fewer forward passes are required for LLM to process user requests; conversely, the lower the prediction accuracy, the more computing resources will be wasted, which will increase the delay of the LLM reasoning system in processing requests.
[0037] Furthermore, the applicant observed that the generation of LLM is localized. Specifically, in different generation segments (or generation intervals, components), word units present different probability distributions. For example, in different parts of a generated text, such as the title, introduction, body, and conclusion, word units present different probability distributions. Especially for fine-tuning large models for specific scenarios, the differences in word unit probability distributions in different generation segments are becoming increasingly large.
[0038] Based on the above observations and analysis, the applicant proposes an improved scheme for accelerating LLM reasoning using speculative sampling. In one embodiment of this improved scheme, different vocabularies are configured for different generated sections of the LLM output text, so that when LLM is used to process user requests, speculative sampling is performed according to the vocabularies corresponding to the generated section to which the current word to be predicted belongs.
[0039] For example, Figure 1 Any Pi and Dj in represents a single word, and P0-P3 represent the input text contained in the user request, D0-D11 represent the output text generated by LLM, D0-D3 in the output text belong to the first generation segment, D4-D8 belong to the second generation segment, and D9-D11 belong to the third generation segment. The word unit distribution in these three generation segments corresponds to vocabulary 1, vocabulary 2 and vocabulary 3 respectively.
[0040] Therefore, by adopting the improved scheme, the hit probability of the predicted generated word can be improved, that is, the accuracy of the speculative sampling algorithm prediction can be improved, thereby effectively reducing the delay of LLM in processing user requests and improving the throughput.
[0041] To help you understand, let's first introduce the preliminary preparations for the improvement plan, including the division of generated segments and the construction of the vocabulary, as well as the fine-tuning of the LLM model.
[0042] 1. Division of generated segments
[0043] It should be understood that the division or pre-setting of generated sections mainly depends on human experience. For different application scenarios, the division method is generally different. For example, in the email generation scenario, the email text output by LLM can be divided into the following generated sections: salutation (the opening salutation, indicating the recipient of the email), introduction (briefly describing the purpose of the email), body (detailed description of the email content), closing remarks (indicating the expected reply or action) and signature (signature and contact information).
[0044] For another example, in a story generation scenario, the story text output by LLM can be divided into the following generated sections: title (summarizing the theme of the story), introduction (describing the background of the story, introducing the main characters and plot), development (unfolding the plot, describing the conflicts and changes of events), climax (the peak of the story, the key point for solving the problem) and ending (ending the story, resolving the conflict or triggering thinking).
[0045] The above is an exemplary introduction to the division of the generated segments.
[0046] 2. Construction of vocabulary
[0047] To help understanding, we first introduce the contents of the vocabulary and then the construction method.
[0048] It should be noted that each of the multiple generated sections in a specific application scenario has an independent vocabulary. Each vocabulary includes multiple key-value pairs, where the key and value in each key-value pair are in the form of word-gram sequences. A word-gram sequence is a sequence of word-grams arranged in chronological order, which may include a single word-gram or multiple word-grams; in addition, a word-gram is the smallest processing unit of LLM for text, which may be a character, word, phrase or special symbol.
[0049] The length of the word sequence corresponding to the key and value can be set as needed. For example, the key and value in a key-value pair both include 1 or 2 words. For another example, the key and value in a key-value pair include 4 and 2 words respectively. In addition, the length of the word sequence corresponding to the key (or value) in any two key-value pairs can be the same or different.
[0050] The vocabulary also records the probability (or frequency) that the value in each key-value pair follows the key. It can be understood that, generally, in a vocabulary, for multiple key-value pairs containing the same key, the sum of the associated multiple probabilities is 1.
[0051] Exemplarily, in a business email generation scenario, the vocabulary corresponding to the title part includes the contents shown in Table 1 below, and the contents in the symbol "" are a word element.
[0052] Table 1
[0053] key value Probability "project" "Progress" - "Report" 0.5 "project" "Status" - "Update" - "Notification" 0.3 "project" "Postponement" - "Notice" 0.1 "project" "Summarize" 0.1 "Meeting" - "Convened" "notify" 0.4 ...... ...... ......
[0054] The above describes the key-value pairs and frequency information contained in the vocabulary. The following describes the construction of the vocabulary.
[0055] The construction of the vocabulary requires the use of corpus in the application scenario. Specifically, multiple complete texts in the application scenario can be collected, and each complete text can be divided into multiple parts corresponding to multiple generated segments, so as to construct a segment text set under each generated segment.
[0056] Furthermore, the corresponding vocabulary can be constructed using the segment text set. It should be understood that the construction of the corresponding vocabulary based on the text data can be implemented in an existing manner without limitation. For example, each segment text in the segment text set can be segmented to obtain a corresponding word sequence, which is then segmented to obtain multiple candidate key-value pairs, and then the frequency is counted and the low-frequency candidate key-value pairs are removed, and finally a constructed vocabulary is obtained. It should be understood that the visualization form of the vocabulary can be a vocabulary tree; in addition, the constructed vocabulary may be updated in the subsequent LLM reasoning stage, which will be specifically introduced later.
[0057] The above introduces the construction method and composition of the vocabulary.
[0058] 3. Fine-tuning of LLM
[0059] It can be understood that the pre-trained LLM can be directly used for reasoning. Generally, in order to improve the reasoning accuracy of LLM in a specific application scenario, the pre-trained LLM can be fine-tuned using fine-tuning samples in the corresponding scenario. The fine-tuning samples include sample input and label text, and the label text corresponds to a label word sequence.
[0060] As mentioned above, when using LLM to process user requests, it is necessary to determine the generation segment to which the current word to be predicted belongs, so as to perform speculative sampling based on the corresponding vocabulary. The determination of the generation segment needs to be judged based on the generated word. Specifically, each generation segment has a pre-configured special start word and end word, so that the generation segment corresponding to the current word to be generated can be judged based on the start word / end word included in the generated word. This requires LLM to have the ability to generate start words and end words. In this regard, the applicant proposes to insert start words and end words in the fine-tuning sample, so that LLM can learn to identify and output special start and end words.
[0061] Specifically, for any fine-tuning sample, the original label word-gram sequence can be marked to correspond to the division position of each generated segment, so as to insert the start word-gram and the end word-gram corresponding to each generated segment to obtain the transformed label word-gram sequence. It should be noted that the marking of the division position and the insertion of special words can be done manually or by a pre-trained LLM model. For example, prompt words can be designed to describe the content characteristics of each generated segment and the corresponding start and end words, so as to obtain the transformed label word-gram sequence output by the LLM.
[0062] For example, the original tag word sequence is: "project" - "progress" - "report" - "dear" - ..., and the transformed tag word sequence is: -"Project" -"Progress" -"Report" -"※" -"Dear" -.......
[0063] It should be understood that in computers, tokens are generally represented as integers, which have a one-to-one correspondence with characters, words, phrases, symbols, etc. with explicit semantics. For special start and end tokens, they can be configured as special integers that do not correspond to regular characters, words, and phrases. Therefore, when LLM performs actual reasoning output, the computer is aware of the start and end tokens, but the user is not. In addition, in the above example, and “※” are only used as intuitive examples. In fact, the start word unit and the end word unit may not correspond to any words or symbols with well-known semantics in the real world.
[0064] From the above, we can get the fine-tuned LLM. For the sake of brevity, the fine-tuned LLM will be referred to as LLM in the following.
[0065] After completing the preliminary preparations in the above improvement scheme, the vocabulary, starting word unit and ending word unit corresponding to each of the multiple generated sections of the output text, and the fine-tuned LLM can be obtained. Based on this, the LLM can be called to process user requests.
[0066] Figure 2 This is one of the flowcharts of the request processing method of the large language model disclosed in the embodiments of this specification. The execution subject of the processing method can be any device, platform, server or equipment cluster with computing and processing capabilities. Figure 2 As shown, the method comprises the following steps:
[0067] Step S210: for the first request, when it is determined that the request is in the decoding stage, a corresponding first generation segment is determined according to the first word-gram sequence included therein. Step S220: based on the first vocabulary corresponding to the first generation segment and the large language model, a speculative sampling process is performed on the first word-gram sequence to obtain a target word-gram sequence following the first word-gram sequence as a processing result of the first request.
[0068] The following is an expanded description of the above steps:
[0069] First, in step S210, for the first request, when it is determined that the request is in the decoding stage, a corresponding first generation segment is determined according to the first word sequence included in the request.
[0070] It should be noted that the "first" in "first request" and "first word sequence", as well as similar terms such as "second" and "third" elsewhere in the text, are all for distinguishing things of the same kind and do not have other limiting functions such as sorting.
[0071] The first request may be any request currently to be processed. For example, a batch of requests participating in the current batch calculation may be scheduled based on the current request queue, and each of the requests may be used as the first request.
[0072] The first request to be processed may be in the prefill stage or the decode stage. To help understanding, these two stages are briefly introduced first.
[0073] LLM mainly includes the pre-filling stage and the decoding stage during inference, and these two stages are executed serially. The pre-filling stage includes the following processing: After the user inputs the data to be processed into the LLM, the LLM self-encodes the data to be processed and generates attention information for the data to be processed, for example, generates KV cache data of the user input word in the data to be processed, and generates the first word. The KV cache data and the first word are used in the decoding stage. The decoding stage refers to the process in which the LLM actually generates the text to be output. In this stage, given the context information generated in the pre-filling stage, including the KV cache data and the first word, the model reasoning continues to be run, and the words of the subsequent text to be output are generated one by one.
[0074] There are many ways to determine whether the first request is in the pre-filling stage or the decoding stage, and only an exemplary description is given. For example, it is determined whether the first request initiates a reset of the KV cache data. If so, it is determined that the first request is in the pre-filling stage, otherwise it is in the decoding stage. For another example, it can be determined whether the first request includes a prompt word (prompt) input by the user. If it is included, it indicates that the first request is in the pre-filling stage, otherwise it is in the decoding stage.
[0075] Further, when it is determined that the first request is in the decoding stage, the corresponding first generation section is determined according to the first word-gram sequence included therein. It should be understood that the first word-gram sequence may include one or more word-grams that have been generated. For example, the first word-gram sequence may only include new word-grams generated in the previous step, such as the first word-gram generated in the pre-filling stage, or the word-gram most recently generated in the decoding stage.
[0076] Specifically, the first generated segment corresponding to the first word element sequence can be determined according to pre-configured rules. The configuration basis of the rules includes the starting word element and the ending word element corresponding to each of the multiple generated segments, and the configured rules are feasible and not unique. Exemplarily, the most recently generated first starting word element can be extracted from the generated sequence corresponding to the first word element sequence, and then the first generated segment corresponding to the first starting word element can be determined according to the mapping relationship between the generated segment and the starting word element. For another example, the first starting word element that does not yet contain a paired ending word element can be extracted from the generated sequence corresponding to the first word element sequence, and then the corresponding first generated segment can be determined according to the mapping relationship between the generated segment and the starting word element.
[0077] In addition, considering that the LLM model is not limited to processing requests under the above-mentioned specific application scenarios (such as email generation), at this time, there are cases where the word sequence corresponding to some requests does not correspond to any generated segments. For these requests, they can be processed based on a pre-configured catch-all word list that does not correspond to any generated segments. The processing process can be simply inferred by referring to other content in the article and will not be elaborated on.
[0078] Based on the above, the first generated section corresponding to the first request can be determined.
[0079] Step S220: Based on the first vocabulary corresponding to the first generated segment and the large language model, perform speculative sampling processing on the first word sequence to obtain a target word sequence following the first word sequence as a processing result of the first request.
[0080] It should be noted that the speculative sampling process in this step can be implemented using an existing speculative sampling algorithm.
[0081] In addition, the applicant also proposed to optimize the speculative sampling processing in this step. Specifically, the mapping relationship between the generation segment and the speculative sampling algorithm can be pre-configured, so that in this step, the first speculative sampling algorithm corresponding to the first generation segment is determined according to the mapping relationship, and the first speculative sampling algorithm is used to perform speculative sampling processing. Exemplarily, the various speculative sampling algorithms involved in the mapping relationship may include: single branch prediction (or single candidate sequence prediction), multi-branch prediction (or multi-candidate sequence prediction), and depth-first search algorithm (Depth-First Search, referred to as DFS) and breadth-first search algorithm (Breadth-First Search, referred to as BFS) under multi-branch prediction. In this way, it is possible to flexibly support multiple sampling modes and multiple speculative sampling algorithms, and further optimize the performance of latency and throughput.
[0082] To help understand, the following Figure 3 , Figure 4 and Figure 5 , an exemplary introduction to the implementation of this step is given, wherein Figure 3 The sub-steps of this step include: Figure 4 is based on Figure 3 A specific example of , which uses a single branch prediction algorithm, Figure 5 The sampling results of multi-branch prediction are shown in FIG. Figure 3 As shown, the implementation of this step S220 may include the following sub-steps S31-S34:
[0083] Step S31, query sampling is performed on the first word list according to the first word sequence to obtain a plurality of predicted word sequences following the first word sequence. It should be understood that the term "several" in the text refers to one or more.
[0084] Specifically, the first key may be determined based on the first word sequence. In one embodiment, the first word sequence may be directly used as the first key. In another embodiment, the last word or the last two words in the first word sequence may be used as the first key.
[0085] Further, several key-value pairs containing the first key can be queried in the first vocabulary, and the several predicted word-unit sequences can be determined based on the mapping values in a predetermined number of key-value pairs whose frequencies are ranked in the front range. It should be understood that the length n of the predicted word-unit sequence can be set as needed. In addition, obtaining several predicted word-unit sequences may go through one or more layers of sampling. For example, when n=2, the first word in the predicted word-unit sequence may be determined based on the first key, and then the first word is used as the key to continue sampling to obtain the second word in the predicted word-unit sequence.
[0086] In one embodiment, query sampling can adopt a sampling mode under single-branch prediction, that is, only the key-value pair with the highest frequency is selected from the vocabulary. In this case, the sampled predicted word-unit sequences are a single predicted word-unit sequence. This mode can efficiently and parallelly perform model decoding inference using fewer computing resources when the prediction accuracy of speculative sampling is high.
[0087] For example, see Figure 4 , where the first word sequence R 0 For: R T0 -R T1 -R T2 -R T3 -R T4 -R T5 , based on R 0 The predicted word sequence obtained by performing single-branch prediction in the first word list is: R (T6) -R (T7) -R (T8) -R (T9) .
[0088] In another embodiment, considering that when the accuracy of speculative sampling prediction is low, if the length of the predicted generated sequence is long, when a certain word in the predicted sequence does not match the actual reasoning result, the subsequent calculations will be wasted, and at this time, the parallel speculative sampling has a poor effect on the reasoning acceleration of the decoding stage. It is proposed that query sampling can also adopt a sampling mode under multi-branch prediction (such as DFS or BFS), that is, multiple groups of key-value pairs with higher frequencies are selected from the vocabulary. At this time, the several predicted word-unit sequences obtained by sampling are multiple predicted word-unit sequences, which can improve the fault tolerance of the prediction results, thereby further improving the accuracy of speculative sampling prediction.
[0089] For example, see Figure 5 , if the current generated word is 1, the look-ahead predicts multiple groups of candidate words (2, 3, 4); at the same time, in the second layer of the tree, if the current word is 2, the look-ahead predicts multiple groups of candidate words (5, 6); if the current word is 3, the look-ahead predicts candidate word 7, .... Thus, the predicted candidate word sequences can be obtained, which are 1-2-5-9, 1-2-6, 1-3-7 and 1-4-8 respectively.
[0090] From the above, one or more predicted word-gram sequences following the first word-gram sequence can be obtained.
[0091] Step S32: construct multiple inputs for a large language model (LLM) based on the plurality of predicted word-gram sequences and the first word-gram sequence.
[0092] Specifically, due to the autoregressive method of LLM, each predicted word-unit sequence needs to be split and then concatenated to the end of the first word-unit sequence one by one as the input sequence of LLM.
[0093] Taking any first prediction sequence among several prediction word-unit sequences as an example, it can be split into multiple subsequences, each of which contains a different number of consecutive word-units starting from the first word-unit in the first prediction sequence. For example, assuming that the first prediction word-unit sequence includes n words, the first prediction word-unit sequence can be split into n subsequences. Furthermore, the multiple subsequences can be spliced after the first word-unit sequence to obtain multiple spliced sequences, so that the first word-unit sequence and the multiple spliced sequences are classified as the multiple inputs.
[0094] For example, see Figure 4 , where the first word sequence R is shown 0 And the four concatenated sequences constructed from the predicted word sequence: R 1 , R 2 , R 3 and R 4 .
[0095] From the above, we can get multiple input sequences for LLM.
[0096] Step S33, calling the large language model LLM to process the constructed multiple inputs to obtain multiple inference results.
[0097] It should be understood that this step can process multiple inputs in parallel, thereby giving full play to the parallel capabilities of GPU hardware and increasing service throughput. In addition, the memory access operation during LLM batch inference is not described here.
[0098] For example, see Figure 4 , call LLM on the sequence R 0 , R 1 , R 2 , R 3 and R 4 By performing parallel processing, we can obtain the result of a single forward calculation, or the inference result: R′ 0 , R′ 1 , R′ 2 , R′ 3 and R′ 4 .
[0099] In this way, the serial limitation brought by LLM autoregression can be overcome, parallel processing can be achieved, and multiple inference results can be obtained.
[0100] Step S34: compare and verify the predicted word-gram sequences using the multiple inference results to obtain a target word-gram sequence.
[0101] Specifically, starting from the first word position after the first word sequence, the inference result and the word at the same position in the predicted word sequence can be compared. When the inference word and the predicted word are the same, it means that the predicted word is valid and reliable.
[0102] And so on, until the inference word unit and the predicted word unit are different, then:
[0103] Subsequent predicted word units starting from different word units are discarded, and the sequence composed of credible predicted word units is used as the target word unit sequence. Figure 4 , through comparison and verification, we can get the effective predicted word as R (T6) and R (T7) , so R (T6)- R (T7) As the target word-gram sequence, or, the concatenated sequence of the first word-gram sequence and the credible predicted word-gram is taken as the target word-gram sequence.
[0104] Alternatively, the subsequent predicted word-grams starting from different word-grams are corrected (or regenerated) using LLM, so that the predicted word-grams that have been verified to be credible and the regenerated predicted word-grams are concatenated to obtain the above-mentioned target word-gram sequence.
[0105] In addition, there is a situation where, assuming that each predicted word has passed verification and is considered credible, the inference word predicted based on the predicted word sequence in the inference result is also considered credible. In this case, the longest inference sequence can be directly used as the target word sequence. For example, assuming Figure 4 All four predicted words in are verified. At this time, the inference sequence R′ is considered 4 The last word R′ in (T10) It is also credible, and R′ can be directly 4 as the target word sequence.
[0106] From the above, speculative sampling processing can be implemented on the first word sequence in the first request based on the first word list, so as to obtain a target word sequence following the first word sequence as the processing result of the first request.
[0107] Combination of the above Figure 3 , Figure 4 and Figure 5 ,right Figure 2 The execution of step S220 is exemplarily introduced.
[0108] Back to Figure 2 According to another aspect of the embodiment, after executing step S220, the method may further include: when the last word (i.e., the last word) of the target word sequence is the terminal word corresponding to the first generated segment, updating the first vocabulary based on the target word sequence.
[0109] Specifically, the key-value pairs and associated frequencies in the first vocabulary may be updated according to the segment sequence between the generated word-gram and the termination word-gram corresponding to the first generated segment in the generated sequence where the target word-gram sequence is located. It is understandable that after the update, the frequencies of some key-value pairs increase, the frequencies of some key-value pairs decrease, and in addition, the key-value pairs whose statistical frequencies are lower than the frequency threshold may be eliminated.
[0110] According to another embodiment, after executing step S220, the method may further include: when the last word of the target word sequence is not a text end word, or when the length of the generated sequence up to the target word sequence does not reach a length threshold, a new request to be processed can be constructed based on the target word sequence; otherwise, a model service result returned to the user is constructed based on the target word sequence.
[0111] Exemplarily, a new pending request may be added to a request queue.
[0112] In summary, the improved scheme disclosed in the embodiments of this specification is adopted to construct multiple independent probability prediction intervals or multiple independent word lists according to the generated features of LLM at different stages. Each interval can be configured with different speculative sampling algorithms according to actual needs to accelerate speculative sampling and improve the hit probability of the speculative sampling algorithm. Specifically:
[0113] 1) Reduce latency based on speculative sampling. This solution first guesses a set of results for each request in the batch request by querying the vocabulary, increasing the amount of data in one round of calculation, making full use of the parallel computing capabilities of the GPU, and using the LLM model to verify the results. This solution can significantly improve service throughput and reduce end-to-end latency.
[0114] 2) Based on single candidate or multiple candidate search, construct independent probability prediction intervals to improve the prediction accuracy of the speculative sampling method and reduce the number of forward calculations required for the decoding stage. In some experiments, the number of forward calculations for some generated segments was reduced from 7.4 to 2.5; the hit rate of speculative sampling was greatly improved, and the average time consumption was reduced to 67.4% of the original.
[0115] 3) Flexible support for multiple sampling modes and multiple speculative sampling algorithms. This solution supports multi-branch LookAhead algorithm, single-branch LookAhead algorithm; BFS search, DFS search; and supports the selection of single sampling length according to business needs. Users can adjust the service focus indicators to obtain shorter latency or higher throughput. If there are no focus indicators, the inference system can also adaptively adjust the speculative sampling strategy to ultimately achieve better latency and throughput performance.
[0116] It should be noted that in the above introduction to the improvement scheme, the prediction and generation of word-grams adopts the query sampling method, specifically, query sampling is performed on the first word table according to the first word-gram sequence in the first request, so as to obtain several predicted word-gram sequences.
[0117] Based on this, the applicant proposes that the vocabulary is essentially a corpus for query sampling (or simply the query sampling corpus), and the form of the query sampling corpus is not limited to the vocabulary, for example, it can also be the segment text set used to construct the vocabulary, or the candidate text set obtained by performing a sliding window process of a predetermined length on each segment text in the segment text set. Accordingly, the first vocabulary in the above steps S220 and S31 is replaced by the first query sampling corpus.
[0118] In one embodiment, assuming that the first query sampling corpus is the first candidate text set, the query sampling in step S31 can be implemented as follows: extracting several candidate texts containing the first word-gram sequence from the first candidate text set, and intercepting the text portion located after the first word-gram sequence in each candidate text, and classifying it into the several predicted word-gram sequences. It can be understood that word segmentation processing is involved from the intercepted text to the corresponding predicted word-gram sequence, which will not be described in detail.
[0119] As described above, query sampling can be implemented based on query corpora in other forms other than vocabulary.
[0120] On the other hand, the applicant also proposed that the prediction of generated word units is not limited to the query sampling method, but can also be a model prediction method. Specifically, different preliminary models (draft models) are pre-configured for different generated segments, where each preliminary model can learn the independent word unit probability distribution under the corresponding generated segment through pre-training.
[0121] For the first preliminary model corresponding to any first generated segment, its training sample can be constructed based on the corresponding first segment text set. For example, for any segment text in the first segment text set, it can be divided into two parts, the first part is used as sample input, and the second part is used as sample label. In addition, the model structures of any two preliminary models can be the same or different, and there is no limitation on this.
[0122] Further, the first vocabulary in the above step S220 can be replaced by the first preliminary model (see Figure 6 ), the query sampling of the first vocabulary according to the first word sequence in step S31 can be replaced by: using the first preliminary model to perform several prediction processes on the first word sequence.
[0123] It should be noted that for the contents not otherwise described in the above-mentioned expansion method of the improvement scheme, those skilled in the art are capable of fully implementing it by referring to the relevant contents in the aforementioned embodiments.
[0124] Corresponding to the above-mentioned request processing method, the embodiment of this specification also provides a request processing device. Figure 7 This is one of the structural schematic diagrams of the request processing device of the large language model disclosed in the embodiments of this specification. The output text of the large language model is pre-set as multiple generation segments, wherein each generation segment is configured with a corresponding starting word and an ending word, and is configured with a corresponding query sampling corpus. Figure 7 The processing device 700 shown in FIG. 7 includes the following functional modules:
[0125] The generated segment determination module 710 is configured to determine the first generated segment corresponding to the first request according to the first word-gram sequence included in the request when it is determined that the request is in the decoding stage. The speculative sampling module 720 is configured to perform speculative sampling processing on the first word-gram sequence based on the first query sampling corpus corresponding to the first generated segment and the large language model, and obtain a target word-gram sequence following the first word-gram sequence as the processing result of the first request.
[0126] In some embodiments, the generated segment determination module 710 is specifically configured to: determine the most recently generated starting word in the first word sequence, and thereby determine the generated segment corresponding to the starting word as the first generated segment.
[0127] In some embodiments, the large language model undergoes the following fine-tuning processing: obtaining a fine-tuning sample, which includes an original first label word-meta sequence; inserting the start word and the end word corresponding to each generated segment according to the division position of the label word-meta sequence corresponding to each generated segment, to obtain a modified second label word-meta sequence; using the fine-tuning sample in which the first label word-meta sequence is replaced with the second label word-meta sequence, the pre-trained large language model is fine-tuned.
[0128] In some embodiments, each generation segment is configured with a corresponding speculative sampling algorithm; the speculative sampling module 720 is specifically configured as: using the first speculative sampling algorithm corresponding to the first generation segment to perform the speculative sampling, the first speculative sampling algorithm is one of the following: single branch prediction, depth-first search algorithm under multi-branch prediction, breadth-first search algorithm under multi-branch prediction.
[0129] In some embodiments, the speculative sampling module 720 is specifically configured as follows: query sampling is performed on the first query sampling corpus according to the first word-gram sequence to obtain a plurality of predicted word-gram sequences following the first word-gram sequence; the large language model is called to process the plurality of inputs constructed based on the plurality of predicted word-gram sequences to obtain a plurality of inference results; and the plurality of predicted word-gram sequences are verified using the plurality of inference results to obtain the target word-gram sequence.
[0130] Further, in some specific embodiments, the first query sampling corpus includes multiple key-value pairs, and the key and value in each key-value pair are in the form of word sequences, and have frequency information that the value is continued after the key; wherein, the query sampling performed by the speculative sampling module 720 includes: determining a first key based on the first word sequence; querying a number of key-value pairs containing the first key in the first query sampling corpus, and determining the number of predicted word sequences based on the values in a predetermined number of key-value pairs whose frequencies are ranked in the front range.
[0131] In some specific embodiments, the construction of the multiple inputs performed by the speculative sampling module 720 includes: for any first prediction sequence among the several prediction word-unit sequences, splitting it into multiple sub-sequences, each of the multiple sub-sequences containing a different number of consecutive words starting from the first word-unit in the first prediction sequence; splicing the multiple sub-sequences after the first word-unit sequence respectively to obtain multiple spliced sequences; and classifying the first word-unit sequence and the multiple spliced sequences as the multiple inputs.
[0132] In some embodiments, the processing device 700 further includes: a query sampling corpus updating unit 730, configured to update the first query sampling corpus based on the target word-gram sequence when the last word-gram of the target word-gram sequence is the terminal word-gram corresponding to the first generated segment.
[0133] In some embodiments, the processing device 700 also includes: a new request construction unit, configured to construct a new request to be processed based on the target word sequence when the last word of the target word sequence is not a text end word, or when the length of the generated sequence up to the target word sequence does not reach a length threshold.
[0134] In some embodiments, the processing device 700 further includes: a batch request scheduling unit configured to schedule a batch of requests participating in this batch calculation based on a current request queue, wherein any request is used as the first request.
[0135] Figure 8 This is the second structural diagram of the request processing device of the large language model disclosed in the embodiment of this specification, wherein the output text of the large language model is pre-set into multiple generation segments, wherein each generation segment is configured with a corresponding starting word and an ending word, and is configured with a corresponding preliminary model. Figure 8 The processing device 800 shown in FIG. 8 includes the following functional modules:
[0136] The generated segment determination module 810 is configured to determine the first generated segment corresponding to the first request according to the first word-gram sequence included in the request when it is determined that the request is in the decoding stage. The speculative sampling module 820 is configured to perform speculative sampling processing on the first word-gram sequence based on the first preliminary model corresponding to the first generated segment and the large language model, and obtain a target word-gram sequence following the first word-gram sequence as the processing result of the first request.
[0137] In some embodiments, the generated segment determination module 810 is specifically configured to: determine the most recently generated starting word in the first word sequence, and thereby determine the generated segment corresponding to the starting word as the first generated segment.
[0138] In some embodiments, the large language model undergoes the following fine-tuning processing: obtaining a fine-tuning sample, which includes an original first label word-meta sequence; inserting the start word and the end word corresponding to each generated segment according to the division position of the label word-meta sequence corresponding to each generated segment, to obtain a modified second label word-meta sequence; using the fine-tuning sample in which the first label word-meta sequence is replaced with the second label word-meta sequence, the pre-trained large language model is fine-tuned.
[0139] In some embodiments, the speculative sampling module 820 is specifically configured as follows: inputting the first word-gram sequence into the first preliminary model several times to obtain several predicted word-gram sequences following the first word-gram sequence; calling the large language model to process the multiple inputs constructed based on the several predicted word-gram sequences to obtain multiple inference results; and using the multiple inference results to verify the several predicted word-gram sequences to obtain the target word-gram sequence.
[0140] Further, in some specific embodiments, the construction of the multiple inputs performed by the speculative sampling module 720 includes: for any first prediction sequence among the several prediction word-unit sequences, splitting it into multiple sub-sequences, each of the multiple sub-sequences containing a different number of consecutive words starting from the first word-unit in the first prediction sequence; splicing the multiple sub-sequences after the first word-unit sequence respectively to obtain multiple spliced sequences; and classifying the first word-unit sequence and the multiple spliced sequences as the multiple inputs.
[0141] In some embodiments, the processing device 800 further includes: a preliminary model retraining unit 830, configured to retrain the first preliminary model based on the target word-gram sequence when the last word-gram of the target word-gram sequence is the terminal word-gram corresponding to the first generated segment.
[0142] In some embodiments, the processing device 800 also includes: a new request construction unit, configured to construct a new request to be processed based on the target word sequence when the last word of the target word sequence is not a text end word, or when the length of the generated sequence up to the target word sequence does not reach a length threshold.
[0143] In some embodiments, the processing device 800 further includes: a batch request scheduling unit configured to schedule a batch of requests participating in this batch calculation based on a current request queue, wherein any request is used as the first request.
[0144] It should be noted that for the introduction of the above functional modules, reference can also be made to the relevant introduction of the process method in the above embodiments.
[0145] According to another embodiment, there is also provided a computer-readable storage medium having a computer program stored thereon, which, when executed in a computer, causes the computer to execute Figure 2 or Figure 3 or Figure 6 The method described.
[0146] According to another embodiment, a computing device is provided, including a memory and a processor, wherein the memory stores executable code, and when the processor executes the executable code, Figure 2 or Figure 3 or Figure 6 The method described.
[0147] Those skilled in the art should be aware that in one or more of the above examples, the functions described in the present invention can be implemented by hardware, software, firmware or any combination thereof. When implemented by software, these functions can be stored in a computer-readable medium or transmitted as one or more instructions or codes on a computer-readable medium.
[0148] The specific implementation methods described above further illustrate the objectives, technical solutions and beneficial effects of the present invention in detail. It should be understood that the above description is only a specific implementation method of the present invention and is not intended to limit the scope of protection of the present invention. Any modifications, equivalent substitutions, improvements, etc. made on the basis of the technical solution of the present invention should be included in the scope of protection of the present invention.
Claims
1. A method for processing requests of a large language model, the implementation of the method is based on a plurality of generated segments pre-set for the output text of the large language model, wherein each generated segment is configured with a corresponding start word and end word, and is configured with a corresponding query sampling corpus; the method comprises: For the first request, when it is determined that the request is in a decoding stage, determining a corresponding first generation segment according to a first word unit sequence included in the request; Based on the first query sampling corpus corresponding to the first generation segment and the large language model, speculative sampling processing is performed on the first word-gram sequence to obtain a target word-gram sequence subsequent to the first word-gram sequence as a processing result of the first request.
2. The method according to claim 1, wherein: Determining a corresponding first generation segment according to the first word-gram sequence included therein includes: The most recently generated starting word-gram in the first word-gram sequence is determined, so that the generation segment corresponding to the starting word-gram is determined as the first generation segment.
3. The method according to claim 1, wherein: The large language model is fine-tuned as follows: Get a fine-tuning sample, which includes the original first-label word sequence; According to the division positions of the first label word-unit sequence corresponding to the respective generated segments, inserting the start word-unit and the end word-unit corresponding to the respective generated segments to obtain a modified second label word-unit sequence; The pre-trained large language model is fine-tuned using a fine-tuning sample in which the first label word sequence is replaced with the second label word sequence.
4. The method according to claim 1, wherein: Each of the generated sections is configured with a corresponding speculative sampling algorithm; wherein, based on the first vocabulary corresponding to the first generated section, speculative sampling is performed on the first word sequence, including: The speculative sampling is performed using a first speculative sampling algorithm corresponding to the first generation segment, where the first speculative sampling algorithm is one of the following: a single branch prediction, a depth-first search algorithm under multi-branch prediction, or a breadth-first search algorithm under multi-branch prediction.
5. The method according to claim 1, wherein: The speculative sampling process includes: performing query sampling on the first query sample corpus according to the first word-gram sequence to obtain a plurality of predicted word-gram sequences following the first word-gram sequence; Calling the large language model to process the multiple inputs constructed based on the plurality of predicted word-unit sequences to obtain multiple inference results; The plurality of predicted word-gram sequences are verified using the plurality of inference results to obtain the target word-gram sequence.
6. The method according to claim 5, wherein: The first query sample corpus includes a plurality of key-value pairs, wherein the key and the value in each key-value pair are in the form of a word sequence, and have frequency information of the value being connected after the key; The query sampling includes: determining a first key based on the first word-gram sequence; The first query sample corpus is queried for a number of key-value pairs containing the first key, and the number of predicted word-unit sequences are determined based on values in a predetermined number of key-value pairs whose frequencies are ranked in a front range.
7. The method according to claim 5, wherein: The construction of the plurality of inputs comprises: For any first prediction sequence among the plurality of prediction word-element sequences, split it into a plurality of sub-sequences, each of the plurality of sub-sequences comprising a different number of consecutive word-element starting from the first word-element in the first prediction sequence; concatenating the plurality of subsequences to the first word sequence respectively to obtain a plurality of concatenated sequences; The first word-gram sequence and a plurality of concatenated sequences are classified as the plurality of inputs.
8. The method according to claim 1, wherein after obtaining the target word-gram sequence, the method further comprises: When the last word of the target word-gram sequence is the terminal word-gram corresponding to the first generated segment, the first query sample corpus is updated based on the target word-gram sequence.
9. The method according to claim 1, wherein: After obtaining the target word sequence, the method further includes: When the last word of the target word-gram sequence is not a text-ending word-gram, or when the length of the generated sequence up to the target word-gram sequence does not reach a length threshold, a new request to be processed is constructed based on the target word-gram sequence.
10. The method according to claim 1, wherein: The method further comprises: A batch of requests participating in this batch calculation are scheduled based on the current request queue, and any one of the requests is used as the first request.
11. A method for processing requests of a large language model, the implementation of the method being based on a plurality of pre-set generation segments for an output text of the large language model, wherein each generation segment is configured with a corresponding start word and end word, and is configured with a corresponding preliminary model; the method comprising: For the first request, when it is determined that the request is in a decoding stage, determining a corresponding first generation segment according to a first word unit sequence included in the request; Based on the first preliminary model and the large language model corresponding to the first generated segment, speculative sampling processing is performed on the first word-gram sequence to obtain a target word-gram sequence following the first word-gram sequence as a processing result of the first request.
12. The method according to claim 11, wherein: The speculative sampling process includes: Inputting the first word-gram sequence into the first preliminary model to obtain a predicted word-gram sequence following the first word-gram sequence; Calling the large language model to process the multiple inputs constructed based on the predicted word-unit sequence to obtain multiple inference results; The predicted word-gram sequence is verified using the multiple inference results to obtain the target word-gram sequence.
13. A request processing device for a large language model, wherein the output text of the large language model is pre-set into a plurality of generation segments, wherein each generation segment is configured with a corresponding start word and end word, and is configured with a corresponding query sample corpus; the device comprises: A generated segment determination module is configured to determine, for the first request, a corresponding first generated segment according to a first word sequence included in the first request when it is determined that the request is in a decoding stage; The speculative sampling module is configured to perform speculative sampling processing on the first word sequence based on the first query sampling corpus corresponding to the first generation segment and the large language model, and obtain a target word sequence following the first word sequence as a processing result of the first request.
14. A request processing device for a large language model, wherein the output text of the large language model is pre-set into a plurality of generation segments, wherein each generation segment is configured with a corresponding start word and end word, and is configured with a corresponding preliminary model; the device comprises: A generated segment determination module is configured to determine, for the first request, a corresponding first generated segment according to a first word sequence included in the first request when it is determined that the request is in a decoding stage; The speculative sampling module is configured to perform speculative sampling processing on the first word sequence based on the first preliminary model corresponding to the first generated segment and the large language model, and obtain a target word sequence following the first word sequence as a processing result of the first request.
15. A computer-readable storage medium having a computer program stored thereon, wherein: When the computer program is executed in a computer, the computer is caused to execute the method according to any one of claims 1 to 12.
16. A computing device comprising a memory and a processor, wherein: The memory stores executable codes, and when the processor executes the executable codes, the method according to any one of claims 1 to 12 is implemented.
Citation Information
Cited By
Training method and training system of end-to-end speech translation model
CN121072785A