MTP speculation sampling method and system for large model

By prefixing meaningless characters in the data processing request in the MTP speculative sampling system, and using the hidden layer state data of the target model and the draft model for multiple rounds of sampling, the problems of information loss and error accumulation are solved, and the accuracy and efficiency of the system are improved.

CN120278243APending Publication Date: 2025-07-08BEIJING SILICONFLOW TECHNOLOGY CO LTD
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510287026.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-12
Publication Date
2025-07-08

AI Technical Summary

Technical Problem

The traditional MTP speculative sampling system has problems such as low accuracy during operation, loss of input information and accumulation of errors, resulting in a reduced acceptance rate of speculative tokens.

Method used

By prefixing meaningless characters in the data processing request, adjusting the initial input data prefilled by the draft model, and using the hidden layer state data of the target model and the draft model for multiple rounds of sampling to eliminate information loss and error accumulation.

Benefits of technology

It improves the accuracy of the MTP speculative sampling system, ensures that the input information is not lost, and reduces error accumulation, and improves the inference accuracy of the large model.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120278243A_ABST
    Figure CN120278243A_ABST
Patent Text Reader

Abstract

The invention relates to an MTP speculation sampling method and system for a large model. The method comprises the following steps: receiving a data processing request of a user, prefixing a meaningless character at the head of the data processing request, and thus obtaining a first round of target model input token set of the data processing request after the prefix is added; a target model initial speculation sampling processing step: carrying out a first round of target model pre-filling processing and a first round of target model sampling processing on the target model to obtain a 0th new verification token; n preset times of initial speculation sampling are carried out, n verification draft tokens are obtained, and each time of speculation sampling comprises the nth time of draft model pre-filling processing and the nth time of draft model sampling processing after the nth time of draft model pre-filling processing; and for a second round of target model input token set, executing the second round of target model pre-filling processing of the target model and sampling in sequence, and based on the obtained n verification tokens, sampling to obtain an (n + 1) th new verification token.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the field of computer information processing, and in particular, to an MTP speculative sampling method and system for large models. Background Art

[0002] With the rapid development of natural language processing (NLP) technology, large language models (such as GPT, Qwen, etc.) have demonstrated excellent performance in tasks such as text generation, question answering, and translation. However, the high performance of these models often comes with huge computational overhead and resource requirements. Against this background, in order to balance performance and efficiency, speculative sampling technology has emerged.

[0003] Speculative sampling is a strategy to optimize generation efficiency. Before sampling using the target model, a much smaller draft model in terms of the number of parameters is used to predict one or more consecutive draft token sequences in advance. Then the target model takes these draft tokens as input, and in the sampling stage, multiple tokens are accepted according to the coincidence degree. In this way, multiple tokens can be generated at once. A draft model with a smaller number of parameters (Draft Model) predicts the speculative tokens or draft tokens that the target model (Target Model) may generate, and inputs these speculative tokens into the target model for verification. If verified, they are accepted and called verified tokens.

[0004] Conventional speculative sampling is divided into 2 steps: first speculation and then sampling, while the difference between MTP (Multi-token Prediction) speculative sampling and conventional speculation is that it first samples and then speculates. However, the traditional MTP speculative sampling system has defects such as low accuracy and loss of information in the input request during operation. There is also a situation where MTP speculation is directly performed using two adjacent verified tokens formed by the target model, which introduces errors and causes the acceptance rate of the speculated or draft tokens to decrease geometrically, resulting in a reduction in the accuracy of large model inference. Therefore, there is a need to obtain an MTP speculative sampling system and method that does not lose the original input information and improves inference accuracy.

[0005] The above information disclosed in the background art section is only used to enhance the understanding of the background of the present application, and thus it may include information that does not constitute the prior art known to those of ordinary skill in the art. Summary of the Invention

[0006] In view of this, the present disclosure provides an MTP speculative sampling method and system for large models. By adjusting the structure of the input request and the initial input data pre-filled by the draft model, it is possible to eliminate the situation where information is lost and the error accumulates excessively, resulting in an excessive cumulative reduction in the acceptance rate.

[0007] Other features and advantages of the present disclosure will become apparent from the following detailed description or will be partly learned through the practice of the present disclosure.

[0008] According to one aspect of the present disclosure, a MTP speculative sampling method for large models is proposed, including: a data processing request preprocessing step of receiving a user's data processing request and prefixing a meaningless character to the header of the data processing request to obtain a first-round target model input token set of the data processing request after adding the prefix; a target model initial speculative sampling processing step of performing a first-round target model pre-filling processing and a first-round target model sampling processing based on the tokens in the first-round target model input token set to obtain the 0th new verification token; a draft model initial speculative sampling processing step of using the hidden layer state data of each token in the first-round target model input token set and the first-round draft model input token set of the draft model formed by discarding the first token in the first-round target model input token set and adding the 0th new verification token as the initial input, starting from the 0th time, performing a predetermined n times of initial speculative sampling, and obtaining n verification draft tokens, where each speculative sampling includes the nth draft model pre-filling processing and the subsequent nth draft model sampling processing; a target model verification sampling processing step of, for the second-round target model input token set formed by merging the 0th new verification token obtained from the first-round target model sampling processing and the n new speculative tokens obtained from the draft model initial speculative sampling processing, performing a second-round target model pre-filling processing of the target model to obtain the hidden layer state data of each token in the second-round target model input token set and the (n + 1) target model prediction heads of the second-round target model input token set, and sequentially sampling based on the raw scores of the (n + 1) tokens in the second-round target model input token set formed by the (n + 1) target model prediction heads, verifying the last n tokens among the (n + 1) tokens, and when obtaining n verification tokens, sampling based on the obtained n verification tokens to obtain the (n + 1)th new verification token, so as to use the hidden layer state data of each token in the second-round target model input token set as the hidden layer state data of each token in the first-round target model input token set and use the second-round draft model input token set of the draft model formed by the finally obtained (n + 1) new verification tokens as the first-round draft model input token set of the draft model as the input, and repeat to start executing the 0th draft model pre-filling processing step and the subsequent steps.

[0009] The MTP speculative sampling method for large models according to the present disclosure, wherein in the first-round target model pre-filling process step, based on the tokens in the first-round target model input token set, the first-round target model pre-filling process of the target model is performed to obtain the hidden layer state data of each token in the first-round target model input token set and the target model prediction head of the first-round target model input token set; and in the first-round target model sampling process step, the first-round target model sampling process is performed on the raw scores (logits) of the last token in the first-round target model input token set formed based on the target model prediction head of the first-round target model input token set to obtain the 0th new verification token.

[0010] The MTP speculative sampling method for large models according to the present disclosure, wherein the initial speculative sampling is performed n times starting from the 0th time, including: the 0th draft model pre-filling process step, using the hidden layer state data of each token in the first-round target model input token set and the first-round draft model input token set of the draft model formed by discarding the first token in the first-round target model input token set and adding the 0th new verification token as the initial input, performing the 0th draft model pre-filling process in the first-round draft model pre-filling process including a predetermined n times of draft model pre-filling processes to obtain the hidden layer state data of the last token in the first-round draft model input token set as the input value of the 1st draft model pre-filling process and obtaining the 0th draft model prediction head of the first-round draft model input token set; and the 0th draft model sampling process step, performing the 0th draft model sampling process in the first-round draft model sampling process including a predetermined n times of draft model sampling processes on the raw scores (logits) of the last token in the first-round draft model input token set formed based on the 0th draft model prediction head to obtain the 1st new speculative token as the input value of the 1st draft model pre-filling process; and repeating the speculation step, using the new speculative token output from the (k - 1)th draft model sampling process and the hidden layer state data of the token obtained from the (k - 1)th draft model pre-filling process as the input, performing the kth draft model pre-filling process in the remaining (n - 1) draft model pre-filling processes to obtain the hidden layer state data of the kth new speculative token and the kth draft model prediction head of the kth new speculative token, and performing the kth draft model sampling process in the remaining (n - 1) draft model sampling processes on the raw scores (logits) of the kth new speculative token formed based on the kth draft model prediction head to obtain the (k + 1)th new speculative token until k is equal to (n - 1).

[0011] The MTP speculative sampling method for large models according to the present disclosure, wherein the data processing request preprocessing step further includes: when the large model is in a pipelined parallel state, obtaining the maximum number of tokens m per data processing system for processing data requests and the number of tokens p of the obtained data processing request; and comparing the magnitudes of the numbers m and the number (p - q), and splitting the tokens of the data processing request into one or more sub-requests in ascending order according to the smaller of the two numbers, and only prefixing a meaningless character to the head of the first sub-request.

[0012] The MTP speculative sampling method for large models according to the present disclosure, wherein q is at least 2 and less than p.

[0013] The MTP speculative sampling method for large models according to the present disclosure, wherein n is 2 - 4.

[0014] The MTP speculative sampling method for large models according to the present disclosure, wherein in the 0th draft model pre-filling processing step, the hidden layer state data of each token in the first-round target model input token set input is a vector matrix with the number of tokens as the dimension, and the first-round draft model input token set of the input draft model generates a vector matrix with the same number of dimensions in the word embedding layer. By splicing the two vector matrices according to the dimension, an input vector matrix for pre-filling processing is formed.

[0015] According to the present disclosure, there is also provided an MTP speculative sampling system for a large model, including: a request preprocessing component, an MTP target model component, and an MTP draft model component. The request preprocessing component receives a user's data processing request and prefixes a meaningless character to the header of the data processing request to obtain a first-round target model input token set of the data processing request after adding the prefix. The MTP target model component, after obtaining the tokens in the first-round target model input token set, performs a first-round target model pre-filling process of the target model to obtain a 0th new verification token. The MTP draft model component uses the hidden layer state data of each token in the first-round target model input token set and the first-round draft model input token set of the draft model formed by discarding the first token in the first-round target model input token set and adding the 0th new verification token as the initial input, starts initial speculative sampling a predetermined n times from the 0th time, and obtains n verification draft tokens. Each speculative sampling includes an nth draft model pre-filling process and an nth draft model sampling process thereafter. The MTP target model component performs a second-round target model pre-filling process of the target model for the second-round target model input token set formed by merging the 0th new verification token obtained from the first-round target model sampling process and the n new speculative tokens obtained from the n draft model sampling processes, obtains the hidden layer state data of each token in the second-round target model input token set and the (n + 1) target model prediction heads of the second-round target model input token set, and sequentially samples based on the raw scores (logits) of the (n + 1) tokens in the second-round target model input token set formed by the (n + 1) target model prediction heads, verifies the last n tokens among the (n + 1) tokens, and when n verification tokens are obtained, samples to obtain a (n + 1)th new verification token based on the obtained n verification tokens, so as to use the hidden layer state data of each token in the second-round target model input token set as the hidden layer state data of each token in the first-round target model input token set and use the second-round draft model input token set of the draft model formed by the finally obtained (n + 1) new verification tokens as the first-round draft model input token set of the draft model as the input of the MTP draft model component, and repeat the process starting from the 0th draft model pre-filling process.

[0016] The MTP speculative sampling system for large models according to the present disclosure, wherein in the first-round target model pre-filling process, the MTP target model component performs the first-round target model pre-filling process based on the tokens in the first-round target model input token set, obtains the hidden layer state data of each token in the first-round target model input token set and the target model prediction head of the first-round target model input token set; and in the first-round target model sampling process, the MTP target model component performs the first-round target model sampling process on the raw scores (logits) of the last token in the first-round target model input token set formed based on the target model prediction head of the first-round target model input token set, and obtains the 0th new verification token.

[0017] The MTP draft model component of the MTP speculative sampling system for large models according to the present disclosure starts the predetermined n initial speculative samplings from the 0th time, including: performing the 0th draft model pre-filling process, using the hidden layer state data of each token in the first-round target model input token set and the first-round draft model input token set of the draft model formed by discarding the first token in the first-round target model input token set and adding the 0th new verification token as the initial input, performing the 0th draft model pre-filling process in the first-round draft model pre-filling process including the predetermined n draft model pre-filling processes, obtaining the hidden layer state data of the last token in the first-round draft model input token set as the input value of the 1st draft model pre-filling process and obtaining the 0th draft model prediction head of the first-round draft model input token set; performing the 0th draft model sampling process, performing the 0th draft model sampling process in the first-round draft model sampling process including the predetermined n draft model sampling processes on the raw scores (logits) of the last token in the first-round draft model input token set formed based on the 0th draft model prediction head, and obtaining the 1st new speculative token as the input value of the 1st draft model pre-filling process; and performing repeated speculation, using the new speculative token output by the (k-1)th draft model sampling process and the hidden layer state data of the token obtained by the (k-1)th draft model pre-filling process as the input, performing the kth draft model pre-filling process in the remaining (n-1) draft model pre-filling processes, obtaining the hidden layer state data of the (k+1)th new speculative token and the kth draft model prediction head of the (k+1)th new speculative token, and performing the kth draft model sampling process in the remaining (n-1) draft model sampling processes on the raw scores (logits) of the (k+1)th new speculative token formed based on the kth draft model prediction head, and obtaining the (k+1)th new speculative token until k is equal to (n-1).

[0018] The MTP speculative sampling system for large models according to the present disclosure, wherein the request preprocessing component, when the large model is in a pipelined parallel state, obtains the maximum number of tokens m per data processing system for processing data requests and the number of tokens p of the obtained data processing requests, compares the magnitudes of the number m and the number (p - 2), and divides the tokens of the data processing requests into one or more sub-requests in ascending order according to the smaller of the two numbers, and only prefixes a meaningless character to the header of the first sub-request.

[0019] The MTP speculative sampling system for large models according to the present disclosure, wherein q is at least 2 and less than p.

[0020] The MTP speculative sampling system for large models according to the present disclosure, where n is 2 - 4.

[0021] The MTP speculative sampling system for large models according to the present disclosure, wherein when the MTP draft model component performs the 0th draft model pre-filling process, the hidden layer state data of each token in the first-round target model input token set received is a vector matrix with the number of tokens as the dimension, and enables the first-round draft model input token set of the input draft model to generate a vector matrix with the same number of dimensions in the word embedding layer. By splicing the two vector matrices according to the dimension, an input vector matrix for the pre-filling process to be processed is formed.

[0022] According to the MTP speculative sampling method and system for large models of the present disclosure, by adding a meaningless token before the initial input request, the defect that in traditional MTP speculative sampling, the input for the MTP draft model discards the first token, resulting in missing input information, is eliminated. And by using the hidden layer state data of each token in the previous-round target model input token set instead of directly using the kv cache of the tokens verified by the previous-round target model in each round of draft model pre-filling process, the error accumulation transfer speed in traditional MTP speculative sampling can be eliminated. Moreover, by splitting the input request into at least two parts and processing them in parallel mode, the possible conflicting situations between pre-filling and decoding or between speculation and sampling can be eliminated.

[0023] It should be understood that the above general description and the following detailed description are only exemplary and do not limit the present disclosure. BRIEF DESCRIPTION OF THE DRAWINGS

[0024] By describing in detail its exemplary embodiments with reference to the drawings, the above and other objectives, features, and advantages of the present disclosure will become more apparent. The following described drawings are only some embodiments of the present disclosure. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.

[0025] Figure 1 It is a flowchart of an embodiment of the MTP speculative sampling method for large models shown according to an exemplary embodiment.

[0026] Figure 2 It is a schematic diagram of a pre-filling process example of an embodiment of the MTP speculative sampling method for large models shown according to an exemplary embodiment.

[0027] Figure 3 It is a schematic diagram of a decoding process example of an embodiment of the MTP speculative sampling method for large models shown according to an exemplary embodiment.

[0028] Figure 4 It is a schematic diagram of a vector concatenation process example in an embodiment of the MTP speculative sampling method for large models shown according to an exemplary embodiment.

[0029] Figure 5 It is a schematic diagram of a case processing of an input data request in the MTP speculative sampling method in the traditional pipelined parallel mode for large models.

[0030] Figure 6 It is a schematic diagram of the preprocessing of an input data request in the MTP speculative sampling method for the pipelined parallel mode shown according to an exemplary embodiment of the present disclosure.

[0031] Figure 7 It is a block diagram of an embodiment of the MTP speculative sampling for large models shown according to an exemplary embodiment. Detailed implementation manners

[0033] Example embodiments will now be described more fully with reference to the accompanying drawings. However, the example embodiments can be implemented in various forms and should not be construed as limited to the embodiments set forth herein; rather, these embodiments are provided so that this disclosure will be thorough and complete, and will fully convey the concept of the example embodiments to those skilled in the art. Like reference numerals in the figures denote like or similar parts, and thus their repetitive description will be omitted.

[0034] In addition, the described features, structures, or characteristics may be combined in any suitable manner in one or more embodiments. In the following description, numerous specific details are provided to give a thorough understanding of the embodiments of the present disclosure. However, those skilled in the art will realize that the technical solutions of the present disclosure can be practiced without one or more of the specific details, or other methods, components, devices, steps, etc. may be used. In other cases, well-known methods, devices, implementations, or operations are not shown or described in detail to avoid obscuring aspects of the present disclosure.

[0035] The block diagrams shown in the accompanying drawings are merely functional entities and do not necessarily correspond to physically independent entities. That is, these functional entities can be implemented in software form, or implemented in one or more hardware modules or integrated circuits, or implemented in different networks and / or processor devices and / or microcontroller devices.

[0036] The flowcharts shown in the accompanying drawings are only illustrative descriptions and do not necessarily include all the content and operations / steps, nor do they have to be executed in the order described. For example, some operations / steps can be decomposed, while some operations / steps can be combined or partially combined, so the actual execution order may change according to the actual situation.

[0037] It should be understood that although terms such as first, second, and third may be used herein to describe various components, these components should not be limited by these terms. These terms are used to distinguish one component from another. Therefore, the first computing device described below can be referred to as the second computing device without departing from the teachings of the present disclosure concept. As used herein, the term "and / or" includes any one of the associated listed items and all combinations of one or more of them.

[0038] Those skilled in the art can understand that the accompanying drawings are only schematic diagrams of exemplary embodiments, and the modules or processes in the accompanying drawings are not necessarily essential for implementing the present disclosure, so they cannot be used to limit the protection scope of the present disclosure.

[0039] Figure 1 is a flowchart of an embodiment of the MTP speculative sampling method for large models shown according to an exemplary embodiment. As Figure 1 shown, in the method 100 shown, at step S110, when the large model processing system receives a user's data processing request, the request preprocessing component 510 (see below Figure 5Description) Prefix a meaningless character to the header of the data processing request to obtain the first-round target model input token set of the data processing request after the prefix is added. For example, if the user input text is: "You ate my pear", and then it is converted into a model input token sequence through a tokenizer. If the token sequence is [1, 2, 3, 4, 5, 6], for traditional MTP speculative sampling, token1 will be discarded, corresponding to the word "You" in the request text. In this way, the word "You" in the request text will be invisible to the MTP speculative sampling process, which leads to the problem of information loss. To this end, in the present disclosure, for all user input texts, a meaningless character is prefixed to the very front header. For example, a special string "<|begin▁of▁sentence|>" is used (other methods can also be used, which will not be elaborated here one by one, and this kind of character belongs to the conventional characters in this field). The string corresponds to token[0], so that after all inputs pass through the tokenizer, a 0 or other meaningless token will be added at the beginning. In this way, for MTP speculative sampling, although the first token is formally discarded, the input information will not be lost in essence.

[0040] Subsequently, at step S120, perform the initial speculative sampling process of the target model. Based on the tokens in the first-round target model input token set, perform the first-round target model pre-filling process and the first-round target model sampling process of the target model to obtain the 0th new verification token. For example, perform the first-round target model pre-filling process at step S121. Specifically, based on the tokens in the first-round target model input token set, perform the first-round target model pre-filling process of the target model to obtain the hidden layer state data of each token in the first-round target model input token set and the target model prediction head of the first-round target model input token set. This step is actually in the pre-filling process of MTP speculative sampling. Figure 2 is an example schematic diagram of the pre-filling process of an embodiment of the MTP speculative sampling method for a large model shown according to an exemplary embodiment. As Figure 2As shown, assume that the initial input token sequence of the user is [0, 1, 2, 3, 4], a total of 5 tokens. Then these 5 tokens will first be input into the target model, and then the raw scores (logits) of the token [4] output by the prediction head of the target model will be obtained. The prediction head is the "decision terminal" of the model, which transforms abstract features into the concrete results required by the task and directly affects the model output quality. The raw scores (logits) output by the language model represent the "confidence" of the model in each candidate word (token), but have not been converted into probability values. The raw scores can be converted into a probability distribution through the softmax function. The raw scores (logits) retain the original judgment of the model, which is convenient for subsequent adjustments (such as comparing the logits differences between the large and small models in speculative sampling). Directly using probabilities may amplify errors due to normalization, while the raw scores (logits) can avoid this problem. At the same time, the hidden layer state data of each token in the token set input to the target model in the first round, that is, the hidden layer state of the target model corresponding to token [0, 1, 2, 3, 4], needs to be saved as the input for the initial prefill stage of the subsequent MTP draft model, that is, as one of the inputs for the 0th draft model prefill process.

[0041] Subsequently, as Figure 1 shown, at step S122, the first-round target model sampling processing step is executed. Specifically, the first-round target model sampling processing is performed on the raw scores (logits) of the last token in the first-round target model input token set formed by the prediction head of the target model based on the first-round target model input token set, and the 0th new verification token is obtained. As Figure 2 shown, it is to perform sampling processing based on the raw scores (logits) of token [4] to obtain the 0th new verification token, that is, the new token [5]. In this way, the initial prefill stage of the entire prefill process at the target model is completed.

[0042] Then, as Figure 1As shown, at step S130, initial speculative sampling processing of the draft model is performed. Specifically, using the hidden layer state data of each token in the first-round target model input token set and the first-round draft model input token set of the draft model formed by discarding the first token in the first-round target model input token set and adding the 0th new verification token as the initial input, starting from the 0th time, perform a predetermined n times of initial speculative sampling, and obtain n verification draft tokens. Each speculative sampling includes the nth draft model pre-filling process and the subsequent nth draft model sampling process. For example, first at step S131, perform the 0th draft model pre-filling process. Specifically, using the hidden layer state data of each token in the first-round target model input token set and the first-round draft model input token set of the draft model formed by discarding the first token in the first-round target model input token set and adding the 0th new verification token as the initial input, perform the 0th draft model pre-filling process in the first-round draft model pre-filling process including a predetermined n times of draft model pre-filling processes, obtain the hidden layer state data of the last token in the first-round draft model input token set as the input value for the 1st draft model pre-filling process, and obtain the 0th draft model prediction head of the first-round draft model input token set. As Figure 2 shown, the number of runs of the MTP draft model is determined by a parameter n. In the example, it is assumed that n is 3, that is, the MTP draft model will run in a loop 3 times (each time including one pre-filling and one sampling). The 0th draft model pre-filling process, that is Figure 2 the first MTP draft model pre-filling in the 3 times in, its input is the hidden state of the target model corresponding to tokens [0, 1, 2, 3, 4] (that is, the hidden layer state data of each token in the first-round target model input token set) and the token sequence [1, 2, 3, 4, 5] (the user's initial input token sequence minus the first token plus the new token [5] generated by the target model), that is, the first-round draft model input token set of the draft model formed by discarding the first token in the first-round target model input token set and adding the 0th new verification token.

[0043] Subsequently, at step S132, the 0th draft model sampling process is performed. Specifically, the 0th draft model sampling process in the first-round draft model sampling process including a predetermined n draft model sampling processes is performed on the raw score (logits) of the last token in the first-round draft model input token set formed based on the 0th draft model prediction head, and the 1st new speculative token is obtained as the input value for the 1st draft model pre-filling process. As Figure 2 shown, the raw score (logits) corresponding to token[5] (i.e., the last token in the first-round draft model input token set formed based on the 0th draft model prediction head) is sampled to obtain draft token[6] (the 1st new speculative token). Draft token[6] will be used as the input for the next draft model pre-filling process.

[0044] Return to reference Figure 1 . Subsequently, at step S133, (n - 1) times of alternating draft model pre-filling processes and draft model sampling processes are executed. This is also referred to as the repeated speculation step here. Specifically, using the new speculative token output from the (k - 1)th draft model sampling process and the hidden layer state data of the token obtained from the (k - 1)th draft model pre-filling process as the input, the kth draft model pre-filling process in the remaining (n - 1) draft model pre-filling processes is performed to obtain the hidden layer state data of the kth new speculative token and the kth draft model prediction head of the kth new speculative token, and the kth draft model sampling process in the remaining (n - 1) draft model sampling processes is performed on the raw score (logits) of the kth new speculative token formed based on the kth draft model prediction head to obtain the (k + 1)th new speculative token until k is equal to (n - 1). When k = 1, it is actually the 2nd time in the predetermined n draft model pre-filling processes. When k = 2, it is actually the 3rd time in the predetermined n draft model pre-filling processes, and so on. For example, as Figure 2As shown, the hidden layer states of draft token [6] and token [5] corresponding to the MTP draft model are used as the input of the second MTP draft model (i.e., the second time in the pre-filling process of the draft model for a predetermined n times when k = 1). Then, draft token [7] is obtained by sampling the logits corresponding to draft token [6]. Next, the hidden layer states of draft token [7] and draft token [6] corresponding to the MTP draft model are used as the input of the third MTP draft model. Finally, draft token [8] is obtained by sampling the logits corresponding to draft token [7]. It is better that n is 2 - 4.

[0045] Return to Figure 1 , finally, at step S140, the target model verification sampling process is performed. Specifically, for the second-round target model input token set formed by combining the 0th new verification token obtained from the first-round target model sampling process and the n new speculative tokens obtained from the initial speculative sampling process of the draft model, the second-round target model pre-filling process of the target model is executed to obtain the hidden layer state data of each token in the second-round target model input token set and the (n + 1) target model prediction heads of the second-round target model input token set, and sampling is sequentially performed based on the raw scores (logits) of the (n + 1) tokens in the second-round target model input token set formed by the (n + 1) target model prediction heads to verify the last n tokens among the (n + 1) tokens. And when n verification tokens are obtained, based on the obtained n verification tokens, the (n + 1)th new verification token is sampled, so as to use the hidden layer state data of each token in the second-round target model input token set as the hidden layer state data of each token in the first-round target model input token set and use the second-round draft model input token set of the draft model formed by the finally obtained (n + 1) new verification tokens as the first-round draft model input token set of the draft model as the input, and repeat to start executing the 0th draft model pre-filling process step and the subsequent steps. Figure 3It is a schematic diagram of an example of the decoding process of the MTP speculative sampling method for large models shown according to an exemplary embodiment. It details a specific example of the target model verification sampling process. After a predetermined n initial speculative samplings in the initial speculative sampling process of the draft model, new tokens are obtained, including token[5] and three draft tokens[6, 7, 8]. Then, these 4 tokens[5, 6, 7, 8] are input into the target model as the second-round target model input token set. Subsequently, the output logits of these 4 tokens corresponding to the target model are input into the sampling module, and the sampling module samples these 4 logits respectively to obtain 4 new tokens[6`, 7`, 8`, 9]. Then, the draft token sequence[6, 7, 8] and the sampling sequence[6`, 7`, 8`] in the second-round target model input token set are compared in order from left to right. Assuming they all match here, that is, all 3 draft token sequences[6, 7, 8] can be accepted, then the token[9] sampled from the logits corresponding to token[8] will be used as the newly generated token, that is, the (n + 1)th new verification token.

[0046] Obviously, the target model verification sampling process here is different from the traditional MTP speculative sampling. In MTP speculative sampling, assume the initial input token sequence is token0 - 26. After directly running the prefill of the target model once to generate token27 initially, the MTP draft model sampling will be run next, inputting token1 - 27 to generate 28, and then running the MTP draft model sampling two more times (usually 1 - 3 times) to generate token29 and token30. Then, in the next run of the target model, it will run 27 - 30, generating the hidden states of these four tokens, and at the same time determining how many of the previously generated draft tokens 28, 29, 30 can be accepted. Assuming all are accepted, then an additional token 31 will be generated.

[0047] Subsequently, in traditional MTP speculative sampling, when running MTP in the next round, usually since the kv cache has already calculated tokens 28 and 29, at this time the MTP or draft model speculative sampling will directly start from 30. The reason is that the MTP draft model is accurate because the input combines the hidden states of the target model, and the reason MTP can perform continuous inference is that the input of the subsequent MTP uses the hidden states of the previous round of MTP to replace the hidden states of the target model, and this replacement has errors. As described above in the present disclosure, the MTP or draft model speculative sampling still starts from token28.

[0048] For conventional draft model speculation, the acceptance rate of the nth draft token is p^n, while the acceptance rate of the nth draft token in MTP is p^(2*n - 1), which decays at a squared rate. The fundamental reason is that using the hidden state of MTP introduces errors. For example, first define the acceptance probability of the nth draft token as P(n). The acceptance probability of the first draft token is p, and the similarity between the hidden states of the target model and MTP is also p. Then the calculation method of P(n) is as follows: The probability that the (n - 1)th draft token is accepted * p * the similarity between the hidden states of the target model and MTP, that is, P(n) = P(n - 1) * p * p P(1) = p P(2) = P(1) * p * p = p^3 P(3) = P(2) * p * p = p^5 … Thus, it is deduced that P(n) = p^(2*n - 1).

[0049] That is to say, there are errors in using the hidden state of the draft model for traditional MTP speculative sampling. Since there are errors in using the hidden state of the MTP draft model, the hidden state of the target model should be used as much as possible. Therefore, as described above, in each round of running MTP or draft model speculative sampling in the present disclosure, it starts from the first draft token generated in the previous round (such as token 28 in the above example) and runs again, regenerating the kv cache for this part, so as to minimize this error accumulation as much as possible.

[0050] Subsequently, at step S150, it is judged whether the speculative sampling ends. This judgment process can adopt the current conventional judgment method, and the judgment method is not the improvement target of this application, so it will not be elaborated here one by one.

[0051] If it is judged at step S150 that the speculative sampling has not ended, then the hidden state of the target model corresponding to the verification tokens [5, 6, 7, 8] and the verification tokens [6, 7, 8, 9] are used as the input of the MTP draft model. Subsequently, the previous speculative sampling steps of the draft model are repeated, that is, the same 3 calls to the MTP draft model are made to obtain 3 draft tokens [10, 11, 12]. Then [9, 10, 11, 12] are used as the input of the target model. In this way, the target model sampling and the MTP draft model are alternately called until the stop token is generated. Therefore, as Figure 3 shown, the speculative sampling in the subsequent decoding process and Figure 2The biggest difference in speculative sampling lies in the sampling module, which samples the logits corresponding to all output tokens of the target model and determines how many of the draft tokens generated by the previous-round MTP draft model can be accepted. Then, the accepted draft tokens and a newly sampled token are used as the input to the next-round MTP draft model. In this way, the target model and the draft model are alternately called until generation stops.

[0052] Figure 4 It is a schematic diagram of the vector splicing processing process instance in the embodiment of the MTP speculative sampling method for large models shown according to an exemplary embodiment. As Figure 3 shown, the input of the MTP draft model has two parts. One is the token sequence [6, 7, 8, 9], and the other is the hidden layer state corresponding to the tokens [5, 6, 7, 8] of the target model. The hidden layer state is a matrix with a dimension of [4, 7168], and each row corresponds to the hidden layer vector of a token. Then, the token sequence [6, 7, 8, 9] is input into the word embedding layer to generate a matrix with a dimension of [4, 7168]. Then, the two matrices are spliced according to the column dimension to obtain a matrix with a dimension of [4, 7168 x 2], which is used as the input to the subsequent layers. That is, the embedding vector of token 6 is spliced with the hidden layer state vector of token 5 corresponding to the target model, and so on. The embedding vector of token 7 is spliced with the hidden layer state vector of token 6 corresponding to the target model, the embedding vector of token 8 is spliced with the hidden layer state vector of token 7 corresponding to the target model, and the embedding vector of token 9 is spliced with the hidden layer state vector of token 8 corresponding to the target model. This is exactly why the initial input of the MTP draft model starts from the first token. The main reason is that each input token vector needs to be spliced with the hidden layer state vector of the previous token corresponding to the target model or the MTP draft model. In this way, compared with the conventional draft model with only tokens as input, the input information is richer, and the probability of the generated draft tokens being accepted will be much higher. That is, for the present disclosure, in the 0th draft model pre-filling processing step, the hidden layer state data of each token in the first-round target model input token set is a vector matrix with the number of tokens as the dimension, and the first-round draft model input token set input into the draft model generates a vector matrix with the same dimension number in the word embedding layer. By splicing the two vector matrices according to the dimension, an input vector matrix for the pre-filling processing to be performed is formed.

[0053] Optionally, during the MTP speculative sampling process, in the case of pipelining parallelism, in the present disclosure, when preprocessing the data processing request in S110, when the large model is in a pipelining parallel state, the maximum number of tokens m per data processing system for processing the data request and the number of tokens p of the obtained data processing request are obtained; and the magnitudes of the number m and the number (p - q) are compared, and the tokens of the data processing request are split into one or more sub-requests in order according to the smaller of the two, and only a meaningless character is prefixed to the head of the first sub-request.

[0054] For example, for the processing of a data processing request, currently, the running process of the model is artificially divided into two stages, prefill and decode. If the current input of the model does not contain the last token, then it is in the prefill state, otherwise it is in the decode state. In the scenario of using MTP speculative sampling, the target model and the MTP draft model run alternately, and the state changes of these two models need to be kept consistent, otherwise there will be problems. Various situations will occur during the running process. For example, assume that the user's initial input is 0 to 26, a total of 27 tokens. There will be three situations: Situation 1: The initial input of the target model is all tokens 0 to 26. At this time, it is in the decode state to generate a new token 27, and then the MTP model processes 1 to 27 and is in the decode state to generate a draft token 28.

[0055] Situation 2: The initial input of the target model is 0 to 25 (system capacity is insufficient). At this time, it is in the prefill state because there is still a 26 not processed and no new token will be generated. Then the MTP model processes 1 to 26 and is in the decode state and generates a draft token 27 before the target model.

[0056] Situation 3: The initial input of the target model is 0 to 24. At this time, it is in the prefill state and no new token will be generated. Then the MTP model processes 1 to 25 and is in the prefill state and no new draft token will be generated. Then the target model processes 25 to 26, is in the decode state and generates a new token 27, and then the MTP model processes 26 to 27, is in the decode state and generates a new draft token 28.

[0057] If there are no artificial restrictions during the system running process, the above three situations may all occur. For situation 2, in the scenario of pipelining parallelism, it will cause the system state to go wrong. Figure 5It is a schematic diagram for handling a case of input data request in the MTP speculative sampling method under the traditional pipeline parallel mode for large models. As Figure 5 shown, assume that the target model is deployed in pipeline parallel on 4 devices now (that is, the model parameters are evenly distributed to 4 devices in the order of layers), the MTP draft model is directly deployed on device 1, and then the target module is deployed on devices 1 - 4. As Figure 5 shown, in the above case 2, first, the parameters of the target model are evenly distributed to 4 devices in the order of layers, and then the MTP draft model is deployed on device 1. Assume that the user input tokens are [0–26], that is, 27 tokens, and assume that the current system capacity limit can run at most 26 tokens at a time. The specific processing process is as follows: At time step 1, Target_P1 runs [0–25] and sends the intermediate layer output to device 2.

[0058] At time step 2, Target_P2 runs [0–25] and sends the intermediate layer output to device 3. At the same time, Target_P1 runs token

[26] and sends the intermediate layer output to device 2.

[0059] At time step 3, Target_P3 runs [0–25] and sends the intermediate layer output to device 4. At the same time, Target_P2 runs token

[26] and sends the intermediate layer output to device 3.

[0060] At time step 4, Target_P4 runs [0–25] and sends the final hidden layer state to device 1. At the same time, Target_P3 runs token

[26] and sends the intermediate layer output to device 4.

[0061] At time step 5, MTP runs [1–26] because it is in the decode state and then samples a draft token 27. At the same time, Target_P4 runs token

[26] and sends the final hidden layer state and logits to device 1.

[0062] At time step 6, the target model has already finished running 0–26 and is in the decode state. It should sample the logits of

[26] , but it is found that a new draft token 27 is added, so the state will change back to prefill, resulting in the inconsistent state between MTP and the target model. The direct reason is that MTP samples before the target model.

[0063] To eliminate this problem, it is necessary to add various conditions for determining whether the draft model has completed the processing of all input tokens of the request. To avoid increasing the complexity of the determination, the present disclosure adopts the method described above. In the pipelined parallel mode, to prevent the decoding process from occurring when the draft model has not been processed completely in the case where the input request is split into two or more parts for any reason, the final draft model processing is directly performed on the part with at least 2 remaining tokens in any possible split situation. Simply put, if the system capacity is not enough to process all tokens, then in the first-round prefill process, at least the last two tokens should be reserved. This way, conflicts will not occur. Figure 6 It is a schematic diagram of the preprocessing of the input data request in the MTP speculative sampling method in the pipelined parallel mode shown according to an exemplary embodiment of the present disclosure. As Figure 6 The processing process shown is as follows: At time step 1, Target_P1 processes [0–24] and sends the intermediate layer output to device 2.

[0064] At time step 2, Target_P2 processes [0–24] and sends the intermediate layer output to device 3. At the same time, Target_P1 processes tokens [25-26] and sends the intermediate layer output to device 2.

[0065] At time step 3, Target_P3 processes [0–24] and sends the intermediate layer output to device 4. At the same time, Target_P2 processes tokens [25-26] and sends the intermediate layer output to device 3.

[0066] At time step 4, Target_P4 processes [0–24] and sends the final hidden layer state to device 1. At the same time, Target_P3 processes tokens [25-26] and sends the intermediate layer output to device 4.

[0067] At time step 5, MTP processes [1–25], and Target_P4 processes tokens [25-26] and sends the final hidden layer state and logits to device 1.

[0068] At time step 6, sample the logits of

[26] of the target model to obtain a new token 27, and then MTP processes [26-27] to sample and obtain a draft token

[28] .

[0069] It can be seen that after reserving at least two tokens in the fill stage, the prefill and decode state changes of the MTP draft model and the target model are consistent, and the situation where the MTP draft model samples earlier than the target model will not occur.

[0070] As described above, q is at least 2 and can also be more, but it is preferably not greater than the maximum number of tokens m per time of the data processing system and less than p.

[0071] Figure 7 It is a block diagram of an embodiment of MTP speculative sampling for a large model shown according to an exemplary embodiment. As Figure 7 shown, the MTP speculative sampling system 700 for a large model according to the present disclosure includes: a request preprocessing component 710, an MTP target model component 720, and an MTP draft model component 730, where The request preprocessing component 710 receives a data processing request from a user, prefixes a meaningless character to the header of the data processing request, so as to obtain a first-round target model input token set of the data processing request after prefix addition. After obtaining the tokens in the first-round target model input token set, the MTP target model component 720 performs a first-round target model pre-filling process of the target model to obtain a 0th new verification token. The MTP draft model component 730 uses the hidden layer state data of each token in the first-round target model input token set and the first-round draft model input token set of the draft model formed by discarding the first token in the first-round target model input token set and adding the 0th new verification token as initial inputs, starts initial speculative sampling a predetermined n times from the 0th time, and obtains n verification draft tokens. Each speculative sampling includes an nth draft model pre-filling process and an nth draft model sampling process thereafter. Among them, the MTP target model component 710 performs a second-round target model pre-filling process of the target model on the 0th new verification token obtained from the first-round target model sampling process and the second-round target model input token set formed by merging the n new speculative tokens obtained from the n draft model sampling processes, obtains the hidden layer state data of each token in the second-round target model input token set and the (n + 1) target model prediction heads of the second-round target model input token set, and sequentially samples based on the raw scores (logits) of the (n + 1) tokens in the second-round target model input token set formed by the (n + 1) target model prediction heads, verifies the last n tokens among the (n + 1) tokens, and in the case of obtaining n verification tokens, samples based on the obtained n verification tokens to obtain a (n + 1)th new verification token, so as to use the hidden layer state data of each token in the second-round target model input token set as the hidden layer state data of each token in the first-round target model input token set and use the second-round draft model input token set of the draft model formed by the finally obtained (n + 1) new verification tokens as the first-round draft model input token set of the draft model as the input of the MTP draft model component, and repeats the process starting from the 0th draft model pre-filling process.

[0072] In the first round of target model pre-filling processing, the MTP target model component 720 performs target model pre-filling processing for the first round based on the tokens in the first-round target model input token set, to obtain the hidden layer state data of each token in the first-round target model input token set and the target model prediction head of the first-round target model input token set; and in the first round of target model sampling processing, the MTP target model component performs the first round of target model sampling processing on the raw scores (logits) of the last token in the first-round target model input token set formed based on the target model prediction head of the first-round target model input token set, to obtain the 0th new verification token.

[0073] The MTP draft model component 730 starts from the 0th time to perform a predetermined n times of initial speculative sampling, including: performing the 0th draft model pre-filling processing, using the hidden layer state data of each token in the first-round target model input token set and the first-round draft model input token set of the draft model formed by discarding the first token in the first-round target model input token set and adding the 0th new verification token as the initial input, performing the 0th draft model pre-filling processing in the first-round draft model pre-filling processing including a predetermined n times of draft model pre-filling processing, to obtain the hidden layer state data of the last token in the first-round draft model input token set as the input value for the 1st draft model pre-filling processing and obtain the 0th draft model prediction head of the first-round draft model input token set; performing the 0th draft model sampling processing, performing the 0th draft model sampling processing in the first-round draft model sampling processing including a predetermined n times of draft model sampling processing on the raw scores (logits) of the last token in the first-round draft model input token set formed based on the 0th draft model prediction head, to obtain the 1st new speculative token as the input value for the 1st draft model pre-filling processing; and performing repeated speculation, using the new speculative token output by the (k - 1)th draft model sampling processing and the hidden layer state data of the token obtained by the (k - 1)th draft model pre-filling processing as the input, performing the kth draft model pre-filling processing in the remaining (n - 1) times of draft model pre-filling processing, to obtain the hidden layer state data of the kth new speculative token and the kth draft model prediction head of the kth new speculative token, and performing the kth draft model sampling processing in the remaining (n - 1) times of draft model sampling processing on the raw scores (logits) of the kth new speculative token formed based on the kth draft model prediction head, to obtain the (k + 1)th new speculative token, until k is equal to (n - 1).

[0074] When the large model is in the pipeline parallel state, the request preprocessing component 710 obtains the maximum number of tokens m per data processing system for processing data requests and the number of tokens p of the obtained data processing request, compares the sizes of the numbers m and (p - 2), and divides the tokens of the data processing request into one or more sub-requests in ascending order according to the smaller of the two numbers, and only prefixes a meaningless character to the head of the first sub-request.

[0075] When the MTP draft model component 730 performs the 0th draft model pre-filling process, the hidden layer state data of each token in the first-round target model input token set received is a vector matrix with the number of tokens as the dimension, and the first-round draft model input token set of the input draft model generates a vector matrix with the same number of dimensions in the word embedding layer. By splicing the two vector matrices according to the dimension, an input vector matrix for the pre-filling process to be processed is formed.

[0076] In summary, compared with the traditional MTP speculative sampling method and system, the system and method of the present disclosure eliminate the defect that the input information is missing because the first token of the input to the MTP draft model is discarded in the traditional MTP speculative sampling by adding a meaningless token before the initial input request. And by using the hidden layer state data of each token in the first-round target model input token set in each round of the draft model pre-filling process, rather than directly using the kvcache of the tokens verified by the previous-round target model, the error accumulation transfer speed in the traditional MTP speculative sampling can be eliminated. Moreover, by splitting the input request into at least two parts and processing them in parallel mode, the possible conflicting situations between pre-filling and decoding or between speculation and sampling can be eliminated.

[0077] Those skilled in the art can understand that the above-mentioned modules can be distributed in the device according to the description of the embodiments, or can be correspondingly changed and distributed in one or more devices different from the present embodiment. The modules of the above embodiments can be combined into one module, or can be further split into multiple sub-modules.

[0078] From the description of the above embodiments, those skilled in the art can easily understand that the exemplary embodiments described herein can be implemented by software or by a combination of software and necessary hardware. Therefore, the technical solutions according to the embodiments of the present disclosure can be embodied in the form of a software product, which can be stored in a non-volatile storage medium (such as a CD-ROM, a USB flash drive, a mobile hard disk, etc.) or on a network, including several instructions to enable a computing device (such as a personal computer, a server, a mobile terminal, or a network device, etc.) to execute the method according to the embodiments of the present disclosure.

[0079] The above specifically shows and describes the exemplary embodiments of the present disclosure. It should be understood that the present disclosure is not limited to the detailed structures, setting manners, or implementation methods described herein; on the contrary, the present disclosure is intended to cover various modifications and equivalent settings included within the spirit and scope of the appended claims.

Claims

1. An MTP speculative sampling method for large models, comprising: A data processing request preprocessing step of receiving a user's data processing request and prefixing a meaningless character to the header of the data processing request, thereby obtaining a first-round target model input token set of the data processing request after adding the prefix; A target model initial speculative sampling processing step of performing a first-round target model pre-filling process and a first-round target model sampling process on the tokens in the first-round target model input token set to obtain the 0th new verification token; A draft model initial speculative sampling processing step of using the hidden layer state data of each token in the first-round target model input token set and the first-round draft model input token set of the draft model formed by discarding the first token and adding the 0th new verification token in the first-round target model input token set as the initial input, starting from the 0th time, performing a predetermined n times of initial speculative sampling, and obtaining n verification draft tokens. Each speculative sampling includes the nth draft model pre-filling process and the subsequent nth draft model sampling process; A target model verification sampling processing step of performing a second-round target model pre-filling process on the second-round target model input token set formed by merging the 0th new verification token obtained from the first-round target model sampling process and the n new speculative tokens obtained from the draft model initial speculative sampling process, obtaining the hidden layer state data of each token in the second-round target model input token set and the (n + 1) target model prediction heads of the second-round target model input token set, and sequentially sampling based on the original scores of the (n + 1) tokens in the second-round target model input token set formed by the (n + 1) target model prediction heads, verifying the last n tokens among the (n + 1) tokens, and in the case of obtaining n verification tokens, sampling to obtain the (n + 1)th new verification token, so as to use the hidden layer state data of each token in the second-round target model input token set as the hidden layer state data of each token in the first-round target model input token set and using the second-round draft model input token set of the draft model formed by the finally obtained (n + 1) new verification tokens as the first-round draft model input token set of the draft model as the input, repeating and starting to execute the 0th draft model pre-filling process step and the subsequent steps.

2. The MTP speculative sampling method for large models according to claim 1, wherein The first-round target model pre-filling process step, based on the tokens in the first-round target model input token set, performs a first-round target model pre-filling process on the target model, obtaining the hidden layer state data of each token in the first-round target model input token set and the target model prediction heads of the first-round target model input token set; and In the first-round target model sampling processing step, the first-round target model sampling processing is performed on the original score of the last token in the first-round target model input token set formed by the target model prediction head based on the first-round target model input token set, to obtain the 0th new verification token.

3. The MTP speculative sampling method for large models according to claim 2, wherein the initial speculative sampling starting from the 0th time for a predetermined n times includes: The 0th draft model pre-filling processing step, using the hidden layer state data of each token in the first-round target model input token set and the first-round draft model input token set of the draft model formed by discarding the first token in the first-round target model input token set and adding the 0th new verification token as the initial input, performing the 0th draft model pre-filling processing in the first-round draft model pre-filling processing including a predetermined n times of draft model pre-filling processing, obtaining the hidden layer state data of the last token in the first-round draft model input token set as the input value for the 1st draft model pre-filling processing and obtaining the 0th draft model prediction head of the first-round draft model input token set; The 0th draft model sampling processing step, performing the 0th draft model sampling processing in the first-round draft model sampling processing including a predetermined n times of draft model sampling processing on the original score of the last token in the first-round draft model input token set formed based on the 0th draft model prediction head, to obtain the 1st new speculative token as the input value for the 1st draft model pre-filling processing; and The repeated speculation step, using the new speculative token output by the (k - 1)th draft model sampling processing and the hidden layer state data of the token obtained by the (k - 1)th draft model pre-filling processing as the input, performing the kth draft model pre-filling processing in the remaining (n - 1) times of draft model pre-filling processing, obtaining the hidden layer state data of the kth new speculative token and the kth draft model prediction head of the kth new speculative token, and performing the kth draft model sampling processing in the remaining (n - 1) times of draft model sampling processing on the original score of the kth new speculative token formed based on the kth draft model prediction head, to obtain the (k + 1)th new speculative token, until k is equal to (n - 1).

4. The MTP speculative sampling method for large models according to claim 3, wherein the data processing request preprocessing step further includes: In the case where the large model is in a pipeline parallel state, obtaining the maximum number of tokens m per time for each data processing system for processing data requests and the number of tokens p of the obtained data processing request; Comparing the size of the quantity m and the quantity (p - q), and splitting the tokens of the data processing request into one or more sub-requests in order according to the smaller of the two quantities, and only prefixing a meaningless character to the head of the first sub-request.

5. The MTP speculative sampling method for large models according to claim 4, wherein q is at least 2 and less than p.

6. The MTP speculative sampling method for large models according to any one of claims 1-5, wherein n is 2-4.

7. The MTP speculative sampling method for large models according to any one of claims 3-5, wherein in the 0th draft model pre-filling processing step, the hidden layer state data of each token in the first-round target model input token set input is a vector matrix with the number of tokens as the dimension, and the first-round draft model input token set of the input draft model generates a vector matrix with the same number of dimensions in the word embedding layer. By splicing the two vector matrices according to the dimension, an input vector matrix to be processed for pre-filling is formed.

8. An MTP speculative sampling system for large models, comprising: A request preprocessing component, an MTP target model component, and an MTP draft model component, wherein the request preprocessing component receives a data processing request from a user, prefixes a meaningless character to the header of the data processing request, so as to obtain the first-round target model input token set of the data processing request after adding the prefix; the MTP target model component, after obtaining the tokens in the first-round target model input token set, performs the first-round target model pre-filling processing on the target model to obtain the 0th new verification token; the MTP draft model component uses the hidden layer state data of each token in the first-round target model input token set and the first-round draft model input token set of the draft model formed by discarding the first token in the first-round target model input token set and adding the 0th new verification token as the initial input, starts from the 0th time to perform a predetermined n times of initial speculative sampling, and obtains n verification draft tokens. Each speculative sampling includes the nth draft model pre-filling processing and the subsequent nth draft model sampling processing; wherein the MTP target model component performs the second-round target model pre-filling processing on the second-round target model input token set formed by merging the 0th new verification token obtained from the first-round target model sampling processing and the n new speculative tokens obtained from the n times of draft model sampling processing, obtains the hidden layer state data of each token in the second-round target model input token set and the (n + 1) target model prediction heads of the second-round target model input token set, and sequentially samples based on the original scores of the (n + 1) tokens in the second-round target model input token set formed by the (n + 1) target model prediction heads, verifies the last n tokens among the (n + 1) tokens, and when n verification tokens are obtained, based on the obtained n verification tokens, samples to obtain the (n + 1)th new verification token, so as to use the hidden layer state data of each token in the second-round target model input token set as the hidden layer state data of each token in the first-round target model input token set and use the second-round draft model input token set of the draft model formed by the finally obtained (n + 1) new verification tokens as the first-round draft model input token set of the draft model as the input of the MTP draft model component, and repeats the processing starting from the 0th draft model pre-filling processing.

9. The MTP speculative sampling system for large models as claimed in claim 8, wherein in the first-round target model pre-filling process of the MTP target model component, based on the tokens in the first-round target model input token set, the first-round target model pre-filling process of the target model is performed to obtain the hidden layer state data of each token in the first-round target model input token set and the target model prediction head of the first-round target model input token set; and in the first-round target model sampling process of the MTP target model component, the first-round target model sampling process is performed on the raw score of the last token in the first-round target model input token set formed based on the target model prediction head of the first-round target model input token set, to obtain the 0th new verification token.

10. The MTP speculative sampling system for large models as claimed in claim 9, wherein the MTP draft model component starts the predetermined n times of initial speculative sampling from the 0th time, including:[[]] performing the 0th draft model pre-filling process, using the hidden layer state data of each token in the first-round target model input token set and the first-round draft model input token set of the draft model formed by discarding the first token in the first-round target model input token set and adding the 0th new verification token as the initial input, performing the 0th draft model pre-filling process in the first-round draft model pre-filling process including the predetermined n times of draft model pre-filling processes, to obtain the hidden layer state data of the last token in the first-round draft model input token set as the input value of the 1st draft model pre-filling process and to obtain the 0th draft model prediction head of the first-round draft model input token set; performing the 0th draft model sampling process, performing the 0th draft model sampling process in the first-round draft model sampling process including the predetermined n times of draft model sampling processes on the raw score of the last token in the first-round draft model input token set formed based on the 0th draft model prediction head, to obtain the 1st new speculative token as the input value of the 1st draft model pre-filling process; and performing repeated speculation, using the new speculative token output by the (k - 1)th draft model sampling process and the hidden layer state data of the token obtained by the (k - 1)th draft model pre-filling process as the input, performing the kth draft model pre-filling process in the remaining (n - 1) times of draft model pre-filling processes, to obtain the hidden layer state data of the kth new speculative token and the kth draft model prediction head of the kth new speculative token, and performing the kth draft model sampling process in the remaining (n - 1) times of draft model sampling processes on the raw score of the kth new speculative token formed based on the kth draft model prediction head, to obtain the (k + 1)th new speculative token, until k is equal to (n - 1).

11. The MTP speculative sampling system for large models as claimed in claim 10, wherein when the large model is in a pipelined parallel state, the request preprocessing component obtains the maximum number of tokens m per data processing system for processing data requests and the number of tokens p of the obtained data processing request, compares the magnitudes of the number m and the number (p - 2), and divides the tokens of the data processing request into one or more sub-requests in ascending order according to the smaller of the two numbers, and prefixes only the first sub-request with a meaningless character at the head.

12. The MTP speculative sampling system for large models as claimed in claim 11, wherein the q is at least 2 and less than p.

13. The MTP speculative sampling system for large models as claimed in any one of claims 10 - 12, wherein the n is 2 - 4.

14. The MTP speculative sampling system for large models as claimed in any one of claims 10 - 12, wherein when the MTP draft model component performs the 0th draft model pre-filling process, the hidden layer state data of each token in the first-round target model input token set received is a vector matrix with the number of tokens as the dimension, and enables the first-round draft model input token set of the input draft model to generate a vector matrix with the same number of dimensions in the word embedding layer, and forms an input vector matrix for the pre-filling process to be performed by concatenating the two vector matrices according to the dimension.

Citation Information

Cited By

  • Large-model end-cloud collaborative reasoning method and system giving consideration to efficiency and privacy

    CN121919915A