Prospective sampling method and device based on draft model token screening

The candidate tokens are generated through the draft model and the scoring mechanism is used to filter, which solves the problem of low generation efficiency when the depth is large in speculative sampling, and realizes efficient and accurate token generation, which is suitable for language models and application scenarios of different scales.

CN119940548APending Publication Date: 2025-05-06BEIJING SILICONFLOW TECHNOLOGY CO LTD

Patent Information

Application Number
CN202510048164.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-13
Publication Date
2025-05-06

AI Technical Summary

Technical Problem

In speculative sampling, the generation efficiency of tokens with larger depths is significantly reduced, and the acceptance rate drops rapidly, making it difficult to guarantee the accuracy and coherence of the generation results.

Method used

The draft model performs inference calculation based on the input sequence, generates multiple candidate tokens, and uses the scoring mechanism to dynamically adjust the speculative sampling path, eliminate invalid or repeated tokens, and accurately filter the path of high-scoring tokens.

Benefits of technology

Significantly improve the probability that the draft model generation token is accepted by the target model, overcome the depth bottleneck problem, optimize the generation structure, improve the inference efficiency, ensure the accuracy and coherence of the generated results, and reduce the calculation cost.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119940548A_ABST
    Figure CN119940548A_ABST
Patent Text Reader

Abstract

The invention provides a speculation sampling method and device based on draft model token screening. The method comprises the following steps: a draft model performs reasoning calculation based on an input sequence to generate a plurality of candidate tokens; calculating a prediction score of each candidate token, wherein the prediction scores are used for representing the probability that the candidate tokens are accepted by the target model; screening the plurality of candidate tokens through the prediction score; inputting the screened candidate tokens into the target model for reasoning calculation again; and sampling a calculation result to generate a target token for subsequent reasoning calculation. According to the speculation sampling method and device based on draft model token screening, the sampling performance can be optimized while the reasoning result quality is ensured, the speculation sampling efficiency during large model reasoning optimization is effectively improved, and the application range is wide.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the field of computer information processing, and in particular to a speculative sampling method and device based on draft model token screening. Background Art

[0002] With the rapid development of natural language processing (NLP) technology, large-scale language models (such as GPT, Qwen, etc.) have demonstrated excellent performance in tasks such as text generation, question answering, and translation. However, the high performance of these models is often accompanied by huge computing overhead and resource requirements. In this context, in order to balance performance and efficiency, speculative sampling technology came into being.

[0003] Speculative sampling is a strategy to optimize generation efficiency. The core idea is to predict the tokens (called draft tokens) that may be generated by the target model through a draft model with a small number of parameters, and input these tokens into the target model for verification.

[0004] The acceptance rate is the proportion of draft tokens that are accepted by the target model. In speculative sampling, the acceptance rate of a token is determined by the draft model's match and the layer depth. When generating tokens with a larger depth, the acceptance rate drops rapidly. For example, a draft model with a match of 63% has an acceptance rate of only 0.63 at a depth of 7. 7 ≈3.94%, resulting in a significant decrease in generation efficiency.

[0005] The above information disclosed in the Background section is only for enhancement of understanding of the background of the present application and therefore it may contain information that does not constitute the prior art that is already known to a person of ordinary skill in the art. Summary of the invention

[0006] In view of this, the present application provides a speculative sampling method and device based on draft model token screening, which has the following advantages: The scoring mechanism is used to dynamically adjust the speculative sampling path, significantly improving the probability that the token generated by the draft model is accepted by the target model, overcoming the depth bottleneck problem.

[0007] Through scoring and screening, invalid or duplicate tokens are eliminated, the generation structure is optimized, and the reasoning efficiency is improved.

[0008] Accurately screen high-scoring token paths to ensure the accuracy and consistency of generated results while reducing computational costs.

[0009] This method is compatible with a variety of draft models and speculative sampling methods, and is suitable for language models and application scenarios of different scales.

[0010] Other features and advantages of the present application will become apparent from the following detailed description, or may be learned in part by the practice of the present application.

[0011] According to one aspect of the present application, a speculative sampling method based on draft model token screening is proposed, the method comprising: the draft model performs inference calculation based on an input sequence to generate multiple candidate tokens; calculates a prediction score for each candidate token respectively, the prediction score being used to represent the probability that the candidate token is accepted by a target model; the multiple candidate tokens are screened using the prediction score; the screened candidate tokens are input into the target model for inference calculation again; and the calculation results are sampled to generate a target token for subsequent inference calculation.

[0012] In an exemplary embodiment of the present application, the draft model performs reasoning calculation based on an input sequence to generate multiple candidate tokens, including: inputting the input sequence into the draft model multiple times; the draft model performs multiple reasoning and sampling to generate the multiple candidate tokens.

[0013] In an exemplary embodiment of the present application, the draft model performs multiple reasoning and sampling to generate the multiple candidate tokens, including: the draft model generates n candidate tokens after d reasoning and sampling, where d is the maximum depth of the token tree structure.

[0014] In an exemplary embodiment of the present application, the prediction score of each candidate token is calculated respectively, including: generating the corresponding prediction score based on the estimated probability and speculative probability of the candidate token obtained during speculative reasoning.

[0015] In an exemplary embodiment of the present application, generating the corresponding prediction score based on the estimated probability and speculative probability of the candidate token obtained during speculative reasoning includes: in, Score the predictions, is the current token, is the estimated probability of the parent node of the current token, It is the speculation probability of the current token.

[0016] In an exemplary embodiment of the present application, the multiple candidate tokens are screened using the predicted scores, including: arranging the multiple candidate tokens from large to small according to the predicted scores, and extracting candidate tokens of target ranking according to the arrangement order; screening the multiple candidate tokens based on the predicted scores of the candidate tokens of the target ranking; and eliminating candidate tokens located at leaf nodes.

[0017] In an exemplary embodiment of the present application, the multiple candidate tokens are screened based on the predicted score of the candidate token of the target ranking, including: taking the predicted score of the candidate token of the target ranking as a screening threshold; and retaining the candidate tokens among the multiple candidate tokens that are greater than the screening threshold.

[0018] In an exemplary embodiment of the present application, the plurality of candidate tokens are screened by the prediction score, and further comprises: arranging the screened candidate tokens in the order of DFS.

[0019] In an exemplary embodiment of the present application, the calculation results are sampled to generate a target token for subsequent reasoning calculations, including: sampling the calculation results, accepting some tokens in the calculation results to generate a target token; performing the next reasoning calculation through the target token, and recalculating the uncached target token in the next reasoning calculation.

[0020] According to one aspect of the present application, a speculative sampling device based on draft model token screening is proposed, and the device includes: a draft module, which is used for the draft model to perform inference calculation based on an input sequence to generate multiple candidate tokens; a scoring module, which is used to calculate the predicted score of each candidate token respectively, and the predicted score is used to represent the probability that the candidate token is accepted by a target model; a screening module, which is used to screen the multiple candidate tokens according to the predicted score; a target module, which is used to input the screened candidate tokens into the target model for inference calculation again; and a sampling module, which is used to sample the calculation results to generate a target token for subsequent inference calculation.

[0021] According to one aspect of the present application, an electronic device is proposed, which includes: one or more processors; a storage device for storing one or more programs; when the one or more programs are executed by the one or more processors, the one or more processors implement the method as described above.

[0022] According to one aspect of the present application, a computer-readable medium is provided, on which a computer program is stored. When the program is executed by a processor, the method described above is implemented.

[0023] According to the speculative sampling method and device based on draft model token screening of the present application, the draft model is used to perform inference calculation based on the input sequence to generate multiple candidate tokens; the prediction score of each candidate token is calculated separately, and the prediction score is used to represent the probability of the candidate token being accepted by the target model; the multiple candidate tokens are screened according to the prediction score; the screened candidate tokens are input into the target model for inference calculation again; the calculation results are sampled to generate target tokens for subsequent inference calculations. This method can optimize the sampling performance while ensuring the quality of the inference results, effectively improve the efficiency of speculative sampling during large model inference optimization, and has a wide range of adaptability.

[0024] It should be understood that the foregoing general description and the following detailed description are exemplary only and are not restrictive of the present application. BRIEF DESCRIPTION OF THE DRAWINGS

[0025] The above and other objects, features and advantages of the present application will become more apparent by describing in detail the exemplary embodiments thereof with reference to the accompanying drawings. The accompanying drawings described below are only some embodiments of the present application, and it is clear to a person skilled in the art that other accompanying drawings can be obtained from these accompanying drawings without creative effort.

[0026] Figure 1 The present invention is a flowchart of a speculative sampling method based on draft model token screening according to an exemplary embodiment.

[0027] Figure 2 The present invention is a flowchart of a speculative sampling method based on draft model token screening according to an exemplary embodiment.

[0028] Figure 3 is a schematic diagram of a speculative sampling method based on draft model token screening according to another exemplary embodiment.

[0029] Figure 4 is a schematic diagram of a speculative sampling method based on draft model token screening according to another exemplary embodiment.

[0030] Figure 5 It is a block diagram of a speculative sampling device based on draft model token screening according to an exemplary embodiment.

[0031] Figure 6 It is a block diagram of an electronic device according to an exemplary embodiment.

[0032] Figure 7It is a block diagram of a computer-readable medium according to an exemplary embodiment. DETAILED DESCRIPTION

[0033] Example embodiments will now be described more fully with reference to the accompanying drawings. However, example embodiments can be implemented in many forms and should not be construed as limited to the embodiments set forth herein; rather, these embodiments are provided so that this application will be comprehensive and complete and will fully convey the concepts of the example embodiments to those skilled in the art. The same reference numerals in the figures represent the same or similar parts, and thus their repeated description will be omitted.

[0034] In addition, described feature, structure or characteristic can be combined in one or more embodiments in any suitable manner. In the following description, many specific details are provided to provide a full understanding of the embodiments of the present application. However, those skilled in the art will appreciate that the technical scheme of the present application can be put into practice without one or more of the specific details, or other methods, components, devices, steps, etc. can be adopted. In other cases, known methods, devices, realizations or operations are not shown or described in detail to avoid blurring the various aspects of the application.

[0035] The block diagrams shown in the accompanying drawings are merely functional entities and do not necessarily correspond to physically independent entities. That is, these functional entities may be implemented in software form, or in one or more hardware modules or integrated circuits, or in different networks and / or processor devices and / or microcontroller devices.

[0036] The flowcharts shown in the accompanying drawings are only exemplary and do not necessarily include all the contents and operations / steps, nor must they be executed in the order described. For example, some operations / steps can be decomposed, and some operations / steps can be combined or partially combined, so the actual execution order may change according to actual conditions.

[0037] It should be understood that although the terms first, second, third, etc. may be used herein to describe various components, these components should not be limited by these terms. These terms are used to distinguish one component from another component. Therefore, the first component discussed below can be referred to as the second component without departing from the teachings of the concepts of the present application. As used herein, the term "and / or" includes any one of the associated listed items and all combinations of one or more.

[0038] Those skilled in the art will appreciate that the drawings are merely schematic diagrams of example embodiments, and the modules or processes in the drawings are not necessarily necessary for implementing the present application, and therefore cannot be used to limit the scope of protection of the present application.

[0039] The technical abbreviations involved in this application are explained as follows: NLP (Natural Language Processing), natural language processing, the interaction technology between computers and human languages, is applied to text generation, translation, question and answer and other fields.

[0040] Token is the smallest unit in natural language processing, usually a word, subword or character, used as input or output of a language model.

[0041] Draft Model, a language model with a small number of parameters, is used to quickly generate candidate tokens as a preliminary reference for target model reasoning.

[0042] Medusa Model, a multi-head prediction model, is a speculative sampling technology model that generates multiple candidate tokens through multiple parallel prediction heads and is suitable for efficient multi-path reasoning.

[0043] Target Model refers to the large-scale language model that actually performs the final reasoning task, which is used to verify and generate the final high-quality output.

[0044] EOS (End of Sequence) is the end mark of the sequence, indicating that the generated text sequence has reached the end point. The model stops generating after encountering the EOS token.

[0045] BFS (Breadth-First Search) is an algorithm that traverses nodes from top to bottom in a tree structure.

[0046] DFS (Depth-First Search), depth-first search, is an algorithm that deeply traverses nodes along subpaths in a tree structure. It is often used to optimize path selection in the reasoning process.

[0047] Speculative Acceptance Rate, the proportion of candidate tokens accepted by the target model, reflects the efficiency and accuracy of speculative sampling.

[0048] Figure 1 The flowchart of a speculative sampling method based on draft model token screening according to an exemplary embodiment is shown. The speculative sampling method 20 based on draft model token screening at least includes steps S102 to S110.

[0049] like Figure 1As shown, in S102, the draft model performs inference calculation based on the input sequence to generate multiple candidate tokens. For example, the input sequence can be input into the draft model multiple times; the draft model performs multiple inferences and sampling to generate the multiple candidate tokens.

[0050] More specifically, the draft model generates n candidate tokens after d inferences and samplings, where d is the maximum depth of the token tree structure.

[0051] The input sequence is input into the draft model multiple times. The draft model performs multiple inferences and samplings, builds a token tree structure, and generates multiple candidate tokens. Assuming that the maximum depth of the tree is d, the number of candidate tokens that can be generated at each layer is limited by the pruning strategy. Each parent node generates several child nodes to form a tree structure. In a specific embodiment, assuming that the root node generates 3 candidate tokens, each node in the first layer generates 2, and the second layer generates 1, the final number of tokens generated depends on the maximum depth and pruning rules.

[0052] In S104, a prediction score of each candidate token is calculated, and the prediction score is used to represent the probability of the candidate token being accepted by the target model. The prediction score corresponding to the candidate token can be generated based on the estimated probability and the speculative probability of the candidate token obtained during speculative reasoning.

[0053] During the Tree Decoding process of the draft model, the model originally stored two core member variables: DraftParents: used to record the parent node relationship of each node in the tree structure. Draft Levels: used to record the level of each node in the tree structure.

[0054] In this application, a new member variable can be added: Draft Credit, which is used to store the predicted score of each Draft Token, indicating the estimated probability of the Token being accepted in the target model (TargetModel).

[0055] Since the probability of the target model cannot be obtained before sampling, there is no way to know the true probability of a token being accepted in advance during the calculation process. In this application, the probability calculated by speculation (such as the draft model) is used as an approximation of the acceptance probability, and the error of the approximation depends on the degree of matching.

[0056] In a specific embodiment, generating the corresponding prediction score based on the estimated probability and speculative probability of the candidate token obtained during speculative reasoning includes: in, Score the predictions, is the current token, is the estimated probability of the parent node of the current token, It is the speculation probability of the current token.

[0057] Through recursive calculation, the predicted score of a node can be expressed as the product of the speculative probabilities of all parent nodes on its path. This calculation adopts the idea of ​​dynamic programming, caching the credit value of the parent node to avoid repeated calculation, thereby improving calculation efficiency.

[0058] Since Draft Parents has recorded the connection information between nodes, Draft Credit only needs to correspond to DraftTokens one by one. This design avoids repeated storage of redundant data, making the reasoning of tree structures more efficient. This improvement lays the foundation for subsequent pruning optimization and sampling strategies. Through predictive scoring, high-quality DraftTokens can be screened more accurately, improving generation quality and reducing reasoning computation overhead.

[0059] In S106, the multiple candidate tokens are screened by the predicted scores. The multiple candidate tokens may be arranged from large to small according to the predicted scores, and the candidate tokens of the target ranking are extracted according to the arrangement order; the multiple candidate tokens are screened based on the predicted scores of the candidate tokens of the target ranking; and the candidate tokens at the leaf nodes are eliminated.

[0060] In S108, the selected candidate tokens are input into the target model for further reasoning and calculation. The selected tokens are used as new inputs of the target model, and the target model is used for further reasoning and generation.

[0061] After eliminating invalid candidate tokens, the computational complexity of the target model is significantly reduced.

[0062] In S110, the calculation result is sampled to generate a target token for subsequent reasoning calculation. The calculation result can be sampled, and some tokens in the calculation result are accepted to generate a target token; the next reasoning calculation is performed through the target token, and the target token that is not cached is recalculated in the next reasoning calculation.

[0063] In a specific embodiment, k candidate tokens in the tree structure may be accepted, and the kv cache information of p tokens may be cached; and the kv cache information of kp tokens that are not cached in the partial candidate tokens may be discarded. More specifically, k candidate tokens in the tree structure may be accepted; and the kv cache of the first consecutive p tokens in the k candidate tokens may be cached.

[0064] The kp uncached tokens and a new token are added to the input sequence to generate the current sequence, and the kp uncached tokens are calculated again to perform the next inference calculation.

[0065] For kp tokens that are not cached, recalculate their kv cache information and add it to the inference calculation. This dynamic recalculation mechanism can ensure the correctness of the generated sequence, while effectively utilizing cache resources and avoiding storing redundant information.

[0066] According to the speculative sampling method based on draft model token screening of the present application, the draft model is used to perform inference calculation based on the input sequence to generate multiple candidate tokens; the prediction score of each candidate token is calculated separately, and the prediction score is used to represent the probability of the candidate token being accepted by the target model; the multiple candidate tokens are screened according to the prediction score; the screened candidate tokens are input into the target model for inference calculation again; the calculation results are sampled to generate target tokens for subsequent inference calculations. This method can optimize the sampling performance while ensuring the quality of the inference results, effectively improve the efficiency of speculative sampling during large model inference optimization, and has a wide range of adaptability.

[0067] Figure 2 The present invention is a flowchart of a speculative sampling method based on draft model token screening according to an exemplary embodiment. Figure 2 The process 20 shown is for Figure 1 A detailed description of S106 “screening the multiple candidate tokens by using the prediction scores” in the process shown.

[0068] like Figure 2 As shown, in S202, the multiple candidate tokens are arranged from large to small according to the predicted scores, and the candidate tokens of the target ranking are extracted according to the arrangement order.

[0069] Sort the generated multiple candidate tokens from large to small according to the draft credit. The draft credit indicates the probability of a token being accepted by the target model and is an important metric for candidate tokens.

[0070] The purpose of sorting is to prioritize tokens that are more likely to be accepted by the target model, thereby improving generation efficiency and quality.

[0071] Determine the screening threshold, that is, the number of candidate tokens (such as k) that the target model allows to enter the next step of reasoning, and extract the top k tokens from the sorted candidate tokens, which are called target ranking candidate tokens. These tokens will become the basis for subsequent screening.

[0072] In S204, the plurality of candidate tokens are screened based on the predicted scores of the candidate tokens of the target ranking. Given a bunch of draft tokens and draft credits, as many draft tokens with larger draft credits as possible are retained.

[0073] The predicted score of the candidate token of the target ranking may be used as a screening threshold; and the candidate tokens with a score greater than the screening threshold among the multiple candidate tokens are retained.

[0074] More specifically, the predicted score of the last token in the target ranking candidate tokens can be used as the screening threshold. If the allowed number of target tokens k=3, the predicted scores of the top three tokens are 0.8, 0.7, and 0.6 respectively.

[0075] The screening threshold is 0.6.

[0076] Traverse all candidate tokens, retain tokens with predicted scores greater than or equal to the screening threshold, and remove tokens with scores lower than the threshold. This process dynamically adjusts the number of candidate tokens to ensure that the tokens entering the next step of reasoning are high-quality candidates.

[0077] It is worth mentioning that we cannot directly remove the draft tokens ranked fourth and later based on the sorting. Because in the calculation, we cannot skip a node and directly take the next node. If a token exists, its parent node must also exist. Although the probability is definitely less than or equal to 1, that is, the credit of the child node will not exceed the credit of the parent node. However, under the influence of parameters such as top-k and top-p, the probability may be 1, that is, it is possible that the credit of a child node is exactly the same as its parent node.

[0078] Here is an example: "You ate", "I", "of", "pear".

[0079] Suppose the draft credits corresponding to three draft tokens ["我", "的", "梨"] are [0.4, 0.4, 0.1] respectively, and only one draft token can be selected for calculation. Then the sorting might be like this ["的", "我", "梨"], because the credits of "的" and "我" are the same, and either one could come first or second.

[0080] If "的" is used but "我" is not, then since the parent node "我" of "的" is pruned, an error will occur directly.

[0081] In S206, candidate tokens located at leaf nodes are removed.

[0082] In this application, a leaf node can refer to a node that has reached the maximum sequence length specified by the user or the machine hardware limit. In this application, a leaf node can also refer to a node containing the EOS (End of Sequence) marker, and such nodes mark the end of the sequence and cannot be extended further.

[0083] Tokens at leaf nodes cannot generate subsequent tokens, so it will increase the inference cost without improving the generation effect. Traverse all candidate tokens and remove all leaf node tokens. This removal operation can reduce the inference calculation burden of the target model and avoid generating redundant invalid sequences.

[0084] If a leaf node is used as a draft token, then it has two results: 1) It is accepted, and then since it reaches here, it is directly returned to the user without generating subsequent tokens.

[0085] 2) It is not accepted. Subsequent tokens cannot be generated.

[0086] Using one token as a draft token is to generate several more tokens during one round of decoding, while leaf nodes cannot bring subsequent tokens. Therefore, whether there are leaf nodes or not does not affect the maximum accepted length, but having leaf nodes will increase the calculation amount of the target model's inference. So all leaf nodes can be removed.

[0087] Take an example: If there is only one draft token "." after "你吃了", then there are three possible results, accepting or not accepting ".", and the possible results are: 1) Do not accept ".", and generate 1 token: ["你吃了", "苹果"] 2) Accept ".", and generate 2 tokens: ["你吃了", ".", "EOS"] 3) Accept "." and generate 2 tokens: ["you ate", ".", "I"] If “you ate” is followed by two draft tokens, “.”, “EOS”, then there are three possible results: 1) "." is not accepted, and 1 token is generated: ["You ate", "Apple"] 2) Accept EOS and generate 2 tokens: [“you ate”, “.”, “EOS”] 3) Do not accept EOS, generate 2 tokens: ["you ate", ".", "I"] In S208, the selected candidate tokens are arranged in the order of DFS. The selected candidate tokens are still in a tree structure (Draft Tree). In general, they are sorted in the BFS manner. In this application, the draft tokens can be arranged in the DFS manner instead of the BFS arrangement.

[0088] DFS sorting can organize tokens in the order of the path from the root node to the leaf node, so that subsequent reasoning can fully and continuously utilize the screening results.

[0089] The DFS sorting process is as follows: Starting from the root node of the Draft Tree, visit each child node in turn and recursively go deeper. After completing the traversal of a branch, return to visit other branches until the entire tree structure is traversed. The arranged token sequence will be passed as input to the target model to ensure the orderliness and coherence of the reasoning process.

[0090] It should be clearly understood that the present application describes how to form and use specific examples, but the principles of the present application are not limited to any details of these examples. On the contrary, based on the teaching of the content disclosed in the present application, these principles can be applied to many other embodiments.

[0091] Figure 3 is a schematic diagram of a speculative sampling method based on draft model token screening according to another exemplary embodiment. Figure 3 The process shown is Figure 1 The process shown is described in detail.

[0092] First, the input sequence is input into the draft model. The draft model performs inference calculation based on the input sequence to generate multiple candidate tokens. More specifically, the input sequence can be input into the draft model multiple times; the draft model performs multiple inferences and sampling to generate the multiple candidate tokens.

[0093] During the reasoning process of the draft model, token filtering is performed after each sampling, and the remaining tokens after token filtering are input into the draft model again for the next speculative calculation.

[0094] In the process of token screening, the token sorting can be optimized after each screening. More specifically, the token sorting can be optimized using the DFS method to increase the speed of subsequent speculative sampling.

[0095] After repeating the speculation process of the draft model many times, the obtained token sequence is sent to the target model for further model reasoning. During the reasoning process of the target model, the KV value of each token needs to be calculated and cached.

[0096] Afterwards, the token sequence is sampled. During the sampling process, some tokens are accepted and some tokens are cached. The first P tokens of these tokens are continuous with the original sequence.

[0097] Finally, the uncached tokens from the previous step and the newly accepted p tokens are cached, and the remaining tokens and a new token are added to the sequence to continue reasoning.

[0098] Under this screening method of the present application, the conditions for the generation of the draft tree can be appropriately relaxed. Compared with the method where each parent node only generates one draft token (sequential speculation) and each parent node generates s drafttokens (s=2,3,4, tree speculation), the maximum number of children (that is, degree) that each layer of node can generate can be customized in the present application. For example max_degree_per_levels=3,2,1 That is, a tree of up to 3 layers is generated. The root node can only generate up to 3 draft tokens, each node in the first layer can only generate up to 2 draft tokens, and each node in the second layer can only generate up to 1 draft token.

[0099] Assuming that the maximum number of draft tokens that can be generated by the sequence is 6, the possible screening situations are shown as follows: Figure 4 As shown: The possible BFS sequence before entering the target model inference is: [“you ate” “his” “my” “pear” “of” “of” “apple”] The sequence after DFS sorting in this application is: [("You ate", "he", "s", "apple", "my", "s", "pear")] A total of 15 draft token candidates were generated, and the degrees of each layer of nodes were 3, 2, and 1 respectively. Finally, 6 draft tokens were selected for the inference of the target model.

[0100] It is worth mentioning that the sequence sorted by BFS can also be used for token screening in this application, which can also speed up the token speculative sampling.

[0101] More specifically, some tokens in the second layer have very low probabilities and are not adopted at all. When generating the draft tokens in the third layer, there is no need to generate subsequent tokens for the draft tokens that are unlikely to be adopted.

[0102] Since only 6 tokens are finally accepted, among the 9 tokens in the first layer plus the second layer, the 3 with the lowest probabilities will definitely be excluded. According to the algorithm, among ["he", "I", "pear", "s", "eat", "s", "eat", "了", "。"], the tokens tied for 6th are "了" and "。", and their credits are both 0.1.

[0103] According to the algorithm, "eat" (0.03) and "eat" (0.02) will be excluded, and there is no need to generate the subsequent draft tokens "s" (0.03) and "s" (0.02) in the third layer. This can reduce some of the computational effort in the inference of the draft model.

[0104] When the subsequent tokens in the third layer: ["apple" (0.12), "apple" (0.06), "。" (0.1), "EOS" (0.1)] are inferred, we know through the algorithm that among these new and old tokens, the token ranked 6th is "apple" (0.12) in the third layer. This means that "了" (0.1) and "。" (0.1) in the second layer also need to be excluded.

[0105] However, in this process, the previous layers will not be excluded before the inference of all draft models is completed. Because once excluded, it means a change in position, which not only requires major modifications to the tree structure (draft parents), but also the kv cache of the draft model will become invalid. So only the newly generated tokens in the third layer with credits less than 0.12 are excluded: ["apple" (0.06), "。" (0.1), "EOS" (0.1)].

[0106] Those skilled in the art will appreciate that all or part of the steps for implementing the above embodiments are implemented as a computer program executed by a CPU. When the computer program is executed by the CPU, the above functions defined by the above method provided in the present application are performed. The program may be stored in a computer-readable storage medium, which may be a read-only memory, a disk, or an optical disk, etc.

[0107] In addition, it should be noted that the above figures are only schematic illustrations of the processes included in the method according to the exemplary embodiments of the present application, and are not intended to be limiting. It is easy to understand that the processes shown in the above figures do not indicate or limit the time sequence of these processes. In addition, it is also easy to understand that these processes can be performed synchronously or asynchronously, for example, in multiple modules.

[0108] The following is an embodiment of the device of the present application, which can be used to execute the embodiment of the method of the present application. For details not disclosed in the embodiment of the device of the present application, please refer to the embodiment of the method of the present application.

[0109] Figure 5 is a block diagram of a speculative sampling device based on draft model token screening according to an exemplary embodiment. Figure 5 As shown, the speculative sampling device 50 based on draft model token screening includes: a draft module 502 , a scoring module 504 , a screening module 506 , a target module 508 , and a sampling module 510 .

[0110] The draft module 502 is used for the draft model to perform reasoning calculations based on the input sequence to generate multiple candidate tokens; the draft module 502 is also used to input the input sequence into the draft model multiple times; the draft model performs multiple reasoning and sampling to generate the multiple candidate tokens.

[0111] The scoring module 504 is used to calculate the predicted score of each candidate token respectively, and the predicted score is used to represent the probability of the candidate token being accepted by the target model; the scoring module 504 is also used to generate the corresponding predicted score based on the estimated probability and speculative probability of the candidate token obtained during speculative reasoning.

[0112] The screening module 506 is used to screen the multiple candidate tokens according to the predicted scores; the screening module 506 is also used to arrange the multiple candidate tokens from large to small according to the predicted scores, and extract the candidate tokens of the target ranking according to the arrangement order; the multiple candidate tokens are screened based on the predicted scores of the candidate tokens of the target ranking; and the candidate tokens located at the leaf nodes are eliminated.

[0113] The target module 508 is used to input the screened candidate tokens into the target model for further reasoning calculation; The sampling module 510 is used to sample the calculation results to generate a target token for subsequent reasoning calculations. The sampling module 510 is also used to sample the calculation results, accept some tokens in the calculation results to generate a target token; perform the next reasoning calculation through the target token, and recalculate the target token that is not cached in the next reasoning calculation.

[0114] According to the speculative sampling device based on draft model token screening of the present application, the draft model is used to perform inference calculation based on the input sequence to generate multiple candidate tokens; the prediction score of each candidate token is calculated separately, and the prediction score is used to represent the probability of the candidate token being accepted by the target model; the multiple candidate tokens are screened according to the prediction score; the screened candidate tokens are input into the target model for inference calculation again; the calculation results are sampled to generate the target token for subsequent inference calculation. This method can optimize the sampling performance while ensuring the quality of the inference results, effectively improve the efficiency of speculative sampling during large model inference optimization, and has a wide range of adaptability.

[0115] Figure 6 It is a block diagram of an electronic device according to an exemplary embodiment.

[0116] Refer to the following Figure 6 The electronic device 600 according to this embodiment of the present application is described. Figure 6 The electronic device 600 shown is merely an example and should not bring any limitation to the functions and scope of use of the embodiments of the present application.

[0117] like Figure 6 As shown, the electronic device 600 is in the form of a general computing device. The components of the electronic device 600 may include but are not limited to: at least one processing unit 610, at least one storage unit 620, a bus 630 connecting different system components (including the storage unit 620 and the processing unit 610), a display unit 640, etc.

[0118] The storage unit stores a program code, which can be executed by the processing unit 610, so that the processing unit 610 performs the steps described in this specification according to various exemplary embodiments of the present application. For example, the processing unit 610 can perform the following steps: Figure 1 , Figure 2 , Figure 4 Follow the steps shown in .

[0119] The storage unit 620 may include a readable medium in the form of a volatile storage unit, such as a random access memory unit (RAM) 6201 and / or a cache memory unit 6202 , and may further include a read-only memory unit (ROM) 6203 .

[0120] The storage unit 620 may also include a program / utility 6204 having a set (at least one) of program modules 6205, such program modules 6205 including but not limited to: an operating system, one or more application programs, other program modules and program data, each of which or some combination may include the implementation of a network environment.

[0121] Bus 630 may represent one or more of several types of bus structures, including a memory unit bus or memory unit controller, a peripheral bus, an accelerated graphics port, a processing unit, or a local bus using any of a variety of bus architectures.

[0122] The electronic device 600 may also communicate with one or more external devices 600' (e.g., keyboards, pointing devices, Bluetooth devices, etc.) so that a user can communicate with the device that interacts with the electronic device 600, and / or the electronic device 600 can communicate with any device that communicates with one or more other computing devices (e.g., routers, modems, etc.). Such communication may be performed through an input / output (I / O) interface 650. In addition, the electronic device 600 may also communicate with one or more networks (e.g., local area networks (LANs), wide area networks (WANs), and / or public networks, such as the Internet) through a network adapter 660. The network adapter 660 may communicate with other modules of the electronic device 600 through a bus 630. It should be understood that, although not shown in the figure, other hardware and / or software modules may be used in conjunction with the electronic device 600, including but not limited to: microcode, device drivers, redundant processing units, external disk drive arrays, RAID systems, tape drives, and data backup storage systems, etc.

[0123] Through the above description of the implementation, it is easy for those skilled in the art to understand that the example implementation described here can be implemented by software, or by combining software with necessary hardware. Figure 7 As shown, the technical solution according to the implementation mode of the present application can be embodied in the form of a software product, which can be stored in a non-volatile storage medium (which can be a CD-ROM, a USB flash drive, a mobile hard disk, etc.) or on a network, and includes a number of instructions to enable a computing device (which can be a personal computer, a server, or a network device, etc.) to execute the above method according to the implementation mode of the present application.

[0124] The software product may be any combination of one or more readable media. The readable medium may be a readable signal medium or a readable storage medium. The readable storage medium may be, for example, but not limited to, a system, device or component of electricity, magnetism, light, electromagnetic, infrared, or semiconductor, or any combination thereof. More specific examples (non-exhaustive list) of readable storage media include: an electrical connection with one or more wires, a portable disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination thereof.

[0125] The computer readable storage medium may include a data signal propagated in a baseband or as part of a carrier wave, wherein a readable program code is carried. This propagated data signal may take a variety of forms, including but not limited to an electromagnetic signal, an optical signal, or any suitable combination of the above. The readable storage medium may also be any readable medium other than a readable storage medium, which may send, propagate, or transmit a program for use by an instruction execution system, an apparatus, or a device or used in combination with it. The program code contained on the readable storage medium may be transmitted with any appropriate medium, including but not limited to wireless, wired, optical cable, RF, etc., or any suitable combination of the above.

[0126] Program code for performing the operations of the present application may be written in any combination of one or more programming languages, including object-oriented programming languages, such as Java, C++, etc., and conventional procedural programming languages, such as "C" or similar programming languages. The program code may be executed entirely on the user computing device, partially on the user device, as a separate software package, partially on the user computing device and partially on a remote computing device, or entirely on a remote computing device or server. In the case of a remote computing device, the remote computing device may be connected to the user computing device through any type of network, including a local area network (LAN) or a wide area network (WAN), or may be connected to an external computing device (e.g., using an Internet service provider to connect through the Internet).

[0127] The computer-readable medium carries one or more programs. When the one or more programs are executed by a device, the computer-readable medium implements the following functions: the draft model performs inference calculation based on the input sequence to generate multiple candidate tokens; calculates the prediction score of each candidate token respectively, and the prediction score is used to represent the probability that the candidate token is accepted by the target model; the multiple candidate tokens are screened by the prediction score; the screened candidate tokens are input into the target model for inference calculation again; and the calculation results are sampled to generate a target token for subsequent inference calculation.

[0128] Those skilled in the art will appreciate that the above modules can be distributed in the device according to the description of the embodiment, or can be changed accordingly and only used in one or more devices different from the embodiment. The modules of the above embodiments can be combined into one module, or further divided into multiple sub-modules.

[0129] Through the description of the above embodiments, it is easy for those skilled in the art to understand that the example embodiments described here can be implemented by software, or by software combined with necessary hardware. Therefore, the technical solution according to the embodiment of the present application can be embodied in the form of a software product, which can be stored in a non-volatile storage medium (which can be a CD-ROM, a USB flash drive, a mobile hard disk, etc.) or on a network, and includes several instructions to enable a computing device (which can be a personal computer, a server, a mobile terminal, or a network device, etc.) to execute the method according to the embodiment of the present application.

[0130] The exemplary embodiments of the present application are specifically shown and described above. It should be understood that the present application is not limited to the detailed structures, configurations or implementations described herein; on the contrary, the present application is intended to cover various modifications and equivalent configurations included in the spirit and scope of the appended claims.

Claims

1. A speculative sampling method based on draft model token screening, characterized in that: include: The draft model performs inference calculations based on the input sequence and generates multiple candidate tokens; Calculate the prediction score of each candidate token respectively, which is used to indicate the probability that the candidate token is accepted by the target model; Screening the multiple candidate tokens by using the prediction scores; Input the filtered candidate tokens into the target model for further reasoning calculation; The calculation results are sampled to generate target tokens for subsequent reasoning calculations.

2. The method according to claim 1, characterized in that The draft model performs inference calculations based on the input sequence and generates multiple candidate tokens, including: inputting the input sequence into the draft model multiple times; The draft model performs multiple inferences and samplings to generate the multiple candidate tokens.

3. The method according to claim 2, characterized in that The draft model performs multiple inferences and sampling to generate the multiple candidate tokens, including: The draft model generates n candidate tokens after d inferences and samplings, where d is the maximum depth of the token tree structure.

4. The method according to claim 1, characterized in that Calculate the prediction score of each candidate token separately, including: The corresponding prediction score is generated based on the estimated probability and speculative probability of the candidate token obtained during speculative reasoning.

5. The method according to claim 4, characterized in that Generating the corresponding prediction score based on the estimated probability and speculative probability of the candidate token obtained during speculative reasoning, including: in, Score the predictions, is the current token, is the estimated probability of the parent node of the current token, It is the speculation probability of the current token.

6. The method according to claim 1, characterized in that Screening the multiple candidate tokens by using the prediction scores includes: Arrange the multiple candidate tokens from large to small according to the predicted scores, and extract the candidate token of the target ranking according to the arrangement order; Screening the multiple candidate tokens based on the predicted scores of the candidate tokens of the target ranking; Eliminate candidate tokens at leaf nodes.

7. The method according to claim 6, characterized in that The plurality of candidate tokens are screened based on the predicted scores of the candidate tokens of the target ranking, including: Using the predicted score of the candidate token of the target ranking as the screening threshold; The candidate tokens among the multiple candidate tokens that are greater than the screening threshold are retained.

8. The method according to claim 6, characterized in that Screening the multiple candidate tokens by using the prediction score also includes: Arrange the filtered candidate tokens in the order of DFS.

9. The method according to claim 1, characterized in that Sampling the calculation results to generate target tokens for subsequent reasoning calculations, including: Sample the calculation results and accept some tokens in the calculation results to generate the target token; The next inference calculation is performed using the target token, and the target token that is not cached is recalculated in the next inference calculation.

10. A speculative sampling device based on draft model token screening, characterized in that: include: The draft module is used for the draft model to perform inference calculations based on the input sequence and generate multiple candidate tokens; A scoring module is used to calculate the prediction score of each candidate token, where the prediction score is used to indicate the probability that the candidate token is accepted by the target model. A screening module, used to screen the multiple candidate tokens according to the prediction score; The target module is used to input the filtered candidate tokens into the target model for further reasoning calculation; The sampling module is used to sample the calculation results to generate target tokens for subsequent reasoning calculations.

Citation Information

Patent Citations

  • Large model reasoning acceleration method and system combining machine learning and speculation sampling

    CN118657220A

  • Large language model reasoning optimization method based on cascade and speculative decoding strategy

    CN119047579A

  • Speculation decoding method and device for large language model and medium

    CN119150848A

  • Text generation method and apparatus, electronic device and medium

    WO2023207690A1

Cited By

  • Large model distributed reasoning acceleration method and device based on speculation sampling

    CN120373477A

  • Large model distributed inference acceleration method and device based on speculative sampling

    CN120373477B

  • Large model speculation reasoning optimization method and system based on online reinforcement learning

    CN121328745A