Medusa model-based token screening method and device
Through the scoring mechanism and path pruning technology, the repetition and incoherence problems in the generation of tokens by Medusa model are solved, and the efficiency and quality of token generation are optimized, which is suitable for long prediction models and speculative sampling technologies.
Patent Information
- Application Number
- CN202510048170.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-13
- Publication Date
- 2025-05-06
- Estimated Expiration
- 2045-01-13
AI Technical Summary
The Medusa model is prone to repetition or incoherence when generating tokens, which affects the generation efficiency and quality.
Through scoring mechanism and path pruning technology, the possibility that the draft tokens generated by Medusa is accepted by the target model is improved, and duplicate tokens and low scoring paths are eliminated to ensure that the generated token sequence is more coherent and accurate.
The token generation efficiency is optimized, the amount of invalid calculation is reduced, and the quality of the generated token sequence is ensured. It is suitable for the Medusa model and other long prediction models and speculative sampling technologies.
Smart Images

Figure CN119940549A_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the field of computer information processing, and in particular to a method and device for token screening based on a medusa model. Background Art
[0002] With the widespread application of large-scale language models (such as GPT, Qwen, etc.), text generation tasks have an important position in the field of natural language processing. However, the high computational overhead of these models limits their application in real-time response and resource-constrained scenarios. In order to reduce the cost of inference and improve generation efficiency, speculative sampling technology has been widely studied and applied.
[0003] Speculative sampling is an optimization method that reduces the computational effort of the target model through prediction. The core idea is to use a smaller-scale model (such as the Medusa model) to generate candidate tokens (draft tokens), and then verify the acceptance probability of these tokens through the target model to improve reasoning efficiency.
[0004] The Medusa model generates multiple draft tokens in parallel through multiple prediction heads (Multi-head), and has the ability to generate multiple candidate tokens at one time, thereby accelerating the reasoning process. However, due to the lack of association between the prediction heads, it is easy to generate repeated or incoherent tokens, affecting the generation efficiency and quality.
[0005] Therefore, a new token screening method and device based on the medusa model is needed.
[0006] The above information disclosed in the Background section is only for enhancement of understanding of the background of the present application and therefore it may contain information that does not constitute the prior art that is already known to a person of ordinary skill in the art. Summary of the invention
[0007] In view of this, the present application provides a token screening method and device based on the medusa model, which has the following technical advantages: Using scoring mechanism and path pruning, we can improve the possibility of Medusa-generated draft tokens being accepted by the target model and optimize the generation efficiency. Remove duplicate tokens through repeated pruning technology to reduce path complexity and invalid calculations; Eliminate low-scoring and incoherent paths to ensure that the final generated token sequence is more coherent and accurate; This method is not only applicable to the Medusa model, but can also be extended to other multi-head forecasting models and speculative sampling techniques, and has wide applicability.
[0008] Other features and advantages of the present application will become apparent from the following detailed description, or may be learned in part by the practice of the present application.
[0009] According to one aspect of the present application, a token screening method based on a medusa model is proposed, the method comprising: performing reasoning calculations on an input sequence using a target model to generate multiple candidate tokens; sampling the multiple candidate tokens to extract some candidate tokens; inputting the some candidate tokens into a medusa model for speculative reasoning to generate multiple speculative tokens in a tree structure; screening the multiple speculative tokens using a pruning method based on the tree structure to generate a target token; performing subsequent reasoning calculations based on the target token to generate an output result.
[0010] In an exemplary embodiment of the present application, a target model is used to perform inference calculation on an input sequence to generate multiple candidate tokens, including: inputting the input sequence into the target model; the target model calculates multiple candidate tokens based on inference in the input sequence; and caching kv cache information corresponding to the multiple candidate tokens.
[0011] In an exemplary embodiment of the present application, the multiple candidate tokens are sampled and some candidate tokens are extracted, including: accepting k candidate tokens from the multiple candidate tokens and caching the kv cache information of p tokens; and discarding the kv cache information of kp tokens that are not cached from the some candidate tokens.
[0012] In an exemplary embodiment of the present application, the part of candidate tokens is input into the medusa model for speculative reasoning to generate multiple speculative tokens in a tree structure, including: each prediction head of the medusa model generates a group of speculative tokens respectively, and organizes them into a tree structure hierarchically.
[0013] In an exemplary embodiment of the present application, the part of candidate tokens is input into the medusa model for speculative reasoning to generate multiple speculative tokens in a tree structure, and further includes: arranging the multiple speculative tokens in the order of DFS.
[0014] In an exemplary embodiment of the present application, the multiple speculative tokens are screened by pruning based on a tree structure to generate a target token, including: screening the multiple speculative tokens by predictive scoring pruning based on the tree structure to generate a target token; and / or screening the multiple speculative tokens by leaf node pruning based on the tree structure to generate a target token; and / or screening the multiple speculative tokens by repeated pruning based on the tree structure to generate a target token.
[0015] In an exemplary embodiment of the present application, the multiple speculative tokens are screened by using a prediction score pruning method based on a tree structure to generate a target token, including: obtaining a prediction score for each speculative token, wherein the prediction score is used to represent the probability that the speculative token is accepted by the target model; and screening the multiple candidate tokens by using the prediction score to generate a target token.
[0016] In an exemplary embodiment of the present application, the plurality of speculative tokens are screened by leaf node pruning based on a tree structure to generate a target token, including: removing speculative tokens located at leaf nodes based on the tree structure to generate a target token.
[0017] In an exemplary embodiment of the present application, the multiple speculative tokens are screened by repeated pruning based on a tree structure to generate a target token, including: removing repeated speculative tokens based on the tree structure to generate a target token.
[0018] According to one aspect of the present application, a token screening device based on a medusa model is proposed, and the device includes: a target module, which is used to use a target model to perform inference calculations on an input sequence to generate multiple candidate tokens; a sampling module, which is used to sample the multiple candidate tokens and extract some candidate tokens; a speculative module, which is used to input the part of the candidate tokens into the medusa model for speculative reasoning to generate multiple speculative tokens in a tree structure; a pruning module, which is used to screen the multiple speculative tokens by pruning based on the tree structure to generate a target token; and an inference module, which is used to perform subsequent inference calculations based on the target token to generate an output result.
[0019] According to one aspect of the present application, an electronic device is proposed, which includes: one or more processors; a storage device for storing one or more programs; when the one or more programs are executed by the one or more processors, the one or more processors implement the method as described above.
[0020] According to one aspect of the present application, a computer-readable medium is provided, on which a computer program is stored. When the program is executed by a processor, the method described above is implemented.
[0021] According to the medusa model-based token screening method and device of the present application, multiple candidate tokens are generated by using the target model to perform inference calculations on the input sequence; the multiple candidate tokens are sampled and some candidate tokens are extracted; the some candidate tokens are input into the medusa model for speculative reasoning to generate multiple speculative tokens in a tree structure; the multiple speculative tokens are screened by pruning based on the tree structure to generate a target token; subsequent inference calculations are performed based on the target token to generate an output result. This method can optimize the token sampling performance while ensuring the quality of the inference results, effectively improve the efficiency of speculative sampling during large model inference optimization, and has a wide range of adaptability.
[0022] It should be understood that the foregoing general description and the following detailed description are exemplary only and are not restrictive of the present application. BRIEF DESCRIPTION OF THE DRAWINGS
[0024] The above and other objects, features and advantages of the present application will become more apparent by describing in detail the exemplary embodiments thereof with reference to the accompanying drawings. The accompanying drawings described below are only some embodiments of the present application, and it is clear to a person skilled in the art that other accompanying drawings can be obtained from these accompanying drawings without creative effort.
[0025] Figure 1 The figure is a flowchart of a token screening method based on the medusa model according to an exemplary embodiment.
[0026] Figure 2 It is a flowchart of a token screening method based on the medusa model according to another exemplary embodiment.
[0027] Figure 3 It is a schematic diagram of a token screening method based on the medusa model according to another exemplary embodiment.
[0028] Figure 4 It is a schematic diagram of a token screening method based on the medusa model according to another exemplary embodiment.
[0029] Figure 5 It is a schematic diagram of a token screening method based on the medusa model according to another exemplary embodiment.
[0030] Figure 6It is a schematic diagram of a token screening method based on the medusa model according to another exemplary embodiment.
[0031] Figure 7 It is a block diagram of a token screening device based on a medusa model according to an exemplary embodiment.
[0032] Figure 8 It is a block diagram of an electronic device according to an exemplary embodiment.
[0033] Fig. 9 It is a block diagram of a computer-readable medium according to an exemplary embodiment. DETAILED DESCRIPTION
[0035] Example embodiments will now be described more fully with reference to the accompanying drawings. However, example embodiments can be implemented in many forms and should not be construed as limited to the embodiments set forth herein; rather, these embodiments are provided so that this application will be comprehensive and complete and will fully convey the concepts of the example embodiments to those skilled in the art. The same reference numerals in the figures represent the same or similar parts, and thus their repeated description will be omitted.
[0036] In addition, described feature, structure or characteristic can be combined in one or more embodiments in any suitable manner. In the following description, many specific details are provided to provide a full understanding of the embodiments of the present application. However, those skilled in the art will appreciate that the technical scheme of the present application can be put into practice without one or more of the specific details, or other methods, components, devices, steps, etc. can be adopted. In other cases, known methods, devices, realizations or operations are not shown or described in detail to avoid blurring the various aspects of the application.
[0037] The block diagrams shown in the accompanying drawings are merely functional entities and do not necessarily correspond to physically independent entities. That is, these functional entities may be implemented in software form, or in one or more hardware modules or integrated circuits, or in different networks and / or processor devices and / or microcontroller devices.
[0038] The flowcharts shown in the accompanying drawings are only exemplary and do not necessarily include all the contents and operations / steps, nor must they be executed in the order described. For example, some operations / steps can be decomposed, and some operations / steps can be combined or partially combined, so the actual execution order may change according to actual conditions.
[0039] It should be understood that although the terms first, second, third, etc. may be used herein to describe various components, these components should not be limited by these terms. These terms are used to distinguish one component from another component. Therefore, the first component discussed below can be referred to as the second component without departing from the teachings of the concepts of the present application. As used herein, the term "and / or" includes any one of the associated listed items and all combinations of one or more.
[0040] Those skilled in the art will appreciate that the drawings are merely schematic diagrams of example embodiments, and the modules or processes in the drawings are not necessarily necessary for implementing the present application, and therefore cannot be used to limit the scope of protection of the present application.
[0041] The technical abbreviations involved in this application are explained as follows: NLP (Natural Language Processing), natural language processing, the interaction technology between computers and human languages, is applied to text generation, translation, question and answer and other fields.
[0042] Token is the smallest unit in natural language processing, usually a word, subword or character, used as input or output of a language model.
[0043] Draft Model, a language model with a small number of parameters, is used to quickly generate candidate tokens as a preliminary reference for target model reasoning.
[0044] Medusa Model, a multi-head prediction model, is a speculative sampling technology model that generates multiple candidate tokens through multiple parallel prediction heads and is suitable for efficient multi-path reasoning.
[0045] Target Model refers to the large-scale language model that actually performs the final reasoning task, which is used to verify and generate the final high-quality output.
[0046] EOS (End of Sequence) is the end mark of the sequence, indicating that the generated text sequence has reached the end point. The model stops generating after encountering the EOS token.
[0047] BFS (Breadth-First Search), breadth-first search, is an algorithm that traverses nodes from top to bottom in a tree structure.
[0048] DFS (Depth-First Search), depth-first search, is an algorithm that deeply traverses nodes along subpaths in a tree structure, and is often used to optimize path selection in the reasoning process.
[0049] Speculative Acceptance Rate, the proportion of candidate tokens accepted by the target model, reflects the efficiency and accuracy of speculative sampling.
[0050] Figure 1 The flowchart of a token screening method based on the medusa model is shown according to an exemplary embodiment. The token screening method based on the medusa model 10 at least includes steps S102 to S110.
[0051] like Figure 1 As shown, in S102, the target model is used to perform inference calculation on the input sequence to generate multiple candidate tokens. The input sequence can be input into the target model; the target model calculates multiple candidate tokens based on the inference in the input sequence; and the kv cache information corresponding to the multiple candidate tokens is cached.
[0052] During the inference process, the target model generates intermediate calculation results and stores them as kv cache (Key-Value cache). Kv cache contains key data of the attention mechanism, which is used to reduce repeated calculations in subsequent inference and improve inference efficiency. These kv caches will be partially retained or discarded based on the selection of candidate Tokens.
[0053] The calculation results may be sampled, and some tokens in the calculation results may be accepted to generate a target token; the next inference calculation may be performed using the target token, and the target token that is not cached may be recalculated in the next inference calculation.
[0054] In S104, the plurality of candidate tokens are sampled to extract some candidate tokens, for example, k candidate tokens from the plurality of candidate tokens are accepted, and kv cache information of p tokens is cached; and kv cache information of kp tokens that are not cached from the partial candidate tokens is discarded.
[0055] In a specific embodiment, k candidate tokens in the tree structure may be accepted, and the kvcache information of p tokens may be cached; and the kv cache information of kp tokens that are not cached in the candidate tokens may be discarded. More specifically, k candidate tokens in the tree structure may be accepted; and the kvcache of the first consecutive p tokens in the k candidate tokens may be cached.
[0056] The kp uncached tokens and a new token are added to the input sequence to generate the current sequence, and the kp uncached tokens are calculated again to perform the next inference calculation.
[0057] For kp tokens that are not cached, recalculate their kv cache information and add it to the inference calculation. This dynamic recalculation mechanism can ensure the correctness of the generated sequence, while effectively utilizing cache resources and avoiding storing redundant information.
[0058] In S106, the part of the candidate tokens is input into the medusa model for speculative reasoning to generate multiple speculative tokens in a tree structure. Each prediction head of the medusa model generates a set of speculative tokens and organizes them into a tree structure in a hierarchical manner.
[0059] More specifically, the sampled candidate tokens are input into the Medusa model for speculative reasoning. The multi-head structure of the Medusa model allows multiple candidate tokens to be generated in parallel, and these tokens are organized into a tree structure. Each prediction head generates a new set of tokens based on the tokens of the previous layer. The speculatively generated tokens are organized in a tree structure, and each layer of tokens corresponds to the parent node of the previous layer.
[0060] In one embodiment, the method further includes: arranging the plurality of speculative tokens in a DFS order. In general, the sequence is sorted in a BFS manner. In this application, the arrangement of draft tokens in BFS can be replaced with a DFS arrangement.
[0061] DFS sorting can organize tokens in the order of the path from the root node to the leaf node, so that subsequent reasoning can fully and continuously utilize the screening results.
[0062] The DFS sorting process is as follows: Starting from the root node of the Draft Tree, visit each child node in turn and recursively go deeper. After completing the traversal of a branch, return to visit other branches until the entire tree structure is traversed. The arranged token sequence will be passed as input to the target model to ensure the orderliness and coherence of the reasoning process.
[0063] In S108, the plurality of speculative tokens are screened by pruning based on the tree structure to generate a target token.
[0064] In one embodiment, for example, the multiple speculative tokens can be screened using a prediction score pruning method based on a tree structure to generate a target token; more specifically, for example, a prediction score of each speculative token can be obtained, and the prediction score is used to represent the probability of the speculative token being accepted by the target model; the multiple candidate tokens are screened using the prediction score to generate a target token.
[0065] In a specific application, during the Tree Decoding process of the draft model, the model originally stores two core member variables: Draft Parents: used to record the parent node relationship of each node in the tree structure. DraftLevels: used to record the level of each node in the tree structure.
[0066] In this application, a new member variable can be added: Draft Credit, which is used to store the predicted score of each Draft Token, indicating the estimated probability of the Token being accepted in the target model (TargetModel).
[0067] Since the probability of the target model cannot be obtained before sampling, there is no way to know the true probability of a token being accepted in advance during the calculation process. In this application, the probability calculated by speculation (such as the draft model) is used as an approximation of the acceptance probability, and the error of the approximation depends on the degree of matching.
[0068] In a specific embodiment, generating the corresponding prediction score based on the estimated probability and speculative probability of the candidate token obtained during speculative reasoning includes: in, Score the predictions, is the current token, is the estimated probability of the parent node of the current token, It is the speculation probability of the current token.
[0069] Through recursive calculation, the predicted score of a node can be expressed as the product of the speculative probabilities of all parent nodes on its path. This calculation adopts the idea of dynamic programming, caching the credit value of the parent node to avoid repeated calculation, thereby improving calculation efficiency.
[0070] Since Draft Parents has recorded the connection information between nodes, Draft Credit only needs to correspond to DraftTokens one by one. This design avoids repeated storage of redundant data, making the reasoning of tree structures more efficient. This improvement lays the foundation for subsequent pruning optimization and sampling strategies. Through predictive scoring, high-quality DraftTokens can be screened more accurately, improving generation quality and reducing reasoning computation overhead.
[0071] More specifically, the multiple candidate tokens may be arranged from large to small according to the predicted scores, and the candidate tokens of the target ranking may be extracted according to the arrangement order; the multiple candidate tokens may be screened based on the predicted scores of the candidate tokens of the target ranking; and the candidate tokens located at the leaf nodes may be eliminated.
[0072] In one embodiment, for example, the multiple speculative tokens may be screened by leaf node pruning based on a tree structure to generate a target token; more specifically, for example, speculative tokens located at leaf nodes may be removed based on a tree structure to generate a target token.
[0073] In this application, a leaf node may refer to a node that has reached the maximum sequence length specified by the user or the machine hardware limit. In this application, a leaf node may also refer to a node containing an EOS (End of Sequence) marker, which marks the end of the sequence and cannot be expanded.
[0074] Leaf node tokens cannot generate subsequent tokens, so they increase the inference cost but do not improve the generation effect. Traverse all candidate tokens and remove all leaf node tokens. This removal operation can reduce the inference calculation burden of the target model and avoid generating redundant invalid sequences.
[0075] If a leaf node is used as a draft token, it has two results: 1) It is accepted and then directly returned to the user without generating subsequent tokens.
[0076] 2) Not accepted. Subsequent tokens cannot be generated.
[0077] The purpose of using a token as a draft token is to generate more tokens in one round of decoding. Leaf nodes cannot bring subsequent tokens, so the presence or absence of leaf nodes does not affect the maximum length accepted. On the contrary, the presence of leaf nodes will increase the computational complexity of the target model inference. Therefore, all leaf nodes can be removed.
[0078] Here is an example: If there is only one draft token "." after "you ate", then there are three possible results: accept or reject ".". The possible results are: 1) "." is not accepted, and 1 token is generated: ["You ate", "Apple"] 2) Accept "." and generate 2 tokens: ["You ate", ".", "EOS"] 3) Accept "." and generate 2 tokens: ["you ate", ".", "I"] If “you ate” is followed by two draft tokens, “.”, “EOS”, then there are three possible results: 1) "." is not accepted, and 1 token is generated: ["You ate", "Apple"] 2) Accept EOS and generate 2 tokens: [“you ate”, “.”, “EOS”] 3) Do not accept EOS, generate 2 tokens: ["you ate", ".", "I"] In one embodiment, for example, the multiple speculative tokens may be screened by repeated pruning based on a tree structure to generate a target token. More specifically, for example, repeated speculative tokens may be eliminated based on a tree structure to generate a target token.
[0079] More specifically, since there is no direct connection between Medusa's headers, it is common for different headers to infer the same token. Therefore, duplicate word pruning is introduced in this application. If a draft token of Medusa contains the current token in its ancestors (i.e., the parent node and the upstream node path of the parent node), the draft token will be deleted.
[0080] In S110, subsequent reasoning calculations are performed based on the target token to generate output results. The pruned target token is considered to be the optimal token set, which is passed to the target model for subsequent reasoning calculations. The high quality of the target token ensures the accuracy of reasoning and the efficiency of generation.
[0081] According to the token screening method based on the medusa model of the present application, a plurality of candidate tokens are generated by performing reasoning calculations on an input sequence using a target model; the plurality of candidate tokens are sampled and some candidate tokens are extracted; the plurality of candidate tokens are input into the medusa model for speculative reasoning to generate a plurality of speculative tokens in a tree structure; the plurality of speculative tokens are screened by pruning based on the tree structure to generate a target token; subsequent reasoning calculations are performed based on the target token to generate an output result. This method can optimize the token sampling performance while ensuring the quality of the reasoning result, effectively improve the efficiency of speculative sampling during large model reasoning optimization, and has a wide range of adaptability.
[0082] It should be clearly understood that the present application describes how to form and use specific examples, but the principles of the present application are not limited to any details of these examples. On the contrary, based on the teaching of the content disclosed in the present application, these principles can be applied to many other embodiments.
[0083] Figure 2 It is a schematic diagram of a token screening method based on the medusa model according to another exemplary embodiment. Figure 2 The process shown is Figure 1 The process shown is described in detail.
[0084] First, the input sequence enters the target model. During the reasoning process of the target model, the KV value of each token needs to be calculated and cached.
[0085] Afterwards, the token sequence is sampled. During the sampling process, some tokens are accepted and some tokens are cached. The first P tokens of these tokens are continuous with the original sequence.
[0086] After sampling, the target model generates multiple candidate tokens, which are then input into the medusa model and speculated again to generate multiple tokens.
[0087] The token sorting can also be optimized. More specifically, the token sorting can be optimized using the DFS method to increase the speed of subsequent speculative sampling.
[0088] Afterwards, multiple tokens are screened, and one or more methods such as prediction score pruning, leaf node pruning, and repeated pruning are used to screen the tokens to generate the target token.
[0089] Finally, cache the tokens that were not cached in the previous step and the newly received p tokens, add the remaining tokens and a new token to the sequence, and add the newly obtained target token to the sequence to continue the inference.
[0090] In the scenario of tree-shaped speculation, the maximum degree per layer takes draft tokens according to 3, 2, 1. The general approach is a full connection, that is, all draft tokens in each layer are connected to all tokens in the lower layer to form the tokens in the lower layer. This can also adopt the Credit Tree and the pruning method based on Credit.
[0091] For example: Use medusa speculation to generate draft tokens for "你吃了" (You have eaten). The first layer is "他" (he) (0.3), "我" (I) (0.2), "苹果" (apple) (0.2). The second layer is "的" (de) (0.6), "了" (le) (0.4). The third layer is "苹果" (apple) (0.3). Then the draft tree and the result selected according to credit are as Figure 3 shown.
[0092] [“你吃了”,“他”(0.3),“我”(0.2),“苹果”(0.2),“的”(0.18),“了”(0.12),“的”(0.12),“的”(0.12)] It can be seen that the sequence is not necessarily smooth under the Cartesian product of multiple heads in medusa. For example, "了" in the second layer obviously cannot be connected to "他" and "我" in the first layer. Another example is that "苹果" in the third layer obviously cannot be connected to "了" in the second layer.
[0093] “你吃了我了苹果” (You have eaten me le apple) is unreasonable. This is the shortcoming of medusa and one of the reasons for the low acceptance rate of multi-head medusa.
[0094] Since there is no direct connection between the heads of Medusa, it often happens that different heads infer the same token.
[0095] Just like the word "苹果" (apple) appeared in both the first layer and the third layer in the previous example.
[0096] The repeated tokens that appear later are very likely not to be accepted. ("你吃苹果了" (You have eaten an apple) is possible, but "你吃苹果了苹果" (You have eaten an apple apple) is impossible.) We have made statistics and the probability of consecutive identical tokens being accepted is extremely low: The following table is a situation table of medusa generating 2 draft tokens: Number of occurrences No repeated draft tokens Repeating the draft token The second draft token is accepted 712 7 The second draft token is not accepted 1407 228 Use the chi-square distribution to verify the correlation between the two random variables of whether there are duplicate draft tokens and whether the second draft token is accepted. There is a 99.9% certainty that these two variables are related.
[0097] When there are no duplicate draft tokens, the probability that the second draft token is accepted is 33.6%. However, when there are duplicate draft tokens, the probability that the second draft token is accepted drops to 3.0%. This is sufficient to show that duplicate draft tokens will significantly reduce the acceptance rate.
[0098] Duplicate word pruning can be performed. If a draft token of medusa has the current token in its ancestors (that is, the path of the parent node and the upstream nodes of the parent node), then this draft token is not added.
[0099] In the above example, the duplicate "apple" will be crossed out, resulting in Figure 4 the result.
[0100] In another specific embodiment, medusa often has a similar distribution for certain tokens in different heads.
[0101] Using medusa to speculatively generate draft tokens for "你吃" (Ni Chi), as Figure 5 shown: The first layer is "了" (0.3), "苹果" (0.2), "的" (0.2), The second layer is "了" (0.6), "的" (0.4), The third layer is "了" (0.3) The result obtained after duplicate pruning is as Figure 6 shown. Under this duplicate pruning, it can be actually measured that the acceptance rate of medusa has been improved to a certain extent (about 6.5%. The acceptance rate has increased from 11.87% to 12.65% with 7 draft tokens). Since many useless token combinations have been reduced, the throughput has increased (about 4%).
[0102] Those skilled in the art will appreciate that all or part of the steps for implementing the above embodiments are implemented as a computer program executed by a CPU. When the computer program is executed by the CPU, the above functions defined by the above method provided in the present application are performed. The program may be stored in a computer-readable storage medium, which may be a read-only memory, a disk, or an optical disk, etc.
[0103] In addition, it should be noted that the above figures are only schematic illustrations of the processes included in the method according to the exemplary embodiments of the present application, and are not intended to be limiting. It is easy to understand that the processes shown in the above figures do not indicate or limit the time sequence of these processes. In addition, it is also easy to understand that these processes can be performed synchronously or asynchronously, for example, in multiple modules.
[0104] The following is an embodiment of the device of the present application, which can be used to execute the embodiment of the method of the present application. For details not disclosed in the embodiment of the device of the present application, please refer to the embodiment of the method of the present application.
[0105] Figure 7 FIG. 1 is a block diagram of a token screening device based on a medusa model according to another exemplary embodiment. Figure 7 As shown, the token screening device 70 based on the medusa model includes: a target module 702, a sampling module 704, a speculation module 706, a pruning module 708, and an inference module 710.
[0106] The target module 702 is used to use the target model to perform inference calculations on the input sequence to generate multiple candidate tokens; the target module 702 is also used to input the input sequence into the target model; the target model calculates multiple candidate tokens based on inference from the input sequence; and caches the kvcache information corresponding to the multiple candidate tokens.
[0107] The sampling module 704 is used to sample the multiple candidate tokens and extract some candidate tokens; the sampling module 704 is also used to accept k candidate tokens from the multiple candidate tokens and cache the kv cache information of p tokens; and discard the kv cache information of kp tokens that are not cached from the partial candidate tokens.
[0108] The speculation module 706 is used to input the candidate tokens into the medusa model for speculative reasoning, and generate multiple speculative tokens in a tree structure; each prediction head of the medusa model generates a set of speculative tokens, and organizes them into a tree structure in a hierarchical manner. The speculation module 706 is also used to arrange the multiple speculative tokens in the order of DFS.
[0109] The pruning module 708 is used to filter the multiple speculative tokens by pruning based on the tree structure to generate a target token; the pruning module 708 is also used to filter the multiple speculative tokens by predictive scoring pruning based on the tree structure to generate a target token; the pruning module 708 is also used to filter the multiple speculative tokens by leaf node pruning based on the tree structure to generate a target token; the pruning module 708 is also used to filter the multiple speculative tokens by repeated pruning based on the tree structure to generate a target token.
[0110] The reasoning module 710 is used to perform subsequent reasoning calculations based on the target token to generate an output result.
[0111] According to the medusa model-based token screening device of the present application, a plurality of candidate tokens are generated by performing inference calculations on an input sequence using a target model; the plurality of candidate tokens are sampled and some candidate tokens are extracted; the some candidate tokens are input into the medusa model for speculative inference to generate a plurality of speculative tokens in a tree structure; the plurality of speculative tokens are screened by pruning based on the tree structure to generate a target token; subsequent inference calculations are performed based on the target token to generate an output result. This method can optimize the token sampling performance while ensuring the quality of the inference result, effectively improve the efficiency of speculative sampling during large model inference optimization, and has a wide range of adaptability.
[0112] Figure 8 It is a block diagram of an electronic device according to an exemplary embodiment.
[0113] Refer to the following Figure 8 hereinafter describes an electronic device 800 according to this embodiment of the present application. Figure 8 The electronic device 800 shown is merely an example and should not bring any limitation to the functions and scope of use of the embodiments of the present application.
[0114] like Figure 8 As shown, the electronic device 800 is in the form of a general computing device. The components of the electronic device 800 may include, but are not limited to: at least one processing unit 810, at least one storage unit 820, a bus 830 connecting different system components (including the storage unit 820 and the processing unit 810), a display unit 840, etc.
[0115] The storage unit stores a program code, which can be executed by the processing unit 810, so that the processing unit 810 performs the steps described in this specification according to various exemplary embodiments of the present application. For example, the processing unit 810 can perform the following steps: Figure 1 , Figure 2 Follow the steps shown in .
[0116] The storage unit 820 may include a readable medium in the form of a volatile storage unit, such as a random access memory unit (RAM) 8201 and / or a cache memory unit 8202 , and may further include a read-only memory unit (ROM) 8203 .
[0117] The storage unit 820 may also include a program / utility 8204 having a set (at least one) of program modules 8205, such program modules 8205 including but not limited to: an operating system, one or more application programs, other program modules and program data, each of which or some combination may include the implementation of a network environment.
[0118] Bus 830 may represent one or more of several types of bus structures, including a memory unit bus or memory unit controller, a peripheral bus, an accelerated graphics port, a processing unit, or a local bus using any of a variety of bus architectures.
[0119] The electronic device 800 may also communicate with one or more external devices 800' (e.g., keyboards, pointing devices, Bluetooth devices, etc.) so that a user can communicate with the device that interacts with the electronic device 800, and / or the electronic device 800 can communicate with any device (e.g., routers, modems, etc.) that communicates with one or more other computing devices. Such communication may be performed through an input / output (I / O) interface 850. In addition, the electronic device 800 may also communicate with one or more networks (e.g., local area networks (LANs), wide area networks (WANs), and / or public networks, such as the Internet) through a network adapter 860. The network adapter 860 may communicate with other modules of the electronic device 800 through a bus 830. It should be understood that, although not shown in the figure, other hardware and / or software modules may be used in conjunction with the electronic device 800, including but not limited to: microcode, device drivers, redundant processing units, external disk drive arrays, RAID systems, tape drives, and data backup storage systems, etc.
[0120] Through the above description of the implementation, it is easy for those skilled in the art to understand that the example implementation described here can be implemented by software, or by combining software with necessary hardware. Fig. 9As shown, the technical solution according to the implementation mode of the present application can be embodied in the form of a software product, which can be stored in a non-volatile storage medium (which can be a CD-ROM, a USB flash drive, a mobile hard disk, etc.) or on a network, and includes a number of instructions to enable a computing device (which can be a personal computer, a server, or a network device, etc.) to execute the above method according to the implementation mode of the present application.
[0121] The software product may be any combination of one or more readable media. The readable medium may be a readable signal medium or a readable storage medium. The readable storage medium may be, for example, but not limited to, a system, device or component of electricity, magnetism, light, electromagnetic, infrared, or semiconductor, or any combination thereof. More specific examples (non-exhaustive list) of readable storage media include: an electrical connection with one or more wires, a portable disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination thereof.
[0122] The computer readable storage medium may include a data signal propagated in a baseband or as part of a carrier wave, wherein a readable program code is carried. This propagated data signal may take a variety of forms, including but not limited to an electromagnetic signal, an optical signal, or any suitable combination of the above. The readable storage medium may also be any readable medium other than a readable storage medium, which may send, propagate, or transmit a program for use by an instruction execution system, an apparatus, or a device or used in combination with it. The program code contained on the readable storage medium may be transmitted with any appropriate medium, including but not limited to wireless, wired, optical cable, RF, etc., or any suitable combination of the above.
[0123] Program code for performing the operations of the present application may be written in any combination of one or more programming languages, including object-oriented programming languages, such as Java, C++, etc., and conventional procedural programming languages, such as "C" or similar programming languages. The program code may be executed entirely on the user computing device, partially on the user device, as a separate software package, partially on the user computing device and partially on a remote computing device, or entirely on a remote computing device or server. In the case of a remote computing device, the remote computing device may be connected to the user computing device through any type of network, including a local area network (LAN) or a wide area network (WAN), or may be connected to an external computing device (e.g., using an Internet service provider to connect through the Internet).
[0124] The above-mentioned computer-readable medium carries one or more programs. When the above-mentioned one or more programs are executed by a device, the computer-readable medium realizes the following functions: using the target model to perform reasoning calculations on the input sequence to generate multiple candidate tokens; sampling the multiple candidate tokens to extract some candidate tokens; inputting the part of the candidate tokens into the medusa model for speculative reasoning to generate multiple speculative tokens in a tree structure; screening the multiple speculative tokens by pruning based on the tree structure to generate a target token; performing subsequent reasoning calculations based on the target token to generate an output result.
[0125] Those skilled in the art will appreciate that the above modules can be distributed in the device according to the description of the embodiment, or can be changed accordingly and only used in one or more devices different from the embodiment. The modules of the above embodiments can be combined into one module, or further divided into multiple sub-modules.
[0126] Through the description of the above embodiments, it is easy for those skilled in the art to understand that the example embodiments described here can be implemented by software, or by software combined with necessary hardware. Therefore, the technical solution according to the embodiment of the present application can be embodied in the form of a software product, which can be stored in a non-volatile storage medium (which can be a CD-ROM, a USB flash drive, a mobile hard disk, etc.) or on a network, and includes several instructions to enable a computing device (which can be a personal computer, a server, a mobile terminal, or a network device, etc.) to execute the method according to the embodiment of the present application.
[0127] The exemplary embodiments of the present application are specifically shown and described above. It should be understood that the present application is not limited to the detailed structures, configurations or implementations described herein; on the contrary, the present application is intended to cover various modifications and equivalent configurations included in the spirit and scope of the appended claims.
Claims
1. A token screening method based on the medusa model, characterized in that: include: Use the target model to perform inference calculations on the input sequence and generate multiple candidate tokens; Sampling the multiple candidate tokens and extracting some candidate tokens; Inputting the candidate tokens into the medusa model for speculative reasoning to generate multiple speculative tokens in a tree structure; Screening the multiple speculative tokens by pruning based on the tree structure to generate a target token; Subsequent reasoning calculations are performed based on the target token to generate output results.
2. The method according to claim 1, characterized in that The target model is used to perform inference calculations on the input sequence to generate multiple candidate tokens, including: inputting the input sequence into the target model; The target model calculates multiple candidate tokens based on reasoning in the input sequence; The kv cache information corresponding to the multiple candidate tokens is cached.
3. The method according to claim 1, characterized in that Sampling the multiple candidate tokens and extracting some candidate tokens includes: Accept k candidate tokens from the multiple candidate tokens, and cache the kv cache information of the p tokens; The kv cache information of the kp tokens that are not cached in the part of candidate tokens is discarded.
4. The method according to claim 1, characterized in that The candidate tokens are input into the medusa model for speculative reasoning to generate multiple speculative tokens in a tree structure, including: Each prediction head of the medusa model generates a set of speculative tokens and organizes them hierarchically into a tree structure.
5. The method according to claim 4, characterized in that Inputting the candidate tokens into the medusa model for speculative reasoning to generate multiple speculative tokens in a tree structure, further comprising: Arrange the multiple speculative tokens in the order of DFS.
6. The method according to claim 1, characterized in that The plurality of speculative tokens are screened by pruning based on the tree structure to generate a target token, including: Screening the multiple speculative tokens using a prediction score pruning method based on a tree structure to generate a target token; and / or Screening the multiple speculative tokens by using leaf node pruning based on the tree structure to generate a target token; and / or The multiple speculative tokens are screened by repeated pruning based on the tree structure to generate a target token.
7. The method according to claim 6, characterized in that The plurality of speculative tokens are screened by using a prediction score pruning method based on a tree structure to generate a target token, including: Get the prediction score of each speculative token, where the prediction score is used to represent the probability that the speculative token is accepted by the target model; The multiple candidate tokens are screened by the prediction scores to generate a target token.
8. The method according to claim 6, characterized in that The plurality of speculative tokens are screened by leaf node pruning based on the tree structure to generate a target token, including: Based on the tree structure, the speculative tokens at the leaf nodes are removed to generate the target token.
9. The method according to claim 6, characterized in that The multiple speculative tokens are screened by repeated pruning based on the tree structure to generate a target token, including: Based on the tree structure, duplicate speculative tokens are removed to generate the target token.
10. A token screening device based on medusa model, characterized in that: include: The target module is used to use the target model to perform inference calculations on the input sequence and generate multiple candidate tokens; A sampling module, used to sample the multiple candidate tokens and extract some candidate tokens; A speculation module, used for inputting the candidate tokens into the medusa model for speculative reasoning, and generating multiple speculative tokens in a tree structure; A pruning module, used for screening the multiple speculative tokens by pruning based on a tree structure to generate a target token; The reasoning module is used to perform subsequent reasoning calculations based on the target token to generate an output result.
Citation Information
Patent Citations
Large language model reasoning acceleration method and related device
CN118333172A
Improvement method of speech synthesis system, electronic equipment and storage medium
CN119169990A
Text generation method and apparatus, electronic device and medium
WO2023207690A1