Efficient draft language model for speculative decoding in autoregressive generative artificial intelligence models

WO2026199095A1PCT designated stage Publication Date: 2026-10-01QUALCOMM INC +10
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
PCT/CN2025/084292
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Filing Date
2025-03-24
Publication Date
2026-10-01

Smart Images

  • Figure CN2025084292_01102026_PF_FP_ABST
    Figure CN2025084292_01102026_PF_FP_ABST
Patent Text Reader

Abstract

Certain aspects of the present disclosure provide techniques and apparatus for machine learning. In an example method, an input is received at a first generative machine learning (ML) model associated with a first vocabulary. Draft token (s) are generated based on the input using the first model. Generating the draft token (s) includes, for each draft token, (i) generating a first token identifier for the draft token based on processing the input using the first model, (ii) mapping the first token identifier to a second token identifier within a second vocabulary for a second generative ML model associated with the first model, where the first vocabulary is a proper subset of the second vocabulary, and (iii) associating the draft token with the second token identifier. A response to the input is generated, using the first model, based on receiving an indication that the draft token(s) have been verified by the second model.
Need to check novelty before this filing date? Find Prior Art

Description

EFFICIENT DRAFT LANGUAGE MODEL FOR SPECULATIVE DECODING IN AUTOREGRESSIVE GENERATIVE ARTIFICIAL INTELLIGENCE MODELSINTRODUCTION

[0001] Aspects of the present disclosure relate to machine learning. More specifically, aspects of the present disclosure relate to speculative decoding in generative artificial intelligence models.

[0002] A wide variety of machine learning model architectures have been trained to perform an assortment of diverse tasks, including computer vision tasks, language tasks, classification and regression tasks, and the like. Recently, research has yielded substantial success in using large language models (LLMs) , large vison models (LVMs) , and / or large multimodal models (LMMs) to process and generate output data. Often, machine learning models (especially LLMs, LVMs, and LMMs) have many parameters (e.g., millions or even billions) , resulting in significant model size, as well as substantial computational expense and time to generate output using the model.

[0003] Some recent efforts to mitigate the computational expense of such generative models include speculative decoding, where a less computationally expensive model (referred to in some aspects as a “draft model” ) can be used to generate a subset of the tokens in the output (rather than using the larger model, often referred to as the “target model, ” for all tokens) . BRIEF SUMMARY

[0004] Certain aspects of the present disclosure provide a processor-implemented method for machine learning. The processor-implemented method includes receiving an input at a first generative machine learning model. The first generative machine learning model is associated with a first vocabulary. The processor-implemented method also includes generating one or more draft tokens based on the input using the first generative machine learning model. Generating the one or more draft tokens includes, for each draft token of the one or more draft tokens: generating a first token identifier for the draft token based on processing the input using the first generative machine learning model; mapping the first token identifier to a second token identifier within a second vocabulary for a second generative machine learning model associated with the first generative machine learning model, wherein the first vocabulary is a proper subset of the second vocabulary; and associating the draft token with the second token identifier. The processor-implemented method also includes generating a response to the input, using the first generative machine learning model, based on receiving an indication that the one or more draft tokens have been verified by the second generative machine learning model.

[0005] Certain aspects of the present disclosure provide a processing system for machine learning. The processing system includes one or more memories comprising processor-executable instructions and one or more processors coupled to the one or more memories. The one or more processors are configured to execute the processor-executable instructions and cause the processing system to: receive an input at a first generative machine learning model, the first generative machine learning model being associated with a first vocabulary; generate one or more draft tokens based on the input using the first generative machine learning model, wherein to generate the one or more draft tokens, the one or more processors are configured to execute the processor-executable instructions and cause the processing system to, for each draft token of the one or more draft tokens: generate a first token identifier for the draft token based on processing the input using the first generative machine learning model; map the first token identifier to a second token identifier within a second vocabulary for a second generative machine learning model associated with the first generative machine learning model, wherein the first vocabulary is a proper subset of the second vocabulary; and associate the draft token with the second token identifier; and generate a response to the input, using the first generative machine learning model, based on receiving an indication that the one or more draft tokens have been verified by the second generative machine learning model.

[0006] Certain aspects of the present disclosure provide a processing system for machine learning. The processing system includes means for receiving an input at a first generative machine learning model, the first generative machine learning model being associated with a first vocabulary. The processing system also includes means for generating one or more draft tokens based on the input using the first generative machine learning model. The means for generating includes, for each draft token of the one or more draft tokens: means for generating a first token identifier for the draft token based on processing the input using the first generative machine learning model; means for mapping the first token identifier to a second token identifier within a second vocabulary for a second generative machine learning model associated with the first generative machine learning model, wherein the first vocabulary is a proper subset of the second vocabulary; and means for associating the draft token with the second token identifier. The processing system further includes means for generating a response to the input, using the first generative machine learning model, based on means for receiving an indication that the one or more draft tokens have been verified by the second generative machine learning model.

[0007] Other aspects provide processing systems configured to perform the aforementioned methods as well as those described herein; non-transitory, computer-readable media comprising instructions that, when executed by one or more processors of a processing system, cause the processing system to perform the aforementioned methods as well as those described herein; a computer program product embodied on a computer-readable storage medium comprising code for performing the aforementioned methods as well as those further described herein; and a processing system comprising means for performing the aforementioned methods as well as those further described herein.

[0008] The following description and the related drawings set forth in detail certain illustrative features of one or more aspects.BRIEF DESCRIPTION OF THE DRAWINGS

[0009] The appended figures depict only certain aspects of this disclosure and are therefore not to be considered limiting of the scope of this disclosure.

[0010] FIG. 1 depicts an example workflow for using an efficient (or at least reduced computationally) draft model to perform speculative decoding in generative artificial intelligence models, according to certain aspects of the present disclosure.

[0011] FIG. 2 is a flow diagram depicting an example method for generating a draft vocabulary for an efficient (or at least reduced computationally) draft model for speculative decoding in generative artificial intelligence models, according to certain aspects of the present disclosure.

[0012] FIG. 3 is a flow diagram depicting an example method for machine learning, according to certain aspects of the present disclosure.

[0013] FIG. 4 depicts an example processing system configured to perform various aspects of the present disclosure.

[0014] FIG. 5 depicts an example processing system configured to perform various aspects of the present disclosure.

[0015] To facilitate understanding, identical reference numerals have been used, where possible, to designate identical elements that are common to the drawings. It is contemplated that elements and features of one aspect may be beneficially incorporated in other aspects without further recitation.DETAILED DESCRIPTION

[0016] Aspects of the present disclosure provide apparatus, methods, processing systems, and computer-readable mediums for providing improved machine learning. Specifically, in certain aspects of the present disclosure, techniques for improved speculative decoding via efficient (or at least reduced computationally) draft models are provided.

[0017] Many model architectures, such as large language models (LLMs) and large vision models (LVMs) , have shown great promise in generating useful output data. However such generative models are often slow at inference time (e.g., taking substantial time to generate output tokens) and are similarly computationally expensive (e.g., consuming substantial memory, as well as processor time and energy, and resulting in substantial heat generation) . As a result, a variety of techniques have been developed to accelerate the token generation rate and / or reduce the computational expense of the token generation. One such technique includes speculative decoding, where a drafting phase is performed to produce “draft” tokens using a relatively less computationally expensive and / or quicker draft model. These draft tokens can then undergo a verification phase to determine whether the draft tokens are “accepted” or “rejected” (e.g., by the target model) to produce the final set of output tokens.

[0018] In some speculative decoding techniques, the draft model may be a pruned version of the target model chosen such that the draft model and target model have similar probability distributions. In other speculative decoding techniques, the draft model may be a smaller version of the target model (e.g., trained on millions of tokens, instead of hundreds of millions or billions of tokens) .

[0019] In certain speculative decoding techniques, the draft model is implemented in part with a language model (LM) head, which is a linear layer that maps the output of one or more hidden layers (e.g., hidden transformer layers) to a vocabulary, predicting the next token (s) in a sequence. In certain cases, the LM head size may result in the drafting phase incurring a significant computational expense (e.g., consuming substantial memory, as well as processor time and energy, and resulting in substantial heat generation) , mitigating the efficiency and performance of speculative decoding techniques.

[0020] Certain aspects of the present disclosure provide techniques for improved speculative decoding based in part on using an efficient (or at least reduced computationally) LM head of a draft model. As described herein, a vocabulary associated with the draft model may be trimmed (or reduced) , such that a resulting size of the trimmed vocabulary for the draft model is less than a size of the vocabulary for the target model. By reducing the size of the draft model’s vocabulary, certain aspects can significantly reduce the size of the draft model’s LM head, thereby reducing the computational expense that may be incurred by the draft model’s LM head during the drafting phase.

[0021] In certain aspects, generating the trimmed vocabulary for the draft model may involve maintaining, within the trimmed vocabulary, a first set of tokens that satisfies one or more conditions and removing, from the trimmed vocabulary, at least a second set of remaining tokens that does not satisfy the one or more conditions. In some aspects, the one or more conditions may be based on each token in the respective set of tokens having a frequency of occurrence in one or more datasets above a threshold. As described further herein, such datasets may include raw datasets (e.g., unprocessed datasets that have not been processed by a machine learning model, such as the draft model and / or target model) , datasets generated by the draft model, datasets generated by the target model, and / or datasets having at least one portion generated by the draft model and at least another portion generated by the target model, as illustrative examples.

[0022] In some aspects, as used herein, a “target model” may generally refer to a generative machine learning model from which output generated data is desired. For example, the target model may correspond to an LLM being used to generate outputs. In some aspects, the target model may incur substantial latency and / or computational expense to generate output tokens (e.g., due to the size of the model) . Similarly, as used herein, a “draft model” may generally refer to a smaller and / or less complex generative machine learning model that can be used as a surrogate for the target model in some cases (e.g., generating similar output, potentially with somewhat reduced accuracy or quality) . Generally, the draft model may incur less latency and / or computational expense to generate output tokens, as compared to the target model, due to the smaller size and / or lower complexity of the draft model relative to the target model.

[0023] Generally, the particular architecture or techniques used to implement the draft model may vary depending on the particular implementation. For example, in some aspects, the draft model may comprise a separate or discrete model from the target model, or may be implemented as a subset of the target model’s layers (e.g., where the draft model corresponds to every other layer of the target model) . In some aspects, the draft model may be implemented using extra parameters in the target model itself to generate draft tokens. In some aspects, the draft model may be implemented by a single transformer layer. Example Speculative Decoding in Generative Artificial Intelligence Models

[0024] Generally, autoregressive token generation (e.g., in LLMs) may take historical tokens as an input in order to generate an output. That is, autoregressive token generation may be represented by the expression: xt ~ p (x|x0, x1, …, xt-1) → xt+1 ~ p (x|x0, x1, …, xt-1, xt) where xt represents a sequence of tokens generated at time t, having a conditional  probability p conditioned on the selection of tokens x0 through xt-1, and xt+1 represents a sequence of tokens generated at time t + 1, having a conditional probability p conditioned on on the selection of tokens x0 through xt. Generally, a single token may be generated each time an autoregressive model is executed, which means that N inferences may be performed to generate a sequence of N tokens. As discussed above, speculative decoding techniques can be used to accelerate token generation by using a draft model, smaller in size than the target model, that speculatively generates tokens, with the target model being used to verify the tokens (speculatively) generated by the draft model.

[0025] In a speculative decoding pipeline, the draft model may speculatively generate n tokens autoregressively, according to the expression: where t corresponds to a point in time and corresponds to the conditional  probability distribution associated with a selected token x at time t conditioned on the selection of tokens x0 through xt-1.

[0026] The target model takes the generated n tokens and processes the n tokens in parallel to generate probability distributions for each of the n tokens, according to the expression: where k corresponds to a token index relative to the generated n tokens.

[0027] The target model can then verify the tokens generated by the draft model by comparing distributions from the draft model and target model to determine whether a token is accepted or rejected. A given token may be accepted when  for some function f and some threshold α (also known as an acceptance rate) . Otherwise, the token may be rejected. The final token may then be generated at the first rejection position or at the last position n based on some function

[0028] Speculative decoding, with an acceptance rate of α, may result in cost reductions relative to using a single autoregressive model to generate tokens iteratively. Inference cost savings, relative to iterative token generation, may be represented by the expression:

[0029] Consider the example, for N = 1000, Ctarget=10, Cdraft=1, n=4, α=3, wherein N corresponds to a number of tokens, CAR corresponds to a computational cost using an acceptance rate of α, Ctarget corresponds to a computational cost of generating a set of tokens using the target model, CSD corresponds to a computational cost of speculative decoding, Cdraft corresponds to a computational cost of generating a set of tokens using the draft model, and n corresponds to a number of tokens generated speculatively generated tokens generated through a single pass through an autoregressive model. In such an example, speculative decoding may result in a 35%reduction in computational expense relative to autoregressive iterative token generation alone.

[0030] However, as discussed, in some speculative decoding techniques, the reduction in computational expense (relative to autoregressive iterative token generation) that may be achieved by speculative decoding may be mitigated (at least somewhat) by the computational expense incurred by the draft model’s LM head during the drafting phase. For example, the draft model’s LM head (e.g., linear layer) may have a size of hidden dimension (ddraft) x vocabulary size (vdraft) , where ddraft is the size (or dimensionality) of the hidden states within the draft model’s layers and vdraft is the total number of unique tokens that the draft model can recognize and use. In many speculative decoding techniques, while ddraft may be lower than the hidden dimension (dtarget) for the target model (where dtarget is the size (or dimensionality) of the hidden states within the target model’s layers) , vdraft may be approximately the same as the vocabulary size (vtarget) , where vtarget is the total number of unique tokens that the target model can recognize and use. Stated differently, in many cases, ddraft < dtarget and vdraft ≈ vtarget. Additionally, because the draft model’s LM head computation may occur each time a draft token is generated, the draft model’s LM head computation may occur draft length times, which can lead to a significant computational expense. For example, in cases with large vocabularies (e.g., certain LLMs may have a vocabulary size of several hundreds of thousands of tokens or more) , the latency of the draft model’s LM head can account for a significant portion of the total latency of the drafting phase. Efficient Workflow for Using an Efficient Draft Model to Perform Speculative Decoding

[0031] FIG. 1 depicts an example workflow 100 for using an efficient (or at least reduced computationally) draft model to perform speculative decoding, according to certain aspects of the present disclosure. In certain aspects, the workflow 100 is performed by a speculative decoding system. That is, the workflow 100 may be performed by any computing system configured to perform speculative decoding, as discussed above and herein. In certain aspects, the workflow 100 may be performed by multiple discrete systems. For example, one computing system (e.g., a portable device) may implement the draft generative model, while another computing system (e.g., a server) may implement the target generative model. In other aspects, the same computing system may implement both the draft generative model and the target generative model.

[0032] In the illustrated example, a prompt 105 (also referred to in some aspects as a “query” ) is accessed by a draft model 110 and a target model 115. As used herein, “accessing” data may generally include receiving, retrieving, obtaining, collecting, generating, or otherwise gaining access to the data. For example, the prompt 105 may be received from a user or entity that uses the machine learning model (s) to generate output. Generally, the particular content and format of the prompt 105 may vary depending on the particular implementation. For example, in some cases, the prompt 105 may comprise textual information (e.g., natural language text) for use as input to an LLM to generate textual output (e.g., a natural language text response) .

[0033] In the depicted workflow 100, the target model 115 (also referred to in some aspects as the “primary model, ” the “primary machine learning model, ” and / or the “primary generative machine learning model” ) generally corresponds to a generative machine learning model used to generate output data based on input prompts. For example, as discussed above, the target model 115 may correspond to an LLM, LVM, LMM, or the like. In some aspects, the target model 115 may be relatively computationally expensive (e.g., incurring substantial latency and / or computational resource usage) during runtime. Further, in the illustrated example, the draft model 110 (also referred to in some aspects as the “secondary model, ” the “secondary machine learning model, ” and / or the “secondary generative machine learning model” ) generally corresponds to a generative machine learning model that can also be used to generate output data based on input prompts. In some aspects, as discussed above, the draft model 110 may be relatively less computationally expensive than the target model 115 (e.g., incurring less latency and / or computational resource usage, as compared to the target model 115) during runtime.

[0034] In some aspects, the draft model 110 may be somewhat less accurate or reliable than the target model 115. However, not all tokens in a generated output are equally “difficult” to generate. Therefore, in some aspects, the draft model 110 may be reliably used to generate some tokens, even if the target model 115 is relied upon for other tokens. Knowing the “difficulty” of a given token (before drafting the token) is, generally speaking, nearly impossible. In some aspects, therefore, the depicted workflow 100 can be used to implement speculative decoding. Generally, as discussed above, speculative decoding involves using the target model 115 to generate some set of tokens in the output, while allowing the draft model 110 to generate another set of one or more tokens. These tokens generated by the draft model 110 (referred to as “draft tokens” in some aspects) can then be evaluated (e.g., using the target model 115) to verify the draft tokens. That is, the target model 115 may accept one or more of the draft tokens, and the target model 115 may reject one or more of the draft tokens (e.g., because the token is too dissimilar from what the target model 115 would have generated) .

[0035] In some aspects, because the draft model 110 is less computationally complex than the target model 115, the set of draft tokens can be generated more rapidly and with less computational expense, as compared to generating the set using the target model 115 alone. Further, because the target model 115 may evaluate multiple draft tokens in parallel (e.g., evaluating the entire sequence of draft tokens at once) , the target model 115 can be used to quickly verify the draft tokens. By combining these features, the workflow 100 can substantially accelerate the output generation process.

[0036] In the illustrated example, the target model 115 may process the prompt 105 to generate a set of one or more tokens 120. As used herein, a “token” corresponds to a unit of output of the model (s) . Generally, the particular format and content of a token may vary depending on the particular implementation. For example, in some aspects, the tokens 120 comprise words, phrases, alphanumeric characters and / or symbols, and the like. Generally, the particular number of tokens 120 to be generated using the target model 115 may vary depending on the particular implementation. For example, in some aspects, the computing system may generate a single token 120 using the target model 115 in between each drafting phase (e.g., before using the draft model 110) . As another example, the computing system may use the target model 115 to generate two or more tokens 120 between drafting phases.

[0037] In certain aspects, the tokens 120 may be generated in part based on a target vocabulary 160 associated with the target model 115. For example, each of the tokens 120 may be associated with a respective token identifier (ID) within the target vocabulary 160. While target vocabulary 160 is depicted within the target model 115 for conceptual clarity, in certain aspects, the target vocabulary 160 may not be included as part of the target model 115 and may be implemented elsewhere, e.g., on the same or different computing device as the target model 115.

[0038] In the illustrated workflow 100, these initial tokens 120 are provided, along with the prompt 105, to the draft model 110, to generate a new set of one or more tokens 125. In some aspects, as discussed above, the tokens 125 may be referred to as “draft” tokens to indicate that the tokens 125 were generated by the draft model 110. Generally, the particular number of tokens 125 to be generated using the draft model 110 may vary depending on the particular implementation.

[0039] In certain aspects, the draft model 110 may be implemented with one or more embedding layers 170, one or more hidden layers 175, and an LM head 180 (e.g., linear layer) . In certain aspects, the tokens 125 may be generated in part based on a draft vocabulary 150 associated with the draft model 110. For example, the LM head 180 may map the output of the hidden layer (s) 175 to token (s) 125 within the draft vocabulary 150. Each of the tokens 125 may be associated with a respective token ID within the draft vocabulary 150. In certain aspects, the draft vocabulary 150 may include a smaller number of tokens / token IDs than the target vocabulary 160. In such aspects, the draft vocabulary 150 may be referred to herein as a trimmed vocabulary.

[0040] In certain aspects described in greater detail herein, the draft vocabulary 150 may be generated at least in part from the target vocabulary 160. For example, the draft vocabulary 150 may be a proper subset of the target vocabulary 160. In some aspects, the draft vocabulary 150 may be generated in part by identifying a set of tokens within the target vocabulary 160 that satisfies one or more conditions, identifying the corresponding set of tokens within the draft vocabulary 150, maintaining the corresponding set of tokens within the draft vocabulary 150, and removing the remaining set of tokens within the draft vocabulary 150 from the draft vocabulary 150, as described further herein. While draft vocabulary 150 is depicted within the draft model 110 for conceptual clarity, in certain aspects, the draft vocabulary 150 may not be included as part of the draft model 110 and may be implemented elsewhere, e.g., on the same or different computing device as the draft model 110.

[0041] In the illustrated workflow 100, the tokens 125 are accessed by a mapping component 165 to determine token IDs within the target vocabulary 160 corresponding to the token IDs for the tokens 125 within the draft vocabulary 150. For example, because the draft vocabulary 150 may be a proper subset of the target vocabulary 160, the draft vocabulary 150 may include a different set of token IDs than the target vocabulary 160 for tokens 125 within the draft vocabulary 150. To address this, the mapping component 165 includes a dictionary mapping 145, which includes mappings of token IDs in the draft vocabulary 150 to token IDs in the target vocabulary 160. The mapping component 165 may generally use the dictionary mapping 145 to map token IDs within the draft vocabulary 150 for the tokens 125 to token IDs within the target vocabulary 160 for the tokens 125. The mapping component 165 may output tokens 155 that correspond to the tokens 125 and that are associated with the mapped token IDs within the target vocabulary 160.

[0042] By way of example, the target vocabulary 160 may have a respective token ID indicating the actual output token for each index in the output tensor (e.g., token ID 0 =“cat, ” token ID 1 = “dog, ” token ID 2 = “whale, ” and so on) . When the draft vocabulary 150 is generated, e.g., by removing one or more tokens from the draft vocabulary 150 that are also within the target vocabulary 160, the token IDs for the remaining tokens within the draft vocabulary 150 may be updated. For instance, assuming “dog, ” is removed, the token IDs for the draft vocabulary 150 may be updated to token ID 0 = “cat, ” token ID 1 = “whale, ” and so on. In this instance, the dictionary mapping 145 may include mapping information (e.g., mapping function / transform) that indicates that token ID 0 in draft vocabulary 150 corresponds to token 0 in target vocabulary 160, token ID 1 in draft vocabulary 150 corresponds to token ID 2 in target vocabulary 160, and so on. Thus, continuing with this example, assuming the draft model 110 outputs a token 125 with a token ID 1 (within the draft vocabulary 150) , the mapping component 165 may use the dictionary mapping 145 to output a token 155 (corresponding to the token 125) that has a token ID 2 (within the target vocabulary 160) .

[0043] In the illustrated workflow 100, the token (s) 155 are accessed by a stopping component 130 to determine whether to stop the drafting phase. As used herein, stopping the drafting phase (also referred to in some aspects as “exiting” from the drafting phase and / or from the draft model) generally corresponds to determining to use the target model 115 for the next one or more token (s) of the output, rather than using the draft model 110 for the next token (s) . The stopping component 130 may use a variety of criteria to determine when to exist the drafting phase. For example, the stopping component 130 may determine whether the current sequence of draft tokens (generated by the draft model 110 during the current drafting phase) meets a defined maximum draft length hyperparameter (e.g., a defined maximum number of draft tokens that should be generated in a given drafting phase) . If the current sequence of draft tokens meets the defined maximum draft length hyperparameter, then the stopping component 130 may stop the drafting phase; otherwise, the stopping component 130 may continue the drafting phase, as illustrated by the dotted arrow 135.

[0044] In the illustrated workflow 100, if the stopping component 130 determines to continue the drafting phase, as illustrated by the dotted arrow 135, then the draft model 110 may be prompted to generate a next token 125 (e.g., using the prompt 105, the tokens 120, and / or the previously generated token 155 as input) . This process can then be repeated until the one or more stopping criteria are satisfied.

[0045] As illustrated, if the stopping component 130 determines to stop the drafting phase, then the set of generated draft tokens 140 (corresponding to tokens 155) is provided for verification to the target model 115. That is, for each iteration of drafting, the draft model 110 may generate and add a new token to the sequence of draft tokens 140 (e.g., adding the highest-scored token at each iteration) . Each draft token 140 may include a respective token ID within the target vocabulary 160. When the drafting phase terminates, this sequence of draft tokens 140 may be provided to the target model 115.

[0046] In some aspects, as discussed above, the target model 115 may be used to verify the draft tokens 140. For example, the prompt 105, any tokens 120 that have already been drafted by the target model 115, and / or the set of draft tokens 140 may be provided as input, allowing the target model 115 to verify (e.g., accept) or reject each token of the sequence of draft tokens 140. For example, the target model 115 may generate a score for each token in the sequence of draft tokens 140, accepting tokens having a score above a threshold (e.g., indicating a higher probability that the target model 115 would have generated the same token) and rejecting tokens having a score below the threshold (e.g., indicating a low probability that the target model 115 would have generated the same token) .

[0047] In some aspects, this verification process also results in generation of a new next token 120 to be added to the sequence of (accepted) tokens. As illustrated, this updated sequence of tokens can then be provided to the draft model 110 to begin the next drafting phase, as discussed above. Although the illustrated example suggests that the target model 115 is used to generate the first token (s) in the output sequence (followed by alternating between the draft model 110 and the target model 115) , in some aspects, the draft model 110 may be used to generate the first token (s) (followed by alternating between the draft model 110 and the target model 115) .

[0048] Advantageously, using the efficient (or at least reduced computationally) draft model described herein may improve the computational efficiency of the speculative decoding system, enabling improved model output generation with reduced computational expense and / or latency. For example, maintaining frequently occurring tokens for the draft model’s LM head may allow for maintaining the accuracy of the final LLM while reducing the draft model’s read-only memory (ROM) size and runtime random access memory (RAM) usage. Additionally, the efficient (or at least reduced computationally) draft model may allow for improved (e.g., higher) memory-bound speed-up on downstream tasks as well as reduced power usage when running certain inference algorithms. Example Method for Generating a Draft Vocabulary for an Efficient Draft Model

[0049] FIG. 2 is a flow diagram depicting an example method 200 for generating a draft vocabulary for an efficient (or at least reduced computationally) draft model for speculative decoding, according to certain aspects of the present disclosure. In some aspects, the method 200 is performed by a computing system, such as the computing system discussed with reference to FIG. 4.

[0050] At block 210, the computing system accesses one or more datasets. As used herein, “accessing” data may generally include receiving, retrieving, obtaining, collecting, generating, or otherwise gaining access to the data.

[0051] In certain aspects, the one or more datasets may include one or more raw datasets. Such raw datasets may include unprocessed data that has not been processed by a machine learning model, such as the draft model (e.g., draft model 110) and / or the target model (e.g., target model 115) , as illustrative examples.

[0052] In certain aspects, the one or more datasets may include one or more target model-based datasets. Such target model-based datasets may include datasets generated by the target model.

[0053] In certain aspects, the one or more datasets may include one or more draft model-based datasets. Such draft model-based datasets may include datasets generated by the draft model.

[0054] In certain aspects, the one or more datasets may include one or more combined draft and target model-based datasets. Such combined draft and target model-based datasets may include datasets having portion (s) generated by the draft model and portion (s) generated by the target model.

[0055] In certain aspects, the one or more datasets may include (i) one or more target model-based datasets, (ii) one or more draft model-based datasets, (iii) one or more combined draft and target model-based datasets, or (iv) any combination thereof.

[0056] At block 220, the computing system determines a set of tokens within the dataset (s) that satisfies one or more conditions. In certain aspects, the one or more conditions may include each token of the set of tokens having a frequency of occurrence in the dataset (s) above a threshold. In such aspects, a maximum number of the set of tokens having the frequency of occurrence in the dataset (s) above the threshold may be based at least in part on one or more hardware parameters (e.g., processor speed, memory capacity, storage capacity, among others) of a computing system (e.g., a speculative decoding system) , such as the computing system discussed above with reference to FIG. 1 and the computing system discussed with reference to FIG. 5, as illustrative examples. Thus, the maximum number of the set of tokens having the frequency of occurrence in the dataset (s) above the threshold and / or the threshold may be hyperparameters of the draft model or generation process.

[0057] At block 230, the computing system generates a modified vocabulary that consists of the set of tokens. In some cases, the modified vocabulary may be generated from the target vocabulary (e.g., target vocabulary 160) . For example, the target vocabulary may be modified (or trimmed) to remove sets of tokens that do not satisfy the one or more conditions, such that the modified vocabulary consists of the set of tokens that does satisfy the one or more conditions.

[0058] At block 240, the computing system uses the modified vocabulary as the draft vocabulary (e.g., draft vocabulary 150) .

[0059] At block 250, the computing system generates a dictionary mapping (e.g., dictionary mapping 145) of token IDs within the draft vocabulary to token IDs within the target vocabulary (e.g., target vocabulary 160) .

[0060] At block 260, the computing system stores the dictionary mapping (e.g., in one or more memories and / or any other suitable storage devices) .

[0061] The various operations of methods described above may be performed by any suitable means capable of performing the corresponding functions. The means may include various hardware and / or software component (s) and / or module (s) , including, but not limited to a circuit, an application specific integrated circuit (ASIC) , or a processor. Example Method for Machine Learning

[0062] FIG. 3 is a flow diagram depicting an example method 300 for machine learning, according to some aspects of the present disclosure. In some aspects, the method 300 is performed by a computing system (e.g., a speculative decoding system) , such as the computing system discussed above with reference to FIG. 1 and the computing system discussed with reference to FIG. 5, as illustrative examples.

[0063] At block 310, the computing system may receive (or otherwise access) an input (e.g., prompt 105, tokens 120, and / or draft tokens 140) at a first generative machine learning model (e.g., draft model 110) . The first generative machine learning model may be associated with a first vocabulary (e.g., draft vocabulary 150) .

[0064] At block 320, the computing system may generate one or more draft tokens (e.g., draft tokens 140) based on the input using the first generative machine learning model. Block 320 may involve performing blocks 330, 340, and 350 for each draft token of the one or more draft tokens.

[0065] At block 330, the computing system may generate a first token identifier for the draft token (e.g., token 125) based on processing the input using the first generative machine learning model.

[0066] At block 340, the computing system may map the first token identifier to a second token identifier (e.g., token 155) within a second vocabulary (e.g., target vocabulary 160) for a second generative machine learning model (e.g., target model 115) associated with the first generative machine learning model. The first vocabulary may be a proper subset of the second vocabulary.

[0067] At block 350, the computing system may associate the draft token with the second token identifier.

[0068] At block 360, the computing system may generate a response to the input, using the first generative machine learning model, based on receiving an indication that the one or more draft tokens have been verified by the second generative machine learning model.

[0069] In certain aspects, method 300 may further involve generating the first vocabulary based at least in part on the second vocabulary. The first vocabulary may consist of a subset of tokens within the second vocabulary that satisfies one or more conditions. In such aspects, the one or more conditions may include each token of the subset of tokens being one of a set of tokens having a frequency of occurrence in one or more datasets above a threshold. In some aspects, a maximum number of the set of tokens having the frequency of occurrence in the one or more datasets above the threshold may be based at least in part on one or more hardware parameters of a computing device.

[0070] In some aspects, the one or more datasets may include one or more unprocessed datasets. Such unprocessed datasets may not have been processed by the first generative machine learning model and / or the second generative machine learning model.

[0071] In some aspects, the one or more datasets may include one or more datasets generated by the first generative machine learning model.

[0072] In some aspects, the one or more datasets may include one or more datasets generated by the second generative machine learning model.

[0073] In some aspects, the one or more datasets may include a dataset having at least a first portion generated by the first generative machine learning model and a second portion generated by the second generative machine learning model.

[0074] In certain aspects, mapping the first token identifier to the second token identifier may include mapping the first token identifier to the second token identifier using a dictionary mapping (e.g., dictionary mapping 145) of token identifiers in the first generative machine learning model to token identifiers in the second generative machine learning model.

[0075] In certain aspects, the input may include an input prompt (e.g., prompt 105) or an output token (e.g., token 120) generated by the second generative machine learning model.

[0076] In certain aspects, the first generative machine learning model may be a draft model and the second generative machine learning model may be a target model. In such aspects, the draft model may be smaller than the target model.

[0077] In certain aspects, the first generative machine learning model may be implemented by a single transformer layer.

[0078] The various operations of methods described above may be performed by any suitable means capable of performing the corresponding functions. The means may include various hardware and / or software component (s) and / or module (s) , including, but not limited to a circuit, an application specific integrated circuit (ASIC) , or a processor. Example Processing System for Generating Draft Vocabulary for an Efficient Draft  Model for Speculative Decoding in Generative Artificial Intelligence Models

[0079] FIG. 4 depicts an example processing system 400 for generating a draft vocabulary for an efficient (or at least reduced computationally) draft model (e.g., draft model 110) for speculative decoding in generative artificial intelligence models, such as described herein, for example, with respect to FIG. 2.

[0080] The processing system 400 includes a central processing unit (CPU) 402, which in some examples may be a multi-core CPU. Instructions executed at the CPU 402 may be loaded, for example, from a program memory associated with the CPU 402 or may be loaded from a memory partition (e.g., of a memory 424) .

[0081] The processing system 400 also includes additional processing components tailored to specific functions, such as a graphics processing unit (GPU) 404, a digital signal processor (DSP) 406, a neural processing unit (NPU) 408, and a connectivity component 412.

[0082] An NPU, such as the NPU 408, is generally a specialized circuit configured for implementing control and arithmetic logic for executing machine learning algorithms, such as algorithms for processing artificial neural networks (ANNs) , deep neural networks (DNNs) , random forests (RFs) , and the like. An NPU may sometimes alternatively be referred to as a neural signal processor (NSP) , tensor processing unit (TPU) , neural network processor (NNP) , intelligence processing unit (IPU) , vision processing unit (VPU) , or graph processing unit.

[0083] NPUs, such as the NPU 408, are configured to accelerate the performance of common machine learning tasks, such as image classification, machine translation, object detection, and various other predictive models. In some examples, a plurality of NPUs may be instantiated on a single chip, such as a system on a chip (SoC) , while in other examples, such NPUs may be part of a dedicated neural-network accelerator.

[0084] NPUs may be optimized for training or inference, or in some cases configured to balance performance between both. For NPUs that are capable of performing both training and inference, the two tasks may still generally be performed independently.

[0085] NPUs designed to accelerate training are generally configured to accelerate the optimization of new models, which is a highly compute-intensive operation that involves inputting an existing dataset (often labeled or tagged) , iterating over the dataset, and then adjusting model parameters, such as weights and biases, in order to improve model performance. Generally, optimizing based on a wrong prediction involves propagating back through the layers of the model and determining gradients to reduce the prediction error.

[0086] NPUs designed to accelerate inference are generally configured to operate on complete models. Such NPUs may thus be configured to input a new piece of data and rapidly process this new piece through an already trained model to generate a model output (e.g., an inference) .

[0087] In some implementations, the NPU 408 is a part of one or more of the CPU 402, the GPU 404, and / or the DSP 406. These may be located on a user equipment (UE) in a wireless communication system or another computing device.

[0088] In some examples, the connectivity component 412 may include subcomponents, for example, for third generation (3G) connectivity, fourth generation (4G) connectivity (e.g., Long-Term Evolution (LTE) ) , fifth generation (5G) connectivity (e.g., New Radio (NR) ) , Wi-Fi connectivity, Bluetooth connectivity, and other wireless data transmission standards. The connectivity component 412 may be further coupled to one or more antennas 414.

[0089] The processing system 400 may also include one or more sensor processing units 416 associated with any manner of sensor, one or more image signal processors (ISPs) 418 associated with any manner of image sensor, and / or a navigation processor 420, which may include satellite-based positioning system components (e.g., global positioning system (GPS) or global navigation satellite system (GLONASS) ) as well as inertial positioning system components.

[0090] The processing system 400 may also include one or more input and / or output devices 422, such as screens, touch-sensitive surfaces (including touch-sensitive displays) , physical buttons, speakers, microphones, and the like.

[0091] In some examples, one or more of the processors of the processing system 400 may be based on an advanced reduced instruction set computing (RISC) machine (ARM) or fifth generation of RISC (RISC-V) instruction set.

[0092] The processing system 400 also includes the memory 424, which is representative of one or more static and / or dynamic memories, such as a dynamic random access memory, a flash-based static memory, and the like. In this example, the memory 424 includes computer-executable components, which may be executed by one or more of the aforementioned processors of the processing system 400.

[0093] In particular, in this example, the memory 424 includes a dataset accessing component 424A, a token selecting component 424B, a vocabulary generating component 424C, and a dictionary mapping generating component 424D. The depicted components, and others not depicted, may be configured to perform various aspects of the methods described herein.

[0094] Generally, the processing system 400 and / or components thereof may be configured to perform the methods described herein. Example Processing Systems for Speculative Decoding in Generative Artificial  Intelligence Models

[0095] FIG. 5 depicts an example processing system 500 for generating a response to a query input into a generative artificial intelligence model based on speculative decoding, such as described herein for example with respect to FIG. 3. In some aspects, the processing system 500 may be included in a mobile device, such as a smartphone, laptop, tablet, or wearable device.

[0096] The processing system 500 includes a central processing unit (CPU) 502, which in some examples may be a multi-core CPU. Instructions executed at the CPU 502 may be loaded, for example, from a program memory associated with the CPU 502 or may be loaded from a memory partition (e.g., of a memory 524) .

[0097] The processing system 500 also includes additional processing components tailored to specific functions, such as a graphics processing unit (GPU) 504, a digital signal processor (DSP) 506, a neural processing unit (NPU) 508, and a connectivity component 512.

[0098] An NPU, such as the NPU 508, is generally a specialized circuit configured for implementing control and arithmetic logic for executing machine learning algorithms, such as algorithms for processing artificial neural networks (ANNs) , deep neural networks (DNNs) , random forests (RFs) , and the like. An NPU may sometimes alternatively be referred to as a neural signal processor (NSP) , tensor processing unit (TPU) , neural network processor (NNP) , intelligence processing unit (IPU) , vision processing unit (VPU) , or graph processing unit.

[0099] NPUs, such as the NPU 508, are configured to accelerate the performance of common machine learning tasks, such as image classification, machine translation, object detection, and various other predictive models. In some examples, a plurality of NPUs may be instantiated on a single chip, such as a system on a chip (SoC) , while in other examples such NPUs may be part of a dedicated neural-network accelerator.

[0100] NPUs may be optimized for training or inference, or in some cases configured to balance performance between both. For NPUs that are capable of performing both training and inference, the two tasks may still generally be performed independently.

[0101] NPUs designed to accelerate training are generally configured to accelerate the optimization of new models, which is a highly compute-intensive operation that involves inputting an existing dataset (often labeled or tagged) , iterating over the dataset, and then adjusting model parameters, such as weights and biases, in order to improve model performance. Generally, optimizing based on a wrong prediction involves propagating back through the layers of the model and determining gradients to reduce the prediction error.

[0102] NPUs designed to accelerate inference are generally configured to operate on complete models. Such NPUs may thus be configured to input a new piece of data and rapidly process this new piece through an already trained model to generate a model output (e.g., an inference) .

[0103] In some implementations, the NPU 508 is a part of one or more of the CPU 502, the GPU 504, and / or the DSP 506. These may be located on a user equipment (UE) in a wireless communication system or another computing device.

[0104] In some examples, the connectivity component 512 may include subcomponents, for example, for third generation (3G) connectivity, fourth generation (4G) connectivity (e.g., LTE) , fifth generation (5G) connectivity (e.g., NR) , Wi-Fi connectivity, Bluetooth connectivity, and other wireless data transmission standards. The connectivity component 512 may be further coupled to one or more antennas 514.

[0105] The processing system 500 may also include one or more sensor processing units 516 associated with any manner of sensor, one or more image signal processors (ISPs) 518 associated with any manner of image sensor, and / or a navigation processor 520, which may include satellite-based positioning system components (e.g., GPS or GLONASS) as well as inertial positioning system components.

[0106] The processing system 500 may also include one or more input and / or output devices 522, such as screens, touch-sensitive surfaces (including touch-sensitive displays) , physical buttons, speakers, microphones, and the like.

[0107] In some examples, one or more of the processors of the processing system 500 may be based on an ARM or RISC-V instruction set.

[0108] The processing system 500 also includes the memory 524, which is representative of one or more static and / or dynamic memories, such as a dynamic random access memory, a flash-based static memory, and the like. In this example, the memory 524 includes computer-executable components, which may be executed by one or more of the aforementioned processors of the processing system 500.

[0109] In particular, in this example, the memory 524 includes a token set generating component 524A, a token selecting component 524B, an output generating component 524C, and generative artificial intelligence models 524D. The depicted components, and others not depicted, may be configured to perform various aspects of the methods described herein.

[0110] Generally, the processing system 500 and / or components thereof may be configured to perform the methods described herein. Example Clauses

[0111] Implementation details of various aspects of the present disclosure are described in the following numbered clauses.

[0112] Clause 1: A processor-implemented method for machine learning, the processor-implemented method comprising: receiving an input at a first generative machine learning model, the first generative machine learning model being associated with a first vocabulary; generating one or more draft tokens based on the input using the first generative machine learning model, comprising, for each draft token of the one or more draft tokens: generating a first token identifier for the draft token based on processing the input using the first generative machine learning model; mapping the first token identifier to a second token identifier within a second vocabulary for a second generative machine learning model associated with the first generative machine learning model, wherein the first vocabulary is a proper subset of the second vocabulary; and associating the draft token with the second token identifier; and generating a response to the input, using the first generative machine learning model, based on receiving an indication that the one or more draft tokens have been verified by the second generative machine learning model.

[0113] Clause 2: The processor-implemented method of Clause 1, further comprising generating the first vocabulary based at least in part on the second vocabulary, wherein the first vocabulary consists of a subset of tokens within the second vocabulary that satisfies one or more conditions.

[0114] Clause 3: The processor-implemented method of Clause 2, wherein the one or more conditions comprise each token of the subset of tokens being one of a set of tokens having a frequency of occurrence in one or more datasets above a threshold.

[0115] Clause 4: The processor-implemented method of Clause 3, wherein the one or more datasets comprise one or more unprocessed datasets, the one or more unprocessed datasets not being processed by the first generative machine learning model or the second generative machine learning model.

[0116] Clause 5: The processor-implemented method in accordance with any of Clauses 3-4, wherein the one or more datasets comprise one or more datasets generated by the first generative machine learning model.

[0117] Clause 6: The processor-implemented method in accordance with any of Clauses 3-5, wherein the one or more datasets comprise one or more datasets generated by the second generative machine learning model.

[0118] Clause 7: The processor-implemented method in accordance with any of Clauses 3-6, wherein the one or more datasets comprise a dataset having at least a first portion generated by the first generative machine learning model and a second portion generated by the second generative machine learning model.

[0119] Clause 8: The processor-implemented method in accordance with any of Clauses 3-7, wherein a maximum number of the set of tokens having the frequency of occurrence in the one or more datasets above the threshold is based at least in part on one or more hardware parameters of a computing device.

[0120] Clause 9: The processor-implemented method in accordance with any of Clauses 1-8, wherein mapping the first token identifier to the second token identifier comprises mapping the first token identifier to the second token identifier using a dictionary mapping of token identifiers in the first generative machine learning model to token identifiers in the second generative machine learning model.

[0121] Clause 10: The processor-implemented method in accordance with any of Clauses 1-9, wherein the input comprises an input prompt or an output token generated by the second generative machine learning model.

[0122] Clause 11: The processor-implemented method in accordance with any of Clauses 1-10, wherein: the first generative machine learning model is a draft model; and the second generative machine learning model is a target model, the draft model being smaller than the target model.

[0123] Clause 12: The processor-implemented method in accordance with any of Clauses 1-11, wherein the first generative machine learning model is implemented by a single transformer layer.

[0124] Clause 13: A processing system comprising: a memory comprising processor-executable instructions; and one or more processors coupled to the one or more memories and configured to execute the processor-executable instructions and cause the processing system to perform a method in accordance with any of Clauses 1-12.

[0125] Clause 14: A mobile device comprising the processing system of Clause 13.

[0126] Clause 15: A processing system comprising means for performing a method in accordance with any of Clauses 1-12.

[0127] Clause 16: A non-transitory computer-readable medium comprising computer-executable instructions that, when executed by one or more processors of a processing system, cause the processing system to perform a method in accordance with any of Clauses 1-12.

[0128] Clause 17: A computer program product embodied on a computer-readable storage medium comprising code for performing a method in accordance with any of Clauses 1-12. Additional Considerations

[0129] The preceding description is provided to enable any person skilled in the art to practice the various aspects described herein. The examples discussed herein are not limiting of the scope, applicability, or aspects set forth in the claims. Various modifications to these aspects will be readily apparent to those skilled in the art, and the generic principles defined herein may be applied to other aspects. For example, changes may be made in the function and arrangement of elements discussed without departing from the scope of the disclosure. Various examples may omit, substitute, or add various procedures or components as appropriate. For instance, the methods described may be performed in an order different from that described, and various steps may be added, omitted, or combined. Also, features described with respect to some examples may be combined in some other examples. For example, an apparatus may be implemented or a method may be practiced using any number of the aspects set forth herein. In addition, the scope of the disclosure is intended to cover such an apparatus or method that is practiced using other structure, functionality, or structure and functionality in addition to, or other than, the various aspects of the disclosure set forth herein. It should be understood that any aspect of the disclosure disclosed herein may be embodied by one or more elements of a claim.

[0130] As used herein, the word “exemplary” means “serving as an example, instance, or illustration. ” Any aspect described herein as “exemplary” is not necessarily to be construed as preferred or advantageous over other aspects.

[0131] As used herein, a phrase referring to “at least one of” a list of items refers to any combination of those items, including single members. As an example, “at least one of: a, b, or c” is intended to cover a, b, c, a-b, a-c, b-c, and a-b-c, as well as any combination with multiples of the same element (e.g., a-a, a-a-a, a-a-b, a-a-c, a-b-b, a-c-c, b-b, b-b-b, b-b-c, c-c, and c-c-c or any other ordering of a, b, and c) .

[0132] As used herein, the term “determining” encompasses a wide variety of actions. For example, “determining” may include calculating, computing, processing, deriving, investigating, looking up (e.g., looking up in a table, a database or another data structure) , ascertaining, and the like. Also, “determining” may include receiving (e.g., receiving information) , accessing (e.g., accessing data in a memory) , and the like. Also, “determining” may include resolving, selecting, choosing, establishing, and the like.

[0133] The methods disclosed herein comprise one or more steps or actions for achieving the methods. The method steps and / or actions may be interchanged with one another without departing from the scope of the claims. In other words, unless a specific order of steps or actions is specified, the order and / or use of specific steps and / or actions may be modified without departing from the scope of the claims. Further, the various operations of methods described above may be performed by any suitable means capable of performing the corresponding functions. The means may include various hardware and / or software component (s) and / or module (s) , including, but not limited to a circuit, an application specific integrated circuit (ASIC) , or processor. Generally, where there are operations illustrated in figures, those operations may have corresponding counterpart means-plus-function components with similar numbering.

[0134] The following claims are not intended to be limited to the aspects shown herein, but are to be accorded the full scope consistent with the language of the claims. Within a claim, reference to an element in the singular is not intended to mean “one and only one” unless specifically so stated, but rather “one or more. ” Unless specifically stated otherwise, the term “some” refers to one or more. No claim element is to be construed under the provisions of 35 U.S.C. § 112 (f) unless the element is expressly recited using the phrase “means for” or, in the case of a method claim, the element is recited using the phrase “step for. ” All structural and functional equivalents to the elements of the various aspects described throughout this disclosure that are known or later come to be known to those of ordinary skill in the art are expressly incorporated herein by reference and are intended to be encompassed by the claims. Moreover, nothing disclosed herein is intended to be dedicated to the public regardless of whether such disclosure is explicitly recited in the claims.

Claims

1.A processing system for machine learning, comprising:one or more memories comprising processor-executable instructions; andone or more processors coupled to the one or more memories and configured to execute the processor-executable instructions and cause the processing system to:receive an input at a first generative machine learning model, the first generative machine learning model being associated with a first vocabulary;generate one or more draft tokens based on the input using the first generative machine learning model, wherein to generate the one or more draft tokens, the one or more processors are configured to execute the processor-executable instructions and cause the processing system to, for each draft token of the one or more draft tokens:generate a first token identifier for the draft token based on processing the input using the first generative machine learning model;map the first token identifier to a second token identifier within a second vocabulary for a second generative machine learning model associated with the first generative machine learning model, wherein the first vocabulary is a proper subset of the second vocabulary; andassociate the draft token with the second token identifier; andgenerate a response to the input, using the first generative machine learning model, based on receiving an indication that the one or more draft tokens have been verified by the second generative machine learning model.2.The processing system of claim 1, wherein:the one or more processors are further configured to execute the processor-executable instructions and cause the processing system to generate the first vocabulary based at least in part on the second vocabulary; andthe first vocabulary consists of a subset of tokens within the second vocabulary that satisfies one or more conditions.3.The processing system of claim 2, wherein the one or more conditions comprise each token of the subset of tokens being one of a set of tokens having a frequency of occurrence in one or more datasets above a threshold.4.The processing system of claim 3, wherein the one or more datasets comprise one or more unprocessed datasets, the one or more unprocessed datasets not being processed by the first generative machine learning model or the second generative machine learning model.5.The processing system of claim 3, wherein the one or more datasets comprise one or more datasets generated by the first generative machine learning model.6.The processing system of claim 3, wherein the one or more datasets comprise one or more datasets generated by the second generative machine learning model.7.The processing system of claim 3, wherein the one or more datasets comprise a dataset having at least a first portion generated by the first generative machine learning model and a second portion generated by the second generative machine learning model.8.The processing system of claim 3, wherein a maximum number of the set of tokens having the frequency of occurrence in the one or more datasets above the threshold is based at least in part on one or more hardware parameters of a computing device.9.The processing system of claim 1, wherein to map the first token identifier to the second token identifier, the one or more processors are configured to execute the processor-executable instructions and cause the processing system to map the first token identifier to the second token identifier using a dictionary mapping of token identifiers in the first generative machine learning model to token identifiers in the second generative machine learning model.10.The processing system of claim 1, wherein the input comprises an input prompt or an output token generated by the second generative machine learning model.11.The processing system of claim 1, wherein:the first generative machine learning model is a draft model; andthe second generative machine learning model is a target model, the draft model being smaller than the target model.12.The processing system of claim 1, wherein the first generative machine learning model is implemented by a single transformer layer.13.A mobile device comprising the processing system of claim 1.14.A processor-implemented method for machine learning, the processor-implemented method comprising:receiving an input at a first generative machine learning model, the first generative machine learning model being associated with a first vocabulary;generating one or more draft tokens based on the input using the first generative machine learning model, comprising, for each draft token of the one or more draft tokens:generating a first token identifier for the draft token based on processing the input using the first generative machine learning model;mapping the first token identifier to a second token identifier within a second vocabulary for a second generative machine learning model associated with the first generative machine learning model, wherein the first vocabulary is a proper subset of the second vocabulary; andassociating the draft token with the second token identifier; andgenerating a response to the input, using the first generative machine learning model, based on receiving an indication that the one or more draft tokens have been verified by the second generative machine learning model.15.The processor-implemented method of claim 14, further comprising generating the first vocabulary based at least in part on the second vocabulary, wherein the first vocabulary consists of a subset of tokens within the second vocabulary that satisfies one or more conditions.16.The processor-implemented method of claim 15, wherein the one or more conditions comprise each token of the subset of tokens being one of a set of tokens having a frequency of occurrence in one or more datasets above a threshold.17.The processor-implemented method of claim 16, wherein the one or more datasets comprise one or more unprocessed datasets, the one or more unprocessed datasets not being processed by the first generative machine learning model or the second generative machine learning model.18.The processor-implemented method of claim 16, wherein the one or more datasets comprise one or more datasets generated by the first generative machine learning model or by the second generative machine learning model.19.The processor-implemented method of claim 16, wherein the one or more datasets comprise a dataset having at least a first portion generated by the first generative machine learning model and a second portion generated by the second generative machine learning model.20.A processing system comprising:means for receiving an input at a first generative machine learning model, the first generative machine learning model being associated with a first vocabulary;means for generating one or more draft tokens based on the input using the first generative machine learning model, wherein the means for generating comprises, for each draft token of the one or more draft tokens:means for generating a first token identifier for the draft token based on processing the input using the first generative machine learning model;means for mapping the first token identifier to a second token identifier within a second vocabulary for a second generative machine learning model associated with the first generative machine learning model, wherein the first vocabulary is a proper subset of the second vocabulary; andmeans for associating the draft token with the second token identifier; andmeans for generating a response to the input, using the first generative machine learning model, based on means for receiving an indication that the one or more draft tokens have been verified by the second generative machine learning model.