Speculative decoding in autoregressive generative artificial intelligence models

By using a draft model to infer and generate multiple token sets in a generative artificial intelligence model and then verifying them with the target model, and combining group inference and parallel inference decoding, the problem of high computational resource consumption when generating responses from large language models is solved, achieving high generation rate and throughput.

CN120958464APending Publication Date: 2025-11-14QUALCOMM INC
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202480019732.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2023-10-02
Filing Date
2024-01-31
Publication Date
2025-11-14

AI Technical Summary

Technical Problem

Generative AI models consume high computational resources when generating responses, especially in large language models. The computational overhead of generating tokens for each round increases with the number of parameters, which limits the storage and processing capabilities of devices.

Method used

By employing speculative decoding technology, multiple token sets are generated using a smaller draft model and verified by a larger target model. This approach combines group speculation and parallel speculative decoding techniques to improve generation rate and throughput.

Benefits of technology

It reduces the computational cost of generating responses, increases the rate and throughput of token generation, and, especially in hybrid systems, reduces server load and maintains model accuracy through the collaborative work of edge devices and cloud servers.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120958464A_ABST
    Figure CN120958464A_ABST
Patent Text Reader

Abstract

Certain aspects of the present disclosure provide techniques and apparatus for generating a response to a query input into a generative artificial intelligence model. The method generally includes generating a plurality of token sets based on an input query and a first generative model, each token set of the plurality of token sets corresponding to a candidate response to the input query; outputting the plurality of token sets to a second generative model for verification; receiving an indication of a token set selected from the plurality of token sets from a second generative model based on the input query and the plurality of token sets; and outputting the selected set of tokens as a response to the input query.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] Cross-references to related applications

[0002] This application claims priority to U.S. Patent Application S / N.18 / 479,659, filed October 2, 2023, entitled "Speculative Decoding in an Autoregressive Generative Artificial Intelligence Model," which claims priority and benefit to U.S. Provisional Patent Application S / N.63 / 454,605, filed March 24, 2023, entitled "Speculative Decoding in an Autoregressive Generative Artificial Intelligence Model," both of which are assigned to the assignee of this application and are incorporated herein by reference in their entirety.

[0003] introduction

[0004] This disclosure relates to generative artificial intelligence models, and more specifically to speculative decoding in generative artificial intelligence models.

[0005] Generative AI models can be used in a variety of environments to generate responses to input queries. For example, generative AI models can be used in chatbot applications, where large language models (LLMs) are used to generate answers to input queries, or at least responses to those queries. Other examples of the use of generative AI models include models that generate stable diffusion of images based on input text descriptions of the content of desired images, and decision transformers that predict future actions based on sequences of previous actions in a given environment.

[0006] Generally, using generative AI models to generate responses to queries can be computationally expensive. For example, in a chatbot deployment that uses a large language model to generate responses to queries formatted as text, the response to the query may be generated for each token (e.g., a word or a portion of a word) generated as part of the response using a one-round pass through the large language model. The output of each round pass may be a probability distribution of the token set (e.g., individual words or portions of multiple words), which can be selected from the token set, for example, by sampling or based on maximum likelihood. Since a one-round pass through the large language model is used to generate each word (or tokens) in the query response, the computational overhead can be modeled as the product of the number of words included in the response and the computational resource overhead of performing a one-round pass through the large language model (e.g., in terms of processing power, memory bandwidth, and / or other computational resources used), which generally increases with the number of parameters within the large language model.

[0007] Brief Overview

[0008] Certain aspects of this disclosure provide a method for generating a response to an input query using a generative artificial intelligence model. The method generally includes: generating a plurality of token sets based on the input query and a first generative model, each of the plurality of token sets corresponding to a candidate response to the input query. The plurality of token sets are output to a second generative model for verification. An indication of a token set selected from the plurality of token sets is received from the second generative model based on the input query and the plurality of token sets. The selected token set is output as a response to the input query.

[0009] Certain aspects of this disclosure provide a method for validating responses to an input query generated using a generative artificial intelligence model. The method generally includes: receiving an input query and a plurality of token sets generated by a first generative model, each of the plurality of token sets corresponding to a candidate response to the input query. A probability distribution associated with each corresponding token set in the plurality of token sets is compared with a corresponding probability distribution generated by a second generative model for the corresponding token set. Token sets are selected based on the comparison, and the selected token set is output to the first generative model.

[0010] Certain aspects of this disclosure provide a method for parallel speculative generation of responses to an input query using a generative artificial intelligence model. The method generally includes: generating a first plurality of token sets based on an input query and a first generative model, each token set in the first plurality of token sets corresponding to a first portion of a candidate response to the input query. The first plurality of token sets are output to a second generative model for verification. While awaiting an indication from the second generative model for a token set selected from the first plurality of token sets, a second plurality of token sets are speculatively generated, each token set in the second plurality of token sets corresponding to a second portion of a candidate response to the input query. The indication for the token set selected from the first plurality of token sets is received from the second generative model. Tokens from the second plurality of token sets associated with the selected token set are output to the second generative model for verification, and the selected token set is output as a response to the input query.

[0011] Other aspects include: a processing system configured to perform the foregoing methods and those methods described herein; a non-transient computer-readable medium including instructions that, when executed by one or more processors of the processing system, cause the processing system to perform the foregoing methods and those methods described herein; a computer program product implemented on a computer-readable storage medium including code for performing the foregoing methods and those methods further described herein; and a processing system including means for performing the foregoing methods and those methods further described herein.

[0012] The following description and related figures illustrate certain illustrative features of one or more aspects. Brief description of the attached diagram

[0014] The accompanying drawings depict only certain aspects of this disclosure and are therefore not intended to limit the scope of this disclosure.

[0015] Figure 1 Examples of group inference decoding in generative models based on various aspects of this disclosure are explained.

[0016] Figure 2 An example relationship between the draft model and the target model used in group inference decoding in generative models, according to various aspects of this disclosure, is explained.

[0017] Figure 3 Examples of parallel speculative decoding in generative models based on various aspects of this disclosure are explained.

[0018] Figure 4 An example pipeline for hybrid speculative decoding in generative models, based on various aspects of this disclosure, is described.

[0019] Figure 5 An example of hybrid speculative decoding in a generative model based on various aspects of this disclosure is explained.

[0020] Figure 6 An example of hybrid parallel speculative decoding in a generative model based on various aspects of this disclosure is explained.

[0021] Figure 7 An example of hybrid group inference decoding in a generative model based on various aspects of this disclosure is explained.

[0022] Figure 8 Example operations for using a generative artificial intelligence model to generate a response to an input query, based on various aspects of this disclosure, are explained.

[0023] Figure 9 Example operations for using a generative artificial intelligence model to verify a response to an input query, based on various aspects of this disclosure, are explained.

[0024] Figure 10 Example operations for using a generative artificial intelligence model to infer and generate responses to input queries and to verify responses to input queries, based on various aspects of this disclosure, are explained.

[0025] Figure 11 An example processing system configured to perform various aspects of this disclosure is described.

[0026] To facilitate understanding, the same reference numerals are used wherever possible to designate common elements shared by all figures. Elements and features conceived in one aspect can be beneficially incorporated into other aspects without further elaboration.

[0027] Detailed description

[0028] This disclosure provides apparatus, methods, processing systems, and computer-readable media for efficiently generating responses to input queries using generative artificial intelligence models.

[0029] Generally, generative AI models generate responses to queries input into the model. For example, a large language model (LLM) deployed within a chatbot can use multiple rounds of passage from the large language model to generate responses to queries, where each successive round is based on the query and tokens (or words or parts thereof) generated using previous rounds of passage from the large language model. These large language models can typically include millions, or even billions, of weights or parameters. Due to the size of these models and the operations performed on each token to predict what the next token should be in response to the query and previously generated tokens, deploying large language models on a wide range of devices that may have limited memory, storage, and / or processing power relative to the cloud computing instances on which large language models are typically run. Furthermore, the computational complexity involved in generating responses to queries provided as model input can involve significant energy consumption, processing time, memory utilization, and / or other resource utilization that may prevent computing resources from being used for other tasks.

[0030] To improve the efficiency and throughput of large language models, speculative decoding techniques allow smaller language models (sometimes called draft large language models, also known as draft models) to operate alongside larger language models (sometimes called target large language models, also known as target models). In this scenario, the draft model speculatively generates additional tokens and the probabilities used to sample these additional tokens based on the currently accepted set of tokens. The target model generates tokens based on those generated by the draft model. To produce the final result, the target model can perform rejection sampling on a per-token basis to accept or reject tokens generated by the draft model, ensuring that the draft and target models have similar probability distributions.

[0031] In some respects, the draft model can be a pruned version of the selected target model so that the draft model and the target model have similar probability distributions. In other respects, the draft model can be a smaller version of the target model (e.g., trained on millions of tokens rather than on hundreds of millions or even billions of tokens).

[0032] Various aspects of this disclosure provide techniques for generating responses to queries input into a large language model using speculative decoding on a group-by-group basis. Generally, a draft model can generate one or more sets of tokens as candidate responses to the query. The target model can then perform sample rejection on a per-set basis. Tokens within the group can be selected based on the conditional probabilities of tokens within the group. By performing rejection sampling on a per-set basis, various aspects of this disclosure preserve a close relationship between the probability distributions within the draft and target models, while increasing the throughput (e.g., the number of tokens generated per second) of the draft and target models compared to draft and target models configured to generate responses on a per-token basis.

[0033] This disclosure provides techniques for generating responses to queries input into a large language model using speculative decoding techniques that run in parallel with a draft model and a target model, also referred to herein as "hybrid speculative decoding." In hybrid speculative decoding, the draft model speculatively generates one or more tokens, which are then validated by the target model. The draft model can be executed on an edge device, and the target model can be executed on a server (e.g., locally accessible from the edge device, hosted in a cloud computing environment, etc.). By executing both the draft model and the target model in parallel, the rate of token generation can be maximized or at least increased. Furthermore, in scenarios where the edge device is disconnected from the server, the edge device can continue generating tokens to produce responses to received queries.

[0034] Speculative Decoding in Generative Artificial Intelligence Models

[0035] Generally speaking (e.g., in large language models), autoregressive token generation takes historical tokens as input to generate output. That is, autoregressive token generation can be represented by the following expression:

[0036] x t ~p(x|x0,x1,…,x t-1 )→x t+1 ~p(x|x0,x1,…,x t-1 ,x t )

[0037] Where x t Let p represent the sequence of tokens generated at time t, with conditional probability p for each token x0 to x1. t-1 The choice is a condition, and x t+1 Let p represent the sequence of tokens generated at time t+1, with conditional probability p for each token x0 to x1. t The choice is a condition. Generally, a single token can be generated each time the autoregressive model is executed, which means that N inferences can be performed to generate a sequence of N tokens. As discussed above, speculative decoding techniques can be used to accelerate token generation by using a draft model that speculatively generates tokens, smaller than the target model, where the target model is used to verify the tokens generated (speculously) by the draft model.

[0038] In the speculative decoding pipeline, the draft model can autoregressively and speculatively generate n tokens based on the following expression:

[0039]

[0040] Where t corresponds to the time point, and Corresponding to the pair of tokens x0 to x0 associated with the token x selected at time t. t-1 The choice of is a conditional probability distribution.

[0041] The target model takes the generated n tokens and processes them in parallel according to the following expression to generate a probability distribution for each of the n tokens:

[0042]

[0043] Where k corresponds to the token index of the n generated tokens.

[0044] The target model can then verify the token generated by the draft model by comparing the distributions from the draft model and the target model to determine whether the token is accepted or rejected.

[0045] For a given function f and a certain threshold α (also known as the acceptance rate), when At that time, given a token It can be accepted. Otherwise, the token can be rejected. This can then be based on a function. The final token is generated at either the first rejection position or the last position n.

[0046] Compared to iteratively generating tokens using a single autoregressive model, speculative decoding with an acceptance rate of α may result in lower costs. The cost savings from inference compared to iterative token generation can be expressed by the following formula:

[0047]

[0048] Considering this example, for N = 1000, C 目标 =10, C 草稿 =1, n=4, α=3, where N corresponds to the number of tokens, C AR Corresponding to the computational cost of using the acceptance rate α, C 目标 Corresponding to the computational cost of generating the token set using the target model, C 草稿 This corresponds to the computational cost of generating the token set using the draft model, and n corresponds to the number of tokens speculatively generated through a single round of passage using the autoregressive model. In such examples, speculative decoding can result in a 35% reduction in computational overhead compared to standalone autoregressive iterative token generation.

[0049] However, as discussed, per-token speculative decoding can impose a limitation on the rate of token generation, since the first token can be sampled individually by the draft model and verified by the target model before the next token is sampled by the draft model and verified by the target model. In other words, generating a response to an input query using per-token speculative decoding can involve executing both the draft and target models for each token generated as a part of the response to the input query, which can utilize considerable computational resources (e.g., processor time, memory, memory bandwidth, etc.) to generate the response.

[0050] Example swarm inference decoding in generative artificial intelligence models

[0051] Figure 1 Example 100 of group inference decoding (GSD) in a generative artificial intelligence model according to various aspects of this disclosure is explained.

[0052] As explained, the draft model and the target model can be used in combination (or otherwise together) to perform group speculative decoding of tokens to generate responses to queries received for processing by one or more generative AI models. As discussed further in detail below, group speculative decoding of tokens allows the draft model to speculatively generate multiple sets of tokens for verification by the target model. Because multiple sets of tokens can be generated by the draft model for verification by the target model, group speculative decoding increases the token generation rate of the generative AI model by generating multiple sets of tokens that can be accepted as correct responses, as a larger number of token sets increases the probability that at least one set includes one or more tokens that will be accepted as a response.

[0053] Given the input of a received query, the draft model typically forms a sampled token group, comprising one or more high-probability node groups derived from the probability distribution of the output over a potential set of tokens. Each high-probability node group may include a group selected based on various techniques, such as optimal k-selection (e.g., selecting the k tokens with the highest probability within the probability distribution), kernel-based selection (e.g., selection based on the sum of probabilities satisfying a threshold probability), etc. Tokens not included in one or more groups may be considered as separate set groups (e.g., groups containing a single member). By selecting candidate token groups, the draft model can sample tokens based on the probability distribution; however, each group can be considered as a single selection with a cluster probability (e.g., as the sum of individual token probabilities) such that the draft model matches or at least approximates the target model, such that the computed probability of each group is a sufficient response to the received query. Clustering achieves this match or at least approximation between the probabilities computed by the draft model and the target model because, although the individual token probabilities may not match very well, the clusters of tokens may have probabilities that are likely similar in both the draft and target models.

[0054] When the sampled token group is input into the draft model in the next iteration of the draft model, the tokens in the sampled token group are input at the sampling positions and processed independently. The result may be a tree data structure 110, where the hint serves as the root node 111 of the tree data structure, and subsequent layers within the tree data structure 110 represent different tokens (or token groups), combined with each of the previously selected token combinations. At some point in time (e.g., after generating a tree with a defined depth (corresponding to the maximum length of the sequence generated by the draft model), the draft model may output the generated tree data structure 110 to the target model for further processing. In some aspects, the tree data structure 110 may be output to the target model along with the respective groupings and selection probabilities generated by the draft model.

[0055] In some respects, the number of nodes at each level of the tree data structure 110 and the depth of the tree data structure 110 can be defined a priori. The number of nodes at each level of the tree can be defined globally, on a per-level basis, or in some other way. For example, the number of nodes at any level of the tree data structure 110 can be defined based on the branch factor of the tree data structure 110 at the immediately preceding level. For example, a branch factor of 2 for a node at level n of the tree data structure 110 can result in 2 nodes (tokens) being generated at level n+1 of the tree data structure 110 for each node (token) at level n. Meanwhile, the depth of the tree data structure 110 can be defined based on the maximum number of tokens (e.g., words) generated using any round of the draft model. For example, if the draft model is configured to generate a sequence of tokens with a maximum length of 5 during any instance of speculative generation, the depth of the tree can be 6 (to include the root node corresponding to the input to the draft model at the first level of the tree data structure 110).

[0056] The draft model can be configured to trigger the target model to generate tokens based on various complexity and / or performance criteria. For example, the draft model can trigger the target model to generate tokens based on a complexity or performance criterion associated with the size of the generated tree. In other examples, the draft model can be configured to trigger the target model to generate tokens based on a time criterion associated with the expected amount of time it would take for the target model to generate a set of tokens, to which the generated tree can be compared. Generally, these complexity and / or performance criteria can set an upper limit on the number of tokens generated by the draft model for validation by the target model. In some aspects, this upper limit can be based on the number of nodes in the tree data structure and can be influenced, for example, by a branching factor defined for different layers of the tree data structure to which the sampled tokens are organized, the depth of the tree data structure, etc. The worst-case computational load for the final round of token generation can be configured to be defined by the memory bandwidth at the device on which the draft model executes.

[0057] Based on the generated tree data structure 110, each group, and the selection probabilities received from the draft model, the target model generates the output (probability) distribution for each partial path of the generated tree data structure 110. Each partial path generally starts from the root node of the generated tree data structure 110 and can be a continuous path through the generated tree data structure 110 that starts from the root node and terminates at any node in the generated tree data structure 110. The target model can generate the output distribution for each partial path using a single-round pass through the target model, for example, by including all tree nodes in the generated tree data structure 110 as token inputs and performing masked self-attention and positional encoding on each partial path within the tree data structure 110. Based on the output distribution of each node in the generated tree data structure 110, the target model can select the token set 112 by traversing from the root node of the generated tree data structure 110. If the target model includes multiple tokens from the selected token set 112, tokens can be selected from the group using normalized group token probabilities (e.g., token probabilities conditioned on the selection of a group that includes a token), and the selected tokens can be redefined as the root node of the generated tree data structure 110. The process discussed herein can be repeated, wherein the selected tokens are defined as the root node of the tree data structure 110 to generate subsequent tokens to be included in the selected token set 112.

[0058] If a token set is rejected based on a probability distribution generated by the target model, a final token set can be selected based on a group-level rejection sampling method. The final token can be selected from the selected final token set based on normalized group token probabilities. If all token sets are accepted by the target model, the final token can be sampled based on the final leaf distribution (e.g., the probability distribution generated at the leaf nodes in tree data structure 110 corresponding to the last accepted or verified tokens along any partial path of tree data structure 110), which allows the target model to generate additional tokens relative to those speculatively generated by the draft model. Using the techniques discussed herein, the selected token set 112 can then be passed back to the draft model as input for a subsequent round of token speculation and verified by the target model.

[0059] Figure 2 Example 200 illustrates the relationship between tokens generated by the draft model and the target model used in group speculation decoding in a generative model according to various aspects of this disclosure.

[0060] The tree generated by the draft model generally allows for diverse inferences, where many candidate token sets are provided by the draft model for validation by the target model. This increases the target model's acceptance rate of tokens generated by the target model. As explained, accepting token number 4 from the token group generated by the draft model at time t = 1210 leads to accepting token number 1 from the token group generated by the draft model at time t = 2220. Subsequently, at time t = 3230, two options exist: token number 3 or token number 5. Token number 3 can be accepted as a sampled token, while token number 5 can be rejected as a sampled token, as explained. At time t = 4240, token 4 from the token set inferred based on the selection of token 3 at time t = 3230 can be accepted as a sampled token based on the calculated probabilities of tokens 3 and 4 conditioned on the acceptance of token 3 at time t = 3230. Meanwhile, both tokens 4 and 5, which are derived from the token set speculatively generated based on the selection of token 5 at time t = 3 230, can be rejected based on the rejection of token 5 at time t = 3 230.

[0061] A token tree, rather than a sequence of tokens, can be fed into the target model as input. For example, a tree could be fed into the target model that includes options such as token 3 or 5 being candidate tokens at time t = 3230, token 3 or 4 being candidate tokens at time t = 4240 if token 3 is selected at time t = 3230, and token 4 or 5 being candidate tokens at time t = 4240 if token 5 is selected at time t = 3230. As explained, at time t = 5250, inputting the token tree data structure (e.g., as explained by the token sequence shown from time t = 1 to t = 4) into the target model allows the target model to generate an output including five tokens—one more than the four token sequences included in the tree generated by the draft model—which provides an improved acceptance rate for the token tree and allows for the generation of additional tokens to be included in the responses generated by the draft and target models. Figure 2 The example described herein includes a set of five tokens, which includes token 4 at time t = 1210, token 1 at time t = 2220, token 3 at time t = 3230, token 4 at time t = 4240 (conditional on selecting token 3 at time t = 3230), and token 2 at time t = 5250 (conditional on selecting token 4 at time t = 4240).

[0062] Generally, a new selection of the token group can be performed whenever the draft model generates a new distribution of one or more subsequent tokens to be included in the response to a received query processed by the draft model and the target model. As discussed, using group selection of tokens allows for a closer match between the probability distributions of the draft model and the target model, minimizing or at least reducing the likelihood of group rejection. In some respects, when there is a low level of uncertainty (e.g., when the next token selection has low entropy), the draft model can choose not to form a group to minimize or at least reduce computational resource consumption.

[0063] Generally, when using group speculation decoding (to generate responses to input queries using generative AI models) can be highly efficient, forming token swarms can reduce the computational complexity of both the draft and target models while maintaining high throughput (e.g., the number of tokens generated per second by the draft and target models). Generally, triggering the target model can be influenced by various considerations, such as the maximum tree size to keep model computational performance metrics within bounds, estimates of when longer speculations have diminishing returns (e.g., when the computational cost of generating a larger tree of speculated tokens is unlikely to be worthwhile, taking into account the possibility that the target model will reach later tokens in the tree), and other performance considerations.

[0064] In some respects, the draft model can match the target model (e.g., in terms of probability distribution), but can have faster inference performance on the same hardware. Generally, a smaller model can generate many speculative tokens, but the probability of those tokens being rejected by the target model increases. As discussed above, group speculative decoding addresses this increased probability of token rejection at the cost of increased computational overhead for longer sequences.

[0065] In some respects, at the draft model, the temperature parameter—or a parameter that affects the draft model’s likelihood of selecting a token with a lower probability of being the exact token to be included in the token set—can be tuned or otherwise selected to improve the performance of group speculation decoding.

[0066] In some respects, the draft model can be fine-tuned to match (or at least approximate) the target model and maximize (or at least increase) the probability that speculative tokens generated by the draft model will be accepted as valid tokens by the target model.

[0067] Generally, group speculative decoding can improve token generation performance compared to per-token speculative decoding. That is, given a Kullback-Leibler (KL) divergence between the draft and target models, measuring how the probability distribution of the draft model differs from that of the target model (treating the target model as a reference distribution), group speculative decoding can generate a larger number of tokens per run of the target model compared to per-token speculative decoding. Different swarming strategies (e.g., group size, additional tokens, etc.) can have different computational complexity characteristics; therefore, the choice of swarming strategy can be based on trade-offs between computational complexity and performance, given bounds on the read bandwidth of the draft and target models, and hardware performance.

[0068] Example Parallel Inference Decoding in Generative Artificial Intelligence Models

[0069] As discussed above, group speculative decoding allows the draft model to speculate on combinations of token generation and target model token acceptance to generate a final token set that matches the distribution of the target model, while increasing the rate of token generation. Furthermore, aspects of this disclosure provide parallel speculative decoding, in which the draft model speculates and generates multiple token sets in parallel, which can provide an increased token generation rate while improving the computational efficiency involved in generating responses to queries using generative artificial intelligence.

[0070] Figure 3 Example 300 of parallel speculative decoding (PSD) in a generative artificial intelligence model according to various aspects of this disclosure is explained.

[0071] As explained, the draft model can run multiple instances of the speculative decoding process to generate multiple candidate token sequences. Each iteration can be performed with different parameters or a seed to sample tokens. Performing N iterations using the S tokens generated in each iteration results in a total of N*S tokens, which can be provided to the target model for processing.

[0072] For example, as explained in Example 300, the draft model can initially perform multiple instances of the speculative decoding process to generate a first candidate sequence 302 and a second candidate sequence 304 for the prompts input into the draft model. As explained, N = 2, S = 4; thus, the draft model can generate two sequences, each containing 4 tokens, for a total of 8 tokens. It should be recognized that this is only an example, and any values ​​of N and S can be used to define the number of candidate sequences generated by the draft model and the number of tokens included in each sequence, respectively.

[0073] The target model takes N*S tokens as input and processes N candidate token sequences in a single run. The result of the target model processing N*S tokens may be a set of N potentially accepted sequences of varying lengths, depending on where the target model rejects a token in each of the N candidate token sequences. Various techniques can be used to select the final sequence based on the N potentially accepted sequences. As a non-limiting example, a "greedy" technique may lead to the selection of the longest sequence, where constraints (e.g., between sequences of the same length) are arbitrarily broken, potentially resulting in the highest number of tokens being selected each time the target model runs. In some cases, the longest-length selection may be biased towards an efficient final sampling distribution, which may cause the probability distribution of the target model to converge, at least to some extent, with the probability distribution of the draft model.

[0074] For example, as explained, the target model accepts two of the four tokens from the first candidate sequence 302 and four of the four tokens from the second candidate sequence 304. Because the second candidate sequence 304 accepts more tokens than the first candidate sequence 302, a "greedy" technique might cause the target model to select the second candidate sequence 304 as the sequence on which the draft model generates subsequent candidate sequences 312 and 314. In this case, speculative token generation is not required based on the first candidate sequence 302 because the first candidate sequence 302 has fewer accepted tokens than the second candidate sequence 304. However, in other cases, other techniques can be implemented to select one of the candidate sequences, including sequences that may not have the most accepted tokens.

[0075] Although Figure 3Parallel speculative decoding based on per-token speculative decoding has been explained; however, it should be recognized that N candidate token sequences can be generated additionally or alternatively based on the group speculative decoding techniques discussed above. For example, using the group speculative decoding techniques discussed above, multiple token sets generated by multiple instances of the draft model can be generated for verification by the target model. In some aspects, the token sets generated by instances of the draft model can be organized into tree data structures (as discussed above), resulting in the generation of multiple tree data structures, where each individual tree data structure corresponds to a token set generated by a specific instance executed by the draft model using instance-specific parameters. The target model can attempt to verify the tokens in each token set to determine the selected token set on which subsequent speculative decoding will be performed. The longest path through each tree data structure can be identified by the target model as a candidate sequence. Using a “greedy” technique, the tokens in the candidate sequence with the longest path can be selected as the token set on which subsequent speculative decoding will be performed, and can be output by the target model to the draft model for use as the root node of the subsequent tree data structure (e.g., a tree data structure corresponding to the subsequent tokens generated as part of the response to the input query).

[0076] Hybrid speculative decoding example

[0077] In some scenarios, it may be possible to use autoregressive inference to execute generative AI models in hybrid systems. Examples of such hybrid systems include systems in which these generative models are executed on edge devices (e.g., smartphones, tablets, laptops, desktop computers, etc.) and cloud servers, on edge devices and home servers, on different home servers, and so on. Thus, in some aspects, a draft model can be executed on a local device, while the target model can be executed on a remote device. Executing generative AI models in hybrid systems can speed up the generation of query responses and reduce the computational load on servers (e.g., by offloading some computational processes to another device). Furthermore, the accuracy of the target model can be preserved when the server executing the target model is available, and large language models (or other generative AI models) can remain available (albeit with lower accuracy) at the edge device when the edge device loses communication with the server.

[0078] However, in hybrid systems, cost reductions may be limited when both the draft model and the target model are hosted in the cloud. Furthermore, the sequential nature of inference (e.g., involving speculatively generating tokens using the draft model and subsequently validating them using the target model) may limit the rate at which tokens are generated.

[0079] To further reduce the computational cost involved in generating query responses using generative artificial intelligence models, various aspects of this disclosure provide techniques for speculative decoding in hybrid environments. In these aspects, speculative decoding in hybrid environments can be performed such that (e.g., at an edge device) generation of candidate token sets using a draft model and (e.g., at a server) token verification using a target model can be performed in parallel. Assuming variations in the number of accepted tokens, the draft model can continue to generate subsequent draft token sets, where padding can be added to account for the difference between the number of accepted tokens and the previous round of speculative token generation. The draft model and the target model can be executed substantially in parallel and sequentially, which accelerates the generation of query responses using generative artificial intelligence models.

[0080] Figure 4 An example pipeline 400 for Hybrid Inference Decoding (HSD) in generative models is described according to various aspects of this disclosure.

[0081] As explained, to maximize or at least increase throughput (e.g., the number of tokens generated per second), a draft model 402 executing on edge device 403 (which may be referred to as a local device or a first device) may continuously generate batches (or sets of tokens, corresponding to different responses of different lengths to input prompts, conditioned on any previously accepted tokens) of candidate tokens and output these tokens to a target model 404 executing on server 405 (which may be referred to as a remote device or a second device, in a cloud computing environment, for example). Edge device 403 may receive the set of accepted tokens from target model 404 from server 405. Draft model 402 may prune the sample tree to conform to the set of accepted tokens. In some aspects, when the set of accepted tokens is empty (e.g., when the target model has not accepted any tokens), draft model 402 may backtrack to the last token accepted by target model 404 and restart token generation from the last accepted token. For example, to backtrack, draft model 402 can prune the generated sample token tree to the last accepted token and start speculative generation again based on the pruned tree.

[0082] In example pipeline 400, a draft model 402 executed on edge device 403 may generate a first speculative token set 410 and output the first speculative token set to a target model 404 executed on server 405 for verification. In some aspects, the first speculative token set 410 may include multiple subsets of speculative tokens, where each subset corresponds to one or more sequences of tokens generated by different instances executed by the draft model 402 using different operating parameters (e.g., based on group speculative decoding, parallel speculative decoding, or other techniques).

[0083] In some respects, while the target model 404 validates the first speculative token set 410, the draft model 402 may generate a second speculative token set 414. The second speculative token set 414 may be generated based on the assumption that the target model 404 accepts 1 to S tokens from the first speculative token set 410, where S corresponds to the number of tokens per sequence generated by the draft model 402. The draft model 402 may then receive a selected token set 412 from the target model 404. Based on the selected token set 412, the draft model 402 may identify a subset of the second speculative token set 414 for analysis (e.g., generating tokens conditioned on accepting the correct number of tokens from the first speculative token set 410) and provide this subset to the target model 404 for validation. The process can continue, wherein the draft model 402 generates a third speculative token set 418, while the target model 404 verifies the second speculative token set 414 and returns a selection 416 from the second speculative token set 414 to the draft model 402, wherein the draft model 402 generates a fourth speculative token set 422, while the target model 404 verifies the third speculative token set 418 and returns a selection 420 from the third speculative token set 418 to the draft model 402, and so on. In other aspects, the draft model 402 may wait for the generation of the second speculative token set 414 until the draft model 402 receives the selected token set 412 from the target model 404.

[0084] Some of the processes described herein for speculatively generating tokens using a draft model in parallel or substantially in parallel and validating a previously generated set of speculatively generated tokens using a target model can be executed based on group selection and target model triggering conditions that enable consecutive or at least near-consecutive operations of draft model 402 and target model 404, while minimizing or at least reducing the likelihood that draft model 402 will backtrack from the last accepted token and restart speculative token generation (e.g., minimizing or at least reducing the likelihood that target model 404 will reject each of the speculatively generated tokens generated by draft model 402). Furthermore, group selection and target model triggering conditions can be selected to avoid or at least reduce the likelihood of excessive computational burden on edge devices or servers (e.g., to minimize or at least reduce memory jitter, where data is repeatedly exchanged between on-processor memory and off-processor memory, network bandwidth bottlenecks, etc., within defined performance metrics) and token generation by the edge device and token validation by the server.

[0085] In some aspects, hybrid speculative decoding can be performed on a per-token basis. In such cases, draft model 402 can operate continuously and generate various branching paths based on the generation of new speculative tokens. In some aspects, edge device 403 can execute multiple instances of raw speculative decoding (e.g., using different parameters input to draft model 402) such that draft model 402 generates multiple paths within a tree representing options verifiable by target model 404. Based on some defined triggering criteria, edge device 403 can provide the generated tree to target model 404 running on server 405 (e.g., in a cloud computing environment), which, as discussed above, can select a sequence of tokens corresponding to a potentially valid response to an input query and return the selected sequence to draft model 402 executed on the edge device. Draft model 402 can then prune the generated tree to conform to the selected sequence returned by target model 404. In this example, the token generation rate can correspond to the rate at which draft model 402 generates tokens, where tree pruning is based on target selection. In this example, inference accuracy and speed can be inversely related because a smaller draft model can generate tokens at a higher rate, but the target model 404 is less likely to accept tokens inferred by the draft model 402; meanwhile, a larger draft model can generate tokens at a lower rate, but the target model 404 is more likely to accept tokens inferred by the draft model 402.

[0086] Figure 5 Example 500 of hybrid speculative decoding in a generative model according to various aspects of this disclosure is explained. In hybrid speculative decoding, in the first round of inference, a draft model (e.g., Figure 4 The first set of speculative generation tokens 502 (also known as "round 1 draft tokens") generated by the draft model 402 described in the text can be provided to the target model (e.g., Figure 4 The target model (404) described herein is provided for verification. As discussed, in various respects, the draft model can be executed on the first device, while the target model can be executed on the second device.

[0087] Depending on the aspects, when the first set of speculatively generated tokens 502 is processed by the target model, the draft model generates a second set of draft tokens (also referred to as "round 2 draft tokens") in the second round of inference, where the assumptions are made for a different number of accepted tokens. For example, as explained, the second set of draft tokens may include (1) a first subset 504, which assumes acceptance of the first draft tokens from the first set and may include a set of tokens speculatively generated based on acceptance of the first tokens; (2) a second subset 506, which assumes acceptance of the first and second draft tokens from the first set and includes a set of tokens speculatively generated based on acceptance of the first and second tokens; (3) a third subset 508, which assumes acceptance of the first to third draft tokens from the first set and includes a set of tokens speculatively generated based on acceptance of the first to third tokens; and (4) a fourth subset 510, which assumes acceptance of all four tokens from the first set and includes a set of tokens speculatively generated based on acceptance of all four tokens. For cases where fewer tokens than those included in the first set are assumed to be accepted, padding (e.g., null values, predefined constants, etc.) can be added to ensure that each hypothesis has the same length. As explained, a second draft token set can be generated in batch processing so that an appropriate selection of draft tokens conditioned on a specific set of accepted tokens from the first set of speculatively generated tokens 502 can be provided to the target model for verification.

[0088] Figure 6 Example 600 of hybrid speculative decoding in a generative model according to various aspects of this disclosure is explained. In this example 600 (also referred to as an example of hybrid parallel speculative decoding (HPSD), the parallel speculative decoding aspects discussed above can be used to generate various sets of candidate tokens for validation by a target model. The target model can operate continuously, as if performing speculative decoding on a per-token basis. However, instead of generating tokens on a per-token basis, the draft model can speculatively generate a group of tokens based on a probability distribution that approximates (but is not necessarily equal to) the probability distribution associated with the target model, and generate different tree data structures based on the input query (or hint) and the generated group of tokens, such as... Figure 6 As explained above, based on some defined triggering criteria, an edge device (also referred to as a local device or first device) can provide the generated tree data structure to a target model (also referred to as a remote device or second device) running on a server (e.g., in a cloud computing environment), which, as discussed above, selects a sequence of tokens corresponding to a potentially valid response to an input query and returns the selected sequence to a draft model executed on the edge device. Generally, the selected token sequence may include speculatively generated tokens verifiable by the target model, plus an additional token generated based on the speculatively generated tokens. The draft model can then prune the generated tree to conform to the tokens verified by the target model.

[0089] As explained, different instances of the draft model using different parameters can independently generate a first speculative token set 602 and a second speculative token set 604 based on input prompts. (e.g., executed locally or on a first device) The draft model can provide the first speculative token set 602 and the second speculative token set 604 to a target model (e.g., executed remotely or on a second device) for verification. In parallel or at least substantially parallel with the target model's verification of the first speculative token set 602 and the second speculative token set 604, the draft model can generate a first subsequent token set 612 based on the assumption that the target model has accepted the first speculative token set 602, and a second subsequent token set 614 based on the assumption that the target model has accepted the first speculative token set 604. In this example, as explained by the ellipse, the target model has accepted the second speculative token set 604 beyond the first speculative token set 602 (e.g., based on the above assumption about...). Figure 3 (Similar to the described method). Since the target model has accepted the second speculative token set 604, the draft model can discard the first subsequent token set 612 and provide a second subsequent token set 614 for verification by the target model. Furthermore, although not shown, the draft model can generate an additional token set based on the second subsequent token set 614, while the target model verifies the second subsequent token set 614.

[0090] The process of speculative token generation, pruning (e.g., by discarding the set of unselected tokens) and target model validation is repeated until a termination condition is reached for the input (e.g., until no further tokens need to be generated to generate a correct response to the input query, until the maximum number of tokens has been generated, until the maximum amount of computational resources has been used to generate the response, etc.).

[0091] In this example, the overall token generation rate corresponds to the rate at which the draft model generates tokens. However, when the target model determines that no speculatively generated token in the generated tree is a candidate included in response to a query, previously generated tokens can be discarded because the draft model can use tokens (up to the last verified token) as input to the draft model to restart speculative token generation. Similar to per-token-based hybrid speculative decoding, in this example, inference accuracy and speed may be inversely related because a smaller draft model can generate tokens at a higher rate, but the target model is less likely to accept tokens speculatively generated by the draft model.

[0092] Figure 7Example 700 of Hybrid Group Speculation Decoding (HGSD) in a generative model according to various aspects of this disclosure is described. In this example 700, the group speculation decoding aspects discussed above can be used to generate various sets of candidate tokens for validation by a target model. The target model (e.g., executed on a remote or second device) can operate continuously, as if performing speculation decoding on a per-token basis. However, instead of generating tokens on a per-token basis, a draft model (e.g., executed on a local or first device) can speculatively generate a group of tokens based on a probability distribution equivalent to the probability distribution associated with the target model, and generate a tree (e.g., a tree data structure) based on the input query (or prompt) and the generated group of tokens. For example, based on some defined triggering criteria, the edge device can provide the generated tree to the target model running on a server (e.g., in a cloud computing environment), as discussed above, which can select a sequence of tokens corresponding to a potentially valid response to the input query and return the selected sequence to the draft model executed on the edge device. Generally, the selected token sequence may include speculatively generated tokens verifiable by the target model, plus an additional token generated based on the speculatively generated tokens. The draft model can then prune the generated tree to conform to the tokens verified by the target model.

[0093] For example, as explained, a tree-like data structure can be generated, where the input prompt serves as the root node of the tree data structure. A first token set 702, including multiple groups (represented as different paths through the tree), can be speculatively generated by a draft model and provided to a target model for verification, as discussed above. When the first token set 702 is verified by the target model, the draft model can speculatively generate a second token set 704, where different subsets of tokens in the second token set 704 are generated based on the assumption that different token sets from the first token set 702 are accepted by the target model. In example 700, the selected token set 706 may correspond to the token set verified by the target model, while the draft model generates subsequent token sets. In general, as discussed above, the tree data structure can be pruned based on the target model's verification of various speculatively generated token groups so that computational resources are not wasted on speculatively generating further token groups based on previous token sets that are not accepted by the target model.

[0094] Similar to Figure 6 Example 600 explained in the text is by Figure 7The overall token generation rate in the hybrid group speculative decoding illustrated in Example 700 can correspond to the rate at which the draft model generates tokens. However, when the target model determines that no speculatively generated token in the generated tree is a candidate included in response to a query, previously generated tokens can be discarded because the draft model can use tokens (up to the last verified token) as input to the draft model to restart speculative token generation. Similar to per-token-based hybrid speculative decoding, in this example, inference accuracy and speed may be inversely related because a smaller draft model can generate tokens at a higher rate, but the target model is less likely to accept tokens speculatively generated by the draft model.

[0095] Example operations for speculative decoding in generative artificial intelligence models

[0096] Figure 8 Example operation 800, which can be performed by a computing device according to various aspects of this disclosure to generate a response to an input query using a generative artificial intelligence model, is described. The computing device used to perform operation 800 can be a device on which at least a draft model can be deployed, such as a smartphone, tablet computer, laptop computer, desktop computer, server, cloud computing instance hosted in a distributed computing environment, etc.

[0097] As explained, operation 800 begins at box 810, where the computing device generates multiple sets of tokens based on the input query and a first generative model. Generally, each of the multiple token sets corresponds to a candidate answer to the input query. For example, the input query can be received from the user of the computing device at a prompt where a text query can be entered, by converting natural language speech into a text representation of the input query.

[0098] In box 820, operation 800 continues, where the computing device outputs the multiple token sets to a second generative model for verification (e.g., by another computing device).

[0099] In box 830, operation 800 continues, wherein the computing device receives an instruction from the second generative model for a set of tokens to be selected from the plurality of token sets based on the input query and the plurality of token sets.

[0100] In box 840, operation 800 continues, where the computing device (e.g., to a user or application, to another model, etc.) outputs the selected set of tokens as a response to the input query.

[0101] In some respects, each of the plurality of token sets includes a group of tokens that have the highest probability within the probability distribution associated with the first generative model on the token universe on which the first generative model is based.

[0102] In some respects, each of the plurality of token sets includes a token group selected based on the sum of probabilities associated with the tokens in that token group, which exceeds a threshold probability.

[0103] In some respects, the multiple token sets are represented as a tree data structure. Within this tree data structure, the root node generally corresponds to the input query, and different paths through the tree data structure generally correspond to different token sets among the multiple token sets. In some respects, the depth of the tree data structure may correspond to the maximum number of tokens generated in a single round of passage through the first generative model, and the maximum size of the tree data structure may be set based on a computational complexity metric associated with generating the target token set by the second generative model.

[0104] In some aspects, operation 800 may further include pruning the tree data structure based on the selected set of tokens. Subsequent sets of tokens are generated based on the pruned tree data structure and the input query and output to a second generative model for validation. Indications for subsequent sets of tokens selected from the subsequent sets of tokens are based on the input query, the pruned tree data structure, and the subsequent sets of tokens received from the second generative model. Outputting the selected set of tokens as a response to the input query may include outputting both the selected set of tokens and the selected subsequent sets of tokens as a response to the input query.

[0105] In some respects, each of the multiple token sets can be generated using a unique instance of the first generative model and a unique parameter that serves as input to the unique instance of the first generative model.

[0106] In some aspects, operation 800 may further include generating a plurality of subsequent token sets based on the input query and the plurality of token sets when the second generative model validates the plurality of token sets. Based on the plurality of subsequent token sets and the selected token set, a refined subsequent token set may be generated, and the refined subsequent token set may be output to the second generative model for validation.

[0107] In some respects, the token set in the subsequent multiple token sets may include several tokens padded to account for a number of tokens less than the maximum number (or a pre-configured number of tokens) in the selected token set.

[0108] In some aspects, operation 800 further includes receiving tokens generated by a second generative model based on a selected set of tokens. The received tokens may be output as additional tokens following the selected set of tokens.

[0109] In some respects, the first generative model can correspond to the draft model in the speculative decoding pipeline, and the second generative model can correspond to the target model in the speculative decoding pipeline.

[0110] In some aspects, the first generative model and the second generative model may have equivalent probability distributions. In some aspects, the first generative model may have a probability distribution that approximates the probability distribution associated with the second generative model. The approximation of the probability distribution may be a probability distribution falling within a threshold difference of the probability distribution associated with the second generative model, such that the probability distribution associated with the first generative model does not need to exactly match the probability distribution associated with the second generative model. In some aspects, the first generative model may be executed locally (e.g., on a local device or local system), and the second generative model may be a remotely hosted model (e.g., on a remote system or remote device).

[0111] Figure 9 Example operation 900 for validating a response to an input query using a generative artificial intelligence model, according to various aspects of this disclosure, is described. Operation 900 can be performed by a device (or a system with multiple devices) on which at least a target model can be deployed, such as a server (e.g., Figure 4 Server 405, cloud computing instances hosted in a distributed computing environment, etc.

[0112] As explained, operation 900 in box 910 begins by receiving an input query and a plurality of token sets from a device on which the first generative model runs, each of the plurality of token sets corresponding to a candidate response to the input query.

[0113] In box 920, operation 900 continues to compare the probability distribution associated with each of the plurality of token sets with the corresponding probability distribution generated by the second generative model for the respective token set.

[0114] In box 930, operation 900 continues to select a token set from the multiple token sets based on the comparison.

[0115] In box 940, operation 900 continues to output the indication of the selected set of tokens to the first generative model.

[0116] In some respects, the multiple token sets are represented as a tree data structure. Within this tree data structure, the root node generally corresponds to the input query, and different paths through the tree data structure generally correspond to different token sets among the multiple token sets. In some respects, the depth of the tree data structure may correspond to the maximum number of tokens generated in a single round of passage through the first generative model, and the maximum size of the tree data structure may be set based on a computational complexity metric associated with generating the target token set by the second generative model.

[0117] In some respects, comparing the probability distribution associated with each of the plurality of token sets with the corresponding probability distribution generated by the second generative model for the respective token set includes generating the probability distribution for each corresponding token set based on a single-round pass of the second generative model. The single-round pass of the second generative model can be performed based on masked self-attention and positional encoding in a tree data structure.

[0118] In some aspects, operation 900 further includes using a second generative model to generate additional tokens based on the selected set of tokens. The additional tokens (or their indications) may be combined with the selected set of tokens (or as part of an indication of the selected set of tokens) and output to the first generative model.

[0119] In some respects, the first generative model can correspond to the draft model in the speculative decoding pipeline, and the second generative model can correspond to the target model in the speculative decoding pipeline.

[0120] Figure 10 The present disclosure describes example operation 1000, which can be performed by a computing device to infer a response to an input query using a generative artificial intelligence model and to verify the response to the input query. The computing device for performing operation 800 can be a device on which at least a draft model can be deployed, such as a smartphone, tablet computer, laptop computer, desktop computer, server, cloud computing instance hosted in a distributed computing environment, etc.

[0121] As explained, operation 1000 begins at box 1010, where the computing device generates a first plurality of token sets based on an input query and a first generative model. Generally, each of the first plurality of token sets may correspond to a candidate answer to the input query. For example, the input query may be received from the user of the computing device at a prompt where a text query can be entered, such as by converting natural language speech into a text representation of the input query.

[0122] In box 1020, operation 1000 continues, where the computing device outputs a first set of multiple tokens to a second generative model for verification (e.g., by another computing device).

[0123] In box 1030, operation 1000 continues, wherein the computing device speculatively generates a second plurality of token sets while waiting for an instruction to be received from the second generative model for a token set selected from the first plurality of token sets. Each token set in the second plurality of token sets generally corresponds to a second portion of the candidate response to the input query.

[0124] In box 1040, operation 1000 continues to receive instructions from the second generative model for the set of tokens selected from the plurality of token sets.

[0125] In box 1050, operation 1000 continues to output tokens associated with the selected token set from the second plurality of token sets to the second generative model for verification.

[0126] In box 1060, operation 1000 continues (e.g., to the user or application, to another model, etc.) to output the selected set of tokens as a response to the input query.

[0127] In some aspects, operation 1000 further includes receiving an indication of a second set of tokens associated with the selected set of tokens, chosen from a second plurality of token sets. The selected second set of tokens may be output as another part of the response to the input query (e.g., to a user or application, to another model, etc.).

[0128] In some respects, each of the first plurality of token sets comprises the group of tokens with the highest probability within the probability distribution associated with the first generative model on the token universe.

[0129] In some respects, each of the first plurality of token sets includes a token group selected based on the sum of probabilities associated with the tokens in that token group, which exceeds a threshold probability.

[0130] In some respects, the first plurality of token sets can be represented as a tree data structure. Within this tree data structure, the root node corresponds to the input query. Each path through this tree data structure corresponds to a token set from the first plurality of token sets.

[0131] In some respects, each of the first plurality of token sets is generated using a unique instance of the first generative model and unique parameters that serve as input to the unique instance of the first generative model.

[0132] In some aspects, operation 1000 further includes generating a refined subsequent token set based on the selected token set and a second plurality of token sets. The refined subsequent token set for verification can be output to a second generative model. While awaiting instruction on the selected second token set from the refined subsequent token set, a third plurality of token sets can be speculatively generated. In some aspects, the token sets in this plurality of subsequent token sets include several tokens padded to account for tokens in the selected token set that are less than a maximum number.

[0133] In some respects, the first generative model can correspond to the draft model in the speculative decoding pipeline. The second generative model can correspond to the target model in the same speculative decoding pipeline.

[0134] In some respects, the first generative model may have a probability distribution that approximates the probability distribution associated with the second generative model. This approximation may be a probability distribution falling within a threshold difference of the probability distribution associated with the second generative model, such that the probability distribution associated with the first generative model does not need to precisely match the probability distribution associated with the second generative model.

[0135] In some respects, the first generative model can be executed locally (e.g., on a local device or local system), and the second generative model can be a model remotely hosted (e.g., on a remote system or remote device).

[0136] Example processing system for speculative decoding in generative artificial intelligence models

[0137] Figure 11 An example processing system 1100 for generating responses to queries input into a generative artificial intelligence model based on group inference decoding is described, such as those described herein. Figures 8 to 10 As described.

[0138] Processing system 1100 includes a central processing unit (CPU) 1102, which in some examples may be a multi-core CPU. Instructions executed at CPU 1102 may be loaded, for example, from program memory associated with CPU 1102 or from a memory partition (e.g., memory 1124).

[0139] The processing system 1100 also includes additional processing components tailored for specific functions, such as a graphics processing unit (GPU) 1104, a digital signal processor (DSP) 1106, a neural processing unit (NPU) 1108, and a connectivity component 1112.

[0140] NPUs (such as the NPU 1108) are generally configured to implement dedicated circuitry for implementing control and arithmetic logic for executing machine learning algorithms, such as those for processing artificial neural networks (ANNs), deep neural networks (DNNs), random forests (RFs), etc. NPUs are sometimes alternatively referred to as neural signal processors (NSPs), tensor processing units (TPUs), neural network processors (NNPs), intelligent processing units (IPUs), vision processing units (VPUs), or graphics processing units.

[0141] NPUs (such as the NPU 1108) are configured to accelerate the execution of common machine learning tasks, such as image classification, machine translation, object detection, and various other predictive models. In some examples, multiple NPUs may be instantiated on a single chip (such as a system-on-a-chip (SoC)), while in other examples, such NPUs may be part of a dedicated neural network accelerator.

[0142] An NPU can be optimized for training or inference, or in some cases configured to balance performance between the two. For an NPU capable of performing both training and inference, these two tasks can generally still be performed independently.

[0143] NPUs designed to accelerate training are typically configured to speed up the optimization of new models. This involves taking an existing dataset (usually labeled or sublabeled), iterating over the dataset, and subsequently tuning model parameters (such as weights and biases) to improve model performance—a highly computationally intensive operation. Generally, optimization based on incorrect predictions involves backtracking through the layers of the model and determining gradients to reduce prediction errors.

[0144] NPUs designed to accelerate inference are typically configured to operate on the full model. Such NPUs can thus be configured to take new data segments as input and rapidly process those segments through an already trained model to generate model outputs (e.g., inference).

[0145] In some implementations, the NPU 1108 is part of one or more of the CPU 1102, GPU 1104, and / or DSP 1106. These can reside in the user equipment (UE) of a wireless communication system or on another computing device.

[0146] In some examples, connectivity component 1112 may include sub-components for, for example, third-generation (3G) connectivity, fourth-generation (4G) connectivity (e.g., Long Term Evolution (LTE)), fifth-generation (5G) connectivity (e.g., New Radio (NR)), Wi-Fi connectivity, Bluetooth connectivity, and other wireless data transmission standards. Connectivity component 1112 may be further coupled to one or more antennas 1114.

[0147] The processing system 1100 may also include one or more sensor processing units 1116 associated with any type of sensor, one or more image signal processors (ISPs) 1118 associated with any type of image sensor, and / or a navigation processor 1120, which may include satellite-based positioning system components (e.g., GPS or GLONASS) and inertial positioning system components.

[0148] The processing system 1100 may also include one or more input and / or output devices 1122, such as screens, touch-sensitive surfaces (including touch-sensitive displays), physical buttons, speakers, microphones, etc.

[0149] In some examples, one or more processors of the processing system 1100 may be based on the ARM or RISC-V instruction set.

[0150] The processing system 1100 also includes a memory 1124, which represents one or more static and / or dynamic memories, such as dynamic random access memory, flash-based static memory, etc. In this example, the memory 1124 includes computer-executable components that can be executed by one or more of the aforementioned processors of the processing system 1100.

[0151] Specifically, in this example, the memory 1124 includes a token set generation component 1124A.

[0152] The token selection receiving component 1124B, the token selection component 1026C, and the output generation component 1124D, the generative model 1124E, and optionally include a comparison component 1124F. The depicted components, and other undepicted components, can be configured to perform various aspects of the methods described herein.

[0153] Generally, the processing system 1100 and / or its components can be configured to perform the methods described herein.

[0154] Example Terms

[0155] Implementation details of each aspect of this disclosure are described in the following numbered clauses:

[0156] Clause 1: A processor-implemented method comprising: generating a plurality of token sets based on an input query and a first generative model, each of the plurality of token sets corresponding to a candidate response to the input query; outputting the plurality of token sets to a second generative model for verification; receiving from the second generative model an indication of a token set selected from the plurality of token sets based on the input query and the plurality of token sets; and outputting the selected token set as a response to the input query.

[0157] Clause 2: The method of Clause 1, wherein each of the plurality of token sets comprises the group of tokens with the highest probability in the probability distribution associated with the first generative model on the token universe.

[0158] Clause 3: The method of Clause 1 or 2, wherein each of the plurality of token sets comprises a token group selected based on the sum of probabilities associated with the tokens in the token group, the sum exceeding a threshold probability.

[0159] Clause 4: The method of any of Clauses 1 to 3, wherein: the plurality of token sets are represented as a tree data structure, the root node of which corresponds to the input query, and each path through the tree data structure corresponds to a token set from the plurality of token sets.

[0160] Clause 5: The method of Clause 4, wherein the depth of the tree data structure corresponds to the maximum number of tokens generated by a single round of passage through the first generative model.

[0161] Clause 6: The method of Clause 4 or 5, wherein the maximum size of the tree data structure is set based on a computational complexity metric associated with the generation of the target token set by the second generative model.

[0162] Clause 7: The method of any of Clauses 4 to 6 further includes: pruning the tree data structure based on the selected set of tokens; generating a plurality of subsequent token sets based on the pruned tree data structure and the input query; outputting the plurality of subsequent token sets to a second generative model for verification; and receiving from the second generative model an indication of a subsequent token set selected from the plurality of subsequent token sets based on the input query, the pruned tree data structure and the plurality of subsequent token sets, wherein outputting the selected token set as a response to the input query includes: outputting the selected token set and the selected subsequent token set as a response to the input query.

[0163] Clause 8: The method of any of Clauses 1 to 7, wherein each of the plurality of token sets is generated using a unique instance of the first generative model and a unique parameter as input to the unique instance of the first generative model.

[0164] Clause 9: The method of any of Clauses 1 to 8 further includes: generating a plurality of subsequent token sets based on the input query and the plurality of token sets when the second generative model validates the plurality of token sets; generating a refined subsequent token set based on the plurality of subsequent token sets and the selected token set; and outputting the refined subsequent token set to the second generative model for validation.

[0165] Clause 10: The method of Clause 9, wherein the token set in the subsequent plurality of token sets includes a number of tokens padded to account for tokens in the selected token set that are less than the maximum number.

[0166] Clause 11: The method of any of Clauses 1 to 10 further includes: receiving a token generated by a second generative model based on a selected set of tokens; and outputting the received token as an additional token following the selected set of tokens.

[0167] Clause 12: The method of any of Clauses 1 to 11, wherein: the first generative model corresponds to the draft model in the speculative decoding pipeline, and the second generative model corresponds to the target model in the speculative decoding pipeline.

[0168] Clause 13: The method of Clause 12, wherein the draft model includes a model trained to have a probability distribution that approximates the corresponding probability distribution of the target model.

[0169] Clause 14: The method of any of Clauses 1 to 13, wherein: the first generative model includes a model executed on a local system, and the second generative model includes a model executed on a remote system.

[0170] Clause 15: A processor-implemented method comprising: receiving an input query and a plurality of token sets generated by a first generative model, each of the plurality of token sets corresponding to a corresponding candidate response to the input query; comparing a probability distribution associated with each of the plurality of token sets with a corresponding probability distribution generated by a second generative model for the corresponding token set; selecting a token set from the plurality of token sets based on the comparison; and outputting an indication of the selected token set to the first generative model.

[0171] Clause 16: The method of Clause 15, wherein: the plurality of token sets are represented as a tree data structure, the root node of which corresponds to the input query, and each path through the tree data structure corresponds to a token set from the plurality of token sets.

[0172] Clause 17: The method of Clause 15 or 16, wherein comparing the probability distribution associated with each of the plurality of token sets with the corresponding probability distribution generated by the second generative model for the corresponding token set includes: generating the probability distribution of each corresponding token set based on a single round of passage of the second generative model.

[0173] Item 18: The method of Item 17, wherein the single-round pass of the second generative model is performed based on masked self-attention and positional encoding in a tree data structure.

[0174] Clause 19: The method of any of Clauses 15 to 18 further includes: generating an additional token based on the selected set of tokens using a second generative model; and outputting the additional token to the first generative model.

[0175] Clause 20: The method of any of Clauses 15 to 19, wherein: a first generative model corresponds to a draft model in the speculative decoding pipeline, and a second generative model corresponds to a target model in the speculative decoding pipeline.

[0176] Clause 21: A processor-implemented method comprising: generating a first plurality of token sets based on an input query and a first generative model, each of the first plurality of token sets corresponding to a first portion of a candidate response to the input query; outputting the plurality of token sets to a second generative model for verification; speculatively generating a second plurality of token sets while awaiting an indication from the second generative model for selecting a token set from the first plurality of token sets, each of the second plurality of token sets corresponding to a second portion of a candidate response to the input query; receiving an indication from the second generative model for selecting a token set from the first plurality of token sets; outputting tokens from the second plurality of token sets associated with the selected token set to the second generative model for verification; and outputting the selected token set as a response to the input query.

[0177] Clause 22: The method of Clause 21 further includes: receiving an indication of a second set of tokens associated with the selected set of tokens selected from a second plurality of token sets; and outputting the selected second set of tokens as another part of the response to the input query.

[0178] Clause 23: The method of Clause 21 or 22, wherein each of the first plurality of token sets comprises the group of tokens with the highest probability in the probability distribution associated with the first generative model on the token universe.

[0179] Clause 24: The method of any of Clauses 21 to 23, wherein each of the first plurality of token sets comprises a token group selected based on the sum of probabilities associated with the tokens in the token group, the sum exceeding a threshold probability.

[0180] Clause 25: The method of any of Clauses 21 to 24, wherein: the first plurality of token sets are represented as a tree data structure, the root node of which corresponds to the input query, and each path through the tree data structure corresponds to a token set from the first plurality of token sets.

[0181] Clause 26: The method of any of Clauses 21 to 25, wherein each of the first plurality of token sets is generated using a unique instance of the first generative model and a unique parameter as input to the unique instance of the first generative model.

[0182] Clause 27: The method of any of Clauses 21 to 26 further includes: generating a refined subsequent set of tokens based on the selected set of tokens and a second plurality of token sets; outputting the refined subsequent set of tokens to a second generative model for verification; and speculatively generating a third plurality of token sets while waiting for an instruction to be received from the second generative model for a second set of tokens selected from the refined subsequent set of tokens.

[0183] Clause 28: The method of Clause 27, wherein the token set in the subsequent plurality of token sets includes a number of tokens padded to account for a number of tokens less than the maximum number in the selected token set.

[0184] Clause 29: The method of any of Clauses 21 to 28, wherein: the first generative model corresponds to the draft model in the speculative decoding pipeline, and the second generative model corresponds to the target model in the speculative decoding pipeline.

[0185] Clause 30: The method of Clause 29, wherein the draft model includes a model trained to have a probability distribution that approximates the corresponding probability distribution of the target model.

[0186] Clause 31: The method of any of Clauses 21 to 30, wherein: the first generative model includes a model executed on a local system, and the second generative model includes a model executed on a remote system.

[0187] Clause 32: A processing system comprising: a memory storing executable instructions; and one or more processors configured to execute the executable instructions to cause the processing system to perform a method as described in any of Clauses 1 to 31.

[0188] Clause 33: A processing system comprising: means for performing a method as described in any of Clauses 1 to 31.

[0189] Clause 34: A computer-readable medium having instructions stored thereon, which, when executed by one or more processors, cause the one or more processors to perform any of the methods described in Clauses 1 to 31.

[0190] Additional considerations

[0191] The foregoing description is provided to enable any person skilled in the art to practice the various aspects described herein. The examples discussed herein are not intended to limit the scope, applicability, or aspects set forth in the claims. Various modifications to these aspects will be readily apparent to those skilled in the art, and the general principles defined herein may be applied to other aspects. For example, changes may be made to the function and arrangement of the elements in discussion without departing from the scope of this disclosure. Various procedures or components may be appropriately omitted, substituted, or added to various examples. For example, the described methods may be performed in a different order than described, and various steps may be added, omitted, or combined. Moreover, features described with reference to some examples may be combined in others. For example, any number of aspects set forth herein may be used to implement an apparatus or practice a method. Furthermore, the scope of this disclosure is intended to cover such apparatuses or methods practiced using other structures, functionalities, or structures and functionalities that supplement or differ from the aspects of this disclosure set forth herein. It should be understood that any aspect of this disclosure disclosed herein may be implemented by one or more elements of the claims.

[0192] As used herein, the term “exemplary” means “serving as an example, instance, or illustration.” Any aspect described herein as “exemplary” should not be construed as superior to or better than the others.

[0193] As used herein, the phrase “at least one of” a list of items refers to any combination of those items, including a single member. As an example, “at least one of a, b, or c” is intended to cover: a, b, c, ab, ac, bc, and abc, as well as any combination with multiple identical elements (e.g., aa, aaa, aab, aac, abb, acc, bb, bbb, bbb, bbc, cc, and ccc, or any other ordering of a, b, and c).

[0194] As used herein, the term "determine" encompasses a wide variety of actions. For example, "determine" can include calculation, computation, processing, derivation, research, searching (e.g., looking in a table, database, or other data structure), ascertaining, and similar actions. Furthermore, "determine" can include receiving (e.g., receiving information), accessing (e.g., accessing data in memory), and similar actions. Likewise, "determine" can also include parsing, selecting, choosing, building, and similar actions.

[0195] The methods disclosed herein include one or more steps or actions for implementing the method. These method steps and / or actions may be interchanged without departing from the scope of the claims. In other words, unless a specific order of steps or actions is specified, the order and / or use of specific steps and / or actions may be modified without departing from the scope of the claims. Furthermore, the various operations of the above methods can be performed by any suitable means capable of performing the corresponding functions. These means may include various hardware and / or software components and / or modules, including but not limited to circuits, application-specific integrated circuits (ASICs), or processors. Generally, where operations illustrated in the drawings are present, these operations may have corresponding paired means with similar numbers plus functional components.

[0196] The following claims are not intended to be limited to the aspects shown herein, but should be granted the full scope consistent with the language of the claims. Within the claims, references to singular elements are not intended to mean “one and only one” (unless specifically stated so), but rather “one or more.” Unless specifically stated otherwise, the term “some / a” refers to one or more. No element of the claims should be interpreted in accordance with the provisions of 35 U.S.SC §112(f) unless the element is expressly stated using the phrase “means for…” or, in the case of a method claim, the element is stated using the phrase “steps for…”. Elements of all aspects described throughout this disclosure that are now or hereafter known to a person skilled in the art are expressly incorporated herein by reference and are intended to be covered by the claims. Furthermore, nothing disclosed herein is intended to be donated to the public, whether or not such disclosure is expressly stated in the claims.

Claims

1. A processing system, comprising: A memory, on which executable instructions are stored; as well as One or more processors, the one or more processors being configured to execute the executable instructions to cause the processing system to: Multiple token sets are generated based on the input query and a first generative model, each of the multiple token sets corresponding to a candidate response to the input query; The multiple token sets are output to the second generative model for verification; Based on the input query and the plurality of token sets, the second generative model receives an indication of the token set to be selected from the plurality of token sets; as well as The selected set of tokens is output as a response to the input query.

2. The processing system of claim 1, wherein each of the plurality of token sets comprises a group of tokens with the highest probability in the probability distribution associated with the first generative model on the token universe.

3. The processing system of claim 1, wherein each of the plurality of token sets comprises a token group selected based on the sum of probabilities associated with tokens in the token group, the sum exceeding a threshold probability.

4. The processing system as described in claim 1, wherein: The multiple token sets are represented as a tree data structure. The root node of the tree data structure corresponds to the input query, and Each path in the tree data structure corresponds to a set of tokens from the plurality of token sets.

5. The processing system of claim 4, wherein the depth of the tree data structure corresponds to the maximum number of tokens generated in a single round of passage through the first generative model.

6. The processing system of claim 4, wherein the maximum size of the tree data structure is set based on a computational complexity metric associated with the generation of the target token set by the second generative model.

7. The processing system of claim 4, wherein the one or more processors are further configured to cause the processing system to: The tree data structure is pruned based on the selected set of tokens; Multiple subsequent token sets are generated based on the pruned tree data structure and the input query; The subsequent sets of tokens are output to the second generative model for verification. as well as Based on the input query, the pruned tree data structure, and the subsequent multiple token sets, the second generative model receives an indication of a subsequent token set selected from the subsequent multiple token sets, wherein outputting the selected token set as a response to the input query includes: outputting the selected token set and the selected subsequent token set as a response to the input query.

8. The processing system of claim 1, wherein each of the plurality of token sets is generated using a unique instance of the first generative model and a unique parameter as input to the unique instance of the first generative model.

9. The processing system of claim 1, wherein the one or more processors are further configured to cause the processing system to: When the second generative model validates the multiple token sets, it generates multiple subsequent token sets based on the input query and the multiple token sets; A refined set of subsequent tokens is generated based on the aforementioned multiple sets of tokens and the selected set of tokens; as well as The refined set of subsequent tokens is output to the second generative model for verification.

10. The processing system of claim 9, wherein the token set in the subsequent plurality of token sets includes a plurality of tokens padded to account for tokens in the selected token set that are less than a maximum number.

11. The processing system of claim 1, wherein the one or more processors are further configured to cause the processing system to: Receive tokens generated by the second generative model based on the selected set of tokens; and Output the received tokens as additional tokens following the selected set of tokens.

12. The processing system of claim 1, wherein: The first generative model corresponds to the draft model in the speculative decoding pipeline, and The second generative model corresponds to the target model in the speculative decoding pipeline.

13. The processing system of claim 12, wherein the draft model comprises a model trained to have a probability distribution that approximates the corresponding probability distribution of the target model.

14. The processing system of claim 1, wherein: The first generative model includes models executed on the local system, and The second generative model includes models executed on remote systems.

15. A processing system, comprising: A memory, on which executable instructions are stored; as well as One or more processors, the one or more processors being configured to execute the executable instructions to cause the processing system to: The system receives an input query and multiple token sets from the first generative model on which it runs, each of the multiple token sets corresponding to a corresponding candidate response to the input query; The probability distribution associated with each of the plurality of token sets is compared with the corresponding probability distribution generated by the second generative model for the respective token set; Select a token set from the plurality of token sets based on the comparison; as well as Output an indication of the selected token set to the first generative model.

16. The processing system of claim 15, wherein: The multiple token sets are represented as a tree data structure. The root node of the tree data structure corresponds to the input query, and Each path in the tree data structure corresponds to a set of tokens from the plurality of token sets.

17. The processing system of claim 15, wherein, in order to compare the probability distribution associated with each of the plurality of token sets with the corresponding probability distribution generated by the second generative model for the corresponding token set, the one or more processors are configured to cause the processing system to generate the probability distribution for each corresponding token set based on a single pass of the second generative model.

18. The processing system of claim 17, wherein the single-round passage of the second generative model is performed based on masked self-attention and positional encoding in a tree data structure.

19. The processing system of claim 15, wherein the one or more processors are further configured to cause the processing system to: The second generative model is used to generate additional tokens based on the selected token set; and The additional token is output to the first generative model.

20. The processing system of claim 15, wherein: The first generative model corresponds to the draft model in the speculative decoding pipeline, and The second generative model corresponds to the target model in the speculative decoding pipeline.

21. A processor-implemented method, comprising: Multiple token sets are generated based on the input query and a first generative model, each of the multiple token sets corresponding to a candidate response to the input query; The multiple token sets are output to the second generative model for verification; Based on the input query and the plurality of token sets, the second generative model receives an indication of the token set to be selected from the plurality of token sets; as well as The selected set of tokens is output as a response to the input query.

22. The method of claim 21, wherein each of the plurality of token sets comprises a group of tokens that have the highest probability within a probability distribution associated with the first generative model in a token universe.

23. The method of claim 21, wherein each of the plurality of token sets comprises a token group selected based on the sum of probabilities associated with tokens in the token group, the sum exceeding a threshold probability.

24. The method of claim 21, wherein: The multiple token sets are represented as a tree data structure. The root node of the tree data structure corresponds to the input query, and Each path in the tree data structure corresponds to a set of tokens from the plurality of token sets.

25. The method of claim 24, wherein the depth of the tree data structure corresponds to the maximum number of tokens generated in a single round of passage through the first generative model.

26. The method of claim 24, wherein the maximum size of the tree data structure is set based on a computational complexity metric associated with the generation of the target token set by the second generative model.

27. The method of claim 24, further comprising: The tree data structure is pruned based on the selected set of tokens; Multiple subsequent token sets are generated based on the pruned tree data structure and the input query; The subsequent sets of tokens are output to the second generative model for verification. as well as Based on the input query, the pruned tree data structure, and the subsequent multiple token sets, the second generative model receives an indication of a subsequent token set selected from the subsequent multiple token sets, wherein outputting the selected token set as a response to the input query includes: outputting the selected token set and the selected subsequent token set as the response to the input query.

28. The method of claim 21, wherein each of the plurality of token sets is generated using a unique instance of the first generative model and a unique parameter as input to the unique instance of the first generative model.

29. The method of claim 21, further comprising: When the second generative model validates the multiple token sets, it generates multiple subsequent token sets based on the input query and the multiple token sets; A refined set of subsequent tokens is generated based on the aforementioned multiple sets of tokens and the selected set of tokens; as well as The refined set of subsequent tokens is output to the second generative model for verification.

30. The method of claim 29, wherein the token set in the subsequent plurality of token sets includes a plurality of tokens padded to account for tokens in the selected token set that are less than a maximum number.

31. The method of claim 21, further comprising: Receive tokens generated by the second generative model based on the selected set of tokens; as well as Output the received tokens as additional tokens following the selected set of tokens.

32. The method of claim 21, wherein: The first generative model corresponds to the draft model in the speculative decoding pipeline, and The second generative model corresponds to the target model in the speculative decoding pipeline.

33. The method of claim 32, wherein the draft model comprises a model trained to have a probability distribution that approximates the corresponding probability distribution of the target model.

34. The method of claim 21, wherein: The first generative model includes models executed on the local system, and The second generative model includes models executed on remote systems.

35. A processor-implemented method, comprising: The device running the first generative model receives an input query and a plurality of token sets, each of the plurality of token sets corresponding to a corresponding candidate response to the input query; The probability distribution associated with each of the plurality of token sets is compared with the corresponding probability distribution generated by the second generative model for the respective token set; Select a token set from the plurality of token sets based on the comparison; as well as Output an indication of the selected token set to the first generative model.

36. The method of claim 35, wherein: The multiple token sets are represented as a tree data structure. The root node of the tree data structure corresponds to the input query, and Each path in the tree data structure corresponds to a set of tokens from the plurality of token sets.

37. The method of claim 35, wherein comparing the probability distribution associated with each corresponding token set in the plurality of token sets with the corresponding probability distribution generated by the second generative model for the corresponding token set comprises: The probability distribution of each corresponding token set is generated based on the single-round passage of the second generative model.

38. The method of claim 37, wherein the single-round passage of the second generative model is performed based on masked self-attention and positional encoding in a tree data structure.

39. The method of claim 35, further comprising: The second generative model is used to generate additional tokens based on the selected set of tokens; as well as The additional token is output to the first generative model.

40. The method of claim 35, wherein: The first generative model corresponds to the draft model in the speculative decoding pipeline, and the second generative model corresponds to the target model in the speculative decoding pipeline.