PROCESSING SYSTEMS AND METHODS IMPLEMENTED BY PROCESSORS

Speculative decoding techniques using a preliminary model to generate token sets verified by a target model enhance the efficiency and throughput of generative AI models, addressing resource limitations and computational costs.

BR112025019648A2Pending Publication Date: 2026-07-28QUALCOMM INC
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
BR112025019648
Authority / Receiving Office
BR · BR
Patent Type
Applications
Current Assignee / Owner
Priority Date
2023-10-02
Filing Date
2024-01-31
Publication Date
2026-07-28

Smart Images

  • Figure 00000000_0000_ABST
    Figure 00000000_0000_ABST
Patent Text Reader

Abstract

Certain aspects of the present disclosure provide techniques and apparatus for generating a response to a query input into a generative artificial intelligence model. The method generally includes generating, based on an input query and a first generative model, a plurality of sets of tokens, each set of tokens in the plurality of sets of tokens corresponding to a candidate response to the input query; outputting, to a second generative model, the plurality of sets of tokens for verification; receiving, from the second generative model, an indication of a selected set of tokens from the plurality of sets of tokens based on the input query and the plurality of sets of tokens; and outputting the selected set of tokens as a response to the input query.
Need to check novelty before this filing date? Find Prior Art

Description

1 / 66 “SPECULATIVE DECODING IN AUTOREGRESSIVE MODELS OF GENERATIVE ARTIFICIAL INTELLIGENCE CROSS-REFERENCE TO RELATED DEPOSIT REQUESTS

[0001] This application claims priority to U.S. Patent Application Serial No. 18 / 479,659, entitled Speculative Decoding in Autoregressive Generative Artificial Intelligence Models, filed October 2, 2023, which claims priority to and benefit of U.S. Provisional Patent Application Serial No. 63 / 454,605, entitled Speculative Decoding in Autoregressive Generative Artificial Intelligence Models, filed March 24, 2023, both of which are assigned to the assignee of this document and are incorporated herein by reference in their entirety. INTRODUCTION

[0002] Aspects of this disclosure relate to generative artificial intelligence models and, more specifically, to speculative decoding in generative artificial intelligence models.

[0003] Generative artificial intelligence models can be used in various environments to generate a response to an input query. For example, generative artificial intelligence models can be used in chatbot applications where large language models (LLMs) are used to generate a response, or at least a replica, to an input query. Other examples where generative artificial intelligence models can be used include stable diffusion, where a model generates an image from an input text description of the Petition 870250082970, dated 09 / 15 / 2025, page 75 / 168 2 / 66 content of the desired image, and decision transformers, in which future actions are predicted based on sequences of previous actions within a given environment.

[0004] In general, generating a response to a query using generative artificial intelligence models can be computationally expensive. For example, in a chatbot deployment where a large-scale language model is used to generate a response to a query formatted as a text query, a response to the query can be generated using a pass through the large-scale language model for each token (e.g., word or part of a word) generated as part of the response. The output of each pass can be a probability distribution on a set of tokens (e.g., words or parts of words) from which the next token (e.g., word or part of a word) can be selected, either by sampling or based on maximum likelihood, for example.Since a pass-through of a large-scale language model is used to generate each word (or token(s)) in a query response, the computational cost can be modeled as the product of the number of words included in the response and the cost of computational resources (e.g., in terms of processing power, memory bandwidth, and / or other computing resources used) to perform a pass-through of the large-scale language model, which generally increases as the number of parameters within the large-scale language model increases. BRIEF SUMMARY

[0005] Certain aspects of the present disclosure Petition 870250082970, dated 09 / 15 / 2025, page 76 / 168 3 / 66 provides a method for generating a response to an input query using a generative artificial intelligence model. In general, the method involves generating, based on an input query and a first generative model, a plurality of token sets, where each set of tokens in the plurality of token sets corresponds to a candidate response to the input query. The plurality of token sets is sent to a second generative model for verification. An indication of a selected token set from among the plurality of token sets is received from the second generative model based on the input query and the plurality of token sets. The selected token set is then issued as a response to the input query.

[0006] Certain aspects of this disclosure provide a method for verifying a response to an input query generated using a generative artificial intelligence model. In general, the method involves receiving an input query and a plurality of token sets generated by a first generative model, where each token set in the plurality of token sets corresponds to a candidate response to the input query. A probability distribution associated with each respective token set in the plurality of token sets is compared to a corresponding probability distribution generated by a second generative model for the respective token set. A token set is selected based on the comparison, and the selected token set is output to the first generative model.

[0007] Certain aspects of the present disclosure Petition 870250082970, dated 09 / 15 / 2025, p. 77 / 168 4 / 66 provides a method for parallel speculative generation of a response to an input query using generative artificial intelligence models. In general, the method involves generating, based on an input query and a first generative model, a first plurality of token sets, where each token set in the first plurality of token sets corresponds to a first portion of a candidate response to the input query. The first plurality of token sets is sent to a second generative model for verification. While waiting to receive, from the second generative model, an indication of a token set selected from the first plurality of token sets, a second plurality of token sets is speculatively generated, where each token set in the second plurality of token sets corresponds to a second portion of the candidate response to the input query.The indication of a set of tokens selected from the first plurality of token sets is received from the second generative model. The tokens from the second plurality of token sets associated with the selected set of tokens are emitted to the second generative model for verification, and the selected set of tokens is emitted as a response to the input query.

[0008] Other aspects provide processing systems configured to perform the aforementioned methods, as well as those described in the present invention; non-transient computer-readable means comprising instructions that, when executed Petition 870250082970, dated 09 / 15 / 2025, page 78 / 168 5 / 66 by one or more processors of a processing system, causing the processing system to perform the methods previously mentioned, as well as those described in the present invention; a computer program product embedded in a computer-readable storage medium comprising code to perform the methods previously mentioned, as well as those additionally described in the present invention; and a processing system comprising means to perform the methods previously mentioned, as well as those additionally described in the present invention.

[0009] The following description and related drawings set out in detail certain illustrative attributes of one or more aspects. BRIEF DESCRIPTION OF THE DRAWINGS

[0010] The attached figures depict only certain aspects of this disclosure and, therefore, should not be considered limiting the scope of this disclosure.

[0011] Figure 1 illustrates an example of speculative group decoding in generative models, according to aspects of the present disclosure.

[0012] Figure 2 illustrates an example relationship between a preliminary model and a target model used in group speculative decoding in generative models, according to aspects of the present disclosure.

[0013] Figure 3 illustrates an example of parallel speculative decoding in generative models, according to aspects of the present disclosure.

[0014] Figure 4 illustrates an example pipeline for hybrid speculative decoding in models. Petition 870250082970, dated 09 / 15 / 2025, page 79 / 168 6 / 66 generative, according to aspects of this disclosure.

[0015] Figure 5 illustrates an example of hybrid speculative decoding in generative models, according to aspects of the present disclosure.

[0016] Figure 6 illustrates an example of hybrid parallel speculative decoding in generative models, according to aspects of the present disclosure.

[0017] Figure 7 illustrates an example of hybrid group speculative decoding in generative models, according to aspects of the present disclosure.

[0018] Figure 8 illustrates example operations for generating a response to an input query using generative artificial intelligence models, in accordance with aspects of this disclosure.

[0019] Figure 9 illustrates example operations for verifying a response to an input query using generative artificial intelligence models, in accordance with aspects of this disclosure.

[0020] Figure 10 illustrates example operations for speculatively generating a response to an input query and verifying the response to the input query using generative artificial intelligence models, in accordance with aspects of this disclosure.

[0021] Figure 11 depicts a sample processing system configured to perform various aspects of the present disclosure.

[0022] To facilitate understanding, identical reference numbers have been used, where possible, to designate identical elements that are common to the drawings. It is contemplated that the elements and attributes of an aspect Petition 870250082970, dated 09 / 15 / 2025, p. 80 / 168 7 / 66 can be beneficially incorporated into other aspects without further mention. DETAILED DESCRIPTION

[0023] Aspects of the present disclosure provide computer-readable devices, methods, processing systems and means for efficiently generating responses to input queries using generative artificial intelligence models.

[0024] In general, generative artificial intelligence models generate a response to a query input in the model. For example, a large-scale language model (LLM) deployed in a chatbot might generate a response to a query using multiple passes through the large-scale language model, with each successive pass being based on the query and the tokens (or words, or parts thereof) generated using previous passes through the large-scale language model. In general, these large-scale language models can include millions, or even billions, of weights or parameters in the model.Due to the size of these models and the operations performed on each token to predict which token should be generated next in response to a query, as well as the previously generated tokens, it may not be practical, or even possible, to deploy large-scale language models on a variety of devices that may have limited memory, storage, and / or processing capacity compared to the cloud computing instances on which large-scale language models typically operate. Additionally, there is the computational complexity involved in generating a response to a query provided as input. Petition 870250082970, dated 09 / 15 / 2025, page 81 / 168 Replacing an 8 / 66 model can involve significant energy expenditure, processing time, memory usage, and / or use of other resources, which may prevent computing resources from being used for other tasks.

[0025] To improve the efficiency and throughput of large-scale language models, speculative decoding techniques allow a smaller language model, sometimes known as a preliminary large-scale language model (also called a preliminary model), to run with a larger language model, sometimes known as a target large-scale language model (also known as a target model). In this case, the preliminary model can speculatively generate additional tokens and probabilities used to sample these additional tokens based on a current set of accepted tokens. The target model can generate tokens based on the tokens generated by the preliminary model.To generate a result, the target model can perform token-by-token rejection sampling in order to accept or reject tokens generated by the preliminary model, so that the preliminary model and the target model have similar probability distributions.

[0026] In some respects, the preliminary model may be a pruned version of the target model, chosen so that the preliminary model and the target model have similar probability distributions. In other respects, the preliminary model may be a smaller version of the target model (for example, trained on millions of tokens, rather than hundreds of millions or even billions of tokens). Petition 870250082970, dated 09 / 15 / 2025, page 82 / 168 9 / 66

[0027] Aspects of the present disclosure provide techniques for generating responses to a query input in a large-scale, group-based language model using speculative decoding techniques. In general, the preliminary model can generate one or more sets of tokens as candidate responses to the query. The target model can, in turn, perform set-by-set rejection sampling. Tokens within a group can be selected based on conditional probabilities of the tokens within the group. By performing set-by-set rejection sampling, aspects of the present disclosure can retain a close relationship between the probability distributions within the preliminary and target models, while increasing the throughput of the preliminary and target models (e.g., the number of tokens generated per second) compared to the preliminary and target models configured to generate token-by-token responses.

[0028] Aspects of the present disclosure provide techniques for generating responses to a query input in a large-scale language model using speculative decoding techniques in which a preliminary model and a target model operate in parallel, also referred to in the present invention as hybrid speculative decoding. In hybrid speculative decoding techniques, a preliminary model can speculatively generate one or more tokens while previously speculatively generated tokens are verified by the target model. The preliminary model can be run on an edge device and the target model can be run on a server (e.g., Petition 870250082970, dated 09 / 15 / 2025, page 83 / 168 10 / 66 locally accessible by the edge device, hosted in a cloud computing environment, etc.). By running both the preliminary and target models in parallel, the rate at which tokens are generated can be maximized or at least increased. Additionally, in situations where the edge device is disconnected from the server, the edge device can continue to generate tokens in order to generate a response to an incoming query. Speculative Decoding in Generative Artificial Intelligence Models

[0029] In general, autoregressive token generation (for example, in large-scale language models) can take historical tokens as an input in order to generate an output. That is, autoregressive token generation can be represented by the expression: Xt ~ p(x|x0,x1,^,xt_1) ^ xt+1~ p(x|x0,x1,^,xt_1,xt) where Xt represents a sequence of tokens generated at time t, having a conditional probability p conditioned on the selection of tokens x0a xt_1, and xt+1 represents a sequence of tokens generated at time t + 1, having a conditional probability p conditioned on the selection of tokens x0a xt. In general, a single token can be generated each time an autoregressive model is run, meaning that N inferences can be made to generate a sequence of N tokens. As discussed above, speculative decoding techniques can be used to accelerate token generation by using a preliminary model, smaller in size than the target model, that generates tokens speculatively, with the target model being used to verify the tokens generated (speculatively) by the model. Petition 870250082970, dated 09 / 15 / 2025, page 84 / 168 11 / 66 preliminary.

[0030] In a speculative decoding pipeline, the preliminary model can speculatively generate n tokens autoregressively, according to the expression: draft draft draft ( i \ draft draft draft xt+1 ~ Pt =P'~-'\X\X0,...Xt), Xt+2 ~ Pt+1 , -, *t+n ~ Pt+n-1 Equation (1) where t corresponds to a point in time and prrattcorresponds to the conditional probability distribution associated with a token x at time t conditioned on the selection of token x0a xt-1.

[0031] The target model takes the n generated tokens and processes them in parallel to generate probability distributions for each of the n tokens, according to the expression: rtargettarget target! _ r targetív\vv vdraft draft}] [Pt -Pt+1 ' — 'Pt+n 1 = 1 / 7 (xr0' —xt>xt+1 '—'xt+k Jik=l,2,...,n Equation (2) where k corresponds to a token index in relation to the n tokens generated.

[0032] The target model can then verify the tokens generated by the preliminary model by comparing the distributions of the preliminary model and the target model to determine whether a token is accepted or rejected. A given token x.^1 can be accepted when f (ρ^'ϡ^) <apara alguma função f e algum limiar a (também conhecido como taxa de aceitação). Caso contrário, o token pode ser rejeitado. O token final pode, então, ser gerado na primeira posição de rejeição ou na última posição n com base em alguma função 3(}^0^^1^961.. Petition 870250082970, dated 09 / 15 / 2025, page 85 / 168 12 / 66

[0033] Speculative decoding, with an acceptance rate of a, can result in cost reductions compared to using a single autoregressive model to iteratively generate tokens. The inference cost savings, compared to iterative token generation, can be represented by the expression: CAR, N = NCtarget > CSD= nNC^rafL +^target a + 1 Equation (3)

[0034] Consider the example, for N = 1000, ctarget = 10, Cdra / t = 1, n = 4, a = 3 where N corresponds to a number of tokens, CAR corresponds to a computational cost using an acceptance rate of a, Ctarget corresponds to a computational cost of generating a set of tokens using the target model, Cdra^t corresponds to a computational cost of generating a set of tokens using the preliminary model, and en corresponds to a number of tokens generated speculatively, generated through a single pass through an autoregressive model. In such an example, speculative decoding can result in a 35% reduction in computational cost compared to iterative autoregressive token generation alone.

[0035] However, token-by-token speculative decoding, as discussed, can impose limits on the rate at which tokens are generated, as a first token may be sampled individually by a preliminary model, and then verified by a target model before the next token is sampled by the preliminary model and verified by the target model. That is, generating a response to an input query using decoding techniques. Petition 870250082970, dated 09 / 15 / 2025, page 86 / 168 13 / 66 Speculative token-by-token trading may involve executing the preliminary model and the target model for each token generated as part of a response to an input query, which may use significant amounts of computational resources (e.g., processor time, memory, memory bandwidth, etc.) in order to generate the response. Speculative Decoding in Example Groups in Generative Artificial Intelligence Models

[0036] Figure 1 illustrates an example 100 of group speculative decoding (GSD) in generative artificial intelligence models, according to aspects of the present disclosure.

[0037] As illustrated, a preliminary model and a target model can be used together (or, in other words, together) with each other to perform speculative token group decoding to generate a response to an incoming query for processing by one or more generative artificial intelligence models. As discussed in more detail below, speculative token group decoding can allow multiple sets of tokens to be speculatively generated by a preliminary model for verification by a target model.Since multiple sets of tokens can be generated by a preliminary model for verification by a target model, speculative group decoding can increase the token generation rate of generative artificial intelligence models by generating multiple sets of tokens that can be accepted as a correct answer, as larger numbers of token sets can increase the probability that at least one set includes an "or". Petition 870250082970, dated 09 / 15 / 2025, page 87 / 168 14 / 66 more tokens will be accepted as a response.

[0038] In general, the preliminary model forms a sampled group of tokens including one or more high-probability node groups from a probability distribution for an output in a set of potential tokens, given an input from the received query. The high-probability node groups may include groups selected based on various techniques, such as topk selection (e.g., selection of the k tokens having the highest probabilities within a probability distribution), core-based selection (e.g., selection based on a sum of probabilities that meet a threshold probability), or similar. Tokens not included in one or more groups may be considered unitary set groups (e.g., groups including a single member).When selecting groups of token candidates, the preliminary model can sample tokens based on probability distribution; however, each group can be treated as a single selection with an aggregate group probability (e.g., as a sum of individual token probabilities) so that the preliminary model matches, or at least approximates, the target model such that the calculated probability of each group is a sufficient answer to the query received. Clustering can achieve this match, or at least approximation, between the probabilities calculated by the preliminary and target models because, although individual token probabilities may not match well, the aggregates of probable tokens may have probabilities that are likely similar in the preliminary and target models. Petition 870250082970, dated 09 / 15 / 2025, page 88 / 168 15 / 66

[0039] When a sampled group of tokens is provided as input to the preliminary model in the next iteration of the preliminary model's execution, the tokens in the sampled group of tokens are provided as input at the sample location and treated independently. The result may be a tree data structure 110, with a prompt as a root node 111 of the tree data structure, and subsequent levels within the tree data structure 110 representing different tokens (or groups of tokens), matching each of the previously selected token combinations. At some point in time (for example, after generating a tree with a defined depth, corresponding to a maximum length of a sequence generated by the preliminary model), the preliminary model may output the generated tree data structure 110 to the target model for further processing.The 110-tree data structure can, in some respects, be output to the target model with groupings and selection probabilities generated by the preliminary model.

[0040] In some respects, the number of nodes at each level of the 110-tree data structure and the depth of the 110-tree data structure can be defined a priori. The number of nodes at each level of the tree can be defined globally, level by level, or in some other way. For example, the number of nodes at any level of the 110-tree data structure can be defined based on a branching factor at the immediately preceding level of the 110-tree data structure. For example, a branching factor of 2 for nodes at the nth level of the 110-tree data structure might result in the generation of Petition 870250082970, dated 09 / 15 / 2025, page 89 / 168 16 / 66 nodes (tokens) in the n+1 level of the 110-tree data structure for each node (token) in the nth level of the tree. Meanwhile, the depth of the 110-tree data structure can be defined based on a maximum number of tokens (e.g., words) that can be generated using any pass through the preliminary model. For example, if the preliminary model is configured to generate a sequence with a maximum length of 5 tokens during any speculative generation instance, then the tree depth could be 6 (to include, in the first level of the 110-tree data structure, the root node corresponding to the entry in the preliminary model).

[0041] The preliminary model can be configured to trigger token generation by the target model based on various complexity and / or performance criteria. For example, the preliminary model can trigger token generation by the target model based on a complexity or performance criterion associated with the size of the generated tree. In other examples, the preliminary model can be configured to trigger token generation by the target model based on a time criterion associated with an expected amount of time for the target model to generate a set of tokens against which the generated tree can be compared. In general, these complexity and / or performance criteria can establish an upper limit on the number of tokens generated by the preliminary model for verification by the target model.This upper limit can, in some respects, be based on the number of nodes in the tree data structure and can be influenced, for example, by a branching factor defined for different values. Petition 870250082970, dated 09 / 15 / 2025, pp. 90 / 168 17 / 66 levels of a tree data structure in which the sampled tokens are organized, a depth of the tree data structure, or similar. The worst-case computational load in the last round of speculative token generation can be configured to be limited by the memory bandwidth on the device where the preliminary model is running.

[0042] Based on the generated 110-tree data structure, groupings, and selection probabilities received from the preliminary model, the target model generates an output distribution (probability) for each partial path of the generated 110-tree data structure. Each partial path generally starts from the root node of the generated 110-tree data structure and can be a contiguous path through the generated 110-tree data structure starting from the root node and ending at any node in the generated 110-tree data structure. The target model can generate the output distribution for each partial path using a single pass through the target model by, for example, including all tree nodes in the generated 110-tree data structure as token inputs and performing masked self-attention and positional encodings for each partial path within the 110-tree data structure.Based on the output distribution for each node in the generated tree data structure 110, the target model can select a set of tokens 112 when traversing from the root node of the generated tree data structure 110. If the target model includes multiple tokens in the selected token set 112, a token can be selected from within the selected set. Petition 870250082970, dated 09 / 15 / 2025, page 91 / 168 18 / 66 group using normalized group token probabilities (e.g., token probabilities conditioned on group selection including a token) and the selected token can be redefined as the root node of the generated tree data structure 110. The process discussed in the present invention can be repeated, with the selected token defined as the root node of the tree data structure 110, in order to generate subsequent tokens to include in the selected token set 112.

[0043] If a group of tokens is rejected based on the probability distribution generated by the target model, a final group of tokens can be selected based on a group-level rejection sampling method. A final token can be selected from the final group of tokens selected, based on the normalized group token probabilities. If all groups of tokens are accepted by the target model, a final token can be sampled from a final tree leaf distribution (e.g., a probability distribution generated over the leaf nodes in the 110-tree data structure corresponding to a last accepted or verified token for any partial path through the 110-tree data structure), which may allow the target model to generate an additional token relative to the token(s) speculatively generated by the preliminary model.The selected set of 112 tokens can then be communicated back to the preliminary model to be used as an input for a subsequent round of token speculation and verification by the target model, using the techniques discussed in the present invention.

[0044] Figure 2 illustrates an example 200 of a Petition 870250082970, dated 09 / 15 / 2025, page 92 / 168 19 / 66 relationship between tokens generated by a preliminary model and a target model used in speculative group decoding in generative models, according to aspects of the present disclosure.

[0045] The tree generated by the preliminary model generally allows for diverse speculation in which many candidate token sets are provided by the preliminary model for verification by the target model, which can improve the acceptance rate, by the target model, of tokens generated by the preliminary model. As illustrated, the acceptance of token number 4 from the group of tokens generated by the preliminary model at time t = 1 210 leads to the acceptance of token number 1 from the group of tokens generated by the preliminary model at time t = 2 220. Subsequently, at time t = 3 230, two options may be present: token number 3 or token number 5. Token number 3 may be accepted as the sampled token, while token number 5 may be rejected as the sampled token, as illustrated.At time t = 4240, token 4 from the set of tokens generated speculatively based on the selection of token 3 at time t = 3230 can be accepted as the sampled token, based on the probability calculated for tokens 3 and 4 conditioned on the acceptance of token 3 at time t = 3230. Meanwhile, both tokens 4 and 5 from the set of tokens generated speculatively based on the selection of token 5 at time t = 3230 can be rejected, based on the rejection of token 5 at time t = 3230.

[0046] The target model can be fed a token tree as input, instead of a token sequence. For example, the target model can Petition 870250082970, dated 09 / 15 / 2025, page 93 / 168 20 / 66 being fed a tree including the illustrated options of token 3 or 5 being a token candidate at time t = 3230, the options of token 3 or 4 at time t = 4240 if token 3 is selected at time t = 3230, and the options of token 4 or 5 at time t = 4240 if token 5 is selected at time t = 3230. As illustrated, at time t = 5250, providing as input a token tree data structure (for example, as illustrated by the token sequence shown at times t = 1 to t = 4) to the target model can allow the target model to generate an output including five tokens—one more than the sequence(s) of four tokens included in the tree generated by the preliminary model—which can provide increased acceptance rates for the token tree and allow the generation of an additional token to be included in the response generated by the preliminary and target models. destination.In the example illustrated in Figure 2, the selected set of five tokens includes token 4 at time t = 1 210, token 1 at time t = 2 220, token 3 at time t = 3 230, token 4 at time t = 4 240 (conditioned on the selection of token 3 at time t = 3 230), and token 2 at time t = 5 250 (conditioned on the selection of token 4 at time t = 4 240).

[0047] In general, a new selection of a token group can be performed each time the preliminary model produces a new distribution for one or more subsequent tokens to be included in a response to a query received and processed by the preliminary model and the target model. As discussed, the use of token group selection can allow for a closer match between the probability distributions of the model. Petition 870250082970, dated 09 / 15 / 2025, page 94 / 168 21 / 66 preliminary and the probability distributions of the target model to minimize, or at least reduce, the possibility of group rejection. In some respects, the preliminary model may choose not to form any groups when there is a low level of uncertainty (e.g., when the next token selection has low entropy) to minimize, or at least reduce, the cost of computational resources.

[0048] In general, by forming token groups when the use of speculative group decoding (to generate responses to input queries using generative artificial intelligence models) is likely to have a high level of effectiveness, computational complexity can be reduced for both preliminary and target models, while maintaining a high level of throughput (e.g., the number of tokens generated by the preliminary and target models per second).In general, the firing of the target model can be affected by several considerations, such as a maximum tree size to keep the model's computational performance metrics within a limit, estimates of when longer speculation has diminishing returns (for example, where the computational cost of generating a larger tree of speculated tokens is unlikely to be worthwhile, given the possibility that the target model will reach later tokens in the tree), and other performance considerations.

[0049] In some respects, the preliminary model may correspond to the target model (for example, in terms of a probability distribution), but it may have Petition 870250082970, dated 09 / 15 / 2025, page 95 / 168 22 / 66 offers faster inference performance than the target model on the same hardware. In general, smaller models can generate many speculative tokens, but may have an increased possibility that the generated tokens will be rejected by the target model. Group speculative decoding, as discussed above, can address this increased possibility of token rejection, at the cost of increasing the computational cost for longer sequences.

[0050] In some respects, in the preliminary model, a temperature parameter — or a parameter that influences the possibility of the preliminary model selecting a token with a lower probability of being an accurate token for inclusion in a token set — can be adjusted or selected to improve the performance of speculative group decoding.

[0051] In some respects, the preliminary model may be subjected to fine-tuning to match (or at least approximate) the target model and maximize (or at least increase) the probability that the speculatively generated tokens, generated by the preliminary model, will be accepted as valid tokens by the target model.

[0052] In general, the performance of token generation for speculative group decoding can be increased compared to token-by-token speculative decoding. That is, for a given Kullback-Leibler (KL) divergence between the preliminary model and the target model, by measuring how the probability distribution of the preliminary model differs from the probability distribution of the target model (treating the target model as the distribution of Petition 870250082970, dated 09 / 15 / 2025, pp. 96 / 168 23 / 66 reference), the number of tokens generated for each target model round may be greater for speculative group decoding than for token-by-token speculative decoding. Different grouping strategies (e.g., group size, additional tokens, etc.) may have different computational complexity characteristics; thus, the selection of a grouping strategy may be based on a trade-off between computational complexity and performance, given the limiting parameters of read bandwidth for the primary and target models and hardware performance. Speculative Parallel Decoding of Examples in Generative Artificial Intelligence Models

[0053] Speculative group decoding, as discussed above, can allow the combination of speculative token generation by the preliminary model and token acceptance by the target model to generate a final set of tokens that corresponds to a distribution by the target model, while increasing the rate at which tokens are generated. Furthermore, aspects of this disclosure provide parallel speculative decoding, in which the preliminary model speculatively generates multiple sets of tokens in parallel, which can provide increased token generation rates while improving the computational efficiency involved in generating query responses using generative artificial intelligence models.

[0054] Figure 3 illustrates an example 300 of parallel speculative decoding (PSD) in generative artificial intelligence models, according to aspects of the present disclosure. Petition 870250082970, dated 09 / 15 / 2025, page 97 / 168 24 / 66

[0055] As illustrated, the preliminary model can run multiple instances of a speculative decoding process to generate a plurality of candidate token sequences. Each iteration can be run with different parameters, or seeds, for token sampling. Running N iterations with S tokens generated per iteration can result in the generation of N * S tokens in total, which can be provided to the target model for processing.

[0056] For example, as illustrated in example 300, the preliminary model can initially execute multiple instances of a speculative decoding process to generate a first candidate sequence 302 and a second candidate sequence 304 for a prompt input in the preliminary model. As illustrated, N = 2 and S = 4; thus, the preliminary model can generate two sequences with four tokens in each sequence, for a total of 8 tokens. It should be recognized that this is only an example, and any values ​​of N and S can be used to define the number of candidate sequences generated by a preliminary model and the number of tokens included in each sequence, respectively.

[0057] The target model uses the N * S tokens as input and processes the N candidate token sequences in one round through the target model. The result of the target model processing the N * S tokens can be a set of N potentially accepted sequences with varying lengths, depending on where the target model rejected a token in each of the N candidate token sequences. Several techniques can be used to select Petition 870250082970, dated 09 / 15 / 2025, pp. 98 / 168 25 / 66 the final sequence based on the N potentially accepted sequences. As a non-limiting example, a greedy technique might result in selecting the sequence with the longest length, with loops (e.g., between sequences that have the same length) arbitrarily broken, which could result in the highest number of tokens being selected each time the target model is run. In some cases, selecting the longest length can influence the final effective sampling distribution, which can cause the probability distribution for the target model to converge, at least to some degree, with the probability distribution for the preliminary model.

[0058] For example, as illustrated, the target model accepts two of the four tokens in the first candidate sequence 302 and accepts four of the four tokens in the second candidate sequence 304. Since more tokens are accepted from the second candidate sequence 304 than from the first candidate sequence 302, the greedy technique can result in the target model selecting the second candidate sequence 304 as the sequence on which the preliminary model generates the subsequent candidate sequences 312 and 314. No speculative token generation needs to be performed based on the first candidate sequence 302 in this case, since the first candidate sequence 302 had a smaller number of accepted tokens than the second candidate sequence 304. However, in other cases, other techniques can be implemented to select one of the candidate sequences, including a sequence that may not have the most accepted tokens.

[0059] Although Figure 3 illustrates a Petition 870250082970, dated 09 / 15 / 2025, pp. 99 / 168 26 / 66 Parallel speculative decoding based on token-by-token speculative decoding, it should be recognized that the N candidate token sequences can also, or alternatively, be generated based on the group speculative decoding techniques discussed above. For example, using the group speculative decoding techniques discussed above, multiple sets of tokens, generated by multiple instances of a preliminary model, can be generated for verification by the target model. In some respects, sets of tokens generated by an instance of a preliminary model can be organized into a tree data structure (as discussed above), resulting in the generation of multiple tree data structures, with each individual tree data structure corresponding to sets of tokens generated by a specific instance of a preliminary model running using instance-specific parameters.The target model can attempt to check tokens in each token set to determine the selected token set upon which subsequent speculative decoding should be performed. The longest path through each tree data structure can be identified as a sequence candidate by the target model. Using a greedy technique, the tokens in the sequence candidate with the longest path can be selected as the token set upon which subsequent speculative decoding should be performed, and can be emitted by the target model to the preliminary model for the preliminary model to use as the root node of a subsequent tree data structure (e.g., a data structure in...). Petition 870250082970, dated 09 / 15 / 2025, pages 100 / 168 27 / 66 tree corresponding to subsequent tokens generated as part of a response to an input query). Example Hybrid Speculative Decoding

[0060] In some cases, it may be possible to run generative artificial intelligence models using autoregressive inference on a hybrid system. Examples of such hybrid systems may include systems in which these generative models run on an edge device (e.g., a smartphone, a tablet computer, a laptop computer, a desktop computer, etc.) and a cloud server, on an edge device and a home server, on different home servers, or similar. Thus, in some respects, the preliminary model may run on a local device, while the target model may run on a remote device.Running generative artificial intelligence models in a hybrid system can accelerate the speed at which query responses are generated and can reduce the computational load on a server (for example, by offloading some computational processes to another device). Additionally, the accuracy of the target model can be preserved when the server on which a target model runs is available, and a large-scale language model (or another generative artificial intelligence model) can remain available (albeit with lower accuracy) on an edge device when the edge device is communicatively disconnected from the server.

[0061] In hybrid systems, however, cost reductions may be limited when the model Petition 870250082970, dated 09 / 15 / 2025, pp. 101 / 168 The preliminary 28 / 66 model and the target model are both hosted on a cloud system. Additionally, the sequential nature of the inference (e.g., involving speculatively generating a token using the preliminary model and then verifying the token using the target model) may limit the rate at which tokens are generated.

[0062] To further reduce the computational cost involved in generating a response to a query using generative artificial intelligence models, aspects of this disclosure provide several techniques for speculative decoding in a hybrid environment. In several respects, speculative decoding in a hybrid environment can be performed such that candidate token set generation using the preliminary model (e.g., on an edge device) and token verification using the target model (e.g., on a server) can be performed in parallel. The preliminary model can continue to generate subsequent preliminary token sets, assuming variations in the number of tokens accepted, where padding can be added to account for differences in the number of tokens accepted from a previous round of speculative token generation.The preliminary model and the target model can be executed continuously and substantially in parallel, which can accelerate the generation of query responses using generative artificial intelligence models.

[0063] Figure 4 illustrates an example pipeline 400 for hybrid speculative decoding (HSD) in generative models, according to aspects of the present disclosure. Petition 870250082970, dated 09 / 15 / 2025, pp. 102 / 168 29 / 66

[0064] As illustrated, to maximize, or at least increase, the throughput (e.g., the number of tokens generated per second), the preliminary 402 model running on a 403 edge device (which may be called a local device or a first device) can continuously generate batches of token candidates (or sets of tokens, corresponding to different responses of varying lengths to an input prompt, conditioned on any previously accepted tokens) and send these tokens to the target 404 model running on a 405 server (e.g., in a cloud computing environment) (which may be called a remote device or a second device). The 403 edge device can receive, from the 405 server, a set of accepted tokens from the target 404 model. The preliminary 402 model can prune the sample tree to fit the set of accepted tokens.In some respects, when the set of accepted tokens is the null set (for example, when the target model does not accept any tokens), the preliminary 402 model may backtrack to the last token accepted by the target 404 model and restart the generation of speculative tokens from the last accepted token. For example, in order to backtrack, the preliminary 402 model may prune a generated tree of sample tokens back to the last accepted token and restart speculative generation based on the pruned tree.

[0065] In example pipeline 400, the preliminary model 402 running on edge device 403 can generate a first set of speculative tokens 410 and issue the first set of speculative tokens to the target model 404 running on server 405 to Petition 870250082970, dated 09 / 15 / 2025, pp. 103 / 168 30 / 66 verification. In some respects, the first set of speculative tokens 410 may include multiple subsets of speculative tokens, with each subset corresponding to one or more token sequences generated by different instances of the preliminary model 402 running using different operational parameters (e.g., according to group speculative decoding, parallel speculative decoding, or another technique).

[0066] In some respects, while the target model 404 verifies the first set of speculative tokens 410, the preliminary model 402 can generate a second set of speculative tokens 414. The second set of speculative tokens 414 can be generated based on assumptions that the target model 404 has accepted between 1 and S tokens from the first set of speculative tokens 410, where S corresponds to the number of tokens per sequence generated by the preliminary model 402. The preliminary model 402 can subsequently receive a selected set of tokens 412 from the target model 404. Based on the selected set of tokens 412, the preliminary model 402 can identify a subset of the second set of speculative tokens 414 for analysis (e.g., tokens whose generation was conditional on accepting the correct number of tokens from the first set of speculative tokens 410) and provide this subset to the target model 404 for verification.This process can continue, with the preliminary model 402 generating a third set of speculative tokens 418 while the target model 404 checks the second set of speculative tokens 414 and returns a selection 416 from the second set of speculative tokens 414. Petition 870250082970, dated 09 / 15 / 2025, pp. 104 / 168 31 / 66 for the preliminary model 402, with the preliminary model 402 generating a fourth set of speculative tokens 422 while the target model 404 checks the third set of speculative tokens 418 and returns a selection 420 from the third set of speculative tokens 418 to the preliminary model 402, and so on. In other respects, the preliminary model 402 may wait to generate the second set of speculative tokens 414 until the preliminary model 402 receives the selected set of tokens 412 from the target model 404.

[0067] Some of the processes described in the present invention for generating tokens speculatively using a preliminary model in parallel, or substantially in parallel, with the verification of previously generated sets of speculatively generated tokens using a target model can be carried out according to the group selection and triggering conditions of the target model selected to achieve continuous, or at least near-continuous, operation of the preliminary model 402 and the target model 404, while minimizing, or at least reducing, the possibility that the preliminary model 402 will backtrack and restart speculative token generation from the last accepted token (for example, to minimize, or at least reduce, the possibility that the target model 404 will reject each of the speculatively generated tokens generated by the preliminary model 402).Additionally, the group selection and trigger conditions of the target model can be selected in order to avoid, or at least reduce, the possibility of overloading the computing capabilities on the edge device or in the... Petition 870250082970, dated 09 / 15 / 2025, pp. 105 / 168 32 / 66 server (for example, so that the edge device generates tokens and the server verifies tokens within defined performance metrics that minimize, or at least reduce, the possibility of encountering bottlenecks in the processing system, such as memory thrashing where data is repeatedly swapped between memory on and off the processor, network bandwidth bottlenecks, and the like).

[0068] In some respects, hybrid speculative decoding can be performed token by token. In this case, the preliminary model 402 can operate continuously and generate multiple branching paths based on the generation of new speculative tokens. In some respects, the edge device 403 can execute a plurality of original speculative decoding instances (e.g., using different input parameters in the preliminary model 402), so that the preliminary model 402 generates multiple paths in a tree representing options that the target model 404 can check.Based on some defined triggering criteria, the 403 edge device can provide the generated tree to the 404 target model operating on the 405 server (e.g., in a cloud computing environment), which, as discussed above, can select a token sequence corresponding to a likely valid response to an incoming query and return the selected sequence to the 402 preliminary model running on the edge device. The 402 preliminary model can then prune the generated tree to fit the selected sequence returned by the 404 target model. In this example, the token generation rate can match. Petition 870250082970, dated 09 / 15 / 2025, pp. 106 / 168 33 / 66 refers to the rate at which the preliminary 402 model generates tokens, with tree pruning being based on target selection. Accuracy and inference speed may have an inverse relationship in this example, as a smaller preliminary model may generate tokens at a higher rate, but with a lower probability that the target 404 model will accept the tokens speculatively generated by the preliminary 402 model; meanwhile, a larger preliminary model may generate tokens at a lower rate, but with a higher probability that the target 404 model will accept the tokens speculatively generated by the preliminary 402 model.

[0069] Figure 5 illustrates an example 500 of hybrid speculative decoding in generative models, according to aspects of the present disclosure. In hybrid speculative decoding, a first set of speculatively generated tokens 502 (also called round 1 preliminary tokens) generated by the preliminary model (e.g., the preliminary model 402 illustrated in Figure 4) in a first round of inference can be provided to the target model (e.g., the target model 404 illustrated in Figure 4) for verification. As discussed, in several aspects, the preliminary model can be run on a first device, while the target model can be run on a second device.

[0070] According to several aspects, while the first set of speculatively generated 502 tokens is processed by the target model, the preliminary model generates a second set of preliminary tokens in a second round of inference (also called tokens Petition 870250082970, dated 09 / 15 / 2025, pp. 107 / 168 34 / 66 preliminary round 2 results), with assumptions being made for different numbers of tokens accepted.For example, as illustrated, the second set of preliminary tokens may include (1) a first subset 504 that assumes acceptance of the first preliminary token from the first set and may include a set of tokens generated speculatively based on acceptance of the first token; (2) a second subset 506 that assumes acceptance of the first and second preliminary tokens from the first set and includes a set of tokens generated speculatively based on acceptance of the first and second tokens; (3) a third subset 508 that assumes acceptance of the first to third preliminary tokens from the first set and includes a set of tokens generated speculatively based on acceptance of the first to third tokens; and (4) a fourth subset 510 that assumes acceptance of all four tokens from the first set and includes a set of tokens generated speculatively based on acceptance of all four tokens.For cases where a smaller number of tokens than the number of tokens included in the first set is assumed to be accepted, padding (e.g., null values, predefined constants, etc.) can be added so that each assumption has the same length. As illustrated, the second set of preliminary tokens can be generated in a batch process, so that the appropriate selection of preliminary tokens, conditioned on a specific set of accepted tokens from the first set of speculatively generated 502 tokens, can be provided to the target model for verification.

[0071] Figure 6 illustrates an example 600 of Petition 870250082970, dated 09 / 15 / 2025, pages 108 / 168 35 / 66 Hybrid speculative decoding in generative models, according to aspects of the present disclosure. In this example 600 (also called a hybrid parallel speculative decoding (HPSD) example), the aspects of parallel speculative decoding discussed above can be used to generate multiple candidate sets of tokens for the target model to verify. The target model can operate continuously, as in token-by-token speculative decoding. However, instead of generating tokens one by one, the preliminary model can speculatively generate groups of tokens based on a probability distribution that approximates, but need not be equal to, the probability distribution associated with the target model and generate different tree data structures, as illustrated in Figure 6, based on the input query (or prompt) and the generated token groups.Based on certain defined triggering criteria, an edge device (also called a local device or first device) can provide the generated tree data structures to the target model operating on a server (e.g., in a cloud computing environment) (also called a remote device or second device), which, as discussed above, selects a sequence of tokens corresponding to a likely valid response to an input query and returns the selected sequence to the preliminary model running on the edge device. Generally, the selected token sequence may include the speculatively generated tokens that the target model can verify, plus an additional token generated based on the generated tokens. Petition 870250082970, dated 09 / 15 / 2025, pp. 109 / 168 36 / 66 speculatively. The preliminary model can then prune the generated tree to fit the tokens verified by the target model.

[0072] As illustrated, different instances of a preliminary model, using different parameters, can independently generate a first set of speculative 602 tokens and a second set of speculative 604 tokens based on an input prompt. The preliminary model (e.g., running on a local device or first device) can provide the first set of speculative 602 tokens and the second set of speculative 604 tokens to a target model (e.g., running on a remote device or second device) for verification.In parallel, or at least substantially in parallel, with the target model verifying the first set of speculative tokens 602 and the second set of speculative tokens 604, the preliminary model can generate a first subsequent set of tokens 612 based on an assumption that the target model has accepted the first set of speculative tokens 602 and a second subsequent set of tokens 614 based on an assumption that the target model has accepted the second set of speculative tokens 604. In this example, as indicated by the oval, the target model accepted the second set of speculative tokens 604 instead of the first set of speculative tokens 602 (e.g., in a manner similar to that described above in relation to Figure 3). Since the target model accepted the second set of speculative tokens 604, the preliminary model can discard the first set of tokens. Petition 870250082970, dated 09 / 15 / 2025, pp. 110 / 168 37 / 66 subsequent 612 and provide the second set of subsequent tokens 614 for verification by the target model. Additionally, although not shown, additional sets of tokens may be generated by the preliminary model based on the second set of subsequent tokens 614 while the target model verifies the second set of subsequent tokens 614.

[0073] The process of generating speculative tokens, pruning (e.g., by discarding unselected token sets), and checking the target model can be repeated until a termination condition is reached for an input (e.g., no additional tokens need to be generated in order to generate an adequate response to the input query, a maximum number of tokens have been generated, a maximum amount of computing resources have been used to generate a response, etc.).

[0074] In this example, the total token generation rate may correspond to the rate at which the preliminary model generates tokens. However, previously generated tokens may be discarded when the target model determines that none of the speculatively generated tokens in the generated tree are candidates for inclusion in response to the query, as the preliminary model may restart speculative token generation using tokens up to the last token verified as an input in the preliminary model. As with token-by-token hybrid speculative decoding, accuracy and inference speed may have an inverse relationship in this example, as a smaller preliminary model may generate tokens at a higher rate, but with a lower probability that the target model will come to... Petition 870250082970, dated 09 / 15 / 2025, pp. 111 / 168 38 / 66 accept the tokens generated speculatively by the preliminary model.

[0075] Figure 7 illustrates an example 700 of hybrid group speculative decoding (HGSD) in generative models, according to aspects of the present disclosure. In this example 700, the aspects of group speculative decoding discussed above can be used to generate multiple candidate sets of tokens for the target model to verify. The target model (e.g., running on a remote device or second device) can operate continuously, as in token-by-token speculative decoding.However, instead of generating tokens one by one, the preliminary model (e.g., running on a local device or first device) can speculatively generate groups of tokens based on a probability distribution that is equivalent to the probability distribution associated with the target model and generate a tree (e.g., a tree data structure) based on the input query (or prompt) and the generated token groups. For example, based on some defined triggering criteria, an edge device can provide the generated tree to the target model operating on a server (e.g., in a cloud computing environment), which, as discussed above, can select a sequence of tokens corresponding to a likely valid response to the input query and return the selected sequence to the preliminary model running on the edge device.In general, the selected token sequence may include speculatively generated tokens that are the target model. Petition 870250082970, dated 09 / 15 / 2025, pp. 112 / 168 39 / 66 can verify, plus an additional token generated based on the speculatively generated tokens. The preliminary model can then prune the generated tree to fit the tokens verified by the target model.

[0076] For example, as illustrated, a tree data structure can be generated with the input prompt serving as a root node of the tree data structure. A first set of 702 tokens, including a plurality of groups (represented as different paths through the tree), can be speculatively generated by the preliminary model and provided to the target model for verification, as discussed above. While the first set of 702 tokens is being verified by the target model, the preliminary model can speculatively generate a second set of 704 tokens, with different subsets of tokens in the second set of 704 tokens being generated based on assumptions that different sets of tokens from the first set of 702 tokens are accepted by the target model.In example 700, the selected token set 706 may correspond to a set of tokens verified by the target model while the preliminary model generates subsequent token sets. In general, the tree data structure can be pruned, as discussed above, based on the target model's verification of various speculatively generated token groups, so that computing resources are not wasted on speculatively generating additional token groups based on previous token sets that were not accepted by the target model. Petition 870250082970, dated 09 / 15 / 2025, pp. 113 / 168 40 / 66

[0077] Similar to example 600 illustrated in Figure 6 shows that the total token generation rate in hybrid group speculative decoding, as illustrated by example 700 in Figure 7, may correspond to the rate at which the preliminary model generates tokens. However, previously generated tokens may be discarded when the target model determines that none of the speculatively generated tokens in the generated tree are candidates for inclusion in response to the query, since the preliminary model can restart speculative token generation using tokens up to the last token verified as an input in the preliminary model. As with token-by-token hybrid speculative decoding, accuracy and inference speed may have an inverse relationship in this example, as a smaller preliminary model may generate tokens at a higher rate, but with a lower probability that the target model will accept the tokens speculatively generated by the preliminary model. Example Operations for Speculative Decoding in Generative Artificial Intelligence Models

[0078] Figure 8 illustrates example 800 operations that can be performed by a computing device to generate a response to an input query using generative artificial intelligence models, according to aspects of this disclosure. The computing device for performing the 800 operations can be a device on which at least one preliminary model can be deployed, such as a smartphone, a tablet computer, a laptop computer, a desktop computer, a server, a cloud computing instance hosted in a distributed computing environment, or similar. Petition 870250082970, dated 09 / 15 / 2025, pp. 114 / 168 41 / 66

[0079] As illustrated, operations 800 begin in block 810, with the computing device generating, based on an input query and a first generative model, a plurality of token sets. In general, each token set in the plurality of token sets can correspond to a candidate answer to the input query. The input query can be received, for example, from a user of the computing device at a prompt where a textual query can be provided as input, through the conversion of a natural language expression into a textual representation of the input query, etc.

[0080] In block 820, operations 800 proceed with the computing device emitting, to a second generative model, the plurality of token sets for verification (e.g., by another computing device).

[0081] In block 830, operations 800 proceed with the computing device receiving, from the second generative model, an indication of a set of tokens selected from the plurality of sets of tokens based on the input query and the plurality of sets of tokens.

[0082] In block 840, operations 800 proceed with the computing device emitting (e.g., to a user or application, to another model, etc.) the set of tokens selected as a response to the input query.

[0083] In some respects, each set of tokens in the plurality of sets of tokens comprises a Petition 870250082970, dated 09 / 15 / 2025, pp. 115 / 168 42 / 66 group of tokens having the highest probabilities within a probability distribution associated with the first generative model in a universe of tokens on which the first generative model is trained.

[0084] In some respects, each token set in the plurality of token sets comprises a group of tokens selected based on a sum of probabilities associated with tokens in the token group, where the sum exceeds a threshold probability.

[0085] In some respects, the plurality of token sets is represented as a tree data structure. Within the tree data structure, a root node can generally correspond to the input query, and different paths through the tree data structure can generally correspond to different token sets among the plurality of token sets. The depth of the tree data structure can, in some respects, correspond to a maximum number of tokens generated by a single pass through the first generative model, and a maximum size of the tree data structure can be defined based on a computational complexity metric associated with generating a target token set by the second generative model.

[0086] In some respects, 800 operations may additionally include pruning a tree data structure based on the selected token set. A plurality of subsequent token sets is generated based on the pruned tree data structure and the input query, and emitted to the second generative model for verification. An indication of a token set Petition 870250082970, dated 09 / 15 / 2025, pp. 116 / 168 The 43 / 66 token set selected subsequently from the plurality of subsequent token sets is received from the second generative model based on the input query, the pruned tree data structure, and the plurality of subsequent token sets. Issuing the selected token set as the response to the input query may include emitting both the selected token set and the subsequent selected token set as the response to the input query.

[0087] In some respects, each respective set of tokens in the plurality of sets of tokens can be generated using a unique instance of the first generative model and unique parameters as inputs to the unique instance of the first generative model.

[0088] In some respects, the 800 operations may additionally include generating a subsequent plurality of token sets based on the input query and the plurality of token sets, while the second generative model verifies the plurality of token sets. Based on the subsequent plurality of token sets and the selected token set, a refined subsequent token set may be generated, and the refined subsequent token set may be issued to the second generative model for verification.

[0089] In some respects, token sets in the plurality of subsequent token sets may include padding taking into account that a number of tokens in the selected token set is less than a maximum number of tokens (or a pre-configured number of tokens).

[0090] In some respects, operations 800 Petition 870250082970, dated 09 / 15 / 2025, pp. 117 / 168 44 / 66 additionally includes receiving a token generated by the second generative model based on the selected token set. The received token can be issued as an additional token subsequent to the selected token set.

[0091] In some respects, the first generative model may correspond to a preliminary model in a speculative decoding pipeline, and the second generative model may correspond to a target model in the speculative decoding pipeline.

[0092] In some respects, the first generative model and the second generative model may have equivalent probability distributions.

[0093] In some respects, the first generative model may have a probability distribution that approximates a probability distribution associated with the second generative model. An approximation of a probability distribution may be a probability distribution that falls within a threshold difference with respect to the probability distribution associated with the second generative model, such that the probability distribution associated with the first generative model need not be an exact match with the probability distribution associated with the second generative model.

[0094] In some respects, the first generative model can be run locally (for example, on a local device or a local system), and the second generative model can be a remotely hosted model (for example, on a remote system or a remote device).

[0095] Figure 9 illustrates example 900 operations to verify a response to an input query. Petition 870250082970, dated 09 / 15 / 2025, pp. 118 / 168 45 / 66 using generative artificial intelligence models, according to aspects of this disclosure. Operations 900 can be performed by a device (or a system with multiple devices) on which at least one target model can be deployed, such as a server (for example, server 405 in Figure 4), a cloud computing instance hosted in a distributed computing environment, or similar.

[0096] As illustrated, operations 900 begin in block 910, with the receipt, from a device operating a first generative model, of an input query and a plurality of token sets, where each token set in the plurality of token sets corresponds to a candidate answer to the input query.

[0097] In block 920, operations 900 proceed with the comparison of a probability distribution associated with each respective set of tokens in the plurality of sets of tokens to a corresponding probability distribution generated by a second generative model for the respective set of tokens.

[0098] In block 930, operations 900 proceed with the selection of a set of tokens from among the plurality of sets of tokens, based on comparison.

[0099] In block 940, operations 900 proceed with the issuance of an indication of the token set selected for the first generative model.

[00100] In some respects, the plurality of token sets is represented as a tree data structure. Within the tree data structure, a Petition 870250082970, dated 09 / 15 / 2025, pp. 119 / 168 The 46 / 66 root node can generally correspond to the input query, and different paths through the tree data structure can generally correspond to different sets of tokens from the plurality of token sets. The depth of the tree data structure can, in some respects, correspond to a maximum number of tokens generated by a single pass through the first generative model, and a maximum size of the tree data structure can be defined based on a computational complexity metric associated with generating a target token set by the second generative model.

[00101] In some respects, comparing the probability distribution associated with each respective set of tokens in the plurality of sets of tokens to a corresponding probability distribution generated by a second generative model for the respective set of tokens involves generating probability distributions for each respective set of tokens based on a single pass through the second generative model. The single pass through the second generative model can be performed based on masked self-attention and positional encodings in a tree data structure.

[00102] In some respects, 900 operations additionally include generating an additional token based on the selected token set using the second model. The additional token (or an indication thereof) may be issued for the first generative model, in conjunction with the selected token set (or as part of the indication of the selected token set).

[00103] In some respects, the first model Petition 870250082970, dated 09 / 15 / 2025, pp. 120 / 168 The generative 47 / 66 model may correspond to a preliminary model in a speculative decoding pipeline, and the second generative model may correspond to a target model in the speculative decoding pipeline.

[00104] Figure 10 illustrates 1000 example operations that can be performed by a computing device to speculatively generate a response to an input query and verify the response to the input query using generative artificial intelligence models, according to aspects of the present disclosure. The computing device to perform the 800 operations can be a device on which at least one preliminary model can be deployed, such as a smartphone, a tablet computer, a laptop computer, a desktop computer, a server, a cloud computing instance hosted in a distributed computing environment, or similar.

[00105] As illustrated, operations 1000 begin in block 1010, with the computing device generating, based on an input query and a first generative model, a first plurality of token sets. In general, each token set in the first plurality of token sets can correspond to a candidate answer to the input query. The input query can be received, for example, from a user of the computing device at a prompt where a textual query can be provided as input, through the conversion of a natural language expression into a textual representation of the input query, etc.

[00106] In block 1020, operations 1000 Petition 870250082970, dated 09 / 15 / 2025, pp. 121 / 168 48 / 66 proceed with the computing device emitting, to a second generative model, the first plurality of token sets for verification (e.g., by another computing device).

[00107] In block 1030, operations 1000 proceed with the computing device generating speculatively, while waiting to receive an indication of a token set selected from the first plurality of token sets of the second generative model, a second plurality of token sets. Each token set in the second plurality of token sets generally corresponds to a second portion of the candidate answer to the input query.

[00108] In block 1040, operations 1000 proceed with the receipt, based on the second generative model, of the indication of the set of tokens selected from the plurality of sets of tokens.

[00109] In block 1050, operations 1000 proceed with the issuance, for the second generative model, of tokens from the second plurality of token sets associated with the token set selected for verification.

[00110] In block 1060, operations 1000 proceed with the issuance (e.g., to a user or application, to another model, etc.) of the set of tokens selected as a response to the input query.

[00111] In some respects, operations 1000 additionally include receiving an indication of a second set of tokens selected from the second plurality of token sets associated with the token set. Petition 870250082970, dated 09 / 15 / 2025, pp. 122 / 168 49 / 66 selected. The second set of tokens selected can be issued (e.g., to a user or application, to another model, etc.) as another portion of the response to the input query.

[00112] In some respects, each set of tokens in the first plurality of token sets comprises a group of tokens having the highest probabilities within a probability distribution associated with the first generative model in a universe of tokens.

[00113] In some respects, each set of tokens in the first plurality of sets of tokens comprises a group of tokens selected based on a sum of probabilities associated with tokens in the group of tokens, where the sum exceeds a threshold probability.

[00114] In some respects, the first plurality of token sets can be represented as a tree data structure. In the tree data structure, a root node can correspond to the input query. Each path through the tree data structure corresponds to a set of tokens from the first plurality of token sets.

[00115] In some respects, each respective set of tokens in the first plurality of token sets is generated using a unique instance of the first generative model and unique parameters as inputs to the unique instance of the first generative model.

[00116] In some respects, the 1000 operations additionally include generating a subsequent refined token set based on the token set. Petition 870250082970, dated 09 / 15 / 2025, pp. 123 / 168 50 / 66 selected and in the second plurality of token sets. The subsequent token set refined for verification can be issued to the second generative model. While waiting to receive an indication of a second token set selected from the subsequent refined token set, a third plurality of token sets can be speculatively generated. In some respects, token sets in the subsequent token set plurality include padding taking into account that a number of tokens in the selected token set is less than a maximum number of tokens.

[00117] In some respects, the first generative model may correspond to a preliminary model in a speculative decoding pipeline. The second generative model may correspond to a target model in the speculative decoding pipeline.

[00118] In some respects, the first generative model may have a probability distribution that approximates a probability distribution associated with the second generative model. An approximation of a probability distribution may be a probability distribution that falls within a threshold difference relative to the probability distribution associated with the second generative model, such that the probability distribution associated with the first generative model need not be an exact match with the probability distribution associated with the second generative model.

[00119] In some respects, the first generative model can be run locally (for example, on a local device or a local system), and the second model Petition 870250082970, dated 09 / 15 / 2025, pp. 124 / 168 51 / 66 generative can be a remotely hosted model (for example, on a remote system or a remote device). Example Processing Systems for Speculative Decoding in Generative Artificial Intelligence Models

[00120] Figure 11 depicts an example processing system 1100 for generating a response to a query input in a generative artificial intelligence model based on group speculative decoding, as described in the present invention, for example, in relation to Figures 8 to 10.

[00121] The 1100 processing system includes a central processing unit (CPU) 1102, which, in some examples, may be a multi-core CPU. Instructions executed on CPU 1102 may be loaded, for example, from a program memory associated with CPU 1102, or they may be loaded from a memory partition (for example, from memory 112 4).

[00122] The 1100 processing system also includes additional processing components adapted to specific functions, such as a graphics processing unit (GPU) 1104, a digital signal processor (DSP) 1106, a neural processing unit (NPU) 1108 and a connectivity component 1112.

[00123] An NPU, such as the NPU 1108, is generally a specialized circuit configured to implement the control and arithmetic logic for executing machine learning algorithms, such as algorithms for Petition 870250082970, dated 09 / 15 / 2025, pages 125 / 168 52 / 66 processing artificial neural networks (ANNs), deep neural networks (DNNs), random forests (RFs), and the like. An NPU may sometimes be alternatively called a neural signal processor (NSP), a tensor processing unit (TPU), a neural network processor (NNP), an intelligence processing unit (IPU), a vision processing unit (VPU), or a graphics processing unit.

[00124] NPUs, such as the NPU 1108, are configured to accelerate the performance of common machine learning tasks, such as image classification, machine translation, object detection, and various other predictive models. In some examples, a plurality of NPUs can be instantiated on a single chip, such as a system on a chip (SoC), while in other examples, these NPUs can be part of a dedicated neural network accelerator.

[00125] NPUs can be optimized for training or inference or, in some cases, configured to balance performance between the two. For NPUs capable of performing both training and inference, the two tasks can still generally be performed independently.

[00126] NPUs designed to accelerate training are generally configured to speed up the optimization of new models, which is a computationally intensive operation involving the insertion of Petition 870250082970, dated 09 / 15 / 2025, pp. 126 / 168 53 / 66 entries into an existing dataset (often identified or tagged), iterating through the dataset, and then adjusting model parameters such as weights and biases to improve model performance. Generally, optimization based on a wrong prediction involves backpropagation through the model layers and determining gradients to reduce the prediction error.

[00127] NPUs designed to accelerate inference are generally configured to operate on complete models. Such NPUs can therefore be configured to take new data as input and quickly process that new data through an already trained model to generate a model output (e.g., an inference).

[00128] In some implementations, the NPU 1108 is part of one or more of the CPU 1102, the GPU 1104, and / or the DSP 1106. These may be located in user equipment (UE), in a wireless communication system, or other computing device.

[00129] In some examples, an 1112 connectivity component may include subcomponents, for example, for third-generation (3G) connectivity, fourth-generation (4G) connectivity (e.g., long-term evolution (LTE long-term evolution)), fifth-generation (5G) connectivity (e.g., New Radio (NR)), Wi-Fi connectivity, Bluetooth connectivity, and other wireless data transmission standards. The 1112 connectivity component may be additionally coupled to one or more 1114 antennas. Petition 870250082970, dated 09 / 15 / 2025, pp. 127 / 168 54 / 66

[00130] The processing system 1100 may also include one or more sensor processing units 1116 associated with any type of sensor, one or more image signal processors (ISPs) 1118 associated with any type of image sensor and / or a navigation processor 1120, which may include satellite-based positioning system components (e.g., GPS or GLONASS) as well as inertial positioning system components.

[00131] The processing system 1100 may also include one or more input and / or output devices 1122, such as screens, touch-sensitive surfaces (including touch-sensitive displays), physical buttons, speakers, microphones and the like.

[00132] In some instances, one or more of the processors in the 1100 processing system may be based on an ARM or RISC-V instruction set.

[00133] The 1100 processing system also includes a 1124 memory, which is representative of one or more static and / or dynamic memories, such as dynamic random access memory, flash-based static memory, and the like. In this example, the 1124 memory includes computer executable components, which can be executed by one or more of the previously mentioned processors of the 1100 processing system.

[00134] In particular, in this example, memory 1124 includes a token set generation component 1124A, a token selection reception component 1124B, a token selection component 1026C, and an output generation component 1124D, generative models 1124E, and, Petition 870250082970, dated 09 / 15 / 2025, pages 128 / 168 55 / 66 optionally, a comparison component 1124F. The components depicted, and others not depicted, can be configured to perform various aspects of the methods described in the present invention.

[00135] In general, the 1100 processing system and / or its components can be configured to perform the methods described in the present invention. Example Clauses

[00136] Implementation details of various aspects of this disclosure are described in the following numbered clauses:

[00137] Clause 1: A processor-implemented method comprising: generating, based on an input query and a first generative model, a plurality of token sets, wherein each token set in the plurality of token sets corresponds to a candidate response to the input query; sending, to a second generative model, the plurality of token sets for verification; receiving, from the second generative model, an indication of a token set selected from among the plurality of token sets based on the input query and the plurality of token sets; and sending the selected token set as a response to the input query.

[00138] Clause 2: The method of clause 1, where each set of tokens in the plurality of sets of tokens comprises a group of tokens having the highest probabilities within a probability distribution associated with the first generative model in a universe of Petition 870250082970, dated 09 / 15 / 2025, pp. 129 / 168 56 / 66 tokens.

[00139] Clause 3: The method of clause 1 or 2, wherein each set of tokens in the plurality of sets of tokens comprises a group of tokens selected based on a sum of probabilities associated with tokens in the group of tokens, wherein the sum exceeds a threshold probability.

[00140] Clause 4: The method of any of clauses 1 to 3, wherein: the plurality of token sets is represented as a tree data structure, a root node of the tree data structure corresponds to the input query, and each path through the tree data structure corresponds to a token set from among the plurality of token sets.

[00141] Clause 5: The method of clause 4, where a depth of the tree data structure corresponds to a maximum number of tokens generated by a single pass through the first generative model.

[00142] Clause 6: The method of clause 4 or 5, in which a maximum size of the tree data structure is defined based on a computational complexity metric associated with the generation of a target token set by the second generative model.

[00143] Clause 7: The method of any of clauses 4 to 6, which additionally comprises: pruning the tree data structure based on the selected token set; generating a plurality of subsequent token sets based on the pruned tree data structure and the input query; emitting, to the second generative model, the plurality of subsequent token sets. Petition 870250082970, dated 09 / 15 / 2025, pp. 130 / 168 57 / 66 for verification; and receive, from the second generative model, an indication of a subsequent selected token set from among the plurality of subsequent token sets based on the input query, in the pruned tree data structure and the plurality of subsequent token sets, whereby issuing the selected token set as the response to the input query comprises issuing the selected token set and the subsequent selected token set as the response to the input query.

[00144] Clause 8: The method of any of clauses 1 to 7, wherein each respective set of tokens in the plurality of sets of tokens is generated using a unique instance of the first generative model and unique parameters as inputs to the unique instance of the first generative model.

[00145] Clause 9: The method of any of clauses 1 to 8, which further comprises: while the second generative model checks the plurality of token sets, generating a subsequent plurality of token sets based on the input query and the plurality of token sets; based on the plurality of subsequent token sets and the selected token set, generating a refined subsequent token set; and issuing, to the second generative model, the refined subsequent token set for verification.

[00146] Clause 10: The method of clause 9, in which token sets in the plurality of subsequent token sets include padding taking into account that a number of tokens in the selected token set Petition 870250082970, dated 09 / 15 / 2025, pp. 131 / 168 58 / 66 is less than a maximum number of tokens.

[00147] Clause 11: The method of any of clauses 1 to 10, which additionally comprises: receiving a token generated by the second generative model based on the selected token set; and issuing the received token as an additional token subsequent to the selected token set.

[00148] Clause 12: The method of any of clauses 1 to 11, wherein: the first generative model corresponds to a preliminary model in a speculative decoding pipeline, and the second generative model corresponds to a target model in the speculative decoding pipeline.

[00149] Clause 13: The method of clause 12, wherein the preliminary model comprises a model trained to have a probability distribution that approximates a corresponding probability distribution for the target model.

[00150] Clause 14: The method of any of clauses 1 to 13, wherein: the first generative model comprises a model running on a local system, and the second generative model comprises a model running on a remote system.

[00151] Clause 15: A processor-implemented method comprising: receiving an input query and a plurality of token sets generated by a first generative model, wherein each respective token set in the plurality of token sets corresponds to a respective candidate answer to the input query; comparing a probability distribution Petition 870250082970, dated 09 / 15 / 2025, pages 132 / 168 59 / 66 associated with each respective set of tokens in the plurality of sets of tokens to a corresponding probability distribution generated by a second generative model for the respective set of tokens; select a set of tokens from among the plurality of sets of tokens, based on the comparison; and issue, to the first generative model, an indication of the selected set of tokens.

[00152] Clause 16: The method of clause 15, wherein: the plurality of token sets is represented as a tree data structure, a root node of the tree data structure corresponds to the input query, and each path through the tree data structure corresponds to a token set among the plurality of token sets.

[00153] Clause 17: The method of clause 15 or 16, in which comparing the probability distribution associated with each respective set of tokens in the plurality of sets of tokens to the corresponding probability distribution generated by the second generative model for the respective set of tokens comprises generating probability distributions for each respective set of tokens based on a single pass through the second generative model.

[00154] Clause 18: The method of clause 17, in which the single pass through the second generative model is performed based on masked self-attention and positional encodings in a tree data structure.

[00155] Clause 19: The method of any of clauses 15 to 18, which additionally comprises: generating an additional token based on the selected token set. Petition 870250082970, dated 09 / 15 / 2025, pp. 133 / 168 60 / 66 using the second generative model; and issue the additional token for the first generative model.

[00156] Clause 20: The method of any of clauses 15 to 19, wherein: the first generative model corresponds to a preliminary model in a speculative decoding pipeline, and the second generative model corresponds to a target model in the speculative decoding pipeline.

[00157] Clause 21: A processor-implemented method comprising: generating, based on an input query and a first generative model, a first plurality of token sets, wherein each token set in the first plurality of token sets corresponds to a first portion of a candidate answer to the input query; sending the plurality of token sets to a second generative model for verification; while waiting to receive, from the second generative model, an indication of a token set selected from the first plurality of token sets, speculatively generating a second plurality of token sets, wherein each token set in the second plurality of token sets corresponds to a second portion of the candidate answer to the input query;Receive, from the second generative model, an indication of the set of tokens selected from the first plurality of token sets; emit, to the second generative model, tokens from the second plurality of token sets associated with the set of tokens selected for verification; and emit the selected set of tokens as a response to the input query. Petition 870250082970, dated 09 / 15 / 2025, pp. 134 / 168 61 / 66

[00158] Clause 22: The method of clause 21, which further comprises: receiving an indication of a second set of tokens selected from the second plurality of sets of tokens associated with the selected set of tokens; and issuing the second selected set of tokens as another portion of the response to the input query.

[00159] Clause 23: The method of clause 21 or 22, wherein each set of tokens in the first plurality of sets of tokens comprises a group of tokens having the highest probabilities within a probability distribution associated with the first generative model in a universe of tokens.

[00160] Clause 24: The method of any of clauses 21 to 23, wherein each set of tokens in the first plurality of sets of tokens comprises a group of tokens selected based on a sum of probabilities associated with tokens in the group of tokens, wherein the sum exceeds a threshold probability.

[00161] Clause 25: The method of any of clauses 21 to 24, wherein: the first plurality of token sets is represented as a tree data structure, a root node of the tree data structure corresponds to the input query, and each path through the tree data structure corresponds to a token set from the first plurality of token sets.

[00162] Clause 26: The method of any of clauses 21 to 25, wherein each respective set of tokens in the first plurality of token sets is generated using a unique instance of the first model. Petition 870250082970, dated 09 / 15 / 2025, pp. 135 / 168 62 / 66 generative and unique parameters as inputs in the unique instance of the first generative model.

[00163] Clause 27: The method of any of clauses 21 to 26, further comprising: generating a refined subsequent token set based on the selected token set and the second token set plurality; issuing, to the second generative model, the refined subsequent token set for verification; and while waiting to receive, from the second generative model, an indication of a second token set selected from the refined subsequent token set; speculatively generating a third token set plurality.

[00164] Clause 28: The method of clause 27, wherein token sets in the plurality of subsequent token sets include padding taking into account that a number of tokens in the selected token set is less than a maximum number of tokens.

[00165] Clause 29: The method of any of clauses 21 to 28, wherein: the first generative model corresponds to a preliminary model in a speculative decoding pipeline, and the second generative model corresponds to a target model in the speculative decoding pipeline.

[00166] Clause 30: The method of clause 29, wherein the preliminary model comprises a model trained to have a probability distribution that approximates a corresponding probability distribution for the target model.

[00167] Clause 31: The method of any one Petition 870250082970, dated 09 / 15 / 2025, pages 136 / 168 63 / 66 of clauses 21 to 30, where: the first generative model comprises a model running on a local system, and the second generative model comprises a model running on a remote system.

[00168] Clause 32: A processing system comprising: a memory that has executable instructions stored therein; and one or more processors configured to execute the executable instructions to make the processing system perform the method of any of clauses 1 to 31.

[00169] Clause 33: A processing system, comprising: means for carrying out the method of any of clauses 1 to 31.

[00170] Clause 34: A computer-readable medium containing instructions which, when executed by one or more processors, cause those processors to perform the method described in any of clauses 1 through 31. Additional Considerations

[00171] The preceding description is provided to enable any person skilled in the art to practice the various aspects described in the present invention. The examples discussed in the present invention are not limiting to the scope, applicability, or aspects set forth in the claims. Various modifications of these aspects will be readily apparent to those skilled in the art, and the generic principles defined in the present invention may be applied to other aspects. For example, changes may be made to the function and arrangement of the elements discussed without departing from the scope of the disclosure. Several Petition 870250082970, dated 09 / 15 / 2025, pages 137 / 168 Examples may omit, substitute, or add various procedures or components, as appropriate. For example, the methods described may be performed in a different order than that described, and various steps may be added, omitted, or combined. Furthermore, attributes described in relation to some examples may be combined in some other examples. For example, an apparatus may be implemented, or a method may be practiced, using any number of the aspects set forth in the present invention. In addition, the scope of the disclosure is intended to cover such an apparatus or method that is practiced with the use of another structure, functionality, or structure and functionality in addition to or different from the various aspects of the disclosure set forth in the present invention. It should be understood that any aspect of the disclosure disclosed in the present invention may be incorporated by one or more elements of a claim.

[00172] As used in the present invention, the term example means that it serves as an example, instance, or illustration. Any aspect described in the present invention as exemplary should not necessarily be interpreted as preferential or advantageous in relation to other aspects.

[00173] As used in the present invention, an expression referring to at least one of a list of items refers to any combination of those items, including single members. For example, at least one of: a, b, or c is intended to encompass a, b, c, ab, ac, bc, and abc, as well as any combination with multiples of the same element (e.g., aa, aaa, aab, aac, a Petition 870250082970, dated 09 / 15 / 2025, pp. 138 / 168 65 / 66 bb, acc, bb, bbb, bbc, cc and ccc, or any other ordering of a, b and c).

[00174] As used in the present invention, the term determine encompasses a wide variety of actions. For example, determine may include calculate, compute, process, derive, investigate, search (e.g., searching in a table, a database, or other data structure), verify, and the like. Furthermore, determine may include receive (e.g., receiving information), access (e.g., accessing data in a memory), and the like. Additionally, determine may include solve, select, choose, establish, and the like.

[00175] The methods disclosed in the present invention comprise one or more steps or actions to perform the methods. The steps and / or actions of the method can be interchanged without departing from the scope of the claims. In other words, unless a particular order of steps or actions is specified, the order and / or use of specific steps and / or actions can be modified without departing from the scope of the claims. Additionally, the various operations of the methods described above can be performed by any suitable means capable of performing the corresponding functions. The means may include various hardware and / or software components and / or modules, including, but not limited to, a circuit, an application-specific integrated circuit (ASIC), or a processor.In general, when operations are illustrated in the figures, these operations may have equivalent components in a more functional way, corresponding to the function. Petition 870250082970, dated 09 / 15 / 2025, pages 139 / 168 66 / 66 similar numbering.

[00176] The following claims are not intended to be limited to the aspects shown in the present invention, but are intended to have the full scope consistent with the language of the claims. In a claim, reference to an element in the singular is not intended to mean one and only one, unless specifically stated otherwise, but instead means one or more. Unless specifically stated otherwise, the term any refers to one or more. No element of a claim shall be interpreted in accordance with the provisions of Title 35 of the USC Code § 112(f), unless the element is expressly mentioned using the expression means to or, in the case of a method claim, the element is mentioned using the expression step to.All structural and functional equivalents to the elements of the various aspects described throughout this disclosure that are known or may become known to those skilled in the art are expressly incorporated into the present invention by reference and are intended to be covered by the claims. Furthermore, nothing disclosed in this invention is intended to be exclusive to the public, regardless of whether such disclosure is explicitly mentioned in the claims. Petition 870250082970, dated 09 / 15 / 2025, pp. 140 / 168

Claims

1 / 12 CLAIMS 1. A processing system characterized by comprising: a memory that has executable instructions stored therein; and one or more processors configured to execute the executable instructions to make the processing system: generate, based on an input query and a first generative model, a plurality of token sets, wherein each token set in the plurality of token sets corresponds to a candidate response to the input query; send, to a second generative model, the plurality of token sets for verification; receive, from the second generative model, an indication of a token set selected from among the plurality of token sets based on the input query and the plurality of token sets; and send the selected token set as a response to the input query.

2. Processing system, according to claim 1, characterized in that each set of tokens in the plurality of sets of tokens comprises a group of tokens having the highest probabilities within a probability distribution associated with the first generative model in a universe of tokens.

3. Processing system, according to claim 1, characterized in that each set of tokens in the plurality of sets of tokens comprises a group of tokens selected based on a sum of probabilities associated with tokens in the group of tokens, wherein the sum exceeds a threshold probability.

4. Processing system, according to claim 1, characterized in that: the plurality of token sets being represented as a tree data structure, a root node of the tree data structure corresponding to the input query, and each path through the tree data structure corresponding to a set of tokens from among the plurality of token sets.

5. Processing system, according to claim 4, characterized in that the depth of the tree data structure corresponds to a maximum number of tokens generated by a single pass through the first generative model.

6. Processing system, according to claim 4, characterized in that a maximum size of the tree data structure is defined based on a computational complexity metric associated with the generation of a target token set by the second generative model.

7. Processing system, according to claim 4, characterized in that one or more processors are additionally configured to make the processing system: prune the tree data structure based on the selected token set; generate a plurality of token sets. Petition 870250082970, dated 09 / 15 / 2025, p.142 / 168 3 / 12 subsequent based on the pruned tree data structure and the input query; issue, to the second generative model, the subsequent plurality of token sets for verification; and receive, from the second generative model, an indication of a subsequent selected token set from among the plurality of subsequent token sets based on the input query, the pruned tree data structure, and the plurality of subsequent token sets, wherein issuing the selected token set as the response to the input query comprises issuing the selected token set and the subsequent selected token set as the response to the input query.

8. Processing system, according to claim 1, characterized in that each respective set of tokens in the plurality of sets of tokens is generated using a unique instance of the first generative model and unique parameters as inputs to the unique instance of the first generative model.

9. Processing system, according to claim 1, characterized in that one or more processors are additionally configured to perform the processing system: while the second generative model verifies the plurality of token sets, generate a subsequent plurality of token sets based on the input query and the plurality of token sets; based on the subsequent plurality of token sets and the selected token set, generate a refined subsequent token set; and issue, to the second generative model, the refined subsequent token set for verification.

10. Processing system, according to claim 9, characterized in that token sets in the plurality of subsequent token sets include padding taking into account that a number of tokens in the selected token set is less than a maximum number of tokens.

11. Processing system, according to claim 1, characterized in that one or more processors are additionally configured to make the processing system: receive a token generated by the second generative model based on the selected token set; and issue the received token as an additional token subsequent to the selected token set.

12. Processing system, according to claim 1, characterized in that: the first generative model corresponds to a preliminary model in a speculative decoding pipeline, and the second generative model corresponds to a target model in the speculative decoding pipeline.

13. Processing system, according to claim 12, characterized in that the preliminary model comprises a model trained to have a probability distribution that approximates a corresponding probability distribution for the target model.

14. Processing system, according to Petition 870250082970, dated 09 / 15 / 2025, pp. 144 / 168 5 / 12, claim 1, characterized by: the first generative model comprising a model running on a local system, and the second generative model comprising a model running on a remote system.

15. A processing system characterized by comprising: a memory that has executable instructions stored therein; and one or more processors configured to execute the executable instructions in order to make the processing system: receive, from a system in which a first generative model operates, an input query and a plurality of token sets, where each respective token set in the plurality of token sets corresponds to a respective candidate answer to the input query; compare a probability distribution associated with each respective token set in the plurality of token sets to a corresponding probability distribution generated by a second generative model for the respective token set; select a token set from among the plurality of token sets, based on the comparison; and issue, to the first generative model, an indication of the selected token set.

16. Processing system, according to claim 15, characterized by: Petition 870250082970, dated 09 / 15 / 2025, p. 145 / 168 6 / 12 the plurality of token sets being represented as a tree data structure, a root node of the tree data structure corresponding to the input query, and each path through the tree data structure corresponding to a token set among the plurality of token sets.

17. Processing system, according to claim 15, characterized in that, to compare the probability distribution associated with each respective set of tokens in the plurality of sets of tokens to the corresponding probability distribution generated by the second generative model for the respective set of tokens, one or more processors are configured to make the processing system generate probability distributions for each respective set of tokens based on a single pass through the second generative model.

18. Processing system, according to claim 17, characterized in that the single passage through the second generative model is performed based on masked self-attention and positional encodings in a tree data structure.

19. Processing system, according to claim 15, characterized in that one or more processors are additionally configured to make the processing system: generate an additional token based on the set of tokens selected using the second generative model; and issue the additional token for the first generative model. Petition 870250082970, dated 09 / 15 / 2025, pp. 146 / 168 7 / 12 20. Processing system according to claim 15, characterized in that: the first generative model corresponds to a preliminary model in a speculative decoding pipeline, and the second generative model corresponds to a target model in the speculative decoding pipeline.

21. A method implemented by a processor characterized by comprising: generating, based on an input query and a first generative model, a plurality of token sets, where each token set in the plurality of token sets corresponds to a candidate response to the input query; sending the plurality of token sets to a second generative model for verification; receiving, from the second generative model, an indication of a token set selected from among the plurality of token sets based on the input query and the plurality of token sets; and sending the selected token set as a response to the input query.

22. A method according to claim 21, characterized in that each set of tokens in the plurality of sets of tokens comprising a group of tokens has the highest probabilities within a probability distribution associated with the first generative model in a universe of tokens.

23. Method according to claim 21, Petition 870250082970, dated 09 / 15 / 2025, p. 147 / 168 8 / 12 characterized by each set of tokens in the plurality of sets of tokens comprising a group of tokens selected based on a sum of probabilities associated with tokens in the group of tokens, wherein the sum exceeds a threshold probability.

24. A method according to claim 21, characterized in that: the plurality of token sets being represented as a tree data structure, a root node of the tree data structure corresponding to the input query, and each path through the tree data structure corresponding to a token set among the plurality of token sets.

25. Method, according to claim 24, characterized by a depth of the tree data structure corresponding to a maximum number of tokens generated by a single pass through the first generative model.

26. Method, according to claim 24, characterized in that a maximum size of the tree data structure is defined based on a computational complexity metric associated with the generation of a target token set by the second generative model.

27. A method according to claim 24, characterized by further comprising: pruning the tree data structure based on the selected token set; generating a plurality of subsequent token sets based on the pruned tree data structure and the input query; issuing, to the second generative model, the subsequent plurality of token sets for verification; and receiving, from the second generative model, an indication of a subsequent selected token set from among the plurality of subsequent token sets based on the input query, the pruned tree data structure, and the plurality of subsequent token sets, wherein issuing the selected token set as the response to the input query comprises issuing the selected token set and the subsequent selected token set as the response to the input query.

28. A method according to claim 21, characterized in that each respective set of tokens in the plurality of sets of tokens is generated using a unique instance of the first generative model and unique parameters as inputs to the unique instance of the first generative model.

29. A method according to claim 21, characterized by further comprising: while the second generative model checks the plurality of token sets, generating a subsequent plurality of token sets based on the input query and the plurality of token sets; based on the subsequent plurality of token sets and the selected token set, generating a refined subsequent token set; and emitting, to the second generative model, the refined subsequent token set for verification.

30. Method, according to claim 29, Petition 870250082970, dated 09 / 15 / 2025, pp. 149 / 168 10 / 12 characterized by token sets in the plurality of subsequent token sets including padding taking into account that a number of tokens in the selected token set is less than a maximum number of tokens.

31. A method according to claim 21, characterized by further comprising: receiving a token generated by the second generative model based on the selected token set; and issuing the received token as an additional token subsequent to the selected token set.

32. Method according to claim 21, characterized in that: the first generative model corresponds to a preliminary model in a speculative decoding pipeline, and the second generative model corresponds to a target model in the speculative decoding pipeline.

33. Method, according to claim 32, characterized in that the preliminary model comprises a model trained to have a probability distribution that approximates a corresponding probability distribution for the target model.

34. Method according to claim 21, characterized in that: the first generative model comprises a model running on a local system, and the second generative model comprises a model running on a remote system.

35. Method implemented by processor Petition 870250082970, dated 09 / 15 / 2025, page 150 / 168 11 / 12 characterized by comprising: receiving, from a device in which a first generative model operates, an input query and a plurality of token sets, where each respective token set in the plurality of token sets corresponds to a respective candidate answer to the input query; comparing a probability distribution associated with each respective token set in the plurality of token sets to a corresponding probability distribution generated by a second generative model for the respective token set; selecting a token set from among the plurality of token sets, based on the comparison; and issuing, to the first generative model, an indication of the selected token set.

36. A method according to claim 35, characterized in that: the plurality of token sets being represented as a tree data structure, a root node of the tree data structure corresponding to the input query, and each path through the tree data structure corresponding to a set of tokens from among the plurality of token sets.

37. Method, according to claim 35, characterized by comparing the probability distribution associated with each respective set of tokens in the plurality of sets of tokens to the corresponding probability distribution generated by the second generative model for the respective set of tokens comprises generating probability distributions for each respective set of tokens based on a single pass through the second generative model.

38. Method, according to claim 37, characterized in that the single passage through the second generative model is performed based on masked self-attention and positional encodings in a tree data structure.

39. A method according to claim 35, characterized by further comprising: generating an additional token based on the selected token set using the second generative model; and issuing the additional token to the first generative model.

40. Method according to claim 35, characterized in that: the first generative model corresponds to a preliminary model in a speculative decoding pipeline, and the second generative model corresponds to a target model in the speculative decoding pipeline. Petition 870250082970, dated 09 / 15 / 2025, pp. 152 / 168