Dual inference acceleration method based on draft prediction and compressed memory
Patent Information
- Application Number
- CN202511739509.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-11-25
- Publication Date
- 2026-09-22
- Estimated Expiration
- 2045-11-25
AI Technical Summary
然而,这类方法普遍依赖规则式分割与统一压缩策略,容易造成信息压缩不足或过度压缩,尤其在数值敏感型任务(如数学题)中存在信息丢失风险,导致模型回答正确率下降
1、本发明通过并行生成与上下文压缩的结合,既解决了传统自回归的串行低效问题,又缓解了长序列推理的冗余开销;
Smart Images

Figure CN121543731B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to artificial intelligence technology, and more particularly to a dual reasoning acceleration method based on draft prediction and compressed memory. Background Technology
[0002] Current mainstream methods for accelerating reasoning in large language models can be broadly categorized into two types: parallel generation methods based on draft mechanisms and memory optimization methods based on context compression. The former introduces a lightweight draft model to predict multiple candidate tokens in parallel, which are then verified one by one by the main model, thus improving generation speed to some extent. However, its main drawback is the limited prediction accuracy of the draft model and the large number of "rejection" operations in the verification stage, leading to wasted computational resources, especially in long text reasoning where performance is unstable. Furthermore, draft models are typically trained independently, making it difficult to effectively capture the contextual features of the main model, resulting in a low token acceptance rate and limiting the upper limit of parallel efficiency improvement. The latter, on the other hand, focuses on reducing the burden on the KVCache and improving context utilization efficiency through thought compression. These methods train the model to compress intermediate thought processes into a small number of key point representations, retaining only crucial context for subsequent generation, significantly reducing memory usage. However, these methods generally rely on rule-based segmentation and uniform compression strategies, which can easily lead to insufficient or excessive information compression, especially in numerically sensitive tasks (such as mathematical problems), posing a risk of information loss and causing a decrease in the model's correct answer rate. Furthermore, current methods do not distinguish the importance of reasoning steps and lack an active mechanism for distinguishing between "acceptable information" and "redundant information".
[0003] Therefore, this invention addresses the shortcomings of each of the draft generation and context compression schemes, and proposes for the first time a dual-path reasoning acceleration mechanism that integrates the two approaches. On the one hand, it improves token generation efficiency through a parallel draft mechanism, and on the other hand, it reduces the length of the historical reasoning chain through a semantic compression strategy, thereby achieving a dual improvement in generation speed and storage efficiency. Summary of the Invention
[0004] The technical problem to be solved by the present invention is to provide a dual reasoning acceleration method based on draft prediction and compressed memory, which addresses the shortcomings of the prior art.
[0005] The technical solution adopted by this invention to solve its technical problem is: a dual reasoning acceleration method based on draft prediction and compressed memory, comprising the following steps: 1) At each stage of inference, obtain the input context for model inference. The model has generated a token sequence. ; sequence In this sequence, the subscript 1 indicates the start position of the sequence, the subscript i indicates the end position of the sequence, and the separator : indicates a continuous interval from the start position to the end position. 2) Based on the initial input context of the model inference The model has generated a token sequence. Parallel prediction of subsequent One candidate token is used as a draft token; 3) Organize the draft tokens according to a tree structure; 4) Perform master model validation; 4.1) Main Model Receive a tree-structured sequence of draft tokens, for each draft token... Calculate conditional probability; 4.2) Based on the conditional probability results, a continuous subsequence in the draft token that has passed verification is recorded as an inference block. ; m is the number of consecutive tokens that have passed verification, m≤k; t represents the t-th inference block; 5) Compress each inference block into a compressed token of fixed length s; 6) Context refactoring; Use the newly generated compressed token Update the context set to obtain the updated context. , recorded as ,in for .
[0006] According to the above scheme, in step 2), a semi-autoregressive draft model is used to predict subsequent events in parallel. The candidate tokens are as follows: ; in, This is a semi-autoregressive draft model. Indicates prediction token sequence.
[0007] According to the above scheme, in step 4.1), the main model For the The conditional probability of a draft token is:
[0008] Among them, subscript This represents the set of all learnable parameters in the main model M.
[0009] According to the above scheme, in step 4.2), the continuous subsequence that has passed verification in the draft token is first selected based on the result with the highest conditional probability. According to the above scheme, step 5) specifically includes the following: Extracting the main model for inference processing Intermediate representation generated at time , in, It is the first The hidden state of a token contains semantic information about that token; Through compression mapping function The middle representation of length m is compressed into a compressed token of fixed length s;
[0010] in, For the feature dimensions of the hidden state, To set the number of compressed tokens; This is the set of learnable parameters for the compression mapping function.
[0011] According to the above scheme, in step 5), the compression mapping function is jointly optimized using smoothed L1 loss and KL distillation loss; ; in, For joint optimization functions, and These are the weighting coefficients; Smoothing L1 loss: ensures consistency between the compressed token and the positional information of the original intermediate representation, and avoids the loss of local key features; KL distillation loss: ensuring the main model is based on compressed tokens The generated distribution is consistent with the distribution generated based on the original inference block, ensuring that the global inference logic is not deviated. Adjusting the compression mapping function using double constraints The parameters determine the generated compressed token. Preserve the information of the original inference block to the greatest extent possible.
[0012] The beneficial effects of this invention are: 1. This invention solves the problem of serial inefficiency in traditional autoregression by combining parallel generation with context compression, and also alleviates the redundant overhead of long sequence inference. 2. This method simultaneously accelerates and optimizes the context during token generation. By introducing a collaborative mechanism between compressed memory and draft verification, it constructs a streaming, lightweight, parallelizable, and low-redundancy inference process. This method does not require changes to the main model structure, is compatible with mainstream Transformer architectures, and can be widely applied to tasks requiring long-distance dependencies, such as multi-step inference, mathematical question answering, and code generation. It significantly reduces inference time and peak KV caching, improving overall inference performance without compromising model output quality. Attached Figure Description
[0013] The present invention will be further described below with reference to the accompanying drawings and embodiments. In the accompanying drawings: Figure 1 This is a flowchart of a method according to an embodiment of the present invention. Detailed Implementation
[0014] To make the objectives, technical solutions, and advantages of this invention clearer, the invention will be further described in detail below with reference to embodiments. It should be understood that the specific embodiments described herein are merely illustrative and not intended to limit the invention.
[0015] like Figure 1 As shown, a dual reasoning acceleration method based on draft prediction and compressed memory includes the following steps: 1) At each stage of inference, obtain the input context for model inference. The model has generated a token sequence. ; sequence In this sequence, the subscript 1 indicates the start position of the sequence, the subscript i indicates the end position of the sequence, and the separator : indicates a continuous interval from the start position to the end position. 2) Based on the initial input context of the model inference The model has generated a token sequence. Parallel prediction of subsequent One candidate token; ; in, This is a semi-autoregressive draft model. Indicates prediction The token sequence; that is, the subsequent tokens predicted by the draft model based on the i generated tokens. A draft token; The j-th draft token is denoted as j=1,...,k; 3) Organize the draft tokens according to a tree structure; In this embodiment, the tree structure organization is selected as a k-ary tree; 4) Master model validation; 4.1) Main Model Receive a tree-structured sequence of draft tokens, for each draft token... Calculate conditional probability; The main model for the first The conditional probability of a draft token is:
[0016] Among them, subscript This represents the set of all learnable parameters in the main model M; 4.2) Based on the conditional probability results, a continuous subsequence in the draft token that has passed verification is recorded as an inference block. ; m is the number of consecutive tokens that have passed verification; the subscript t indicates the t-th inference block; In step 4.2), the continuous subsequence that has passed verification in the draft token is selected as the result with the highest conditional probability.
[0017] Alternatively, you can choose the result with the most consecutive tokens from among the ones with the highest conditional probabilities; 5) Compress each inference block into a compressed token of fixed length s; Extracting the main model for inference processing Intermediate representation generated at time , in, It is the first The hidden state of a location token contains semantic information about that token; Through compression mapping function The intermediate representation of length m is compressed into a compressed token of fixed length s. The compressed token is essentially the memory points (semantic summary) of the reasoning block.
[0018] in, For the hidden dimension, To set the number of compressed tokens; The compression mapping function is jointly optimized using smoothed L1 loss and KL distillation loss. ; in, For joint optimization functions, and These are the weighting coefficients; Smoothing L1 loss: ensures consistency between the compressed token and the positional information of the original intermediate representation, and avoids the loss of local key features; KL distillation loss: ensuring the main model is based on compressed tokens The generated distribution is consistent with the distribution generated based on the original inference block, ensuring that the global inference logic is not deviated. Adjusting the compression mapping function using double constraints The parameters determine the generated compressed token. Preserve the information of the original inference block to the greatest extent possible.
[0019] To measure the performance of this method in context compression, the information dependency metric Dependency (Dep) is introduced, defined as:
[0020] in Indicates the currently generated token The size of the set of historical tokens of interest. The smaller the Dep value, the less the model relies on the original historical context, and the more significant the compression effect.
[0021] 6) Context refactoring; Use the newly generated compressed token The updated context set is obtained as follows: , recorded as ,in for ; 7) Use the reconstructed context for reasoning; Based on the updated context set (Proceed to step 2) The draft model is called again to generate k candidate tokens in parallel. The main model verifies the new inference block, compresses it, and updates the context until the generated sequence reaches the preset length or triggers the termination of inference.
[0022] .
[0023] To demonstrate the effectiveness of the method of this invention, we systematically compared it with the following three representative solutions: (1) Token-level speculation baseline. This type of method only improves the generation speed by generating drafts, but has a low acceptance rate and a lot of redundant verification. (2) Full sequence compression baseline, which deletes historical tokens through heuristics or scoring, but does not retain the ability to discriminate reasoning information; (3) Semantic segmentation compression baseline, which supports reducing context by segment compression, but still relies on independent judgment after generation, making it difficult to coordinate with the main model inference path.
[0024] Experiments show that, compared to existing mainstream inference acceleration methods, the draft prediction and compressed memory-based method of this invention achieves better results. Three benchmark datasets were used for evaluation in the experiments: the GSM8K dataset, the BBH-Logical Deduction dataset, and the MATH (Mini) dataset.
[0025] Table 1. Comparison Experiment Results of GSM8K
[0026] Table 2. Comparative Experimental Results of BBH-Logical Deduction
[0027] Table 3. Results of MATH (Mini) Comparison Experiment
[0028] The experimental results in Tables 1 to 3 demonstrate that the proposed method achieves significant performance improvements across several representative complex inference tasks. On datasets such as GSM8K, BBH-Logical Deduction, and MATH (Mini), the proposed method reduces average inference latency by approximately 28.5%, peak KV cache token count by over 65%, and information dependency index by approximately 50% with almost no loss in accuracy, significantly alleviating computational and memory pressure on the model during multi-step inference. Compared to speculative decoding methods that rely solely on draft generation, the proposed method significantly reduces redundant verification times and improves overall throughput efficiency through the collaborative work of structured verification and semantic compression. Compared to AnLLM, which relies on heuristic rule-based context pruning, the proposed method possesses stronger representational capabilities and stability, eliminating the need for manual pruning strategies during inference and dynamically preserving key information. Furthermore, compared to LightThinker, which only performs semantic compression, the proposed method avoids logical interruptions caused by information loss through the linkage between draft-driven generation and compression modules, achieving more stable and efficient inference acceleration and context optimization, demonstrating stronger task generalization ability and engineering practicality.
[0029] It should be understood that those skilled in the art can make improvements or modifications based on the above description, and all such improvements and modifications should fall within the protection scope of the appended claims.
Claims
1. A dual reasoning acceleration method based on draft prediction and compressed memory, characterized in that, Includes the following steps: 1) At each stage of inference, obtain the input context for model inference. The model has generated a token sequence. ; where, input context For text; sequence In this sequence, the subscript 1 indicates the start position of the sequence, the subscript i indicates the end position of the sequence, and the separator : indicates a continuous interval from the start position to the end position. 2) Based on the input context of the model inference and the token sequence already generated by the model. Parallel prediction of subsequent One candidate token is used as a draft token; In step 2), a semi-autoregressive draft model is used to predict subsequent events in parallel. The candidate tokens are as follows: ; in, This is a semi-autoregressive draft model. Indicates prediction token sequence; 3) Organize the draft tokens according to a tree structure; 4) Perform master model validation; details are as follows: 4.1) Main Model Receive a tree-structured sequence of draft tokens, for each draft token... Calculate conditional probability; In step 4.1), the main model For the The conditional probability of a draft token is: Among them, subscript Main model The set of all learnable parameters in the dataset; 4.2) Based on the conditional probability results, a continuous subsequence in the draft token that has passed verification is recorded as an inference block. ; m is the number of consecutive tokens that have passed verification, m≤k; t represents the t-th inference block; Among them, the continuous subsequence that has passed verification in the draft token is selected as the result with the highest conditional probability; 5) Compress each inference block obtained from verification into a compressed token of fixed length s; 6) Context refactoring; Use the newly generated compressed token The updated model has generated a token sequence. Obtain the updated set of contexts.
2. The dual reasoning acceleration method based on draft prediction and compressed memory according to claim 1, characterized in that, In step 5), the specific details are as follows: Extracting the main model for inference processing Intermediate representation generated at time , in, It is the first The hidden state of each token; Through compression mapping function The middle representation of length m is compressed into a compressed token of fixed length s; in, For the feature dimensions of the hidden state, To set the number of compressed tokens; This is the set of learnable parameters for the compression mapping function.
3. The dual reasoning acceleration method based on draft prediction and compressed memory according to claim 2, characterized in that, In step 5), the compression mapping function is jointly optimized using smoothed L1 loss and KL distillation loss. ; in, For joint optimization functions, and These are the weighting coefficients.
4. The dual reasoning acceleration method based on draft prediction and compressed memory according to claim 1, characterized in that, In step 6), the updated context set is obtained as follows: , recorded as ,in for .
5. A computer device comprising a memory and a processor, wherein the memory stores a computer program, characterized in that, When the processor executes the computer program, it implements the steps of the method according to any one of claims 1 to 4.
6. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by a processor, it implements the method described in any one of claims 1 to 4.
Citation Information
Patent Citations
Large language model reasoning acceleration method and system based on progressive grassy tree
CN120654833A
Fast Speculative Decoding Using Multiple Parallel Drafts
US20250209355A1