Self-speculation decoding method and system based on layered quantization KV cache
By hierarchically quantizing the KV cache of the target model, the memory and computing bottlenecks of the KV cache in long-context tasks are solved, the inference efficiency and generation quality are improved, and the efficient alignment of the draft model and the target model is achieved.
Patent Information
- Application Number
- CN202510338366.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-21
- Publication Date
- 2025-06-20
AI Technical Summary
The prior art faces memory and computing bottlenecks in KV caches, alignment problems between draft models and target models, and insufficient compatibility between quantization and speculative decoding, resulting in poor generation quality and efficiency.
The KV cache of the target model is optimized by hierarchical quantization processing, and it is decomposed into high-digit quantization parts and low-digit quantization parts. The draft model generates candidate token sequences based on the high-digit quantization parts. The target model is verified in parallel in the asynchronous queue and updated the hierarchical quantized KV cache.
It significantly reduces the memory usage of KV cache, improves inference efficiency, ensures generation quality, and improves semantic alignment between the draft model and the target model.
Smart Images

Figure CN120181241A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the technical fields of natural language processing (NLP) and deep learning, and relates to a self-speculative decoding method and system based on hierarchical quantization KV cache. Background Art
[0002] Currently, with the wide application of large language models (LLMs), the problems of inference efficiency and memory occupancy in long context tasks are becoming increasingly prominent. Traditional autoregressive decoding methods need to generate tokens one by one, resulting in a linear growth relationship between the generation speed and the context length. For example, generating a thousand-word document may take several minutes, severely limiting application scenarios with high real-time requirements (such as interactive dialogue systems). For this reason, the speculative decoding technology has been proposed. By introducing a lightweight draft model to pre-generate a candidate token sequence and then quickly verified by the target model, the number of calls to the target model can be reduced. However, the existing speculative decoding solutions still face the following key problems in long context tasks:
[0003] Memory and computational bottlenecks of KV cache:
[0004] During the autoregressive generation process, the model needs to cache the key-value vectors (KV cache) of historical tokens to accelerate attention calculation. For long context tasks (such as processing an input of 4096 tokens), the memory occupancy of the KV cache can reach dozens of GB (taking FP16 precision as an example), far exceeding the hardware video memory capacity. Existing technologies use INT8 quantization to compress the KV cache, but the quantization process needs to frequently perform dequantization operations to restore the accuracy, resulting in additional computational overhead. In addition, quantization errors may accumulate, affecting the generation quality, especially more significantly in long sequence generation.
[0005] Alignment problem between the draft model and the target model:
[0006] The core of speculative decoding is that the candidate tokens generated by the draft model need to be highly consistent with the output of the target model. However, existing solutions usually use a draft model with a much smaller number of parameters than the target model (such as 1 / 10 of the number of parameters). Such small models are difficult to fully capture complex semantic dependencies in long context tasks, resulting in a decrease in the acceptance rate of candidate tokens. For example, in long dialogue generation, the draft model may mispredict the turning points of the dialogue logic, forcing the target model to frequently regenerate, which instead reduces the overall speedup ratio.
[0007] Insufficient compatibility between quantization and speculative decoding:
[0008] Existing work has attempted to combine KV cache quantization with speculative decoding, but faces a dilemma: if only the KV cache of the target model is quantized, the draft model still needs to load the full-precision cache, and the memory occupancy is not significantly reduced; if both the draft model and the target model are quantized, the low-precision (e.g., 4-bit) weights will cause the quality of the draft model's generation to deteriorate, further reducing the acceptance rate.
[0009] In addition, traditional quantization schemes (such as uniform quantization) are not optimized for the asynchronous characteristics of speculative decoding, resulting in the quantization / dequantization operations being unable to be efficiently parallelized with the candidate generation and verification processes.
[0010] Unique challenges of long-context tasks:
[0011] Long text generation requires the model to maintain a coherent understanding of long-distance context, which poses higher requirements for the accuracy and capacity of the KV cache. Due to memory limitations, existing speculative decoding schemes often need to truncate or segment long contexts, resulting in inconsistent generated content. For example, when generating a technical document, if the draft model ignores early-defined terms due to insufficient cache capacity, the target model may not be able to correct such errors, and the final output will not meet the requirements.
[0012] Regarding the above problems, the prior art has not provided an effective solution.
[0013] It should be noted that the information disclosed in the above background art section is only used to enhance the understanding of the background of the present application, and thus may include information that does not constitute the prior art known to those of ordinary skill in the art. Summary of the Invention
[0014] To have a basic understanding of some aspects of the disclosed embodiments, a simple summary is given below. The summary is not a general review, nor is it intended to identify key / important components or delineate the protection scope of these embodiments, but rather serves as a preface to the subsequent detailed description.
[0015] The embodiments of the present disclosure provide a self-speculative decoding method and system based on hierarchical quantization of the KV cache, which significantly accelerate the inference process of large language models while maintaining the generation quality by optimizing the management and quantization strategy of the KV cache.
[0016] In some embodiments, the method includes:
[0017] Perform hierarchical quantization processing on the KV cache of the target model, decomposing the original cache data into a high-precision quantization part and a low-precision quantization part;
[0018] The draft model generates a candidate token sequence based on the KV cache of the high-precision quantization part and stores the candidate token sequence in an asynchronous queue;
[0019] The target model retrieves the candidate token sequence from the asynchronous queue, combines the high-precision quantization part and the low-precision quantization part into a full-precision cache, and performs parallel verification on the candidate token sequence;
[0020] According to the verification result, accept the valid candidate token or have the target model regenerate a replacement token, and update the hierarchical quantization KV cache.
[0021] Preferably, the hierarchical quantization process includes:
[0022] Quantize the original KV cache into a 4-bit high-precision representation and a 4-bit low-precision representation, where the 4-bit high-precision representation is used for the inference of the draft model, and the combination of the 4-bit high-precision and the 4-bit low-precision is used as an INT8 precision representation for the verification of the target model.
[0023] Preferably, the weights of the draft model are represented in 4-bit quantization, and the weights of the target model are represented in full-precision 16-bit floating-point numbers.
[0024] Preferably, the asynchronous queue is a first-in-first-out queue, which is used to implement the asynchronous parallel processing of the draft model and the target model. The draft model continuously generates candidate token sequences and writes them into the queue, and the target model reads the candidate sequences from the queue in batches for verification.
[0025] Preferably, the parallel verification includes:
[0026] The target model calculates the probability distribution of multiple candidate token sequences at one time, and determines whether to accept the candidate token according to a preset threshold; if the candidate token is rejected, the target model generates a new token to replace the rejected part.
[0027] Preferably, the update of the hierarchical quantization KV cache includes:
[0028] Re-perform hierarchical quantization on the KV cache corresponding to the new token generated by the target model, and update the high-precision and low-precision quantization parts.
[0029] In some embodiments, the self-speculative decoding system based on the hierarchical quantization KV cache includes:
[0030] Hierarchical quantization module: used to decompose the KV cache of the target model into a high-precision quantization part and a low-precision quantization part;
[0031] Draft model module: load 4-bit quantization weights and generate candidate token sequences based on the high-precision quantization KV cache;
[0032] Asynchronous queue module: store the candidate token sequences generated by the draft model;
[0033] Target model module: Load 16-bit full-precision weights, combine the high and low quantization parts into an INT8 cache, batch verify candidate sequences, and generate new tokens to replace the tokens that fail the verification.
[0034] Cache update module: Update the hierarchical quantization KV cache according to the verification results.
[0035] Preferably, the hierarchical quantization module is further configured to: Generate a 4-bit low-bit representation through quantization error calculation, and dynamically combine it with the 4-bit high-bit to form an INT8 cache.
[0036] In some embodiments, the self-speculative decoding device based on the hierarchical quantization KV cache includes a processor and a memory storing program instructions, and the processor is configured to run the self-speculative decoding method based on the hierarchical quantization KV cache.
[0037] In some embodiments, the computer-readable storage medium stores a computer program, and when the program is executed by a processor, it implements the self-speculative decoding method based on the hierarchical quantization KV cache.
[0038] A self-speculative decoding method and system based on the hierarchical quantization KV cache provided by the embodiments of the present disclosure can achieve the following technical effects:
[0039] By decomposing the KV cache into a 4-bit high-bit and a 4-bit low-bit representation, the draft model only needs to load the high 4-bit part (accounting for 50% of the original INT8 cache), and the target model dynamically combines the high and low bits to restore the INT8 precision, reducing the memory occupancy of the KV cache.
[0040] The draft model and the target model are decoupled through a FIFO queue. The draft model continuously generates candidate tokens, and the target model batch verifies multiple groups of candidate sequences in the queue, eliminating the serial waiting time of "generation-verification" in traditional speculative decoding.
[0041] The target model uses full-precision weights to ensure the accuracy of the probability distribution calculated in the verification stage and avoid the problem of error accumulation introduced by low-precision models. The draft model generates candidate tokens based on the 4-bit high-bit KV cache, and the target model uses the complete INT8 cache during verification. The two share the same cache structure, with a higher semantic alignment degree.
[0042] The above general description and the following description are only exemplary and explanatory, and are not used to limit this application. Description of the Drawings
[0043] One or more embodiments are exemplarily illustrated by corresponding drawings. These exemplary illustrations and the drawings do not constitute a limitation on the embodiments. Elements with the same reference numerals in the drawings are shown as similar elements. The drawings do not constitute a scale limitation, and wherein:
[0044] Figure 1 Data processing block diagram of a self-speculative decoding model based on a hierarchical quantization KV cache;
[0045] Figure 2 Flow block diagram of a self-speculative decoding method based on a hierarchical quantization KV cache;
[0046] Figure 3 It is a schematic diagram of the method flow described in the present invention;
[0047] Figure 4 It is a schematic diagram of the device structure of the present invention. Detailed implementation manners
[0048] In order to be able to understand the features and technical content of the embodiments of the present disclosure in more detail, the implementation of the embodiments of the present disclosure will be described in detail below with reference to the drawings. The attached drawings are for reference and illustration purposes only and are not used to limit the embodiments of the present disclosure. In the following technical descriptions, for the sake of explanation, multiple details are provided to provide a full understanding of the disclosed embodiments. However, one or more embodiments can still be implemented without these details. In other cases, well-known structures and devices can be shown in a simplified manner.
[0049] In the embodiments of the present disclosure, terms such as "first" and "second" in the specification and claims and the above drawings are used to distinguish similar objects and do not necessarily describe a specific order or sequence. It should be understood that the data used in this way can be interchanged under appropriate circumstances so as to implement the embodiments of the present disclosure described herein. In addition, the terms "comprising" and "having" and any variations thereof are intended to cover non-exclusive inclusion.
[0050] Unless otherwise specified, the term "plurality" means two or more.
[0051] In the embodiments of the present disclosure, the character " / " indicates that the objects before and after are in an "or" relationship. For example, A / B means: A or B.
[0052] The term "and / or" is a description of the association relationship of an object and indicates that three relationships can exist. For example, A and / or B means: A or B, or, the three relationships of A and B.
[0053] The term "corresponding" can refer to an association relationship or a binding relationship. A corresponding to B means that there is an association relationship or a binding relationship between A and B.
[0054] As Figures 1-3 , a self-speculative decoding method based on hierarchical quantization of KV cache, the method comprising:
[0055] S1: Perform hierarchical quantization processing on the KV cache of the target model, and decompose the original cache data into a high-bit quantization part and a low-bit quantization part.
[0056] That is, quantize the FP16 KV cache into high and low 4-bit representations, specifically as follows:
[0057] Calculate the high 4-bit representation.
[0058] Calculate the quantization error and quantize it into a low 4-bit representation.
[0059] S2: The draft model generates a candidate token sequence based on the KV cache of the high-bit quantization part and stores the candidate token sequence in an asynchronous queue.
[0060] The draft model only loads the high 4-bit representation.
[0061] S3: The target model obtains the candidate token sequence from the asynchronous queue, combines the high-bit quantization part and the low-bit quantization part into a full-precision cache, and performs parallel verification on the candidate token sequence.
[0062] The target model loads the high and low 4-bit representations and combines them into a KV cache with INT8 precision.
[0063] The parallel verification includes:
[0064] The target model calculates the probability distribution for multiple candidate token sequences at one time and determines whether to accept the candidate token according to a preset threshold; if the candidate token is rejected, the target model generates a new token to replace the rejected part.
[0065] S4: According to the verification result, accept the valid candidate token or have the target model regenerate a replacement token, and update the hierarchical quantization KV cache.
[0066] As a refinement of the above embodiment, the hierarchical quantization processing includes:
[0067] Quantize the original 8-bit KV cache into a high 4-bit representation and a low 4-bit representation, where the high 4-bit representation is used for the inference of the draft model, and the combination of the high 4 bits and the low 4 bits is used for the verification of the target model with INT8 precision.
[0068] As a refinement of the above embodiment, the weights of the draft model are highly quantized and represented in 4 bits, aiming to quickly generate a set of candidate tokens for the target model to verify;
[0069] The weights of the target model are represented by 16-bit floating-point numbers in full precision, aiming to verify the tokens generated by the draft model, regenerate the tokens that fail the verification, and update the corresponding KVCache. The target model can receive multiple token sequences as input at one time, so as to verify multiple candidate token sequences at one time, greatly improving the verification efficiency.
[0070] Since the draft model is highly quantized, the speed of generating candidate tokens by the draft model is several times that of the verification speed of the target model.
[0071] As a refinement of the above embodiment, the asynchronous queue is a first-in-first-out queue (FIFO). The speed of generating candidate tokens by the draft model is several times that of the verification speed of the target model. To improve the inference efficiency and reduce the unnecessary waiting caused by the verification work of the target model, the draft model and the target model work asynchronously in parallel, that is, while the draft model generates candidate tokens, the target model verifies the current candidate tokens. To achieve this goal, a "draft model output FIFO" is set between the draft model and the target model. After the draft model generates a candidate token sequence, it enqueues the candidate token sequence. After the target model completes the previous verification, it inputs all the candidate token sequences stored in the draft model output FIFO into the model for verification. After the verification is completed, these candidate token sequences are dequeued.
[0072] A self-speculative decoding system based on hierarchical quantization KV cache, comprising:
[0073] Hierarchical quantization module: used to decompose the KV cache of the target model into a high-precision quantization part and a low-precision quantization part.
[0074] Draft model module: loads 4-bit quantization weights and generates candidate token sequences based on the high-precision quantization KV cache.
[0075] Asynchronous queue module: stores the candidate token sequences generated by the draft model.
[0076] Target model module: loads 16-bit weights in full precision, combines the high-precision and low-precision quantization parts into an INT8 cache, batch-verifies the candidate sequences, and generates new tokens to replace the tokens that fail the verification.
[0077] Cache update module: updates the hierarchical quantization KV cache according to the verification results.
[0078] Combined with Figure 4As shown, an apparatus 300 for self-speculative decoding based on a hierarchical quantization KV cache according to an embodiment of the present disclosure includes a processor 304 and a memory 301. Optionally, the apparatus may further include a communication interface 302 and a bus 303. Among them, the processor 304, the communication interface 302, and the memory 301 can communicate with each other through the bus 303. The communication interface 302 can be used for information transmission. The processor 304 can call the logical instructions in the memory 301 to execute the self-speculative decoding method based on the hierarchical quantization KV cache in the above embodiment.
[0079] In addition, when the logical instructions in the above-mentioned memory 301 are implemented in the form of software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium.
[0080] The memory 301, as a computer-readable storage medium, can be used to store software programs and computer-executable programs, such as the program instructions / modules corresponding to the method in the embodiment of the present disclosure. The processor 304 executes functional applications and data processing by running the program instructions / modules stored in the memory 301, that is, implements the self-speculative decoding method based on the hierarchical quantization KV cache in the above embodiment.
[0081] The memory 301 may include a program storage area and a data storage area. Among them, the program storage area can store an operating system and application programs required for at least one function; the data storage area can store data created according to the use of the terminal device, etc. In addition, the memory 301 may include a high-speed random access memory and may also include a non-volatile memory.
[0082] An embodiment of the present disclosure provides a computer-readable storage medium storing computer-executable instructions, and the computer-executable instructions are set to execute the self-speculative decoding method based on the hierarchical quantization KV cache.
[0083] The above-mentioned computer-readable storage medium may be a transient computer-readable storage medium or a non-transient computer-readable storage medium.
[0084] The technical solution of the embodiments of the present disclosure can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes one or more instructions for causing a computer device (which may be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the method described in the embodiments of the present disclosure. The foregoing storage medium may be a non-transitory storage medium, including: various media capable of storing program codes such as USB flash drives, mobile hard disks, read-only memories (ROMs), random access memories (RAMs), magnetic disks, or optical discs, or may also be a transitory storage medium.
[0085] The above description and drawings fully illustrate the embodiments of the present disclosure so that those skilled in the art can practice them. Other embodiments may include structural, logical, electrical, process, and other changes. The embodiments merely represent possible variations. Unless explicitly required, individual components and functions are optional, and the order of operations may vary. Parts and features of some embodiments may be included in or substituted for parts and features of other embodiments. Moreover, the terms used in this application are only for describing the embodiments and are not used to limit the claims. As used in the description of the embodiments and the claims, unless the context clearly indicates otherwise, the singular forms "a", "an", and "the" are intended to include the plural forms as well. Similarly, as used in this application, the term "and / or" refers to any and all possible combinations including one or more of the associated listed items. Additionally, when used in this application, the term "comprise" and its variants "comprises" and / or "comprising" etc. mean the presence of the stated features, wholes, steps, operations, elements, and / or components, but do not exclude the presence or addition of one or more other features, wholes, steps, operations, elements, components, and / or groupings thereof. Without further limitation, an element defined by the statement "comprising one..." does not exclude the presence of additional identical elements in the process, method, or device comprising the element. In this document, each embodiment may focus on the differences from other embodiments, and the same or similar parts among the various embodiments may be referred to each other. For the methods, products, etc. disclosed in the embodiments, if they correspond to the method part disclosed in the embodiments, the relevant parts may refer to the description of the method part.
[0086] Those skilled in the art will realize that the units and algorithm steps of each example described in combination with the embodiments disclosed herein can be implemented by electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are executed in hardware or software can depend on the specific application and design constraints of the technical solution. The skilled person can use different methods for each specific application to implement the described functions, but such implementation should not be considered to exceed the scope of the embodiments of the present disclosure. The skilled person can clearly understand that for the convenience and brevity of description, the specific working processes of the systems, devices, and units described above can refer to the corresponding processes in the foregoing method embodiments and will not be elaborated herein.
[0087] In the embodiments disclosed herein, the disclosed methods, products (including but not limited to devices, equipment, etc.) can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the division of the units can be merely a logical function division, and there can be other division methods in actual implementation. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Additionally, the displayed or discussed couplings or direct couplings or communication connections to each other can be through some interfaces. The indirect couplings or communication connections of the devices or units can be in electrical, mechanical, or other forms. The units described as separate components can be or can not be physically separated, and the components displayed as units can be or can not be physical units, that is, they can be located in one place or can be distributed to multiple network units. Some or all of the units can be selected according to actual needs to implement this embodiment. Additionally, in the embodiments of the present disclosure, the functional units can be integrated in one processing unit, or each unit can exist physically separately, or two or more units can be integrated in one unit.
[0088] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to embodiments of the present disclosure. In this regard, each block in the flowchart or block diagram may represent a module, a segment of a program, or a portion of code, which contains one or more executable instructions for implementing a specified logical function. In some alternative implementations, the functions noted in the blocks may occur in a different order than noted in the accompanying drawings. For example, two consecutive blocks may actually be executed substantially in parallel, or they may sometimes be executed in the reverse order, depending on the functions involved. In the description corresponding to the flowcharts and block diagrams in the accompanying drawings, the operations or steps corresponding to different blocks may also occur in a different order than disclosed in the description, and sometimes there is no specific order between different operations or steps. For example, two consecutive operations or steps may actually be executed substantially in parallel, or they may sometimes be executed in the reverse order, depending on the functions involved. Each block in the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, can be implemented by a dedicated hardware-based system that performs the specified functions or actions, or can be implemented by a combination of dedicated hardware and computer instructions.
Claims
1. A self-speculative decoding method based on hierarchical quantization KV cache, characterized in that: The following steps are involved: Perform hierarchical quantization on the KV cache of the target model, decomposing the original cache data into a high-bit quantization part and a low-bit quantization part; The draft model generates a candidate token sequence based on the KV cache of the high-order quantization part, and stores the candidate token sequence in an asynchronous queue; The target model obtains the candidate token sequence from the asynchronous queue, combines the high-order quantized part and the low-order quantized part into a full-precision cache, and performs parallel verification on the candidate token sequence; Based on the verification results, the valid candidate token is accepted or the replacement token is regenerated by the target model, and the hierarchical quantization KV cache is updated.
2. The self-speculative decoding method based on hierarchical quantization KV cache according to claim 1, characterized in that: The hierarchical quantization process comprises: The original KV cache is quantized into a high-order 4-bit representation and a low-order 4-bit representation, wherein the high-order 4-bit representation is used for reasoning of the draft model, and the high-order 4 bits and the low-order 4 bits are combined into an INT8 precision representation for verification of the target model.
3. The self-speculative decoding method based on hierarchical quantization KV cache according to claim 1, characterized in that: The weights of the draft model are represented by 4-bit quantization, and the weights of the target model are represented by full-precision 16-bit floating point numbers.
4. The self-speculative decoding method based on hierarchical quantization KV cache according to claim 1, characterized in that: The asynchronous queue is a first-in-first-out queue, which is used to realize asynchronous parallel processing of the draft model and the target model. The draft model continuously generates candidate token sequences and writes them into the queue, and the target model reads candidate sequences in batches from the queue for verification.
5. The self-speculative decoding method based on hierarchical quantization KV cache according to claim 1, characterized in that: The parallel verification includes: The target model calculates the probability distribution of multiple candidate token sequences at one time, and determines whether to accept the candidate token based on the preset threshold; if the candidate token is rejected, the target model generates a new token to replace the rejected part.
6. The self-speculative decoding method based on hierarchical quantization KV cache according to claim 1, characterized in that: The updating of the hierarchical quantization KV cache includes: Re-quantize the KV cache corresponding to the new token generated by the target model, and update the high-order and low-order quantization parts.
7. A self-speculative decoding system based on hierarchical quantization KV cache using any of the methods of claims 1-6, characterized in that: include: Hierarchical quantization module: used to decompose the KV cache of the target model into a high-bit quantization part and a low-bit quantization part; Draft model module: loads 4-bit quantization weights and generates candidate token sequences based on high-bit quantization KV cache; Asynchronous queue module: stores the candidate token sequence generated by the draft model; Target model module: loads full-precision 16-bit weights, combines high-bit and low-bit quantized parts into INT8 cache, performs batch verification on candidate sequences, and generates new tokens to replace tokens that fail verification; Cache update module: updates the hierarchical quantized KV cache based on the verification results.
8. The system according to claim 7, characterized in that The hierarchical quantization module is further configured as follows: The low-order 4-bit representation is generated by quantization error calculation and dynamically combined with the high-order 4 bits into INT8 cache.
9. A self-speculative decoding device based on hierarchical quantization KV cache, comprising a processor and a memory storing program instructions, characterized in that: The processor is configured to run the self-speculative decoding method based on hierarchical quantization KV cache according to any one of claims 1-6.
10. A computer-readable storage medium, characterized in that: A computer program is stored thereon, and when the program is executed by a processor, the self-speculative decoding method based on hierarchical quantization KV cache as described in any one of claims 1 to 6 above is implemented.
Citation Information
Cited By
KV cache data quantification device and method
CN121031682A
Inference acceleration method, electronic equipment, storage medium and program product
CN121706943A
Machine-learned speculative decoding engines
WO2026050635A1