Quantization method and device for realizing elastic KV cache by computing power through intelligent computing cloud platform

CN122507325BActive Publication Date: 2026-09-29DATACANVAS LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202611001046.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2026-07-07
Publication Date
2026-09-29
Estimated Expiration
2046-07-07

AI Technical Summary

Technical Problem

[0009]本发明提供一种智能计算云平台通过算力实现弹性KV缓存的量化方法及装置,以解决现有技术中存在智能计算云平台中缺少一种同时兼顾“低比特长期保存、局部按需恢复、触发范围可控、纠偏结果可增量融合”的KV缓存的量化方法的问题

Benefits of technology

[0029]在本发明中,步骤S1、将智能计算云平台的历史token划分为多个缓存块,并将所述历史token的Key数据和/或Value数据进行量化,写入对应缓存块中,在所述历史token的高精度KV数据中选取多个位置作为候选锚点;步骤S2、计算所述缓存块的摘要Key数据和量化误差代理数据,作为所述缓存块的块级摘要数据;步骤S3、响应于所述智能计算云平台生成新的目标token,基于所述目标token对所述缓存块进行打分,以生成缓存块评分,并基于所述缓存块评分从多个缓存块中筛选出候选缓存块;步骤S4、计算所述候选缓存块的不确定性指标数据,并根据所述不确定性指标数据从所述候选缓存块中确定待恢复精度的目标缓存块;步骤S5、定位距离所述目标缓存块最近的前序候选锚点作为上游目标锚点,并基于所述上游目标锚点对所述历史token的高精度KV数据进行局部回放,以生成所述目标缓存块的目标高精KV数据。这样,本发明通过步骤S1至S5构建了一套轻量、智能、闭环的KV缓存动态管理机制,在智能计算云平台中实现了高效率与高质量的协同优化:智能计算云平台以低比特量化方式长期保存历史KV缓存并稀疏设置高精度锚点,显著降低显存占用;通过块级摘要与Query驱动的打分快速筛选相关缓存块,并进一步结合不确定性指标精准识别真正需要纠偏的目标缓存块。该方案在不保留全量高精度KV、不引入额外重建模型、不扩大回放范围的前提下,以极小的追加计算代价,有效抑制了低比特量化带来的语义偏差和错误累积,同时保障了智能计算云平台高并发场景下的低时延与资源可调度性,显著提升了超长上下文大模型推理的准确性、稳定性。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122507325B_ABST
    Figure CN122507325B_ABST
Patent Text Reader

Abstract

The application provides a method and device for quantifying elastic KV cache through computing power of an intelligent computing cloud platform, and relates to the technical fields of intelligent computing centers, intelligent computing centers, computing power infrastructure and intelligent computing cloud technology.The method comprises the following steps: S1, dividing historical tokens into multiple cache blocks, quantifying KV data and writing the data into corresponding cache blocks, and selecting multiple candidate anchor points; S2, calculating block-level summary data; S3, in response to a new target token, scoring the cache blocks to generate cache block scores and screening out candidate cache blocks; S4, calculating uncertainty index data and determining a target cache block with to-be-restored precision according to the uncertainty index data; and S5, locating an upstream target anchor point and locally playing back the historical tokens based on the upstream target anchor point to generate target high-precision KV data of the target cache block.The application can greatly improve the quantification effect of KV cache of the intelligent computing cloud platform.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the fields of intelligent computing centers, smart computing centers, computing infrastructure, and smart cloud technologies, specifically to a method and apparatus for quantizing elastic key-value caching through computing power in an intelligent computing cloud platform. Background Technology

[0002] With the rapid development of artificial intelligence technology, "intelligent computing centers" and "smart computing centers" have emerged.

[0003] An "intelligent computing center" refers to a facility that provides the necessary computing power, data, and algorithms for artificial intelligence applications (such as the development, training, and inference of deep learning models) by utilizing large-scale heterogeneous computing resources, including general-purpose and intelligent computing power. Intelligent computing centers encompass facilities, hardware, and software, and can provide full-stack capabilities from underlying computing power to top-level application enablement.

[0004] "Intelligent computing center" includes, but is not limited to, "intelligent computing center".

[0005] "Intelligent computing center" or artificial intelligence computing center is a type of computing infrastructure that provides computing power services, data services, and algorithm services required for artificial intelligence applications, based on artificial intelligence theory and adopting artificial intelligence computing architecture.

[0006] "Computing power" is the core of "intelligent computing center" and "smart computing center". It is the ability of computer equipment or computing / data center to process parameters. It is the ability of computer hardware and software to work together to execute a certain computing requirement. It is the computing power to achieve the target result output by processing parameter data. It is a new type of productivity that integrates parameter computing power, network carrying capacity and data storage capacity. It mainly provides services to society through computing power infrastructure.

[0007] With the large-scale deployment of large language model inference services on intelligent computing cloud platforms, the ability to generate long contexts has become a key driver for improving user experience and expanding application scenarios. However, as the length of contexts continues to grow (e.g., 32K, 100K, or even millions of tokens), the model needs to cache the key and value vectors corresponding to all historical tokens during autoregressive decoding to support attention computation. This results in the KV cache consuming a large amount of GPU memory, severely restricting the concurrency density, energy efficiency, and service cost of intelligent computing cloud platforms. To alleviate storage pressure, existing technologies commonly use low-bit quantization to compress the KV cache. Although this method significantly reduces memory usage, low-precision representation inevitably introduces quantization errors, especially in long sequences, which can easily cause attention weight shifts, thus affecting the output quality of the current generation step. Once low-bit KV has participated in the current inference and caused a deviation, if it is not corrected in time, the error may gradually accumulate through the autoregressive mechanism, leading to semantic drift or even logical collapse.

[0008] It is evident that since the emergence of intelligent computing centers, there has been a lack of a quantization method for KV cache that simultaneously achieves "long-term storage of low-bit values, local on-demand recovery, controllable trigger range, and incremental fusion of correction results." This problem has been an urgent challenge to be solved in this field. Summary of the Invention

[0009] This invention provides a quantization method and apparatus for elastic KV cache in an intelligent computing cloud platform using computing power, in order to solve the problem in the prior art that there is a lack of a KV cache quantization method in intelligent computing cloud platforms that simultaneously takes into account "long-term storage of low bits, local on-demand recovery, controllable trigger range, and incremental fusion of correction results".

[0010] To solve the above problems, the present invention is implemented as follows: In a first aspect, the present invention provides a method for quantizing elastic key-value caching through computing power in an intelligent computing cloud platform, comprising: Step S1: Divide the historical tokens of the intelligent computing cloud platform into multiple cache blocks, quantify the key data and / or value data of the historical tokens, write them into the corresponding cache blocks, and select multiple positions as candidate anchor points in the high-precision KV data of the historical tokens. Step S2: Calculate the digest key data and quantization error proxy data of the cache block as the block-level digest data of the cache block; Step S3: In response to the intelligent computing cloud platform generating a new target token, the cache block is scored based on the target token to generate a cache block score, and candidate cache blocks are selected from multiple cache blocks based on the cache block score; Step S4: Calculate the uncertainty index data of the candidate cache blocks, and determine the target cache block whose accuracy needs to be restored from the candidate cache blocks based on the uncertainty index data; Step S5: Locate the nearest preceding candidate anchor point to the target cache block as the upstream target anchor point, and perform partial replay of the high-precision KV data of the historical token based on the upstream target anchor point to generate the target high-precision KV data of the target cache block.

[0011] In one embodiment, step S1 includes: Step S1.1: Quantize the Key data based on a first quantization granularity and quantize the Value data based on a second quantization granularity, wherein the first quantization granularity is greater than the second quantization granularity.

[0012] In one embodiment, step S1 includes: Step S1.2: For the high-precision KV data of the historical tokens, set a candidate anchor point every preset time interval according to the time sequence; Each candidate anchor point includes at least one of the following data: the hidden state of the predetermined network layer, the compressed representation of the input of the predetermined network layer, and the checkpoint state.

[0013] In one embodiment, step S2 includes: Step S2.1: Calculate the mean squared error of the cache block before and after updating the historical token, and / or calculate the historical energy statistics and / or activity value and / or prior data of the cache block; Step S2.2: Calculate the quantization error proxy data based on the mean square error value and / or the historical energy statistics value and / or the activity value and / or the prior data.

[0014] In one embodiment, step S3 includes: Step S3.1: Based on the similarity between the target token and the cache block across multiple attention heads, calculate the access tendency data of the cache block; Step S3.2: Calculate the cache block score of the cache block based on the similarity, the access tendency data, and the prior data of the cache block.

[0015] In one embodiment, step S3.2 includes: Step S3.2.1: Take the maximum value of the similarity between the target token and the cache block on multiple attention heads as the target similarity, and perform a weighted calculation based on the target similarity, the access tendency data and the prior data of the cache block to generate the cache block score.

[0016] In one embodiment, step S4 includes: Step S4.1: Calculate the score proximity data between the candidate buffer block and its neighboring candidate buffer blocks, calculate the quantization error value of the candidate buffer block, calculate the output distribution entropy of the current decoding step, and calculate the importance fluctuation data of the candidate buffer block; Step S4.2: Perform a weighted calculation based on the score proximity data, the quantization error value, the output distribution entropy, and the importance fluctuation data to generate the uncertainty index data.

[0017] In one embodiment, step S5 includes: Step S5.1: Perform multi-dimensional cropping on the data between the upstream target anchor point and the target cache block to generate the first candidate data between the upstream target anchor point and the target cache block; Step S5.2: Filter out the first candidate data to generate the second candidate data corresponding to the target cache block; Step S5.3: Calculate the attention contribution difference between the second candidate data and the low-bit data; Step S5.4: Correct the basic output data of the target cache block based on the attention contribution difference to obtain the target high-precision KV data.

[0018] In one embodiment, the multi-dimensional cropping includes one or more of the following: Perform time-dimensional pruning on the data between the upstream target anchor point and the target cache block; Perform layer-by-layer dimensional pruning on the data between the upstream target anchor point and the target cache block; Head dimension pruning is performed on the data between the upstream target anchor point and the target cache block; Channel dimension pruning is performed on the data between the upstream target anchor point and the target cache block; The data between the upstream target anchor point and the target cache block is pruned to a precision dimension.

[0019] In one embodiment, step S5 includes: Step S5.5: Based on the upstream target anchor point, perform partial replay of the Key data of the target cache block to generate high-precision target Key data; Step S5.6: In response to the uncertainty index data of the target high-precision key data being greater than the judgment threshold, the value data of the target cache block is partially replayed based on the upstream target anchor point to generate target high-precision value data, and the target high-precision KV data is generated based on the target high-precision value data and the target high-precision key data.

[0020] In one embodiment, step S5 includes: Step S5.7: In response to the uncertainty index data of the target high-precision key data being less than or equal to the judgment threshold, the target high-precision key data is used as the target high-precision KV data.

[0021] In one embodiment, step S5 includes: Step S5.8: Based on the upstream target anchor point, perform partial playback of the target cache block according to the first recovery accuracy to generate candidate high-precision KV data; Step S5.9: In response to the uncertainty index data of the candidate high-precision KV data being greater than the judgment threshold, the target cache block is partially replayed based on the upstream target anchor point according to the second recovery accuracy to generate target high-precision KV data, wherein the second recovery accuracy is greater than the first recovery accuracy.

[0022] In one embodiment, step S5 includes: Step S5.9: In response to the uncertainty index data of the candidate high-precision KV data being less than or equal to the judgment threshold, the candidate high-precision KV data is used as the target high-precision KV data.

[0023] In one embodiment, the method further includes: In response to the target high-precision KV data meeting the release condition, the target high-precision KV data is released.

[0024] In one embodiment, the release condition includes one or more of the following: The target high-precision KV data is used continuously a set number of times; The popularity of the target high-precision KV data is less than the popularity threshold; The existing pressure value of the intelligent computing cloud platform is greater than the pressure threshold.

[0025] Secondly, the present invention also provides a quantization device for an intelligent computing cloud platform to achieve elastic KV caching through computing power, comprising: The partitioning module is used to divide the historical tokens of the intelligent computing cloud platform into multiple cache blocks, quantify the key data and / or value data of the historical tokens, write them into the corresponding cache blocks, and select multiple positions as candidate anchor points in the high-precision KV data of the historical tokens. The calculation module is used to calculate the digest key data and quantization error proxy data of the cache block as the block-level digest data of the cache block; The scoring module is used to respond to the intelligent computing cloud platform generating a new target token, score the cache block based on the target token to generate a cache block score, and filter candidate cache blocks from multiple cache blocks based on the cache block score; A filtering module is used to calculate the uncertainty index data of the candidate cache blocks and determine the target cache block whose accuracy needs to be restored from the candidate cache blocks based on the uncertainty index data. The replay module is used to locate the nearest preceding candidate anchor point to the target cache block as the upstream target anchor point, and to perform partial replay of the high-precision KV data of the historical token based on the upstream target anchor point to generate the target high-precision KV data of the target cache block.

[0026] Thirdly, the present invention also provides an electronic device, including a processor, a memory, and a computer program stored in the memory and executable on the processor. When the computer program is executed by the processor, it implements the steps in the quantization method of the intelligent computing cloud platform for achieving elastic KV caching through computing power as described in the first aspect above.

[0027] Fourthly, the present invention also provides a computer-readable storage medium storing a computer program, which, when executed by a processor, implements the steps in the quantization method of the intelligent computing cloud platform for achieving elastic KV caching through computing power as described in the first aspect above.

[0028] Fifthly, the present invention also provides a computer program product, including computer instructions, which, when executed by a processor, implement the steps in the quantization method of the intelligent computing cloud platform for achieving elastic KV caching through computing power as described in the first aspect above.

[0029] In this invention, step S1 involves dividing the historical tokens of the intelligent computing cloud platform into multiple cache blocks, quantizing the key and / or value data of the historical tokens, writing them into the corresponding cache blocks, and selecting multiple positions as candidate anchor points in the high-precision KV data of the historical tokens; step S2 involves calculating the digest key data and quantization error proxy data of the cache blocks as block-level digest data of the cache blocks; step S3 involves responding to the intelligent computing cloud platform generating a new target token, scoring the cache blocks based on the target token to generate cache block scores, and selecting candidate cache blocks from multiple cache blocks based on the cache block scores; step S4 involves calculating the uncertainty index data of the candidate cache blocks, and determining the target cache block whose precision needs to be restored from the candidate cache blocks based on the uncertainty index data; step S5 involves locating the nearest preceding candidate anchor point to the target cache block as the upstream target anchor point, and performing partial replay of the high-precision KV data of the historical tokens based on the upstream target anchor point to generate the target high-precision KV data of the target cache block. Thus, this invention constructs a lightweight, intelligent, and closed-loop dynamic management mechanism for KV cache through steps S1 to S5, achieving high-efficiency and high-quality collaborative optimization in the intelligent computing cloud platform: the intelligent computing cloud platform uses low-bit quantization to store historical KV cache for a long time and sparsely sets high-precision anchor points, significantly reducing memory usage; relevant cache blocks are quickly filtered through block-level summaries and query-driven scoring, and further, uncertainty indicators are combined to accurately identify the target cache blocks that truly need correction. This solution, without retaining the full high-precision KV, introducing additional reconstruction models, or expanding the playback range, effectively suppresses semantic bias and error accumulation caused by low-bit quantization with minimal additional computational cost, while ensuring low latency and resource schedulability in high-concurrency scenarios of the intelligent computing cloud platform, significantly improving the accuracy and stability of inference for large models with ultra-long contexts. Attached Figure Description

[0030] To more clearly illustrate the technical solution of the present invention, the accompanying drawings used in the description of the present invention will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0031] Figure 1 This is a flowchart of a quantization method for an intelligent computing cloud platform to achieve elastic KV caching through computing power, provided by the present invention. Figure 2 This is a flowchart of another method for quantizing elastic KV caching through computing power in an intelligent computing cloud platform provided by the present invention; Figure 3 This is a flowchart of another method for quantizing elastic KV caching through computing power in an intelligent computing cloud platform provided by the present invention; Figure 4 This is a flowchart of another method for quantizing elastic KV caching through computing power in an intelligent computing cloud platform provided by the present invention; Figure 5 This is a flowchart of another method for quantizing elastic KV caching through computing power in an intelligent computing cloud platform provided by the present invention; Figure 6 This is a flowchart of another method for quantizing elastic KV caching through computing power in an intelligent computing cloud platform provided by the present invention; Figure 7 This is a flowchart of another method for quantizing elastic KV caching through computing power in an intelligent computing cloud platform provided by the present invention; Figure 8 This is a structural diagram of a quantization device for an intelligent computing cloud platform that achieves elastic KV caching through computing power, provided by the present invention. Figure 9 This is a structural diagram of an electronic device provided by the present invention. Detailed Implementation

[0032] The technical solutions of this invention will now be clearly and completely described with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of this invention. All other embodiments obtained by those skilled in the art based on the embodiments of this invention without creative effort are within the scope of protection of this invention.

[0033] The “computing power” mentioned in this invention refers to: the ability of computer equipment or computing / data center to process information; the ability of computer hardware and software to work together to perform a certain computing requirement; the computing power to achieve the target result output by processing information data; and a new type of productivity that integrates information computing power, network carrying capacity, and data storage capacity, mainly providing services to society through computing power infrastructure.

[0034] The "computational power" (CP) described in this invention refers to the ability of a data center server to process data and output results. It is a comprehensive indicator of a data center's computing power, encompassing general computing power, supercomputing power, and intelligent computing power. The commonly used unit of measurement is floating-point operations per second (FLOPS, 1 EFLOPS = 10^18 FLOPS), with higher values ​​indicating stronger overall computing power. It is estimated that 1 EFLOPS is approximately the computing power output of 5 Tianhe-2A supercomputers, 500,000 mainstream server CPUs, or 2 million mainstream laptops. The calculation formula is: CP = CP通用 +CP 智能 +CP 超级 .

[0035] The "Network Power" (NP) mentioned in this invention refers to the performance of data transmission capability of computing facilities, which includes comprehensive capabilities such as network architecture, network bandwidth, transmission latency, intelligent management and scheduling, and involves network transmission within and between data centers. It is a comprehensive indicator for measuring network transmission scheduling capability.

[0036] The "Storage Power" (SP) described in this invention refers to the comprehensive capabilities of a data center in four aspects: data storage capacity, performance, security and reliability, and green and low-carbon operation. It is a comprehensive indicator for measuring the data storage capacity of a data center, including external storage devices such as storage arrays and internal storage devices within servers. The commonly used unit of measurement for storage capacity is exabytes (EB, 1EB = 2^60 bytes), while the commonly used unit of measurement for performance is the number of read / write operations per second (IOPS / TB). Disaster recovery ratio is an important indicator of security and reliability.

[0037] The "computing infrastructure" mentioned in this invention refers to a new type of information infrastructure that integrates information computing power, network carrying capacity, and data storage capacity, enabling centralized computing, storage, transmission, and application of information.

[0038] The "new information infrastructure" mentioned in this invention refers to network infrastructure such as 5G networks, fiber optic broadband networks, backbone networks, international communication networks, and satellite internet; computing infrastructure such as data centers, general computing centers, intelligent computing centers, and supercomputing centers; and new technology facilities such as artificial intelligence, blockchain, and quantum computing.

[0039] The “computing power” mentioned in this invention includes: general computing power, intelligent computing power, and supercomputing power.

[0040] The "general computing power" mentioned in this invention refers to the computing power provided by servers based on CPU (Central Processing Unit) chips, which is used to support basic general computing such as cloud computing and edge computing.

[0041] The "intelligent computing power" mentioned in this invention refers to: a computing platform deployed on a large scale based on dedicated chips such as GPU (Graphics Processing Unit), FPGA (Field Programmable Gate Array), and ASIC (Application Specific Integrated Circuit) for various artificial intelligence innovative applications, such as natural language processing and machine vision.

[0042] The “supercomputing power” mentioned in this invention refers to the computing power provided by high-performance computing clusters such as supercomputers. It utilizes the centralized computing resources of multiple computer systems working in parallel and uses a dedicated operating system to handle extremely complex or data-intensive problems. It is mainly used for computing in cutting-edge scientific fields, such as planetary simulation, drug molecule design, and gene analysis.

[0043] The "intelligent computing center" described in this invention refers to a facility that, through the use of large-scale heterogeneous computing resources, including general-purpose computing power (CPU) and intelligent computing power (GPU, FPGA, ASIC, etc.), primarily provides the necessary computing power, data, and algorithms for artificial intelligence applications (such as the development, training, and inference of deep learning models). The intelligent computing center encompasses facilities, hardware, and software, and can provide full-stack capabilities from underlying computing power to top-level application enablement.

[0044] The "intelligent computing cloud platform" mentioned in this invention, abbreviated as "intelligent computing cloud", refers to a cloud computing platform that integrates hardware and software resources based on an intelligent computing center.

[0045] The "intelligent computing center" mentioned in this invention includes, but is not limited to, "smart computing center".

[0046] The "intelligent computing center" mentioned in this invention, also known as an artificial intelligence computing center, is a type of computing infrastructure that provides computing power services, data services, and algorithm services required for artificial intelligence applications, based on artificial intelligence theory and adopting an artificial intelligence computing architecture.

[0047] The "computing center" mentioned in this invention refers to a facility that is mainly composed of infrastructure such as wind, thermal, hydro, and electricity, and IT hardware and software equipment, and has computing power, carrying capacity, and storage capacity, including general data centers, intelligent computing centers, supercomputing centers, etc.

[0048] The "supercomputing center" mentioned in this invention refers to a supercomputing data center, which is a data center based on supercomputers or large-scale computing clusters. It can provide large-scale computing, storage and network services and is widely used in aerospace, defense, oil exploration, climate modeling and genome sequencing and other application scenarios.

[0049] The “computing resources” mentioned in this invention refer to the technologies and facilities required for the development of the digital society that have the ability to compute, transmit, store and apply information, including but not limited to computing resources such as CPUs and GPUs, network resources such as switches and routers, storage resources such as storage arrays and distributed storage, security resources such as firewalls and intrusion detection systems, and supporting and guaranteeing resources such as wind, fire, water and electricity.

[0050] The "model" mentioned in this invention includes, but is not limited to, "large language model" and "multimodal large model".

[0051] The "large language model" mentioned in this invention refers to a large-scale language model (LLM), which is a language model with a large number of parameters. It is designed to understand and generate human language. It is trained with a large amount of text data and can perform a wide range of tasks, including text summarization, translation, and sentiment analysis.

[0052] The “Multimodal Large Models” mentioned in this invention refer to models that combine multimodal information such as text, images, videos, and audio for training, including but not limited to multimodal large language models.

[0053] It is important to emphasize that in the highly centralized and resource-intensive infrastructure environment of intelligent computing cloud platforms, Large Language Model (LLM) inference services are rapidly evolving towards ultra-long contexts (such as 32K–1M tokens). This trend presents unprecedented challenges to KV cache management: it must support massive concurrent user requests while ensuring millisecond-level response times; it must accommodate extremely long historical contexts without unlimited expansion of GPU memory resources. In this context, traditional "full-precision caching" or "static low-bit compression" solutions are no longer sustainable, necessitating a new KV cache quantization mechanism—an integrated technical loop that simultaneously achieves "long-term low-bit storage, local on-demand recovery, controllable trigger range, and incremental fusion of correction results." Its necessity can be explained from the following dimensions: I. Low-bit long-term storage: Breaking through the storage bottleneck of intelligent computing cloud platforms and enabling high-density multi-tenant deployment.

[0054] The core objective of an intelligent computing center is to maximize the service output per unit of hardware resources. In scenarios with extremely long contexts, the key-value cache for a single request can reach tens of GB (FP16). If 100 concurrent requests run simultaneously, the key-value cache alone will consume a huge amount of existing resources, placing extremely high demands on hardware. At the same time, the multi-tenant shared architecture requires the intelligent computing cloud platform to be able to elastically schedule and isolate resources, and video memory is a hard limiting factor.

[0055] Low-bit quantization (such as INT4 / FP8) can compress cache size by 5–10 times, enabling a single card to support more tenant instances and significantly improving resource utilization and revenue per unit of computing power. Without long-term low-bit storage, intelligent computing centers will be unable to provide long-context services cost-effectively.

[0056] II. Localized on-demand recovery: Avoid "quality collapse" of the intelligent computing cloud platform and ensure the accuracy of key semantics.

[0057] In intelligent computing cloud platforms, not all historical tokens are equally important. Attention mechanisms inherently possess sparse focusing characteristics, and current generation typically relies on only a few key contextual fragments (such as recent instructions and entity definitions).

[0058] While uniform low-bit quantization introduces errors evenly, in tasks such as logical reasoning, code generation, and medical question answering, even minor deviations at critical points can lead to serious errors. Therefore, intelligent computing centers need to provide differentiated SLAs for clients in different industries (e.g., finance requires 99.99% accuracy), and cannot accept "acceptable average quality but failure at critical points." Dynamically restoring high-precision key-value pairs only to "hotspot cache blocks" highly relevant to the query allows for precise repair of high-risk areas with less than 1% additional computational overhead, achieving "great results with minimal cost." Localized, on-demand restoration is the core lever for balancing "cost" and "reliable output."

[0059] 3. Controllable triggering range: meets the deterministic latency requirements of the intelligent computing cloud platform in the cloud-native environment.

[0060] The intelligent computing cloud platform targets enterprise customers and, through a KV caching quantization mechanism that features "low-bit long-term storage, local on-demand recovery, controllable trigger range, and incremental fusion of correction results," can provide predictable and commensurate latency services.

[0061] If coarse-grained recovery strategies such as global replay and full-layer reconstruction are used in existing technologies, the computation latency may suddenly increase from 10ms to 100ms+, causing request queuing, timeouts, or even a cascading failure. In high-concurrency environments, sudden recomputation can also trigger GPU resource contention, affecting the service quality of other tenants.

[0062] This solution uses uncertainty indicators to strictly limit the recovery range, ensuring that additional calculations are completed in microseconds and that resource usage is schedulable.

[0063] IV. Incremental fusion of correction results: Enables seamless, lightweight, and composable quality repair within the intelligent computing cloud platform.

[0064] In pipelined, asynchronous cloud inference engines, any "start from scratch" operation will disrupt execution continuity. Directly replacing low-bit key-value pairs will lead to inconsistent attention weights, requiring a recalculation of the entire context, which violates the principle of efficiency. Incremental fusion mechanisms can correct only the biased parts without recalculating the entire context, thus avoiding gradient oscillations or numerical overflows caused by precision jumps.

[0065] Please see Figure 1 , Figure 1 This is a flowchart of a quantization method for an intelligent computing cloud platform to achieve elastic key-value caching through computing power, as provided by the present invention. Figure 1 As shown, the method includes: Step S1: Divide the historical tokens of the intelligent computing cloud platform into multiple cache blocks, quantify the key data and / or value data of the historical tokens, write them into the corresponding cache blocks, and select multiple positions as candidate anchor points in the high-precision KV data of the historical tokens.

[0066] In this invention, the historical token sequence generated during the autoregressive decoding process is first divided into several consecutive or non-overlapping cache blocks based on a preset time window length (e.g., every 512 tokens, or according to semantic boundaries such as sentences / paragraphs) or the dynamic context importance assessment results. Each cache block corresponds to a local context, serving as the basic unit for subsequent quantization storage and selective recovery.

[0067] Subsequently, low-bit quantization is performed on the key and / or value data (i.e., KV data) corresponding to all tokens within each cache block. Quantization methods can employ symmetric / asymmetric linear quantization (such as INT4, INT8), per-group quantization, or calibration-based non-uniform quantization strategies to significantly compress storage volume with controllable precision loss. The quantized KV data is persistently written to the main cache area as a long-term resident low-precision main cache for regular attention computation.

[0068] Meanwhile, the intelligent computing cloud platform sparsely selects several key positions along the timeline of the token sequence as candidate anchor points. Each candidate anchor point not only records the complete token index at the corresponding time, but also saves its high-precision key-value data, as well as optional auxiliary state information, including but not limited to the hidden state of the current Transformer layer, the statistical characteristics of the attention head (such as mean and variance), and checkpoint snapshots that can support forward inference.

[0069] In this invention, the anchor point selection strategy can be based on any one or a combination of the following: Sampling at fixed intervals (e.g., setting an anchor point every 2048 tokens); Semantic mutation detection (e.g., triggered by a jump in perplexity or a topic switching signal); Attention entropy threshold (a key point is defined as a location where attention is highly concentrated); User instruction boundaries (such as the position before "Please summarize the above text" is automatically set as the anchor point).

[0070] In one possible implementation, let T be the number of historical tokens accumulated in a certain layer of the large language model before time t. This number is then divided into several cache blocks of fixed or dynamic length. (Each layer of the Transformer maintains its own KV Catch, hence "each layer." Historical tokens represent 100 tokens generated during the inference process of the large model, which can be understood as T = 100. The so-called dynamic length refers to a method that dynamically determines the size of each block based on semantics, access frequency, etc.) For the b-th cache block, let its low-bit main cache be denoted as Kˆ. b Vˆ b The low-bit format can be 2 bits, 3 bits, 4 bits, 8 bits, compressed floating-point format, or other compact representations lower than conventional high-precision caches. After each new token is generated, the intelligent computing cloud platform quantizes and writes the key and value of the corresponding layer of the token into the corresponding cache block, and simultaneously updates the digest statistics of the cache block.

[0071] The beneficial technical effects achieved in step S1 are as follows: This invention, through a three-pronged design of cache block partitioning, low-bit main memory, and anchor point snapshots, enables intelligent computing cloud platforms to significantly reduce KV cache memory usage while retaining the ability to perform local high-precision replays starting from the nearest upstream anchor point within any cache block range. This provides a data foundation for subsequent on-demand correction. In the resource-sensitive and service-critical environment of intelligent computing cloud platforms, this design truly enables the KV caching mechanism to achieve the engineering goals of "low-bit long-term storage, local on-demand recovery, controllable trigger range, and incremental fusion of correction results," providing key technical support for large-scale inference of ultra-long context models.

[0072] Step S2: Calculate the digest key data and quantization error proxy data of the cache block as the block-level digest data of the cache block.

[0073] In this invention, to reduce the computational complexity of attention matching, the intelligent computing cloud platform aggregates and compresses all quantized key vectors within each cache block to generate one or more digest key data, serving as the semantic "fingerprint" of that block. In this invention, the digest key data may include, but is not limited to, one or more of the following: Mean / Weighted Mean Key: Takes the average of all keys within the block, preserving the overall semantic direction; Principal Component Key: Extracts the principal component orientation of the Key within the block to capture the main mutation patterns; Cluster center Key: If the semantics within a block are diverse, K cluster centers can be retained (e.g., K=2~4) to improve expressive power; Learnable Projective Summary: Maps intra-block keys to fixed-dimensional summary vectors using a lightweight linear layer.

[0074] The summary key is used in subsequent steps to perform fast coarse-grained similarity calculations (such as dot product or cosine similarity) with the current query, in order to initially filter out potentially relevant "hot cache blocks" and avoid the computational overhead of traversing all historical cache blocks.

[0075] In this invention, since the main cache uses low-bit quantization, information loss is unavoidable. Quantization error proxy data is used to predict the impact of this loss on attention output. The intelligent computing cloud platform simultaneously calculates and records the error proxy index during quantization, serving as a basis for block-level risk assessment.

[0076] The two types of data mentioned above together constitute the block-level digest data for each cache block. This data is extremely small (typically <1KB / block) and can reside permanently in a cache (such as L2 cache or shared memory). During each decoding step, the intelligent computing cloud platform first uses the current query and the digest keys of all cache blocks to quickly calculate the block-level importance score.

[0077] The beneficial technical effects achieved in step S2 are as follows: By constructing block-level summary data with two dimensions—semantic summarization and error proxy—the intelligent computing cloud platform can efficiently predict the "value" and "risk" of each cache block without accessing the original high-precision key-value pairs. This not only significantly reduces the computational overhead of dynamic evaluation but also ensures the accuracy and economy of subsequent recovery actions. It is a key prerequisite for achieving high-quality, low-latency, high-concurrency, long-context reasoning in the intelligent computing cloud platform.

[0078] Step S3: In response to the intelligent computing cloud platform generating a new target token, score the cache blocks based on the target token to generate a cache block score, and select candidate cache blocks from multiple cache blocks based on the cache block score.

[0079] After each step of autoregressive decoding generates a new target token, a cache block score is generated. Based on this score, candidate cache blocks are selected from multiple cache blocks. This allows for the rapid identification of historical context fragments most likely to affect the quality of current and subsequent generation, providing a basis for whether to trigger local recovery. This enables a joint assessment of "relevance + risk" with extremely low overhead, accurately narrowing down the candidate range.

[0080] In this invention, while generating a new target token, the model outputs the corresponding current semantic vector. Then, it iterates through all cache blocks, performing a fast similarity match between the current semantic vector and its digest key data to calculate the generated cache block score.

[0081] Because the digest key is extremely small and pre-stored, the process can be completed in microseconds even with a million token contexts. Finally, based on the cache block scores, the top-K cache blocks that are semantically "potentially relevant" are quickly selected as candidate cache blocks.

[0082] The beneficial technical effects achieved in step S3 are as follows: By using query-driven relevance assessment and multi-dimensional uncertainty weighting, a "risk-value" profile of the cache block is dynamically constructed in each decoding step, achieving efficient focusing from the full history to key segments. It is not only an "intelligent gate" connecting the main cache and partial recovery, but also a core control mechanism that ensures the intelligent computing cloud platform maintains low latency and high generation quality under ultra-long contexts and high concurrency loads. This design enables the system to achieve a fine balance between "resource saving" and "quality assurance," laying a precise decision-making foundation for subsequent controlled playback and incremental fusion.

[0083] Step S4: Calculate the uncertainty index data of the candidate cache blocks, and determine the target cache block whose accuracy needs to be restored from the candidate cache blocks based on the uncertainty index data.

[0084] After identifying several highly important candidate cache blocks, the system does not directly perform recovery on all high-importance blocks. Instead, it further calculates the uncertainty index for each candidate cache block. Based on the uncertainty index, it determines which of these high-importance blocks are truly worth recovering with high precision. Only if a block is not only important but also has unstable current sorting, large quantization error, hesitant model behavior, or significant recent performance fluctuations, is it considered worthwhile to invest additional computing power for recovery.

[0085] It should be noted that uncertainty metrics are a set of lightweight surrogate signals used to quantify the reliability of model output and the risk of historical cache errors in the current inference step. These include, but are not limited to, the entropy of the current output distribution, the quantization projection error of the candidate cache block in the Query direction, the historical attention sensitivity of the block, and the recent score volatility. These metrics together constitute a comprehensive risk score, used to determine whether it is necessary to trigger high-precision local recovery for a specific cache block, thereby ensuring generation quality while avoiding unnecessary computational overhead.

[0086] In this invention, the intelligent computing cloud platform can construct uncertainty indicators from multiple dimensions to form a comprehensive risk score. For example, uncertainty can be constructed from the following dimensions: indicator quantification error risk, output distribution uncertainty, block-level attention sensitivity, and temporal local volatility.

[0087] In this invention, based on the aforementioned indicators, the intelligent computing cloud platform can use a weighted fusion and threshold gating mechanism to determine the final target cache block. A comprehensive uncertainty score can be calculated based on the uncertainty indicator data, and then based on the uncertainty score...

[0088] In this invention, based on the aforementioned indicators, the intelligent computing cloud platform can use weighted fusion to determine the final target cache block. A comprehensive uncertainty score can be calculated based on the uncertainty index data, and then the candidate cache block with the highest uncertainty score is selected as the target cache block.

[0089] The beneficial technical effects achieved by step S4 are as follows: In the selected high-importance cache blocks, a multi-dimensional uncertainty criterion is further introduced to accurately identify blocks with both high relevance and high risk attributes. High-precision recovery is triggered only for local historical content that may actually damage the quality of the current generation. This significantly improves the semantic accuracy and stability of long context reasoning without increasing the average latency. At the same time, it avoids redundant calculations on blocks that are "seemingly important but actually harmless", effectively ensuring the resource efficiency and latency controllability of the intelligent computing cloud platform in high-concurrency scenarios.

[0090] Step S5: Locate the nearest preceding candidate anchor point to the target cache block as the upstream target anchor point, and perform partial replay of the high-precision KV data of historical tokens based on the upstream target anchor point to generate the target high-precision KV data of the target cache block.

[0091] In this invention, the intelligent computing cloud platform does not store all historical high-precision key-value pairs (KVs) indefinitely, but only saves a small number of "anchor points" at intervals. When a cache block truly deserves refinement, it starts from the nearest anchor point, replays only a small segment of necessary calculations, and temporarily restores the high-precision KVs of the target cache block. To support subsequent recovery without permanently retaining all high-precision KVs, this invention sparsely saves multiple anchor point states along the historical sequence. Let the target cache block be b, and its nearest upstream anchor point be ab. The anchor point state can be a hidden state of a predetermined layer, a compressed representation of the input of a predetermined layer, or a checkpoint state that can continue to recover subsequent KVs. Unlike the approach of "saving the information required for reconstruction for all locations," this invention only saves a small number of anchor points at coarser intervals, so high-precision states will not reside on a large scale for a long time. When the target cache block is triggered, the intelligent computing cloud platform starts from the anchor point, performs local forward replay only for the necessary time period between the anchor point and the target cache block, and restores the high-precision target KV data of the target cache block.

[0092] It should be noted that this step is the high-precision repair execution stage in the entire dynamic correction mechanism. It does not recalculate the entire sequence, does not rely on external models, and only starts from the nearest credible starting point to lightly reconstruct the local high-precision state.

[0093] The intelligent computing cloud platform uses the state of the upstream target anchor point as the initial condition and performs forward reasoning only on tokens within the time window covered by the target cache block, but only at necessary levels and components, to achieve lightweight partial replay. It should be noted that this lightweight partial replay can meet the following requirements.

[0094] Only replay the token of the target cache block; Layer selection: If the error mainly comes from certain layers (such as the top semantic layer), only replay those layers; Head selection: Reconstruct only high-response attention heads based on the Query attention distribution.

[0095] The beneficial technical effects achieved by step S5 are as follows: While completely avoiding long-term storage of the full high-precision KV cache, it supports on-demand, accurate, high-precision local recovery capabilities with minimal storage cost (only a few anchor points). When a target cache block is determined to be high-risk, the intelligent computing cloud platform only replays the necessary time period from the nearest anchor point to the block, and this can be limited to critical layers or attention points. This allows for the temporary reconstruction of the target high-precision KV data with controllable computational overhead, ensuring the effectiveness of the correction while eliminating memory overload and global recalculation overhead.

[0096] In this invention, step S1 involves dividing the historical tokens of the intelligent computing cloud platform into multiple cache blocks, quantizing the key and / or value data of the historical tokens, writing them into the corresponding cache blocks, and selecting multiple positions as candidate anchor points in the high-precision KV data of the historical tokens; step S2 involves calculating the digest key data and quantization error proxy data of the cache blocks as block-level digest data of the cache blocks; step S3 involves responding to the intelligent computing cloud platform generating a new target token, scoring the cache blocks based on the target token to generate cache block scores, and selecting candidate cache blocks from multiple cache blocks based on the cache block scores; step S4 involves calculating the uncertainty index data of the candidate cache blocks, and determining the target cache block whose precision needs to be restored from the candidate cache blocks based on the uncertainty index data; step S5 involves locating the nearest preceding candidate anchor point to the target cache block as the upstream target anchor point, and performing partial replay of the high-precision KV data of the historical tokens based on the upstream target anchor point to generate the target high-precision KV data of the target cache block. Thus, this invention constructs a lightweight, intelligent, and closed-loop dynamic management mechanism for KV cache through steps S1 to S5, achieving high-efficiency and high-quality collaborative optimization in the intelligent computing cloud platform: the intelligent computing cloud platform uses low-bit quantization to store historical KV cache for a long time and sparsely sets high-precision anchor points, significantly reducing memory usage; relevant cache blocks are quickly filtered through block-level summaries and query-driven scoring, and further, uncertainty indicators are combined to accurately identify the target cache blocks that truly need correction. This solution, without retaining the full high-precision KV, introducing additional reconstruction models, or expanding the playback range, effectively suppresses semantic bias and error accumulation caused by low-bit quantization with minimal additional computational cost, while ensuring low latency and resource schedulability in high-concurrency scenarios of the intelligent computing cloud platform, significantly improving the accuracy and stability of inference for large models with ultra-long contexts.

[0097] In one embodiment, step S1 includes: Step S1.1: Quantize the Key data based on the first quantization granularity and quantize the Value data based on the second quantization granularity, wherein the first quantization granularity is greater than the second quantization granularity.

[0098] It should be noted that the role of the key data is to participate in similarity calculation. Its absolute value does not directly affect the output content, but rather the attention weight is determined by the relative direction and sorting.

[0099] Attention mechanisms are robust to quantization errors in key data, especially after softmax normalization, where small directional perturbations are often smoothed out. Therefore, in this invention, key data can be quantized with coarser granularity and stored in significantly compressed form, with limited impact on the final output.

[0100] Value data directly constitutes the weighted combination of attention outputs, and its numerical precision directly affects the semantic details and factual accuracy of the generated content. Quantization errors are directly injected into the final hidden state and amplified layer by layer through the autoregressive process, easily leading to factual errors or logical confusion. Therefore, in this invention, Value data adopts a relatively fine quantization granularity to retain key content information.

[0101] The beneficial technical effects achieved by step S1.1 are as follows: Based on the different roles played by the Key and Value in the attention mechanism and their sensitivity to quantization errors, different bit widths are assigned to them, thereby achieving a better balance between overall compression ratio and generation quality, and reducing the computing power expenditure cost of the intelligent computing cloud platform.

[0102] In one embodiment, step S1 includes: Step S1.2: For the high-precision KV data of historical tokens, set a candidate anchor point every preset time interval in chronological order; wherein, each candidate anchor point includes at least one of the following data: the hidden state of the predetermined network layer, the compressed representation of the input of the predetermined network layer, and the checkpoint state.

[0103] In this invention, to support subsequent recovery without permanently retaining the full high-precision key-value pairs, multiple anchor states are sparsely stored along the historical sequence. Let the target cache block be b, and its nearest upstream anchor be denoted as ab. Anchor states can be hidden states of predetermined network layers, compressed representations of predetermined network layer inputs, or checkpoint states. Unlike the approach of "saving the information needed for reconstruction at all locations," this invention only stores a small number of anchors at coarser intervals; therefore, high-precision states are not retained on a large scale for extended periods.

[0104] In this context, the hidden state of a predetermined network layer refers to a pre-selected layer (such as layer 4, layer 8, or the top layer) in a neural network like the Transformer. During the autoregressive generation process, the output vectors of these intermediate layers contain the semantic representation of the current token at that layer. Saving this as anchor content allows it to be used directly as the starting point for forward computation during subsequent local replays, avoiding the need to reason layer by layer from the input token. This significantly accelerates the reconstruction of high-precision key-value pairs while ensuring semantic consistency.

[0105] Compressed representation of the input to a predetermined network layer refers to a compact encoding obtained by lightweight compression of the input data of a specified network layer (usually the output of the previous layer). This can be achieved through methods such as principal component analysis, quantized cluster center indexing, low-rank approximation, or small autoencoders. This representation significantly reduces storage overhead while retaining sufficient information to approximate the original input. In resource-constrained scenarios, it can serve as an alternative to hidden states, decompressing or interpolating to restore them during local playback, achieving a trade-off between controlled accuracy loss and higher storage efficiency.

[0106] Checkpoint states refer to a set of snapshots of intermediate variables within the model, saved to support deterministic forward replay. These typically include, but are not limited to, the running mean and variance of LayerNorm, intermediate tensors of residual connections, the on / off states of activation functions, and seeds for random number generators. These states ensure that when re-executing the forward propagation from the anchor point, the numerical results in the original inference path can be accurately reproduced. This avoids inconsistencies between the recovered key-value pairs and the original context due to floating-point nondeterminism or missing internal states, and is crucial for high-reliability inference scenarios (such as code generation and legal question answering).

[0107] In one embodiment, such as Figure 2 As shown, step S2 includes: Step S2.1: Calculate the mean squared error of the cache block before and after updating the historical token, and / or calculate the historical energy statistics and / or activity value and / or prior data of the cache block.

[0108] Step S2.2: Calculate quantization error proxy data based on the mean square error value and / or historical energy statistics and / or activity value and / or prior data.

[0109] To avoid performing high-precision comparisons of all historical key-value pairs directly in each decoding step, this invention maintains lightweight digest information for each cache block. In traditional autoregressive decoding, generating a new token requires interacting with the high-precision patterns of all tokens in all historical cache blocks, which is extremely costly.

[0110] In this invention, by using the approach of "first recall, then fine sorting", let the digest key of the b-th cache block on the h-th attention head be... , which represents its "average, representative value, or digest value", is not the key of each token in the block, but rather a simplified feature that represents the entire block.

[0111] The quantization error proxy is Q errb This indicates the degree of error risk introduced by quantization compression. It can be an approximation of the mean square error before and after quantization within the block, with historical energy statistics represented by E. bThis indicates the overall "strength," "amplitude," and "activity" of the Key or Value within this block. The location or recent reuse prior is R. b Even without a precise comparison of the current step, the intelligent computing cloud platform can, based on historical experience, prioritize certain blocks more or less. This can be based on location priors (blocks closer to the current moment may have higher priority) or recent reuse priors (a block has been frequently noticed and repeatedly hit in the past few steps). This statistical information does not require precise reconstruction of the key and value of each token within the block, but rather serves to provide a low-cost basis for hotspot identification when the current query arrives.

[0112] The mean squared error (MSE) values ​​before and after updating historical tokens refer to the calculated mean squared error between the key / value vectors before quantization (high precision, such as FP16) and after quantization (low bit, such as INT8) when writing a newly generated token to a cache block. This allows for a comparison of the overall block-level MSE change before and after the write. For example, if the newly added token is semantically critical and has significant quantization distortion, it will lead to a significant increase in the block's MSE. This difference effectively reflects the degree of precision loss introduced by the latest context, serving as a dynamic signal to determine whether the block has become "high-risk" due to new content, providing real-time error awareness for subsequent recovery decisions.

[0113] Historical energy statistics refer to the average or cumulative value of the squared L2 norm of all key or value vectors within a cache block, used to measure the strength of the semantic signal carried by that block. High energy indicates that the vector amplitude of the cache block data is large and its features are significant, possibly corresponding to named entities, keywords, or high-confidence semantic units; low energy may indicate padding characters, stop words, or ambiguous context in the cache block. This statistic can not only be used to normalize quantization errors (e.g., calculating relative MSE = MSE / energy), but also help determine the semantic importance of blocks. Even with the same error, the distortion of high-energy blocks often has a greater impact on attention output.

[0114] Activity value is a dynamic metric that measures the degree to which a cache block has received attention from the attention mechanism over a number of decoding steps. It is typically calculated by accumulating the average attention weight of the tokens within the block under historical queries, or the sum of the similarities between the digest key and historical queries. A high activity value indicates that the block is continuously involved in semantic reasoning (e.g., repeatedly cited contextual facts) and belongs to a "high-frequency usage area." If such cache blocks also have significant quantization errors, their potential harm is even greater. Activity values ​​can be updated online using a sliding window or exponential smoothing, enabling low-overhead tracking of usage activity.

[0115] Prior data refers to static or semi-static weight information pre-defined based on task type, user instructions, context structure, or domain knowledge, used to guide cache management strategies. For example, when a user enters "Please remember the following configuration," the system can mark the cache blocks containing several subsequent tokens as having high prior importance; in code generation scenarios, syntax structures such as function definitions and class declarations can be automatically assigned higher recovery priority; in multilingual dialogues, a more conservative quantification strategy can be adopted for non-native language segments. Prior data does not rely on runtime statistics but can significantly improve the robustness and adaptability of the system in critical scenarios.

[0116] In one embodiment, such as Figure 3 As shown, step S3 includes: Step S3.1: Based on the similarity between the target token and the cached block across multiple attention heads, calculate the access preference data for the cached block.

[0117] Step S3.2: Calculate the cache block score based on similarity, access preference data, and prior data of cache blocks.

[0118] In this invention, when generating a new token, instead of performing a detailed calculation on all historical key-value blocks, a coarse score is first assigned to each "cache block". We select the few most noteworthy blocks. Let the query of the current decoding step be q, and its vector representation at the h-th attention head be q. h .

[0119] In this invention, step S3.2 includes: Step S3.2.1: Take the maximum similarity between the target token and the cache block across multiple attention heads as the target similarity, and perform a weighted calculation based on the target similarity, access tendency data and prior data of the cache block to generate a cache block score.

[0120] The formula for calculating the cache block score is as follows:

[0121] In the formula, The importance of the b-th cache block is determined by three factors: how similar the current query is to this cache block, whether this block has historically been important, and whether this block has any prior advantages in terms of location / recent reuse.

[0122] Where α, β, and γ are adjustable weights. This is used to describe the target similarity between the current query and the b-th cache block at the h-th head, i.e., the similarity between the current query and the block's digest key. This is similar to the idea in standard attention, except that instead of comparing the query with the key of each token in the block, it first compares it with the block-level digest. `max` represents the maximum similarity among all attention heads, which is taken as the target similarity. If any attention head deems the block relevant, then the block is considered important. Access tendency data used to characterize the historically high-intensity access of this cache block. This is the prior data for the cache block.

[0123] In one embodiment, such as Figure 4 As shown, step S4 includes: Step S4.1: Calculate the score proximity data between the candidate buffer block and the adjacent candidate buffer block, calculate the quantization error value of the candidate buffer block, calculate the output distribution entropy of the current decoding step, and calculate the importance fluctuation data of the candidate buffer block.

[0124] Step S4.2: Perform weighted calculations based on score proximity data, quantization error value, output distribution entropy, and importance fluctuation data to generate uncertainty index data.

[0125] After identifying several highly important candidate cache blocks, the intelligent computing cloud platform does not directly perform recovery on all high-scoring blocks. Instead, it further calculates the uncertainty index for each candidate cache block. Based on the uncertainty index, it then determines which of these high-scoring blocks are truly worth recovering with high precision. Only if a block is not only important but also has unstable current sorting, large quantization error, a hesitant model, or significant recent performance fluctuations, is it considered worthwhile to invest additional computing power for recovery.

[0126] In this invention, uncertainty index data can be generated using the following formula:

[0127] in, This represents the score proximity data between the candidate block and its neighboring candidate blocks, used to measure whether the current ranking is stable. Specifically, it measures whether the score difference between the b-th block and the blocks ranked before and after it is too small. If they are very close, it indicates that the current ranking is fragile. This represents the quantization error value of the b-th cache block, used to determine the degree of compression of the cache block. The low-bit result we see now may not be very reliable, as its importance judgment may be distorted by the quantization error. This represents the output distribution entropy of the current decoding step, which reflects the current uncertainty of the model. If the current model output distribution entropy is high, it means that the model is not so certain about the next token, the probability of candidate tokens is relatively dispersed, and the current generation is in a state of hesitation. This indicates the fluctuation range of the candidate block's importance over the most recent steps, if the cached block is not a "stable important block" but a "swinging important block". λ1 to λ4 are weight parameters.

[0128] The intelligent computing cloud platform not only evaluates the semantic importance of candidate cache blocks, but also comprehensively considers the magnitude of their quantization error and the uncertainty of the current model output, thereby accurately identifying and skipping cache blocks that are "highly relevant but low risk", and triggering high-precision recovery only for blocks with the superposition of four risks: "highly relevant, high error, high confusion, and high volatility".

[0129] The beneficial technical effects achieved by step S4.2 are as follows: Based on score proximity data, quantization error value, output distribution entropy and importance fluctuation data, uncertainty index data are calculated, which effectively avoids wasting valuable computing power on cache blocks with high scores but whose low-bit representation is already stable enough, significantly reduces the frequency of invalid local replays, and ensures that high-precision context is injected in a timely manner at the critical moment when the model is most prone to errors. In this way, the accuracy and robustness of long context generation are significantly improved without increasing the average latency, providing a data foundation for intelligent computing cloud platforms to provide high-quality large model inference services under high concurrency and low latency constraints.

[0130] In one embodiment, such as Figure 5 As shown, step S5 further includes: Step S5.1: Perform multi-dimensional cropping on the data between the upstream target anchor point and the target cache block to generate the first candidate data between the upstream target anchor point and the target cache block.

[0131] In this invention, the intelligent computing cloud platform performs multi-dimensional dynamic pruning on the time window between the upstream target anchor point and the target cache block, significantly compressing the scale of the data to be replayed.

[0132] It should be noted that multi-dimensional cropping includes one or more of the following: Perform time-dimensional pruning on the data between the upstream target anchor point and the target cache block; Perform layer-by-layer pruning on the data between the upstream target anchor point and the target cache block; Perform header dimension pruning on the data between the upstream target anchor point and the target cache block; Perform channel dimension pruning on the data between the upstream target anchor point and the target cache block; Perform precision dimension pruning on the data between the upstream target anchor point and the target cache block.

[0133] Among them, the time dimension pruning is to retain only the subsequences that are semantically strongly related to the target cache block (such as by sliding window or by truncating irrelevant prefixes based on attention prior).

[0134] Layer pruning involves replaying only the Transformer layers that have the greatest impact on the current output (such as the top semantic layer), skipping the bottom general feature layers.

[0135] Head-dimensional pruning is based on similarity weights, activating only high-response attention heads, while the remaining heads directly reuse low-bit KV data.

[0136] Channel dimensional pruning involves masking or skipping computations of channels with low sensitivity in the Key / Value vector (such as those with small variance or weak gradient).

[0137] Precision dimensional cutting employs a hybrid precision approach (e.g., FP8 backbone + FP16 critical head) for non-critical parts, balancing speed and fidelity.

[0138] Step S5.2: Filter out the first candidate data to generate the second candidate data corresponding to the target cache block.

[0139] To reconstruct a high-precision key-value pair for the target cache block, the intelligent computing cloud platform needs to replay a historical sequence starting from the upstream target anchor point, for example, from token 8000 to token 10000, while the target cache block covers tokens 9500–10000. The first candidate data generated at this point contains all intermediate tokens and their states within the entire replay window. However, not all tokens within this window belong to the target cache block; some may include: Leading padding tokens (such as the transition segment from the anchor point to the start of the target block); The contents of other low-importance cache blocks; Historical fragments that are semantically irrelevant or have minimal quantization error.

[0140] Performing high-precision calculations and subsequent difference analysis on these non-target areas would result in unnecessary waste of computing power.

[0141] Therefore, filtering out the first candidate data can accurately extract the data subset corresponding to the target cache block.

[0142] In one possible implementation, if the cache block uses a non-contiguous or dynamically partitioned strategy (such as semantic segmentation), then by matching the block ID or metadata tag, it is ensured that only the token belonging to the current target cache block is retained.

[0143] The beneficial technical effect achieved by step S4.2 is that by trimming the first candidate data (replay path) into a dedicated high-precision candidate set for the target cache block, irrelevant historical content between the target cache block is completely removed. This mechanism ensures the granular accuracy and resource economy of local recovery, and greatly reduces the computing power cost of the intelligent computing cloud platform.

[0144] Step S5.3: Calculate the attention contribution difference between the second candidate data and the low-bit data.

[0145] Step S5.4: Correct the basic output data of the target cache block based on the attention contribution difference to obtain the target high-precision KV data.

[0146] In this invention, the correction of the basic output data of the target cache block based on the attention contribution difference can be achieved using the following formula:

[0147]

[0148] in, and These represent the approximate results after restoring the low-bit main cache of all cache blocks. For the selected target cache block set Ω, the intelligent computing cloud platform uses the higher-precision key-value patch obtained by partial playback restoration. For basic output Perform differential correction.

[0149] As can be seen from the above formula, this invention does not recalculate the high-precision attention corresponding to all historical blocks, but only adds the difference between the attention contribution of the target block under high precision and low bit approximation as a patch to the basic output. This ensures the targeted nature of the correction while significantly controlling additional computing power overhead.

[0150] In one embodiment, such as Figure 6 As shown, in this invention, step S5 includes: Step S5.5: Perform partial replay of the Key data of the target cache block based on the upstream target anchor point to generate high-precision target Key data.

[0151] In existing technologies, the Key and Value are typically treated as an inseparable whole, and both are reconstructed synchronously once recovery is triggered. However, based on an in-depth analysis of the error propagation characteristics of the attention mechanism, this invention proposes that the Key and Value have different impact paths on the final output, and their recovery priority and necessity should also be dynamically decoupled.

[0152] Since the Key primarily influences attention weight allocation, while the Value directly determines the output content, this invention proposes a two-stage, condition-triggered replay strategy for the intelligent computing cloud platform: first, the Key is restored, and only when its uncertainty remains high is the more expensive Value further restored.

[0153] The intelligent computing cloud platform first takes the upstream target anchor point as the starting point, performs lightweight forward inference only on the time window covered by the target cache block, but only calculates and saves the key vector (which can be limited to the key layer / head) to generate high-precision key data of the target.

[0154] In one possible implementation, if attention is already stable under a high-precision key, then there is no need to restore the value.

[0155] Step S5.6: In response to the uncertainty index data of the target high-precision key data being greater than the judgment threshold, the value data of the target cache block is partially replayed based on the upstream target anchor point to generate target high-precision value data, and target high-precision KV data is generated based on the target high-precision value data and the target high-precision key data.

[0156] The intelligent computing cloud platform evaluates the uncertainty index of the target high-precision key data. If this index exceeds a preset judgment threshold, it indicates that even with the use of high-precision keys, the model is still in a state of high uncertainty, or that although the attention weights have been corrected, the original low-bit values ​​may have severely distorted the semantic content. In this case, the high-precision values ​​must be restored to truly guarantee the generation quality.

[0157] Therefore, the intelligent computing cloud platform once again started from the same upstream anchor point and executed a second round of partial replay. However, this time it focused on the Value path, generated target high-precision Value data, and combined it with the existing high-precision Key to form complete target high-precision KV data.

[0158] In this invention, in response to the uncertainty index data of the target high-precision key data being less than or equal to the judgment threshold, the target high-precision key data is used as the target high-precision KV data.

[0159] In one embodiment, such as Figure 7 As shown, step S5 includes: Step S5.8: Based on the upstream target anchor point, perform local replay of the target cache block according to the first recovery accuracy to generate candidate high-precision KV data.

[0160] Step S5.9: In response to the uncertainty index data of the candidate high-precision KV data being greater than the judgment threshold, the target cache block is partially replayed based on the upstream target anchor point according to the second recovery accuracy to generate the target high-precision KV data, wherein the second recovery accuracy is greater than the first recovery accuracy.

[0161] In large-scale model inference with ultra-long contexts on intelligent computing cloud platforms, high-precision local replay can improve quality, but its computational cost is still very high. This invention introduces a hierarchical accuracy recovery strategy: The first level of restoration accuracy serves as a low-cost, exploratory repair. The second level of restoration accuracy serves as the final high-fidelity restoration.

[0162] The intelligent computing cloud platform first generates a candidate high-precision key-value (KV) data set with lower precision and assesses whether it is sufficient to eliminate the current uncertainty. Only when the candidate still fails to meet the quality requirements is a second playback with higher precision initiated. This "try first, then decide, and upgrade if necessary" mechanism significantly increases the cost of computing power investment.

[0163] In this invention, in response to the uncertainty index data of the candidate high-precision KV data being less than or equal to the judgment threshold, the candidate high-precision KV data is used as the target high-precision KV data.

[0164] In this invention, the high-precision patches generated through partial playback do not need to be retained indefinitely. The intelligent computing cloud platform can configure lifecycle control strategies for the target key-value patches, such as: automatic invalidation after several consecutive decoding steps; invalidation when block heat drops below a threshold; early release when memory pressure exceeds a predetermined threshold; and batch cleanup when generation enters a new semantic stage. This avoids the long-term accumulation of temporary high-precision patches in the intelligent computing cloud platform. Furthermore, this invention can dynamically adjust anchor density, importance threshold, uncertainty threshold, and partial playback range based on resource budget information such as memory availability, processor utilization, service latency targets, and request priorities. For example, when the GPU is idle, the trigger threshold can be lowered and the partial playback range expanded to achieve higher precision; when the intelligent computing cloud platform is busy, the trigger threshold can be increased, or partial playback can be limited to restoring only a portion of the header or a portion of the channel to ensure overall latency stability.

[0165] In this invention, the target high-precision KV data is released in response to the target high-precision KV data meeting the release condition.

[0166] It should be noted that the release conditions include one or more of the following: Continuously use the target high-precision KV data a set number of times; The popularity of the target high-precision KV data is less than the popularity threshold; The existing pressure value of the intelligent computing cloud platform is greater than the pressure threshold.

[0167] Please see Figure 8 , Figure 8 This is a structural diagram of a quantization device for an intelligent computing cloud platform that achieves elastic KV caching through computing power, as provided by the present invention. Figure 8As shown, the quantization device 800 for elastic KV caching in the intelligent computing cloud platform includes: The partitioning module 801 is used to divide the historical tokens of the intelligent computing cloud platform into multiple cache blocks, quantify the key data and / or value data of the historical tokens, write them into the corresponding cache blocks, and select multiple positions as candidate anchor points in the high-precision KV data of the historical tokens.

[0168] The calculation module 802 is used to calculate the digest key data and quantization error proxy data of the cache block as the block-level digest data of the cache block.

[0169] The scoring module 803 is used to respond to the intelligent computing cloud platform generating a new target token, score the cache blocks based on the target token to generate cache block scores, and filter candidate cache blocks from multiple cache blocks based on the cache block scores.

[0170] The filtering module 804 is used to calculate the uncertainty index data of candidate cache blocks and determine the target cache block whose accuracy needs to be restored from the candidate cache blocks based on the uncertainty index data.

[0171] The replay module 805 is used to locate the nearest preceding candidate anchor point to the target cache block as the upstream target anchor point, and to perform partial replay of the high-precision KV data of historical tokens based on the upstream target anchor point in order to generate the target high-precision KV data of the target cache block.

[0172] In one embodiment, the partitioning module 801 is further configured to: The Key data is quantized based on a first quantization granularity, and the Value data is quantized based on a second quantization granularity, wherein the first quantization granularity is larger than the second quantization granularity.

[0173] In one embodiment, the partitioning module 801 is further configured to: For high-precision key-value data of historical tokens, a candidate anchor point is set at preset time intervals in chronological order; Each candidate anchor point includes at least one of the following data: the hidden state of the predetermined network layer, the compressed representation of the input of the predetermined network layer, and the checkpoint state.

[0174] In one embodiment, the calculation module 802 is further configured to: Calculate the mean squared error of the cache block before and after updating the historical token, and / or calculate the historical energy statistics and / or activity value and / or prior data of the cache block; Quantization error proxy data is calculated based on mean square error and / or historical energy statistics and / or activity values ​​and / or prior data.

[0175] In one embodiment, the scoring module 803 is further configured to: Based on the similarity between the target token and the cached block across multiple attention heads, the access preference data of the cached block is calculated; Based on similarity, access preference data, and prior data of cached blocks, a cache block score is calculated.

[0176] In one embodiment, the scoring module 803 is further configured to: The maximum similarity between the target token and the cached block across multiple attention heads is taken as the target similarity. A weighted calculation is then performed based on the target similarity, access preference data, and prior data of the cached block to generate a cached block score.

[0177] In one embodiment, the scoring module 804 is further configured to: Calculate the score proximity data between the candidate buffer block and its neighboring candidate buffer blocks, calculate the quantization error value of the candidate buffer block, calculate the output distribution entropy of the current decoding step, and calculate the importance fluctuation data of the candidate buffer block; Uncertainty index data is generated by weighting the score proximity data, quantization error value, output distribution entropy, and importance fluctuation data.

[0178] In one embodiment, the playback module 805 is further configured to: The data between the upstream target anchor and the target cache block is pruned in multiple dimensions to generate the first candidate data between the upstream target anchor and the target cache block; The first candidate data is filtered out to generate the second candidate data corresponding to the target cache block; Calculate the difference in attention contribution between the second candidate data and the low-bit data; The basic output data of the target cache block is corrected based on the difference in attention contribution to obtain the target high-precision KV data.

[0179] In one embodiment, multi-dimensional cropping includes one or more of the following: Perform time-dimensional pruning on the data between the upstream target anchor point and the target cache block; Perform layer-by-layer pruning on the data between the upstream target anchor point and the target cache block; Perform header dimension pruning on the data between the upstream target anchor point and the target cache block; Perform channel dimension pruning on the data between the upstream target anchor point and the target cache block; Perform precision dimension pruning on the data between the upstream target anchor point and the target cache block.

[0180] In one embodiment, the playback module 805 is further configured to: Based on the upstream target anchor point, the key data of the target cache block is partially replayed to generate high-precision target key data; In response to the uncertainty index data of the target high-precision key data exceeding the judgment threshold, the value data of the target cache block is partially replayed based on the upstream target anchor point to generate the target high-precision value data, and the target high-precision KV data is generated based on the target high-precision value data and the target high-precision key data.

[0181] In one embodiment, the playback module 805 is further configured to: If the uncertainty index of the target high-precision key data is less than or equal to the judgment threshold, the target high-precision key data will be used as the target high-precision KV data.

[0182] In one embodiment, the playback module 805 is further configured to: Based on the upstream target anchor point, the target cache block is partially replayed according to the first recovery accuracy to generate candidate high-precision KV data; In response to the uncertainty index data of the candidate high-precision KV data being greater than the judgment threshold, the target cache block is partially replayed based on the upstream target anchor point according to the second recovery accuracy to generate the target high-precision KV data, wherein the second recovery accuracy is greater than the first recovery accuracy.

[0183] In one embodiment, the playback module 805 is further configured to: If the uncertainty index of the candidate high-precision KV data is less than or equal to the judgment threshold, the candidate high-precision KV data is used as the target high-precision KV data.

[0184] In one embodiment, the playback module 805 is further configured to: In response to the target high-precision KV data meeting the release conditions, the target high-precision KV data is released.

[0185] In one embodiment, the release condition includes one or more of the following: Continuously use the target high-precision KV data a set number of times; The popularity of the target high-precision KV data is less than the popularity threshold; The existing pressure value of the intelligent computing cloud platform is greater than the pressure threshold.

[0186] The quantization device for the intelligent computing cloud platform to achieve elastic KV caching through computing power provided by the present invention is capable of realizing the various processes of the various embodiments of the above-mentioned intelligent computing cloud platform to achieve elastic KV caching through computing power. The technical features are one-to-one and can achieve the same technical effect. To avoid repetition, it will not be described again here.

[0187] It should be noted that the quantization device for elastic KV caching implemented by the intelligent computing cloud platform in this invention can be a device, or it can be a component, integrated circuit, or chip in an electronic device.

[0188] The present invention also provides an electronic device, see [link to relevant documentation]. Figure 9 , Figure 9 This is a schematic diagram of the structure of an electronic device provided by an embodiment of the present invention. The electronic device includes a memory 901, a processor 902, and a program or instructions stored in the memory 901 that run on the processor. When the program or instructions are executed by the processor 902, they can achieve the following: Figure 1 The corresponding intelligent computing cloud platform achieves the same beneficial effect by implementing any step of the elastic KV cache quantization method embodiment through computing power, and will not be elaborated here.

[0189] The processor 902 can be a CPU, ASIC, FPGA, or GPU.

[0190] Those skilled in the art will understand that all or part of the steps of the above-described intelligent computing cloud platform's method for implementing elastic KV caching through computing power can be accomplished by hardware related to program instructions, and the program can be stored in a readable medium.

[0191] The present invention also provides a readable storage medium on which a computer program is stored, and which, when executed by a processor, can perform the above-described functions. Figure 1 The corresponding intelligent computing cloud platform can achieve any step in the quantization method embodiment of elastic KV caching through computing power and achieve the same technical effect. To avoid repetition, it will not be described again here. The storage medium mentioned is such as read-only memory (ROM), random access memory (RAM), magnetic disk, or optical disk.

[0192] The present invention also provides a computer program product, including computer instructions that, when executed by a processor, implement the above-described... Figure 1 The corresponding intelligent computing cloud platform implements each process of the quantization method embodiment of elastic KV caching through computing power, and can achieve the same technical effect. To avoid repetition, it will not be described in detail here.

[0193] Through the above description of the embodiments, those skilled in the art can clearly understand that the methods of the above embodiments can be implemented by means of software plus necessary general-purpose hardware platforms. Of course, they can also be implemented by hardware, but in many cases the former is a better implementation method. Based on this understanding, the technical solution of the present invention, or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk) and includes several instructions to cause a terminal (which may be a mobile phone, computer, server, air conditioner, or second terminal device, etc.) to execute the methods of the various embodiments of the present invention.

[0194] The embodiments of the present invention have been described above with reference to the accompanying drawings. However, the present invention is not limited to the specific embodiments described above. The specific embodiments described above are merely illustrative and not restrictive. Those skilled in the art can make many other forms under the guidance of the present invention without departing from the spirit and scope of the claims, and all of these forms are within the protection scope of the present invention.

Claims

1. A method for quantizing elastic key-value caching through computing power in an intelligent computing cloud platform, characterized in that, include: Step S1: Divide the historical tokens of the intelligent computing cloud platform into multiple cache blocks, quantify the key data and / or value data of the historical tokens, write them into the corresponding cache blocks, and select multiple positions as candidate anchor points in the high-precision KV data of the historical tokens. Step S2: Calculate the digest key data and quantization error proxy data of the cache block as the block-level digest data of the cache block; Step S3: In response to the intelligent computing cloud platform generating a new target token, the cache block is scored based on the target token to generate a cache block score, and candidate cache blocks are selected from multiple cache blocks based on the cache block score; Step S4: Calculate the uncertainty index data of the candidate cache blocks, and determine the target cache block whose accuracy needs to be restored from the candidate cache blocks based on the uncertainty index data; Step S5: Locate the nearest preceding candidate anchor point to the target cache block as the upstream target anchor point, and perform partial replay of the high-precision KV data of the historical token based on the upstream target anchor point to generate the target high-precision KV data of the target cache block; Step S5 includes: Step S5.1: Perform multi-dimensional cropping on the data between the upstream target anchor point and the target cache block to generate the first candidate data between the upstream target anchor point and the target cache block; Step S5.2: Filter out the first candidate data to generate the second candidate data corresponding to the target cache block; Step S5.3: Calculate the attention contribution difference between the second candidate data and the low-bit data; Step S5.4: Correct the basic output data of the target cache block based on the attention contribution difference to obtain the target high-precision KV data; The multi-dimensional cropping includes one or more of the following: Perform time-dimensional pruning on the data between the upstream target anchor point and the target cache block; Perform layer-by-layer dimensional pruning on the data between the upstream target anchor point and the target cache block; Head dimension pruning is performed on the data between the upstream target anchor point and the target cache block; Channel dimension pruning is performed on the data between the upstream target anchor point and the target cache block; The data between the upstream target anchor point and the target cache block is pruned to a precision dimension.

2. The method according to claim 1, characterized in that, Step S1 includes: Step S1.1: Quantize the Key data based on a first quantization granularity and quantize the Value data based on a second quantization granularity, wherein the first quantization granularity is greater than the second quantization granularity.

3. The method according to claim 1, characterized in that, Step S1 includes: Step S1.2: For the high-precision KV data of the historical tokens, set a candidate anchor point every preset time interval according to the time sequence; Each candidate anchor point includes at least one of the following data: the hidden state of the predetermined network layer, the compressed representation of the input of the predetermined network layer, and the checkpoint state.

4. The method according to claim 1, characterized in that, Step S2 includes: Step S2.1: Calculate the mean squared error of the cache block before and after updating the historical token, and / or calculate the historical energy statistics and / or activity value and / or prior data of the cache block; Step S2.2: Calculate the quantization error proxy data based on the mean square error value and / or the historical energy statistics value and / or the activity value and / or the prior data.

5. The method according to claim 1, characterized in that, Step S3 includes: Step S3.1: Based on the similarity between the target token and the cache block across multiple attention heads, calculate the access tendency data of the cache block; Step S3.2: Calculate the cache block score of the cache block based on the similarity, the access tendency data, and the prior data of the cache block.

6. The method according to claim 5, characterized in that, Step S3.2 includes: Step S3.2.1: Take the maximum value of the similarity between the target token and the cache block on multiple attention heads as the target similarity, and perform a weighted calculation based on the target similarity, the access tendency data and the prior data of the cache block to generate the cache block score.

7. The method according to claim 1, characterized in that, Step S4 includes: Step S4.1: Calculate the score proximity data between the candidate buffer block and its neighboring candidate buffer blocks, calculate the quantization error value of the candidate buffer block, calculate the output distribution entropy of the current decoding step, and calculate the importance fluctuation data of the candidate buffer block; Step S4.2: Perform a weighted calculation based on the score proximity data, the quantization error value, the output distribution entropy, and the importance fluctuation data to generate the uncertainty index data.

8. The method according to claim 1, characterized in that, Step S5 includes: Step S5.5: Based on the upstream target anchor point, perform partial replay of the Key data of the target cache block to generate high-precision target Key data; Step S5.6: In response to the uncertainty index data of the target high-precision key data being greater than the judgment threshold, the value data of the target cache block is partially replayed based on the upstream target anchor point to generate target high-precision value data, and the target high-precision KV data is generated based on the target high-precision value data and the target high-precision key data.

9. The method according to claim 8, characterized in that, Step S5 includes: Step S5.7: In response to the uncertainty index data of the target high-precision key data being less than or equal to the judgment threshold, the target high-precision key data is used as the target high-precision KV data.

10. The method according to claim 1, characterized in that, Step S5 includes: Step S5.8: Based on the upstream target anchor point, perform partial playback of the target cache block according to the first recovery accuracy to generate candidate high-precision KV data; Step S5.9: In response to the uncertainty index data of the candidate high-precision KV data being greater than the judgment threshold, the target cache block is partially replayed based on the upstream target anchor point according to the second recovery accuracy to generate target high-precision KV data, wherein the second recovery accuracy is greater than the first recovery accuracy.

11. The method according to claim 10, characterized in that, Step S5 includes: Step S5.9: In response to the uncertainty index data of the candidate high-precision KV data being less than or equal to the judgment threshold, the candidate high-precision KV data is used as the target high-precision KV data.

12. The method according to claim 1, characterized in that, The method further includes: In response to the target high-precision KV data meeting the release condition, the target high-precision KV data is released.

13. The method according to claim 12, characterized in that, The release conditions include one or more of the following: The target high-precision KV data is used continuously a set number of times; The popularity of the target high-precision KV data is less than the popularity threshold; The memory pressure value of the intelligent computing cloud platform is greater than the pressure threshold.

14. A quantization device for an intelligent computing cloud platform to achieve elastic key-value caching through computing power, characterized in that, include: The partitioning module is used to divide the historical tokens of the intelligent computing cloud platform into multiple cache blocks, quantify the key data and / or value data of the historical tokens, write them into the corresponding cache blocks, and select multiple positions as candidate anchor points in the high-precision KV data of the historical tokens. The calculation module is used to calculate the digest key data and quantization error proxy data of the cache block as the block-level digest data of the cache block; The scoring module is used to respond to the intelligent computing cloud platform generating a new target token, score the cache block based on the target token to generate a cache block score, and filter candidate cache blocks from multiple cache blocks based on the cache block score; A filtering module is used to calculate the uncertainty index data of the candidate cache blocks and determine the target cache block whose accuracy needs to be restored from the candidate cache blocks based on the uncertainty index data. The replay module is used to locate the nearest preceding candidate anchor point to the target cache block as the upstream target anchor point, and to perform partial replay of the high-precision KV data of the historical token based on the upstream target anchor point to generate the target high-precision KV data of the target cache block. The playback module is also used for: The data between the upstream target anchor point and the target cache block is cropped in multiple dimensions to generate the first candidate data between the upstream target anchor point and the target cache block; The first candidate data is filtered out to generate the second candidate data corresponding to the target cache block; Calculate the attention contribution difference between the second candidate data and the low-bit data; The basic output data of the target cache block is corrected based on the attention contribution difference to obtain the target high-precision KV data; The multi-dimensional cropping includes one or more of the following: Perform time-dimensional pruning on the data between the upstream target anchor point and the target cache block; Perform layer-by-layer dimensional pruning on the data between the upstream target anchor point and the target cache block; Head dimension pruning is performed on the data between the upstream target anchor point and the target cache block; Channel dimension pruning is performed on the data between the upstream target anchor point and the target cache block; The data between the upstream target anchor point and the target cache block is pruned to a precision dimension.

15. An electronic device, characterized in that, include: A processor, a memory, and a program stored in the memory and executable on the processor, wherein the program, when executed by the processor, implements the steps of the quantization method for achieving elastic KV caching through computing power by the intelligent computing cloud platform as described in any one of claims 1 to 13.

16. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program, which, when executed by a processor, implements the steps of the quantization method for the intelligent computing cloud platform to achieve elastic KV caching through computing power as described in any one of claims 1 to 13.

17. A computer program product, characterized in that, It includes computer instructions, which, when executed by a processor, implement the steps of the quantization method for the intelligent computing cloud platform to achieve elastic KV caching through computing power as described in any one of claims 1 to 13.

Citation Information

Patent Citations

  • Large language model low-delay reasoning method based on dynamic reasoning graph optimization

    CN121072787A

  • Dynamic route selection method and system, electronic equipment and medium

    CN121396886A