Dynamic cache management method and device based on kv cache, storage medium and equipment

CN122817282APending Publication Date: 2026-09-25BEIJING YIYUAN TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202610749699.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-05-28
Publication Date
2026-09-25

AI Technical Summary

Technical Problem

[0003]本申请提供了一种基于KV Cache的动态缓存管理方法、装置、存储介质及设备,能够解决由于存储与计算分离导致的大模型推理中KV Cache访问带宽需求过高、总线传输成为性能瓶颈的问题

Benefits of technology

[0017]本申请实施例提供的基于KV Cache的动态缓存管理方法、装置、存储介质及设备,能够在主机与存储介质之间增加主控端,由主控端接收主机发送的大模型第N层推理计算所需的当前查询向量和当前会话标识,并从缓冲区或者非易失性存储阵列中读取当前会话标识对应的KV Cache,由主控端根据当前查询向量和读取的KV Cache计算当前注意力,直接将数据量较小的当前注意力反馈给主机,而无需再传输数据量较大的KV Cache,这样不仅减少了主机端的计算量,还大大降低了主机与控制端之间的带宽。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122817282A_ABST
    Figure CN122817282A_ABST
Patent Text Reader

Abstract

The application discloses a dynamic cache management method and device based on KV Cache, a storage medium and equipment. The method is applied to a master control end. The method comprises the following steps: receiving a current query vector and a current session identifier required by a large model N-layer inference calculation sent by a host, wherein N is a positive integer; in the case that a KV Cache corresponding to the current session identifier is pre-stored in a buffer area, reading the KV Cache corresponding to the current session identifier from the buffer area; in the case that the KV Cache corresponding to the current session identifier is not in the buffer area, caching the KV Cache corresponding to the current session identifier from a non-volatile storage array to the buffer area, and reading the KV Cache corresponding to the current session identifier from the buffer area; calculating a current attention output vector according to the current query vector and the KV Cache corresponding to the current session identifier; and feeding back the current attention output vector to the host.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of data caching technology, and more specifically, to a dynamic cache management method, apparatus, storage medium, and device based on KV Cache. Background Technology

[0002] In the inference process of LLM (Large Language Model), the generation of each new token requires accessing the entire KV cache of that session. This mechanism presents significant storage and memory access challenges: the KV cache data volume of a single session can reach hundreds of MB to several GB, and a complete read is required for each token generation. Existing architectures generally adopt a "storage and computation separation" design, storing the KV cache in NAND or DRAM (Dynamic Random Access Memory), while computational tasks are handled by the CPU (Central Processing Unit), GPU (Graphics Processing Unit), or NPU (Neural Network Processing Unit), with all raw data transferred between the two via a bus. As a result, the system's bandwidth requirements are directly proportional to the data volume and access frequency. For example, if a single decoding requires reading 500 MB of KV cache, and the task is to be completed within 50 ms, the required bandwidth is as high as 10 GB / s. It is evident that the fundamental contradiction of traditional solutions lies in the fact that the physical separation of storage and computation forces a large amount of data to be repeatedly moved via the bus, making it difficult to meet the stringent requirements of low latency and high throughput for large model inference. Summary of the Invention

[0003] This application provides a dynamic cache management method, apparatus, storage medium, and device based on KV Cache, which can solve the problem of excessive KV Cache access bandwidth requirements and bus transmission becoming a performance bottleneck in large model inference caused by the separation of storage and computation.

[0004] The specific technical solution is as follows: In a first aspect, embodiments of this application provide a dynamic cache management method based on KV Cache, the method being applied to a main control terminal, the method comprising: The receiver sends the current query vector and current session identifier required for the Nth layer inference computation of the large model, where N is a positive integer. If the KV Cache corresponding to the current session identifier is pre-stored in the buffer, the KV Cache corresponding to the current session identifier is read from the buffer; If the KV Cache corresponding to the current session identifier is not in the buffer, the KV Cache corresponding to the current session identifier is cached from the non-volatile storage array into the buffer, and the KV Cache corresponding to the current session identifier is read from the buffer. Calculate the current attention output vector based on the current query vector and the KV Cache corresponding to the current session identifier; The current attention output vector is fed back to the host.

[0005] In one possible implementation, the method further includes: Receive the current inference level progress of the Nth layer inference calculation of the large model sent by the host, or monitor the current inference level progress of the Nth layer inference calculation of the large model executed by the host; The target session identifier required for the N+Kth layer inference is predicted based on the current inference level progress, wherein the session identifier corresponding to the N+K-1th layer inference is the current session identifier, and the target session identifier is a different session identifier from the current session identifier; If the KV Cache corresponding to the target session identifier is not in the buffer, the KV Cache corresponding to the target session identifier is pre-cached from the non-volatile storage array into the buffer.

[0006] In one possible implementation, the method further includes: The size of K is constrained based on the single-level computation time of the large model and the average read latency of the non-volatile memory array.

[0007] In one possible implementation, the master control unit includes a prefetch scheduler, a controller for a non-volatile memory array, and a prefetch queue; The method of predicting the target session identifier required for the N+K layer inference based on the current inference level progress includes: the prefetch scheduler predicting the target session identifier required for the N+K layer inference based on the current inference level progress; The process of pre-caching the KV Cache corresponding to the target session identifier from the non-volatile storage array to the buffer includes: the pre-read scheduler issuing a pre-read command to the controller of the non-volatile storage array, the pre-read command including the storage address range, data length, and priority of the KV Cache corresponding to the target session identifier; after receiving the pre-read command, the controller reads the KV Cache corresponding to the target session identifier from the non-volatile storage array based on the pre-read command and outputs it to the pre-read queue; the pre-read queue buffers and performs flow control on the received KV Cache, and caches the buffered KV Cache to the buffer.

[0008] In one possible implementation, the method further includes: The main control terminal communicates with the host through a preset low-speed interface.

[0009] Secondly, embodiments of this application provide a dynamic cache management device based on KV Cache, the device being applied to a main control terminal, the device comprising: The receiving unit is used to receive the current query vector and current session identifier required for the Nth layer inference computation of the large model sent by the host, where N is a positive integer; The reading unit is used to read the KV Cache corresponding to the current session identifier from the buffer if the KV Cache corresponding to the current session identifier is pre-stored in the buffer; A caching unit is configured to cache the KV Cache corresponding to the current session identifier from the non-volatile storage array into the buffer when there is no KV Cache corresponding to the current session identifier in the buffer. The reading unit is also used to read the KV Cache corresponding to the current session identifier from the buffer; The calculation unit is used to calculate the current attention output vector based on the current query vector and the KV Cache corresponding to the current session identifier; The feedback unit is used to feed back the current attention output vector to the host.

[0010] In one possible implementation, the device further includes: The progress acquisition unit is used to receive the current inference level progress of the Nth layer inference calculation of the large model sent by the host, or to monitor the current inference level progress of the host performing the Nth layer inference calculation of the large model. The prediction unit is used to predict the target session identifier required for the N+Kth layer inference based on the current inference level progress, wherein the session identifier corresponding to the N+K-1th layer inference is the current session identifier, and the target session identifier is a different session identifier from the current session identifier. The caching unit is further configured to, in the absence of a KV Cache corresponding to the target session identifier in the buffer, pre-cache the KV Cache corresponding to the target session identifier from the non-volatile storage array into the buffer.

[0011] In one possible implementation, the device further includes: A constraint unit is used to constrain the size of K based on the single-level computation time of the large model and the average read latency of the non-volatile memory array.

[0012] In one possible implementation, the master control unit includes a prefetch scheduler, a controller for a non-volatile memory array, and a prefetch queue; The prediction unit is used by the pre-read scheduler to predict the target session identifier required for the N+K layer inference based on the current inference level progress. The caching unit is further configured to have the prefetch scheduler issue a prefetch command to the controller of the non-volatile storage array. The prefetch command includes the storage address range, data length, and priority of the KV Cache corresponding to the target session identifier. After receiving the prefetch command, the controller reads the KV Cache corresponding to the target session identifier from the non-volatile storage array based on the prefetch command and outputs it to the prefetch queue. The prefetch queue performs buffering and flow control on the received KV Cache and caches the buffered KV Cache in the buffer.

[0013] In one possible implementation, the device further includes: A communication unit is used for communication between the main control terminal and the host through a preset low-speed interface.

[0014] Thirdly, embodiments of this application provide a computer-readable storage medium having a computer program stored thereon that, when executed by a processor, implements the method as described in any possible implementation of the first aspect.

[0015] Fourthly, embodiments of this application provide an electronic device, which includes: One or more processors; The processor is coupled to a storage device for storing one or more programs; When one or more programs are executed by one or more processors, the electronic device performs the method as described in any possible implementation of the first aspect.

[0016] Fifthly, embodiments of this application provide a computer program product containing instructions that, when executed on a computer or processor, cause the computer or processor to perform the method described in any possible implementation of the first aspect.

[0017] The dynamic cache management method, apparatus, storage medium, and device based on KV Cache provided in this application embodiment can add a master control terminal between the host and the storage medium. The master control terminal receives the current query vector and current session identifier required for the Nth layer inference calculation of a large model sent by the host, and reads the KV Cache corresponding to the current session identifier from the buffer or non-volatile storage array. The master control terminal calculates the current attention based on the current query vector and the read KV Cache, and directly feeds back the current attention with a small amount of data to the host without transmitting the KV Cache with a large amount of data. This not only reduces the amount of computation on the host side, but also greatly reduces the bandwidth between the host and the control terminal.

[0018] Furthermore, the technical effects that can be achieved by the embodiments of this application include: 1. In this embodiment, the master control terminal predicts the target session identifier required for the N+K layer inference based on the current inference level progress calculated by the Nth layer inference of the large model. If there is no KV Cache corresponding to the target session identifier in the buffer, the KV Cache corresponding to the target session identifier can be cached in advance from the non-volatile storage array to the buffer. Thus, when the large model performs the N+K layer inference, the master control terminal can directly read the required KV Cache from the buffer with higher read efficiency for attention calculation, without having to read from the non-volatile storage array with lower read efficiency. This can greatly reduce the latency of KV Cache data reading and improve the adaptability to real-time inference of the large model.

[0019] 2. In this embodiment, the size of K is constrained based on the single-level computation time of the large model and the average read latency of the non-volatile storage array, rather than blindly setting the size of K. This can prevent the large model from inferring to the N+K layer before the pre-read is completed due to the K value being too small, thus making the pre-read unable to achieve the purpose of reducing latency. On the other hand, it can prevent the layers before the N+K layer from being pre-read due to the K value being too large, thus failing to reduce the data read latency.

[0020] 3. The main control terminal communicates with the host through a preset low-speed interface, which can not only meet the low-bandwidth data transmission requirements, but also reduce interface costs. Attached Figure Description

[0021] To more clearly illustrate the technical solutions in the embodiments of this application or the prior art, the accompanying drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are merely some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without any creative effort.

[0022] Figure 1 A flowchart illustrating a dynamic cache management method based on KV Cache provided in this application embodiment;

[0023] Figure 2 This is a block diagram of a dynamic cache management device based on KV Cache provided in an embodiment of this application. Detailed Implementation

[0024] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only a part of the embodiments of this application, and not all of them. All other embodiments obtained by those skilled in the art based on the embodiments of this application without creative effort are within the scope of protection of this application.

[0025] It should be noted that, unless otherwise specified, the embodiments and features described in this application can be combined with each other. The terms "comprising" and "having," and any variations thereof, in the embodiments and drawings of this application are intended to cover non-exclusive inclusion. For example, a process, method, system, product, or device that includes a series of steps or units is not limited to the listed steps or units, but may optionally include steps or units not listed, or may optionally include other steps or units inherent to these processes, methods, products, or devices.

[0026] This application provides a flowchart illustrating a dynamic cache management method based on KV Cache. The method is applied to the main control unit, such as... Figure 1 As shown, the method includes: S110: Receive the current query vector and current session identifier required for the Nth layer inference computation of the large model sent by the host.

[0027] Where N is a positive integer. The large model can be a Transformer model or other models, and this application does not limit it.

[0028] In the hierarchical inference process of a large model, the self-attention calculation of each layer (e.g., the Nth layer) requires the query vector corresponding to the current token and the session ID used to distinguish different dialogue sessions. In this embodiment, the main control terminal receives the current query vector and the current session ID sent by the host (e.g., CPU / GPU / NPU). The current query vector is used for attention calculation with the KV Cache corresponding to the current session ID, while the current session ID is used to quickly locate historical key-value pairs belonging to that session in the global cache.

[0029] In large model inference, the host obtains the current query vector and current session identifier required for the Nth layer inference computation of the large model in the following way: For the currently generated token, the host starts from the input embedding and performs forward computation layer by layer. At the Nth layer, the query vector is calculated in real time based on the weight matrix of the current layer and the output of the previous layer (or the representation of the current token). The session identifier is assigned and maintained by the inference service when the session is created, and is used to uniquely identify the dialogue.

[0030] After obtaining the current session identifier, the corresponding KV cache can be retrieved from the non-volatile storage array or buffer. This is because the KV cache data for each session is pre-stored in high-capacity storage media (such as a non-volatile storage array) or may be cached in a buffer. When the host sends the current session identifier, the master control unit uses a pre-established mapping table or direct address index to convert the session identifier into the corresponding physical storage address, thereby quickly locating and reading the KV cache data corresponding to that session without scanning the entire storage space.

[0031] In addition, communication between the main control unit and the host is no longer limited to high-speed interface communication; it can also communicate through a preset low-speed interface, such as the eMMC (Embedded Multi Media Card) interface.

[0032] S120: If the KV Cache corresponding to the current session identifier is pre-stored in the buffer, read the KV Cache corresponding to the current session identifier from the buffer.

[0033] During the initialization phase, the host loads large model weights into the inference engine and stores all session KV caches in compressed format in a non-volatile memory array. The master control unit establishes a KV cache index table (including the mapping relationship between session identifiers and storage addresses in the non-volatile memory array). During the prefill phase, the host processes user-input prompts, computes the KV caches of all input tokens in parallel, and writes all KV caches into the non-volatile memory array via the master control unit. Simultaneously, it caches some KV caches in a buffer, such as the recently accessed KV caches. During the decoding phase, the master control unit first checks if the buffer contains a pre-stored KV cache corresponding to the current session identifier. If the buffer does contain such a cache, it directly reads it from the buffer. If the buffer does not contain a KV cache corresponding to the current session identifier, step S130 is executed. The buffer can be a DRAM array or other types of data.

[0034] S130: If there is no KV Cache corresponding to the current session identifier in the buffer, cache the KV Cache corresponding to the current session identifier from the non-volatile storage array into the buffer, and read the KV Cache corresponding to the current session identifier from the buffer.

[0035] If the KV Cache corresponding to the current session identifier is not in the buffer, the master can first cache the KV Cache corresponding to the current session identifier from the non-volatile storage array into the buffer, and then read the KV Cache corresponding to the current session identifier from the buffer. Alternatively, it can directly read the KV Cache corresponding to the current session identifier from the non-volatile storage array.

[0036] The non-volatile memory array serves as the primary storage carrier for the KV cache, storing all KV cache data from concurrent sessions. Specifically, it can be a multi-channel NAND Flash array. A core configuration might include: 16 identical NAND Flash chips forming a 16-channel parallel architecture, with a single chip capacity of 128GB and a total storage capacity of 2TB, capable of storing approximately 4000 KV cache entries with a single session capacity of 500MB; a single chip read bandwidth of 1GB / s, and an aggregated parallel read bandwidth of up to 16GB / s, matching the bandwidth requirements of DRAM data transfer. The total system capacity and read bandwidth can be linearly expanded by increasing or decreasing the number of NAND channels.

[0037] S140: Calculate the current attention output vector based on the KV Cache corresponding to the current query vector and the current session identifier.

[0038] After obtaining the KV Cache corresponding to the current query vector and the current session identifier, the master control unit can calculate the current attention output vector based on the KV Cache. The specific calculation formula includes: ; Where output represents the current attention output vector, query represents the current query vector, K represents the key value in the KV Cache, V represents the value value in the KV Cache, d represents the dimension of the current query vector, T represents the transpose, and softmax function represents the normalized exponential function.

[0039] In calculation When multiplying and accumulating arrays are needed, the architecture of these arrays can be multiplying and accumulating tree arrays, systolic arrays, vectorized arrays, etc.

[0040] The softmax function requires exponential operations and normalization, which makes hardware implementation complex. To achieve fast calculation, methods include, but are not limited to, table lookup + linear interpolation and piecewise linear approximation.

[0041] S150: Feed the current attention output vector back to the host.

[0042] After the master control terminal obtains the current attention output vector, it can feed the current attention output vector back to the host through a preset low-speed interface so that the host can continue to complete the remaining calculations of the Nth layer.

[0043] The dynamic cache management method based on KV Cache provided in this application embodiment can add a master control terminal between the host and the storage medium. The master control terminal receives the current query vector and current session identifier required for the Nth layer inference calculation of the large model sent by the host, and reads the KV Cache corresponding to the current session identifier from the buffer or non-volatile storage array. The master control terminal calculates the current attention based on the current query vector and the read KV Cache, and directly feeds back the current attention with a small amount of data to the host, without having to transmit the KV Cache with a large amount of data. This not only reduces the amount of computation on the host side, but also greatly reduces the bandwidth between the host and the control terminal.

[0044] In one possible implementation, to further reduce the latency of KV Cache data reading and improve adaptability to real-time inference with large models, another embodiment of this application provides the following method: A1. Receive the current inference level progress of the Nth layer inference calculation of the large model sent by the host, or listen to the current inference level progress of the Nth layer inference calculation of the large model executed by the host.

[0045] The host can send the inference progress of each layer of the large model to the master control terminal in real time or periodically, so that the master control terminal can passively receive the inference progress sent by the host. The master control terminal can also actively monitor the inference progress of each layer of the large model during inference. For example, by inserting custom callback functions or hooks before and after the forward propagation of each layer of the model, the master control terminal can actively trigger progress reporting when the calculation of that layer is completed, thereby realizing real-time monitoring of the inference level by the master control terminal.

[0046] A2. Predict the target session identifier required for the N+Kth level of inference based on the current inference level progress.

[0047] In this context, the session identifier corresponding to the N+K-1th layer inference is the current session identifier, and the target session identifier is different from the current session identifier. (Hereinafter referred to as the first constraint)

[0048] The master control unit may include a prefetch scheduler, a controller for the non-volatile memory array, and a prefetch queue. In this step, the prefetch scheduler predicts the target session identifier required for inference at layer N+K based on the current inference layer progress.

[0049] In practice, the master control unit can begin predicting the target session identifier required for the N+K layer inference when the current inference layer progress reaches a preset progress threshold. If the current inference layer progress has not reached the preset progress threshold, this prediction operation will not be performed. The remaining time corresponding to the preset progress threshold can be greater than or equal to the average read latency of the non-volatile memory array. The remaining time corresponding to the preset progress threshold is the time required for the remaining progress corresponding to the preset progress threshold.

[0050] In one possible implementation, the master controller can constrain the size of K based on the single-level calculation time of the large model and the average read latency of the non-volatile memory array. (Hereinafter referred to as the second constraint)

[0051] Specific constraint expressions include, but are not limited to, the following: K ≥ T_read_nand / T_layer; Where K represents the number of read-ahead layers, T_layer represents the computation time of a single layer in a large model, and T_read_nand represents the average read latency of the non-volatile memory array.

[0052] Example parameters: If the calculation time for a single layer is T_layer=40ms and the number of layers to be read in advance is K=2, then the available time window K×T_layer is 80ms, which can completely cover the average read latency of 50ms for NAND Flash, achieving full coverage of latency, and the host and controller will not perceive any additional access latency.

[0053] Therefore, K only needs to satisfy the first and second constraints. When there are multiple K values ​​that satisfy both constraints, any one of them can be chosen, or the minimum value can be selected.

[0054] In another possible implementation, the value of K can also be dynamically adjusted according to the system's runtime state. This dynamic adjustment can be under the combined constraints of the first and second constraints, or solely under the first constraint. Specifically, the master control unit can monitor the current read load of the non-volatile memory array and the KV cache hit rate of the buffer in real time, and dynamically adjust the value of the pre-read advance layer number K accordingly: when the current read load of the non-volatile memory array is less than a first load threshold, and / or the KV cache hit rate of the buffer is less than a first preset hit threshold, the value of K is decreased; when the current read load of the non-volatile memory array is greater than or equal to a second load threshold, and / or the KV cache hit rate of the buffer is greater than or equal to a second preset hit threshold, the value of K is increased. The first load threshold is less than or equal to the second load threshold, and the first preset hit threshold is less than or equal to the second preset hit threshold.

[0055] In layman's terms, when the non-volatile storage array is under low load (i.e., the current read load is less than the first load threshold), its actual read latency is low, and the value of K can be appropriately reduced to lower unnecessary prefetch overhead. When the non-volatile storage array is under high load (i.e., the current read load is greater than or equal to the second load threshold), its actual read latency increases, and the value of K needs to be increased accordingly to ensure that the prefetch operation can be completed before the target layer inference begins, maintaining the latency masking effect. When the KV Cache hit rate of the buffer is less than the first preset hit threshold (i.e., the current hit rate is low), it indicates that most of the data obtained by the prefetch operation has not been reused, and the value of K can be appropriately reduced to avoid invalid prefetches occupying storage bandwidth and buffer space. When the KV Cache hit rate of the buffer is greater than or equal to the second preset hit threshold (i.e., the current hit rate is high), it indicates that the prefetch data can effectively hit subsequent inference requests, and the value of K needs to be increased accordingly to load more cached data of future layers in advance, further improving the latency masking effect. By dynamically adjusting, the system can adaptively balance prefetch effectiveness and resource overhead under different cache hit scenarios, thereby improving the overall efficiency of the system.

[0056] By dynamically adjusting, the system can adaptively balance prefetch timeliness and resource overhead under different loads and / or different cache hit scenarios, thereby improving the overall efficiency of the system.

[0057] A3. If there is no KV Cache corresponding to the target session identifier in the buffer, pre-cache the KV Cache corresponding to the target session identifier from the non-volatile storage array into the buffer.

[0058] In a master control system comprising a prefetch scheduler, a controller for a non-volatile memory array, and a prefetch queue, the prefetch scheduler issues a prefetch command to the controller of the non-volatile memory array. The prefetch command includes the storage address range, data length, and priority of the KV Cache corresponding to the target session identifier. After receiving the prefetch command, the controller reads the KV Cache corresponding to the target session identifier from the non-volatile memory array based on the prefetch command and outputs it to the prefetch queue. The prefetch queue performs buffering and flow control on the received KV Cache, caching the buffered KV Cache into a buffer.

[0059] The storage address range directly specifies the start and end physical addresses of the KV Cache data corresponding to the target session identifier in the non-volatile storage array (such as NAND Flash), which is the core basis for data location. The data length is used to verify the integrity of the read operation, ensuring that the amount of data read from the specified address range is exactly equal to the KV Cache size of the session (e.g., 500MB), preventing over-reading or under-reading. The priority is used to schedule multiple pre-read commands: when the storage array is busy, the controller executes the higher-priority commands first, and then initiates the read according to the address range when it is the turn of this command. Therefore, the three parameters work together to enable the controller to obtain the KV Cache data corresponding to the target session identifier in an orderly and complete manner.

[0060] Specifically, the pre-read queue can be a FIFO (First In First Out) pre-read queue. When the queue is full, it sends a back pressure signal to the controller, and when the queue is empty, it sends a pre-read acceleration notification to the pre-read scheduler.

[0061] Specifically, the buffer can be a DRAM buffer. For example, its configuration can be: capacity configurable from 1 to 8GB, with a minimum configuration of 1GB supporting KV cache for 1 active session + 1 pre-read session, and a recommended configuration of 2-4GB supporting KV cache reuse for 4-8 sessions; the cache replacement strategy supports LRU (Least Recently Used) algorithm or a management strategy based on session lifecycle.

[0062] In addition, for high-priority sessions, their KV Cache is kept in the buffer to ensure zero-latency access; for ordinary-priority sessions, on-demand hierarchical pre-reading is performed; for system idle periods, predictive pre-reading is performed based on historical access characteristics to preload the KV Cache of frequently accessed sessions into the buffer and improve cache hit rate.

[0063] In general, the caching rules for buffers include, but are not limited to, the following: Resident protection rule: The KV cache of an active session currently performing inference must be locked and kept in the buffer, and should not be evicted, to ensure low-latency access during the inference process; Read-ahead reservation rule: Reserve a fixed buffer address space for read-ahead operations to ensure that KVCache data that has been read ahead can be written normally without space conflicts; Eviction rules: When the buffer is full, algorithms such as LRU are used to evict the least accessed inactive session KV Cache data; after session inference is completed, it can be configured to release immediately or retain for a period of time for cache reuse; Capacity configuration rules: The minimum configuration is the KV Cache storage space for a preset number of sessions (e.g., 2) (approximately 1GB). For example, to support basic operation with 1 active session + 1 pre-read session, the recommended configuration is the storage space for 4-8 sessions (2-4GB). This supports cache reuse of KV Cache across multiple sessions, reducing the overhead of repeated reads from non-volatile storage arrays.

[0064] This embodiment predicts the target session identifier required for the N+Kth layer inference by using the master control terminal to calculate the current inference level progress of the Nth layer inference of the large model. If the KV Cache corresponding to the target session identifier is not present in the buffer, it can be pre-cached from the non-volatile storage array into a buffer. Therefore, when the large model performs the N+Kth layer inference, the master control terminal can directly read the required KV Cache from the more efficient buffer for attention calculation, without needing to read from the less efficient non-volatile storage array. This significantly reduces the latency of KV Cache data reading and improves adaptability to real-time inference of the large model. Furthermore, this embodiment constrains the size of K based on the single-layer computation time of the large model and the average read latency of the non-volatile storage array, rather than blindly setting the value of K. This prevents situations where the pre-reading is incomplete before the large model infers to the N+Kth layer due to an excessively small K value, thus failing to reduce latency. Conversely, it prevents situations where the layers before the N+Kth layer cannot be pre-read due to an excessively large K value, thus failing to reduce data read latency.

[0065] Based on the above method embodiments, another embodiment of this application provides a dynamic cache management device based on KV Cache. This device is applied to the main control terminal, such as... Figure 2 As shown, the device includes: The receiving unit 210 is used to receive the current query vector and current session identifier required for the Nth layer inference calculation of the large model sent by the host, where N is a positive integer; The reading unit 220 is used to read the KV Cache corresponding to the current session identifier from the buffer when the buffer pre-stores the KV Cache corresponding to the current session identifier; The caching unit 230 is used to cache the KV Cache corresponding to the current session identifier from the non-volatile storage array into the buffer when there is no KV Cache corresponding to the current session identifier in the buffer. The reading unit 220 is also used to read the KVCache corresponding to the current session identifier from the buffer; The calculation unit 240 is used to calculate the current attention output vector based on the current query vector and the KV Cache corresponding to the current session identifier; Feedback unit 250 is used to feed back the current attention output vector to the host.

[0066] In one possible implementation, the device further includes: The progress acquisition unit is used to receive the current inference level progress of the Nth layer inference calculation of the large model sent by the host, or to monitor the current inference level progress of the host performing the Nth layer inference calculation of the large model. The prediction unit is used to predict the target session identifier required for the N+Kth layer inference based on the current inference level progress, wherein the session identifier corresponding to the N+K-1th layer inference is the current session identifier, and the target session identifier is a different session identifier from the current session identifier. The cache unit 230 is further configured to pre-cache the KV Cache corresponding to the target session identifier from the non-volatile storage array into the buffer when there is no KV Cache corresponding to the target session identifier in the buffer.

[0067] In one possible implementation, the device further includes: A constraint unit is used to constrain the size of K based on the single-level computation time of the large model and the average read latency of the non-volatile memory array.

[0068] In one possible implementation, the master control unit includes a prefetch scheduler, a controller for a non-volatile memory array, and a prefetch queue; The prediction unit is used by the pre-read scheduler to predict the target session identifier required for the N+K layer inference based on the current inference level progress. The cache unit 230 is further configured to have the prefetch scheduler issue a prefetch command to the controller of the non-volatile storage array. The prefetch command includes the storage address range, data length, and priority of the KV Cache corresponding to the target session identifier. After receiving the prefetch command, the controller reads the KV Cache corresponding to the target session identifier from the non-volatile storage array based on the prefetch command and outputs it to the prefetch queue. The prefetch queue performs buffering and flow control on the received KV Cache and caches the buffered KV Cache in the buffer.

[0069] In one possible implementation, the device further includes: A communication unit is used for communication between the main control terminal and the host through a preset low-speed interface.

[0070] The dynamic cache management device based on KV Cache provided in this application embodiment can add a master control terminal between the host and the storage medium. The master control terminal receives the current query vector and current session identifier required for the Nth layer inference calculation of the large model sent by the host, and reads the KV Cache corresponding to the current session identifier from the buffer or non-volatile storage array. The master control terminal calculates the current attention based on the current query vector and the read KV Cache, and directly feeds back the current attention with a small amount of data to the host, without having to transmit the KV Cache with a large amount of data. This not only reduces the amount of computation on the host side, but also greatly reduces the bandwidth between the host and the control terminal.

[0071] Based on the above method embodiments, another embodiment of this application provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the method as described in any of the above embodiments.

[0072] Based on the above method embodiments, another embodiment of this application provides an electronic device or computer device, including: One or more processors; The processor is coupled to a storage device for storing one or more programs; When the one or more programs are executed by the one or more processors, the electronic device or computer device performs the method as described in any of the above embodiments.

[0073] Based on the above embodiments, another embodiment of this application provides a computer program product, which includes instructions that, when executed on a computer or processor, cause the computer or processor to perform the method described in any of the above embodiments.

[0074] The above-described apparatus embodiments correspond to the method embodiments and have the same technical effects. For detailed descriptions, please refer to the method embodiments. The apparatus embodiments are derived from the method embodiments; detailed descriptions can be found in the method embodiments section, and will not be repeated here. Those skilled in the art will understand that the accompanying drawings are merely schematic diagrams of one embodiment, and the modules or processes shown in the drawings are not necessarily essential for implementing this application.

[0075] Those skilled in the art will understand that the modules in the apparatus of the embodiments can be distributed in the apparatus of the embodiments as described in the embodiments, or they can be located in one or more devices different from this embodiment with corresponding changes. The modules of the above embodiments can be combined into one module, or they can be further divided into multiple sub-modules.

[0076] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of this application, and are not intended to limit them. Although this application has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of this application.

Claims

1. A dynamic cache management method based on KV Cache, characterized in that, The method is applied to the main control terminal, and the method includes: The receiver sends the current query vector and current session identifier required for the Nth layer inference computation of the large model, where N is a positive integer. If the KV Cache corresponding to the current session identifier is pre-stored in the buffer, the KV Cache corresponding to the current session identifier is read from the buffer; If the KV Cache corresponding to the current session identifier is not in the buffer, the KV Cache corresponding to the current session identifier is cached from the non-volatile storage array into the buffer, and the KV Cache corresponding to the current session identifier is read from the buffer. Calculate the current attention output vector based on the current query vector and the KV Cache corresponding to the current session identifier; The current attention output vector is fed back to the host.

2. The method according to claim 1, characterized in that, The method further includes: Receive the current inference level progress of the Nth layer inference calculation of the large model sent by the host, or monitor the current inference level progress of the Nth layer inference calculation of the large model executed by the host; The target session identifier required for the N+Kth layer inference is predicted based on the current inference level progress, wherein the session identifier corresponding to the N+K-1th layer inference is the current session identifier, and the target session identifier is a different session identifier from the current session identifier; If the KV Cache corresponding to the target session identifier is not in the buffer, the KV Cache corresponding to the target session identifier is pre-cached from the non-volatile storage array into the buffer.

3. The method according to claim 2, characterized in that, The method further includes: The size of K is constrained based on the single-level computation time of the large model and the average read latency of the non-volatile memory array.

4. The method according to claim 2, characterized in that, The main control unit includes a prefetch scheduler, a controller for the non-volatile memory array, and a prefetch queue; The method of predicting the target session identifier required for the N+K layer inference based on the current inference level progress includes: the prefetch scheduler predicting the target session identifier required for the N+K layer inference based on the current inference level progress; The process of pre-caching the KV Cache corresponding to the target session identifier from the non-volatile storage array to the buffer includes: the pre-read scheduler issuing a pre-read command to the controller of the non-volatile storage array, the pre-read command including the storage address range, data length, and priority of the KV Cache corresponding to the target session identifier; after receiving the pre-read command, the controller reads the KV Cache corresponding to the target session identifier from the non-volatile storage array based on the pre-read command and outputs it to the pre-read queue; the pre-read queue buffers and performs flow control on the received KV Cache, and caches the buffered KV Cache to the buffer.

5. The method according to any one of claims 1-4, characterized in that, The method further includes: The main control terminal communicates with the host through a preset low-speed interface.

6. A dynamic cache management device based on KV Cache, characterized in that, The device is applied to the main control terminal, and the device includes: The receiving unit is used to receive the current query vector and current session identifier required for the Nth layer inference computation of the large model sent by the host, where N is a positive integer; The reading unit is used to read the KV Cache corresponding to the current session identifier from the buffer if the KV Cache corresponding to the current session identifier is pre-stored in the buffer; A caching unit is configured to cache the KV Cache corresponding to the current session identifier from the non-volatile storage array into the buffer when there is no KV Cache corresponding to the current session identifier in the buffer. The reading unit is also used to read the KV Cache corresponding to the current session identifier from the buffer; The calculation unit is used to calculate the current attention output vector based on the current query vector and the KV Cache corresponding to the current session identifier; The feedback unit is used to feed back the current attention output vector to the host.

7. The apparatus according to claim 6, characterized in that, The device further includes: The progress acquisition unit is used to receive the current inference level progress of the Nth layer inference calculation of the large model sent by the host, or to monitor the current inference level progress of the host performing the Nth layer inference calculation of the large model. The prediction unit is used to predict the target session identifier required for the N+Kth layer inference based on the current inference level progress, wherein the session identifier corresponding to the N+K-1th layer inference is the current session identifier, and the target session identifier is a different session identifier from the current session identifier. The caching unit is further configured to, in the absence of a KV Cache corresponding to the target session identifier in the buffer, pre-cache the KV Cache corresponding to the target session identifier from the non-volatile storage array into the buffer.

8. The apparatus according to claim 6, characterized in that, The device further includes: A constraint unit is used to constrain the size of K based on the single-level computation time of the large model and the average read latency of the non-volatile memory array.

9. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the program is executed by the processing unit, it implements the method as described in any one of claims 1-5.

10. An electronic device, characterized in that, The electronic device includes: One or more processing units; The processing unit is coupled to the storage unit, which is used to store one or more programs. When the one or more programs are executed by the one or more processing units, the electronic device performs the method as described in any one of claims 1-5.