Key-value cache management method and device, equipment, storage medium and program product

CN122816544APending Publication Date: 2026-09-25MOORE THREAD INTELLIGENT TECHNOLOGY (HANGZHOU) CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202611105014.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-07-23
Publication Date
2026-09-25

AI Technical Summary

Technical Problem

以输入序列为单位的存储策略会为每个请求独立存储相同前缀的KV数据,导致存储资源的浪费和计算冗余

Benefits of technology

[0012]在此基础上,基于第一词元块的哈希值查询哈希映射表来确定匹配的键值缓存块,其中哈希映射表存储有哈希值与已分配的键值缓存块的标识之间的对应关系,因而能够根据第一词元块的哈希值在哈希映射表中直接定位到可能对应的候选键值缓存块的标识,无需将第一词元块的词元序列与已分配的各个键值缓存块所关联的词元序列逐一进行比对,从而显著减少了匹配过程中的内容比较次数,降低了匹配处理所需的计算量。此外,由于哈希映射表的查找操作所消耗的时间与已分配的键值缓存块的总数量无关,即使已分配的键值缓存块的数量增多,匹配操作所消耗的时间也不会随之显著增长,从而在保证匹配可行性的前提下提高了确定与第一词元块匹配的键值缓存块的效率,进而提高了大语言模型推理系统的吞吐量,并减少了大语言模型的首词元延迟(Time To First Token,TTFT),或者说减少了大语言模型从用户提交输入序列开始,到其生成并输出第一个响应词元为止的所经历的时间。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122816544A_ABST
    Figure CN122816544A_ABST
Patent Text Reader

Abstract

The embodiment of the present disclosure provides a key-value cache management method and device, equipment, a storage medium and a program product, relates to the technical field of computers, in particular to the technical field of large language models, and specifically relates to a key-value cache management method and device, equipment, a storage medium and a program product. The key-value cache management method comprises the following steps: dividing a plurality of word pieces included in a first input sequence of a large language model into a plurality of word piece blocks, wherein a first word piece block in the plurality of word piece blocks comprises word pieces meeting a predetermined quantity condition; determining a key-value cache block matched with the first word piece block based on a hash value of the first word piece block and a hash mapping table, wherein the hash mapping table is used to store a corresponding relationship between the hash value and an identifier of an allocated key-value cache block. Thus, the embodiment of the present disclosure can solve the technical problems of high storage resource occupation and redundant calculation in the related art, and effectively reduce the storage resource occupation and reduce the repeated calculation.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The embodiments disclosed herein relate to the field of computer technology, and more particularly to the field of large language model technology, specifically to key-value cache management methods, key-value cache management devices, electronic devices, readable storage media, and computer program products. Background Technology

[0002] Large Language Models (LLMs) require attention calculations for each token in the input sequence when processing the input prompt and generating output. This calculation involves repeatedly reading the key and value vectors of each token. To avoid redundant calculations, the inference system of a large language model can employ a key-value cache (KV cache) technique to store the calculated key and value vectors and reuse them directly in subsequent processing, thereby significantly reducing computational overhead.

[0003] The input sequence can be segmented into multiple tokens by a tokenizer. In some technical solutions, the KV cache management method stores the key vectors and value vectors of all tokens in the sequence contiguously, treating the entire input sequence as a unit. However, in multi-request concurrency or batch processing scenarios, the input sequences of different requests often have the same prefix (e.g., system prompts, dialogue templates, or common knowledge fragments). The storage strategy based on the input sequence would store the KV data with the same prefix independently for each request, resulting in wasted storage resources and computational redundancy.

[0004] Therefore, there is an urgent need in this field for a solution that can enable different requests to share key-value cache data with the same prefix without increasing too much computational overhead, so as to reduce storage consumption and reduce redundant calculations. Summary of the Invention

[0005] Embodiments of this disclosure provide key-value cache management methods and apparatuses, electronic devices, readable storage media, and computer program products that can at least partially solve the above-described problems or other problems in the art.

[0006] According to a first aspect of this disclosure, a key-value cache management method is provided, the method comprising: dividing a plurality of lexical units included in a first input sequence of a large language model into a plurality of lexical blocks, wherein a first lexical block among the plurality of lexical blocks includes lexical units that meet a predetermined number of conditions; determining a key-value cache block matching the first lexical block based on the hash value of the first lexical block and a hash mapping table, wherein the hash mapping table is used to store the correspondence between the hash value and the identifier of the allocated key-value cache block.

[0007] According to a second aspect of this disclosure, a key-value cache management apparatus is provided, the apparatus comprising: a partitioning unit configured to partition a plurality of lexical units included in a first input sequence of a large language model into a plurality of lexical blocks, wherein a first lexical block among the plurality of lexical blocks includes lexical units that meet a predetermined number of conditions; and a matching unit configured to determine a key-value cache block that matches the first lexical block based on a hash value of the first lexical block and a hash mapping table, wherein the hash mapping table is used to store the correspondence between hash values ​​and identifiers of allocated key-value cache blocks.

[0008] According to a third aspect of this disclosure, an electronic device is provided, the electronic device including a processor, the processor being used to implement the key-value cache management method of the first aspect and any possible implementation thereof.

[0009] According to a fourth aspect of this disclosure, a non-volatile computer-readable storage medium is provided, on which computer program instructions are stored, which, when executed by a processor, implement the key-value cache management method of the first aspect and any possible implementation thereof.

[0010] According to a fifth aspect of this disclosure, a computer program product is provided, comprising a computer program that, when executed by a processor, implements the steps of the key-value cache management method in the first aspect and any possible implementation thereof.

[0011] The key-value cache management method, apparatus, electronic device, readable storage medium, and computer program product provided by the embodiments of this disclosure divide the multiple lexical units included in the first input sequence of a large language model into multiple lexical blocks, thereby reducing the management granularity of the key-value cache from the entire input sequence to lexical blocks, improving the fineness of cache management, and facilitating the orderly allocation of storage resources.

[0012] Based on this, the hash map table is queried based on the hash value of the first word block to determine the matching key-value cache block. The hash map table stores the correspondence between hash values ​​and the identifiers of allocated key-value cache blocks. Therefore, the identifier of a possible candidate key-value cache block can be directly located in the hash map table based on the hash value of the first word block, without having to compare the word sequence of the first word block with the word sequences associated with each allocated key-value cache block one by one. This significantly reduces the number of content comparisons in the matching process and lowers the computational load required for matching. In addition, since the time consumed by the hash map table lookup operation is independent of the total number of allocated key-value cache blocks, even if the number of allocated key-value cache blocks increases, the time consumed by the matching operation will not increase significantly. This improves the efficiency of determining the key-value cache block that matches the first word block while ensuring matching feasibility, thereby improving the throughput of the large language model inference system and reducing the time to first token (TTFT) latency of the large language model, or in other words, reducing the time elapsed from the user submitting the input sequence to the generation and output of the first response word.

[0013] It should be understood that the description in this section is not intended to identify key or important features of the embodiments of this disclosure, nor is it intended to limit the scope of this disclosure. Other features of this disclosure will become readily apparent from the following description.

[0014] Other features and aspects of this disclosure will become clear from the following detailed description of exemplary embodiments with reference to the accompanying drawings. Attached Figure Description

[0015] Other features, objects, and advantages of this disclosure will become more apparent from the following detailed description of non-limiting embodiments with reference to the accompanying drawings: Figures 1 to 2 These are flowcharts of a key-value cache management method provided according to exemplary embodiments of this disclosure; Figures 3 to 4 These are schematic diagrams illustrating the process of a key-value cache management method provided according to exemplary embodiments of this disclosure; Figures 5 to 6 These are flowcharts of a key-value cache management method provided according to exemplary embodiments of this disclosure; Figure 7 This is a schematic diagram illustrating the determination of the hash value of a first word block according to an exemplary embodiment of this disclosure; Figure 8 This is a block diagram of a key-value cache management apparatus provided according to an exemplary embodiment of the present disclosure; Figure 9This is a schematic block diagram of an electronic device provided according to an exemplary embodiment of the present disclosure. Detailed Implementation

[0016] The various exemplary embodiments, features, and aspects of this disclosure will now be described in detail with reference to the accompanying drawings. The same reference numerals in the drawings denote elements that have the same or similar functions. Although various aspects of the embodiments are shown in the drawings, they are not necessarily drawn to scale unless specifically indicated otherwise.

[0017] The term “exemplary” as used herein means “serving as an example, implementation method, or illustration.” Any implementation method described herein as “exemplary” is not necessarily to be construed as superior to or better than other implementation methods.

[0018] Furthermore, to better illustrate this disclosure, numerous specific details are set forth in the following detailed description. Those skilled in the art will understand that this disclosure can be practiced without certain specific details. In some instances, apparatuses, means, elements, and circuits well known to those skilled in the art have not been described in detail in order to highlight the main points of this disclosure.

[0019] Some embodiments of this disclosure provide a key-value cache management method. Figure 1 This is a flowchart of a key-value cache management method 1000 provided according to an exemplary embodiment of the present disclosure.

[0020] like Figure 1 As shown, the key-value cache management method 1000 includes: Step S1: Divide the multiple lexical units included in the first input sequence of the large language model into multiple lexical blocks, wherein the first lexical block in the multiple lexical blocks includes lexical units that meet a predetermined number of conditions.

[0021] Step S2: Based on the hash value of the first word block and the hash mapping table, determine the key-value cache block that matches the first word block, wherein the hash mapping table is used to store the correspondence between the hash value and the identifier of the allocated key-value cache block.

[0022] The input sequence of a large language model can be segmented into multiple tokens by a tokenizer. Each token has a corresponding identifier, with tokens containing the same content sharing the same identifier. When processing the input sequence and generating the output, attention calculation is required for each token in the input sequence. This calculation involves repeatedly reading the key and value vectors of the tokens. To avoid redundant calculations, the inference system of a large language model can employ a key-value cache technique, storing the calculated key and value vectors in the CPU's (Central Processing Unit) memory or the GPU's (Graphics Processing Unit) video memory. These vectors can then be reused in subsequent processing, significantly reducing computational overhead.

[0023] In some technical solutions, KV cache management treats the entire input sequence as a unit, storing the key and value vectors of all words in the sequence contiguously in CPU memory or GPU video memory. However, in multi-request concurrency or batch processing scenarios, the input sequences of different requests often share the same prefix, which can be understood as multiple words at the beginning of the input sequence. Different requests may have completely identical prefixes, such as system prompts, dialogue templates, or common knowledge fragments; therefore, the key and value vectors corresponding to these identical prefixes are also the same. KV cache management based on the input sequence as a unit stores the KV data with the same prefix independently for each request, leading to wasted storage resources and computational redundancy.

[0024] Other technical solutions attempt to divide the input sequence into fixed-size word blocks and manage the KV cache on a block-by-block basis. Specifically, the input sequence can be divided into multiple word blocks according to a fixed number of words, with each word block corresponding to a key-value cache block used to store the key vectors and value vectors of all words within that block. In this way, the allocation and storage of the KV cache no longer require contiguous storage space, but can be distributed across multiple non-contiguous storage units, thereby reducing storage fragmentation and improving the utilization of storage resources. However, in the above technical solutions, even if different requests have word blocks with completely identical content (e.g., word blocks corresponding to the same prefix), each request will still be allocated a separate key-value cache block. In other words, each request stores its own copy of the KV data with the same prefix. While this solution optimizes storage management by dividing word blocks, there is redundancy in storage resources for word blocks with the same content, leaving room for further optimization.

[0025] To at least address this problem, in embodiments of this disclosure, the first input sequence of the large language model is divided into multiple word blocks, wherein the first word block among the multiple word blocks includes word blocks that meet a predetermined number of conditions; based on the hash value of the first word block and a hash mapping table, a key-value cache block matching the first word block is determined, wherein the hash mapping table is used to store the correspondence between the hash value and the identifier of the allocated key-value cache block.

[0026] The first word block includes word elements that meet a predetermined quantity condition, which can be understood as a pre-set criterion used to determine whether the word block is in a full state. Specifically, this predetermined quantity condition manifests as whether the number of word elements contained in the word block reaches a certain predetermined standard, and the specific value of this standard (e.g., 256) can be flexibly determined by those skilled in the art based on hardware configuration or application scenario. In other words, this predetermined quantity condition defines the transition boundary of the word block from "not full" to "full," and is not limited to a fixed specific value, but rather indicates the boundary state of the number of word elements contained in the word block relative to the predetermined condition.

[0027] Based on this, the first word block can be understood as a complete word block, which is in a "filled" state. In addition, multiple word blocks may also include a second word block. If the number of words in the second word block does not reach the predetermined number, it can be understood as an incomplete word block, which is in a "not filled" state, and words can be added later.

[0028] By refining the input sequence into multiple token blocks, the management granularity of the key-value cache is reduced from the entire input sequence to token blocks, thus providing a foundation for prefix sharing across requests. Even if the input sequences of different requests only share some prefixes, multiple token blocks formed based on the same prefixes can share the same key-value cache block. Compared to storage strategies based on the entire input sequence, the token block granularity significantly improves the utilization of storage resources and sharing flexibility.

[0029] A key-value cache block is a contiguous storage unit used to store the key vectors and value vectors of all words within a word block. Each key-value cache block corresponds to a word block, storing the key vector and value vector of each word in that word block for direct reuse during subsequent decoding.

[0030] A hash value is a fixed-length numerical value calculated by a hash function on an input of arbitrary length. It can be used to quickly compare whether data might be the same. For example, a hash value can be a fixed-length numerical value calculated by a hash function on a sequence of words within a word block. It is used to quickly compare whether word sequences within different word blocks might be the same, where a word sequence can be understood as the sequence formed by the sequential arrangement of all words within a word block.

[0031] A hash map is a data structure that uses hash values ​​as keys and storage locations or data identifiers as values, enabling fast lookups. In embodiments of this disclosure, a hash map is used to store the correspondence between hash values ​​and identifiers of allocated key-value cache blocks. It may include multiple pairs of mappings, where each pair may include a hash value and the identifier of the allocated key-value cache block corresponding to that hash value.

[0032] Therefore, in the KV cache management scheme disclosed herein, the identifier of a possible corresponding candidate key-value cache block can be directly located in the hash map table based on the hash value of the first word block. The identifier of the candidate key-value cache block is selected from the identifiers of the allocated key-value cache blocks in the hash map table. This method eliminates the need to compare the word sequence of the first word block with the word sequences associated with each of the allocated key-value cache blocks one by one, thereby significantly reducing the number of content comparisons during the matching process and reducing the computational load required for matching.

[0033] Furthermore, since the time consumed by the lookup operation of the hash map is independent of the total number of allocated key-value cache blocks, even if the number of allocated key-value cache blocks increases, the time consumed by the matching operation will not increase significantly. This improves the efficiency of determining the key-value cache block that matches the first word block while ensuring the feasibility of matching, thereby improving the throughput of the large language model inference system and reducing the TTFT of the large language model.

[0034] Figure 2 This is a flowchart of a key-value cache management method 1000 provided according to an exemplary embodiment of the present disclosure. Figure 3 This is a schematic diagram of the process of a key-value cache management method 1000 provided according to an exemplary embodiment of the present disclosure.

[0035] like Figure 2 As shown, the key-value cache management method 1000 may further include: Step S2-1: Query the hash mapping table to obtain the identifier of the candidate key-value cache block corresponding to the hash value of the first word block; and in response to the fact that the word sequence associated with the candidate key-value cache block is the same as the word sequence of the first word block, determine that the first word block reuses the candidate key-value cache block.

[0036] For example, combined Figure 2 and Figure 3The first input sequence 1, the second input sequence 2, and the third input sequence 3 of the large language model can be divided into multiple word blocks after word segmentation. Taking the first input sequence 1 as an example, it can include a first word block with a hash value of 0xABC123. In addition, the second input sequence 2 can also include a first word block with a hash value of 0xABC123; the third input sequence 3 can also include a first word block with a hash value of 0xABC123.

[0037] The key-value cache block pool may include multiple key-value cache blocks, which may include free key-value cache blocks and key-value cache blocks that have already stored the key vectors and value vectors of the terms. The key-value cache blocks that have already stored the key vectors and value vectors of the terms are the allocated key-value cache blocks. For example, Figure 3 Block 0 shown is an idle key-value cache block; Block 10 and Block 20 are allocated key-value cache blocks.

[0038] A hash table may include multiple pairs of mappings, where each mapping may include a hash value and an identifier of the allocated key-value cache block corresponding to that hash value. Figure 3 The hash map shown includes the mapping between the hash value 0xABC123 and the identifier of the key-value cache block Block10.

[0039] Therefore, taking the first input sequence 1 as an example, after determining that the hash value of a first word block it includes is 0xABC123, the hash mapping table can be queried to obtain the identifier Block10 of the candidate key-value cache block corresponding to the hash value 0xABC123 of the first word block.

[0040] After retrieving the identifier Block10 of the candidate key-value cache block, the associated lexical sequence is obtained and compared with the lexical sequence of the first lexical block. If the lexical sequence associated with the candidate key-value cache block Block10 is the same as that of the first lexical block, it indicates that the key vectors and value vectors stored in the candidate key-value cache block Block10 are consistent with the key vectors and value vectors required by the first lexical block. In this case, it is determined that the first lexical block can reuse the candidate key-value cache block Block10.

[0041] By first quickly locating candidate key-value cache blocks using hash values ​​and then performing precise comparisons using word sequences, the system leverages the efficient lookup capabilities of hash maps while avoiding erroneous matches that might result from hash collisions, thus improving matching efficiency while ensuring correctness. After a successful match, the first word block does not need to recalculate its key and value vectors; it directly reuses the already stored key-value cache block, reducing the computational resources consumed by redundant calculations. Furthermore, first word blocks with identical content in different input sequences (e.g., first input sequence 1, second input sequence 2, and third input sequence 3) can share the same key-value cache block, reducing redundant storage resource usage and contributing to improved overall throughput of large language model inference systems.

[0042] Furthermore, the number of identifiers of candidate key-value cache blocks corresponding to the hash value of the first word block obtained by querying the hash mapping table based on the hash value of the first word block may not be unique. In this case, the word sequence associated with each candidate key-value cache block can be compared with the word sequence of the first word block to determine whether there is a candidate key-value cache block that can be reused by the first word block.

[0043] Figure 4 This is a schematic diagram of the process of a key-value cache management method 1000 provided according to an exemplary embodiment of the present disclosure.

[0044] like Figure 4 As shown, in some embodiments of this disclosure, the key-value cache management method 1000 may further include: querying a hash map table; if the hash value of the first word block is not found, allocating an idle key-value cache block for the first word block, and updating the hash map table based on the hash value of the first word block and the identifier of the key-value cache block allocated to the first word block; or querying a hash map table; if the identifier of a candidate key-value cache block corresponding to the hash value of the first word block is obtained, and the word sequence associated with the candidate key-value cache block is different from the word sequence of the first word block, allocating an idle key-value cache block for the first word block, and updating the hash map table based on the hash value of the first word block and the identifier of the key-value cache block allocated to the first word block.

[0045] Specifically, in one implementation, step 101 can be executed to begin allocating key-value cache blocks, followed by step 102 to input a first input sequence and a list of token identifiers including identifiers of multiple tokens, so that the hash value of the first token block and the token sequence of the first token block can be calculated using the token identifier list in subsequent steps. In step 103, the multiple tokens included in the first input sequence are divided into multiple token blocks, and each token block of the first input sequence is traversed, wherein the multiple token blocks are numbered. i You can choose from the range 0 to N-1, where N is a positive integer greater than 1.

[0046] In step 104, the identifiers of the tokens included in each token block can be obtained, so that in step 105, it can be determined whether the token block is complete based on the identifiers of the tokens in each token block. If the token block is complete, or if the number of tokens in the token block reaches a predetermined condition, step 106 can be executed to determine the hash value of the token block, which will be defined as the first token block for ease of explanation later. If the token block is not complete, or if the number of tokens in the token block does not reach the predetermined condition, step 107 can be executed to set its hash value to -1 and allocate an idle key-value cache block for it, which will be defined as the second token block for ease of explanation later.

[0047] After determining the hash value of the first word block, step 108 can be executed to query the hash mapping table based on the hash value of the first word block; and step 109 can be executed to determine whether there is a hash value for the first word block.

[0048] If the hash map table contains the hash value of the first word block, it indicates the existence of an allocated key-value cache block. The key and value vectors stored in this allocated key-value cache block may match the key and value vectors required by the first word block. At least one such allocated key-value cache block is identified as a candidate key-value cache block. In subsequent step 110, it is confirmed whether the word sequence associated with the candidate key-value cache block is the same as the word sequence of the first word block. If they are the same, it means that the key and value vectors stored in the candidate key-value cache block are consistent with the key and value vectors required by the first word block, and the first word block can reuse the candidate key-value cache block. After a successful match, the first word block does not need to recalculate the key and value vectors; it can directly reuse the already stored key-value cache block, reducing the computational resources consumed by repeated calculations.

[0049] If there is no hash value for the first word block in the hash map table (i.e., no hash value), it means that there is currently no allocated key-value cache block that can be shared with the first word block. Therefore, step 115 can be executed to allocate a free key-value cache block for the first word block and update the hash map table based on the hash value of the first word block and the identifier of the key-value cache block allocated to the first word block.

[0050] Furthermore, in step 110, it is confirmed whether the lexical sequence associated with the candidate key-value cache block is the same as the lexical sequence of the first lexical block. If they are different, it indicates that the key vectors and value vectors stored in the candidate key-value cache block are inconsistent with the key vectors and value vectors required by the first lexical block, and the first lexical block cannot reuse the candidate key-value cache block. Therefore, step 115 can be executed to allocate an idle key-value cache block for the first lexical block, and the hash mapping table is updated based on the hash value of the first lexical block and the identifier of the key-value cache block allocated to the first lexical block.

[0051] When querying the hash map table, if no candidate key-value cache block identifier corresponding to the hash value of the first word block is found, it means that no allocated key-value cache block with the same content as the first word block is currently stored. In this case, an idle key-value cache block can be allocated to the first word block, and the hash map table can be updated with a new mapping relationship based on the hash value of the first word block and the identifier of the allocated key-value cache block. Furthermore, if one or more candidate key-value cache block identifiers corresponding to the hash value of the first word block are found, but after comparing the word sequences, it is found that the word sequences associated with all candidate key-value cache blocks are different from the word sequence of the first word block (i.e., a hash collision occurs), in this case, an idle key-value cache block can also be allocated to the first word block, and the hash map table can be updated with a new mapping relationship based on the hash value of the first word block and the identifier of the allocated key-value cache block.

[0052] Through the above processing, whether the hash value never appears or appears but the content does not match, new storage space can be allocated for the first key-value cache block, and the correspondence between its hash value and the identifier of the allocated key-value cache block can be recorded. This ensures that the first key-value cache block can be correctly hit when it is reused by other requests in the future. This mechanism guarantees the integrity of key-value cache management and avoids processing failures due to mismatches. In addition, by including the identifier of the newly allocated key-value cache block in the hash mapping table, the reusable cache resource pool is continuously expanded, which helps to improve the hit rate of subsequent requests, further reduce redundant calculations, and reduce the overall consumption of storage resources.

[0053] Figure 5 This is a flowchart of a key-value cache management method 1000 provided according to an exemplary embodiment of the present disclosure.

[0054] like Figure 5 As shown, the key-value cache management method 1000 may further include: Step S3: Allocate an idle key-value cache block for the second word block among multiple word blocks, wherein the number of words in the second word block has not reached a predetermined number condition.

[0055] As described above, the first input sequence of the large language model is divided into multiple word blocks, which may include a first word block and a second word block. The first word block contains a predetermined number of words and can be understood as a complete word block, in a "filled" state. The second word block does not meet the predetermined number of words and can be understood as an incomplete word block, in an "unfilled" state, where words can be added later.

[0056] For a first token block that is in a "filled" state, the matching key-value cache block can be determined based on its hash value and hash mapping table. This helps to quickly locate allocated key-value cache blocks with the same content. When a match is successful, the first token block can directly reuse this key-value cache block, thereby reducing the computational resources consumed by repeatedly calculating the key and value vectors of token blocks with the same content. Furthermore, this mechanism allows first token blocks with the same content in different requests to share the same key-value cache block, reducing storage resource consumption.

[0057] For incomplete word blocks (i.e., the second word block) where the number of words has not reached the predetermined number, they are in an "unfilled" state and can be further supplemented with words. Therefore, hashing is not performed on the second word block where the number of words has not reached the predetermined number; instead, an idle key-value cache block can be allocated to it directly. In streaming input or dynamically appended word scenarios, the last word block of the input sequence is usually in an unfilled state, and its content changes frequently as new words are added. If hashing is performed on this second word block and attempts are made to share it, the hash value needs to be recalculated every time a word is appended, resulting in significant computational overhead.

[0058] Therefore, the key-value caching management scheme provided by the embodiments of this disclosure, while ensuring fine-grained sharing at the lexical block level and reducing storage resource consumption, adopts a strategy of direct allocation and temporary non-sharing for incomplete lexical blocks, delaying hash calculation and sharing detection until the lexical block is filled, thereby reducing unnecessary computational overhead, helping to improve the overall efficiency in the dynamic decoding stage, reducing computational load and storage fragmentation while maintaining sharing efficiency, improving the throughput of the large language model inference system, and reducing the first lexical latency of the large language model.

[0059] Refer again Figure 5 The key-value cache management method 1000 may also include: Step S4: Add at least one lexical element from the second input sequence of the large language model to the second lexical block; in response to the number of lexical elements in the second lexical block after adding at least one lexical element reaching a predetermined number condition, determine the hash value of the second lexical block; based on the hash value of the second lexical block and the hash mapping table, obtain the identifier of the candidate key-value cache block corresponding to the hash value of the second lexical block from the identifiers of the allocated key-value cache blocks; in response to the lexical sequence associated with the candidate key-value cache block being the same as the lexical sequence of the second lexical block, determine that the candidate key-value cache block is reused in the second lexical block, and release the free key-value cache block allocated to the second lexical block.

[0060] Specifically, for ease of description, the currently processed input sequence can be referred to as the first input sequence, and the next input sequence processed after the first input sequence can be referred to as the second input sequence. It is understood that the naming of the first and second input sequences is only used to distinguish different input sequences within the same large language model and does not constitute any restriction on the order or number of input sequences.

[0061] The first input sequence is divided into multiple word blocks, including a first word block with a predetermined number of words and a second word block with a less than predetermined number of words (i.e., an incomplete word block). The second input sequence includes at least one word, which can be added to the second word block of the first input sequence. After adding at least one word to the second word block of the first input sequence, the number of words in the second word block reaches the predetermined number, or in other words, the second word block changes from an unfilled state to a filled state. After the second word block becomes filled, the hash value of the second word block can be determined by referring to the processing method of the first word block above. The hash mapping table is queried to obtain the identifier of the candidate key-value cache block corresponding to the hash value, and the word sequence associated with the candidate key-value cache block is compared with the word sequence of the second word block. If the word sequence associated with the candidate key-value cache block is the same as the word sequence of the second word block, it is determined that the filled second word block can reuse the candidate key-value cache block.

[0062] Furthermore, in some embodiments, the key-value cache management method 1000 may further include: adding at least one lexical element included in the second input sequence of the large language model to the second lexical block; determining the hash value of the second lexical block in response to the number of lexical elements in the second lexical block after adding at least one lexical element reaching a predetermined number condition; querying the hash map table, and updating the hash map table based on the hash value of the second lexical block and the identifier of the free key-value cache block allocated to the second lexical block if the hash value of the second lexical block is not found; or querying the hash map table, and updating the hash map table based on the hash value of the second lexical block and the identifier of the free key-value cache block allocated to the second lexical block if the identifier of the candidate key-value cache block corresponding to the hash value of the second lexical block is obtained, and the lexical sequence associated with the candidate key-value cache block is different from the lexical sequence of the second lexical block.

[0063] In other words, through the above method, the initially incomplete second word block in the first input sequence can be filled with the help of the words in the second input sequence, thus transforming it into a complete word block and participating in the query and sharing of the hash map table. This mechanism enables word blocks from different input sequences to be filled collaboratively, reducing the waste caused by incomplete blocks at the end that cannot be shared, improving the word block filling rate and sharing reuse rate, further reducing storage redundancy, and helping to improve the overall efficiency of large language model inference systems.

[0064] Furthermore, in some embodiments of this disclosure, step S1, which divides the multiple lexical units included in the first input sequence of the large language model into multiple lexical blocks, may include: in response to the existence of a stored third lexical block, wherein the number of lexical units in the third lexical block has not reached a predetermined number condition, adding at least one lexical unit from the first input sequence to the third lexical block until the number of lexical units in the third lexical block reaches the predetermined number condition or all lexical units in the first input sequence have been added; and in response to the presence of lexical units that have not yet been added in the first input sequence, dividing the unequal lexical units into multiple lexical blocks.

[0065] In embodiments of this disclosure, the inference system of the large language model may have processed one or more other input sequences before the currently processing first input sequence. When these previously processed input sequences are segmented into lexical blocks, incomplete lexical blocks (i.e., unfilled lexical blocks) may be generated, where the number of lexical units does not meet a predetermined condition. For ease of description, these are referred to as third lexical blocks in the following text. The third lexical block is similar to the second lexical block in the first input sequence, both being in an unfilled state, but the difference is that the third lexical block already existed and was stored in the system before processing the current first input sequence.

[0066] To fully utilize allocated storage resources, the inference system of the large language model checks for the existence of a stored third word block before dividing the multiple words in the first input sequence into multiple word blocks. If a third word block exists, at least one word can be taken from the first input sequence and added to it. This addition process can continue until the number of words in the third word block reaches a predetermined condition (i.e., it becomes full), or until all words in the first input sequence have been added. Through this filling operation, the previously incomplete third word block has the opportunity to become a complete word block, thus participating in subsequent hash map lookups and shared reuse. After the third word block becomes a complete word block, the hash value of the third word block can be determined by referring to the processing method of the first word block above. The hash map is then queried to obtain the identifier of the candidate key-value cache block corresponding to the hash value, and the word sequence associated with the candidate key-value cache block is compared with the word sequence of the third word block. If the lexical sequence associated with the candidate key-value cache block is the same as the lexical sequence of the third lexical block, then the third lexical block in its filled state can reuse the candidate key-value cache block. The specific processing details are the same as those for the first lexical block, and will not be repeated here.

[0067] If there are still remaining words in the first input sequence after filling all the third word blocks, the words that have not been added in the first input sequence are divided into multiple word blocks. The complete word blocks (i.e., the first word blocks) can be processed in the same way as the first word blocks described above; the incomplete word blocks (i.e., the second word blocks) are stored in the system as new unfilled word blocks for use in subsequent input sequences.

[0068] By employing the above method, existing incomplete lexical blocks can be prioritized to carry the lexical units of the current input sequence, reducing storage fragmentation caused by frequent allocation of new lexical blocks and improving the utilization rate of storage resources. Furthermore, as more incomplete lexical blocks are filled into complete lexical blocks, the number of key-value cache blocks that can be shared across requests increases, further reducing redundant computation and storage, and contributing to improved overall efficiency of large language model inference systems.

[0069] In addition, in an implementation where at least one lexical element from the second input sequence of the large language model is added to the second lexical block and the number of lexical elements in the second lexical block after adding at least one lexical element reaches a predetermined number, the hash map is queried. If the hash value of the second lexical block is not found, the hash map can be updated based on the hash value of the second lexical block and the identifier of the free key-value cache block allocated to the second lexical block.

[0070] Furthermore, in this implementation, if the hash map is queried and the identifier of the candidate key-value cache block corresponding to the hash value of the second word block is obtained, and the word sequence associated with the candidate key-value cache block is different from the word sequence of the second word block, the hash map can be updated based on the hash value of the second word block and the identifier of the free key-value cache block allocated to the second word block.

[0071] Specifically, when querying the hash map table, no candidate key-value cache block identifier corresponding to the hash value of the second token block was found. This indicates that a key-value cache block with the same content as the second token block has not yet been stored. In this case, an idle key-value cache block can be allocated for the second token block, and the hash map table can be updated with a new mapping relationship based on the hash value of the second token block and the identifier of the allocated key-value cache block. Through this update operation, if subsequent requests encounter token blocks with the same hash value, they can locate the newly stored key-value cache block through the hash map table, thus having the opportunity to achieve reuse.

[0072] When querying the hash map table, one or more candidate key-value cache block identifiers corresponding to the hash value of the second word block are found. However, after comparing the word sequences, the word sequences associated with all candidate key-value cache blocks are different from the word sequence of the second word block (i.e., a hash collision occurs). In this case, a free key-value cache block can still be allocated to the second word block, and the hash value of the second word block and the identifier of the newly allocated key-value cache block can be added to the hash map table to update the hash map table. This approach ensures that even if a hash collision occurs, the original mapping relationship will not be overwritten, thus guaranteeing that all stored key-value cache blocks can be correctly found.

[0073] Through the two update mechanisms described above, whether the hash value has never appeared or appears but the content does not match, new storage space can be allocated for the filled second key-value block, and its mapping relationship can be recorded in the hash mapping table. This ensures that the second key-value block can be correctly hit when processed by other requests in the future, thereby expanding the reusable key-value cache block resource pool, helping to improve the hit rate of subsequent requests, reduce redundant calculations, and further reduce the consumption of storage resources. In addition, the method for handling hash collisions preserves the original mapping relationship, avoiding cache block loss due to overwriting, which helps to improve the correctness and reliability of the key-value cache management method.

[0074] Refer again Figure 4 In some embodiments of this disclosure, the key-value cache management method 1000 may further include: recording the number of lexical blocks that use the key-value cache block to obtain the reference count corresponding to the key-value cache block; wherein, in response to the reuse of the candidate key-value cache block by the first lexical block or the second lexical block, the reference count corresponding to the candidate key-value cache block is increased; and in response to the release of the reuse of the candidate key-value cache block by the first lexical block or the second lexical block, the reference count corresponding to the candidate key-value cache block is decreased.

[0075] In this embodiment, the key-value cache management method 1000 may introduce a reference counting mechanism to record the number of token blocks currently using the same key-value cache block. Each key-value cache block can be associated with a reference count, the initial value of which can be set to 0. When a token block (e.g., the first token block or the second token block) reuses the key-value cache block, the reference count of the key-value cache block can be increased by 1; when the token block no longer uses the key-value cache block, the reference count of the key-value cache block can be decreased by 1. Only when the reference count decreases to 0 can the key-value cache block be released, that is, returned to the idle key-value cache block pool for subsequent allocation.

[0076] Taking three different input sequences as an example, after word segmentation, the first input sequence 1, the second input sequence 2, and the third input sequence 3 can yield multiple word blocks. The first input sequence 1 may include a first word block, the word sequence of which is the same as the word sequence of a first word block in the second input sequence 2 and the word sequence of a first word block in the third input sequence 3. For ease of description, the first word block with the same word sequence in the first input sequence 1, the second input sequence 2, and the third input sequence 3 will be referred to as the first word block with the same content.

[0077] When a first word block with the same content as the first input sequence 1 reuses a key-value cache block, the reference count of that key-value cache block can change from 0 to 1. When a first word block with the same content as the second input sequence 2 reuses the same key-value cache block, the reference count of that key-value cache block can change from 1 to 2. When a first word block with the same content as the third input sequence 3 reuses the same key-value cache block, the reference count of that key-value cache block can change from 2 to 3.

[0078] As the first input sequence 1, the second input sequence 2, and the third input sequence 3 are processed, the first token block with the same content in each input sequence will no longer use the key-value cache block, and the reference count of the key-value cache block will decrease by 1 in turn. The key-value cache block can only be released when its reference count becomes 0. This reference counting mechanism ensures that the key-value cache block is released only after all users have finished processing, preventing other currently used token blocks from accessing invalid data due to premature release, thus improving the security and reliability of key-value cache management. Based on the reference counting mechanism, as an option, the key-value cache block can be released in response to its reference count becoming zero. Alternatively, the most recent usage time of the key-value cache block can be recorded; and in response to the number of idle key-value cache blocks falling below a threshold, the key-value cache block with the earliest most recent usage time can be selected from multiple key-value cache blocks with a reference count of zero for release.

[0079] Specifically, some embodiments of this disclosure can further optimize the eviction strategy for key-value cache blocks by incorporating the most recent usage time. Specifically, when a candidate key-value cache block is reused in response to a first or second lexical block, and the reference count of the candidate key-value cache block is increased, the most recent usage time of the candidate key-value cache block can be updated to the current time. When the number of idle key-value cache blocks falls below a threshold, the key-value cache block with the earliest most recent usage time can be selected from multiple key-value cache blocks with a reference count of zero for release. Whenever a lexical block reuses a key-value cache block, in addition to increasing the reference count of that key-value cache block, the most recent usage time of that key-value cache block can also be updated to the current time.

[0080] When the number of idle key-value cache blocks in the key-value cache block pool falls below a preset threshold, it indicates a shortage of available key-value cache block resources, necessitating the release of some allocated key-value cache blocks to free up space. Idle key-value cache blocks can be understood as those not currently used by any token blocks and not yet allocated, with a reference count of zero. The preset threshold can be pre-set based on parameters such as the total key-value cache block capacity and current load, for example, set to 10% of the total number of key-value cache blocks or a fixed value. When the number of idle key-value cache blocks falls below this threshold, an eviction operation is triggered: one or more key-value cache blocks with the earliest recent usage time are selected from all key-value cache blocks with a reference count of zero for release. The earliest recent usage time means that the key-value cache block has not been reused for a long time and is considered "cold" data. Prioritizing the release of these "cold" data preserves recently reused "hot" data, thereby improving cache hit rate within limited storage resources and reducing the computational overhead caused by frequent recalculation of key and value vectors.

[0081] The reference counting mechanism and the eviction policy based on recent usage time described above are independent yet complementary. The reference counting mechanism ensures that cache blocks currently in use are not mistakenly released, while the eviction policy proactively cleans up less frequently used, colder data when cache resources are scarce, providing storage space for new lexical blocks. The combination of the two helps improve the utilization efficiency of cache resources and reduce redundant computations while ensuring data security, thereby increasing the overall throughput of large language model inference systems and reducing response latency.

[0082] Figure 6 This is a flowchart of a key-value cache management method 1000 provided according to an exemplary embodiment of the present disclosure.

[0083] like Figure 6 As shown, the key-value cache management method 1000 may further include: Step S2-2: In response to the determination by the Bloom filter that the hash value of the first word block exists in the hash map table, the step of determining the key-value cache block that matches the first word block based on the hash value of the first word block and the hash map table is executed.

[0084] In addition, in response to the Bloom filter determining that the hash value of the first word block does not exist in the hash map, an idle key-value cache block is allocated for the first word block, and the hash map is updated based on the hash value of the first word block and the identifier of the key-value cache block allocated to the first word block.

[0085] A Bloom filter is a probabilistic data structure used to determine whether an element might exist or definitely does not exist in a set. Specifically, in the embodiments of this disclosure, the element can be a hash value, and the set can be a hash map. Whenever a hash value is successfully added to the hash map, it is added to the Bloom filter. Thus, the Bloom filter records information about all hash values ​​that "definitely exist in the hash map." When executing a query, for the hash value being queried, the Bloom filter may return two results: a "definitely does not exist" result, indicating that the hash value is definitely not in the hash map; or a "possibly exists" result, indicating that the hash value might exist in the hash map, but there is a certain probability of a false positive, or a probability that it actually does not exist but is mistakenly judged as possibly existing.

[0086] Therefore, in response to the Bloom filter determining that the hash value of the first word block exists in the hash map table, or in other words, the Bloom filter returns a result of "possibly exists", the step of determining the key-value cache block matching the first word block based on the hash value of the first word block and the hash map table can be performed. In response to the Bloom filter determining that the hash value of the first word block does not exist in the hash map table, or in other words, the Bloom filter returns a result of "definitely does not exist", a free key-value cache block can be allocated for the first word block, and the hash map table can be updated based on the hash value of the first word block and the identifier of the key-value cache block allocated to the first word block.

[0087] In the embodiments of this disclosure, the hash value of the first token block can be input into a Bloom filter for judgment before querying the hash map table. If the Bloom filter determines that the hash value does not exist, it can be directly determined that there is no key-value cache block identifier corresponding to the hash value in the hash map table, thus eliminating the need to perform subsequent hash map table query operations and directly proceeding to the process of allocating a free key-value cache block for the first token block. If the Bloom filter determines that the hash value exists, the step of determining a matching key-value cache block based on the hash value of the first token block and the hash map table can continue.

[0088] By introducing a Bloom filter as a pre-filter, a large number of non-existent query requests can be quickly eliminated, reducing the number of invalid accesses to the hash map table, thereby reducing query latency and improving overall processing efficiency. It should be noted that the Bloom filter is only an exemplary pre-filter implementation, and the embodiments disclosed herein are not limited to using a Bloom filter. Any probabilistic data structure or deterministic filtering mechanism capable of quickly determining whether an element may exist can be applied to the technical solutions of this disclosure.

[0089] Figure 7This is a schematic diagram illustrating the determination of the hash value of a first word block according to an exemplary embodiment of this disclosure.

[0090] like Figure 7 As shown, in some embodiments of this disclosure, the key-value cache management method 1000 may further include determining the hash value of the first lexical block.

[0091] Optionally, the key-value cache management method 1000 may include: in response to a first word block having an adjacent preceding word block, obtaining the hash value of the preceding word block as a prefix hash, and determining the hash value of the first word block based on the prefix hash and the word sequence of the first word block. Optionally, the key-value cache management method 1000 may further include: in response to a first word block not having an adjacent preceding word block, determining the hash value of the first word block based on the word sequence of the first word block. In other words, in the above two embodiments, the hash value of the first word block can be calculated in different ways depending on whether the first word block has an adjacent preceding word block.

[0092] Specifically, when the first word block has an adjacent preceding word block, the hash value already calculated for that preceding word block is used as a prefix hash. This prefix hash is then concatenated with the word sequence of the first word block. The concatenated data is then input into a hash function to calculate the hash value of the first word block. Here, the prefix hash refers to the hash value of the preceding word block. Incorporating information from previous word blocks when calculating the hash value of the current word block distinguishes it from cases where the input sequence contains the same word sequence but with different contexts. When the first word block does not have an adjacent preceding word block, the word sequence of the first word block can be directly input into the hash function to calculate its hash value.

[0093] In practical implementation, efficient non-cryptographic hash functions such as xxhash can be used for calculation. Taking xxhash as an example, an xxhash object is created; if a prefix hash exists, the prefix hash is converted to 8-byte little-endian format and updated to the hash object; then, the token identifiers in the token sequence are converted to byte arrays and updated to the hash object; finally, the hash object's interface is called to obtain the final hash value.

[0094] For example, the process of calculating a hash value based on prefix hashing is illustrated using three consecutive first word blocks. Each first word block may include a predetermined number of words, such as 256 words. The first word block among multiple first word blocks (referred to as first word block 0) does not have an adjacent preceding word block and may be the first word block in the first input sequence. In this case, the word sequence of first word block 0 is directly input into the xxhash hash function to calculate the hash value of first word block 0, denoted as hash_0.

[0095] The second word block in a plurality of first word blocks (referred to as first word block 1) is adjacent to first word block 0 and follows first word block 0. Therefore, first word block 1 has an adjacent preceding word block (i.e., first word block 0). When calculating the hash value of first word block 1, the hash value hash_0 of first word block 0 is first obtained as a prefix hash. This prefix hash is then concatenated with the word sequence of first word block 1. The concatenated data is then input into the xxhash hash function to calculate the hash value of first word block 1, denoted as hash_1.

[0096] The third word block in a plurality of first word blocks (referred to as first word block 2) is adjacent to first word block 1 and is located after first word block 1. Therefore, first word block 2 has an adjacent preceding word block (i.e., first word block 1). When calculating the hash value of first word block 2, the hash value hash_1 of first word block 1 is first obtained as a prefix hash. This prefix hash is then concatenated with the word sequence of first word block 2. The concatenated data is then input into the xxhash hash function to calculate the hash value of first word block 2, denoted as hash_2.

[0097] Through the above method, a transitive chain is formed between hash_0, hash_1, and hash_2: the generation of hash_0 depends only on the word sequence of the first word block 0 itself; the generation of hash_1 depends on hash_0 and the word sequence of the first word block 1 itself, thus implicitly containing the information of the first word block 0; the generation of hash_2 depends on hash_1 and the word sequence of the first word block 2 itself, thus implicitly containing the information of the first word block 0 and the first word block 1. Therefore, the hash value of each word block includes the information of all the word blocks preceding it, ensuring the context relevance of the hash value. Even if a word block in two different input sequences contains the exact same word sequence, if their prefix contexts are different (i.e., the content of the preceding word blocks is different), the calculated hash values ​​will be different, thus avoiding the erroneous sharing of the same key-value cache block in different contexts. This mechanism improves the uniqueness of hash values ​​and the correctness of shared matching.

[0098] In other words, through the aforementioned prefix hashing propagation method, the hash value of the current word block not only depends on its own word sequence but also implicitly contains information from all preceding word blocks. This mechanism guarantees the uniqueness and contextual relevance of hash values, improving the correctness and security of sharing. Furthermore, using efficient hash functions such as xxhash can generate hash values ​​with low collision rates while maintaining computational speed, reducing the overhead of additional content comparisons caused by hash collisions.

[0099] Furthermore, in some embodiments of this disclosure, an asynchronous pre-computation mechanism can be employed to further reduce the impact of hash value calculation on the main processing flow. Specifically, during the process of querying the hash map table based on the hash value of the m-th first word block among multiple first word blocks, the hash value of the (m+n)-th first word block among the multiple first word blocks is calculated, where m and n are both positive integers. In other words, the process in which the inference system can query the hash map table based on the hash value of one of the multiple first word blocks can be understood as the main flow processing the current first word block. Additionally, the inference system can also compute the hash value of another first word block among the multiple first word blocks in parallel in a background thread, such as the hash value of the next first word block to be processed or the hash value of a predicted first word block that may be accessed. This computation process is executed in parallel with the query and allocation operations of the main flow and does not block the main flow.

[0100] In this way, when the main process finishes processing the current first word block and moves on to the next first word block, the hash value of that next first word block may have already been pre-calculated. This allows it to be directly used to query the hash mapping table without waiting for the hash value calculation. Asynchronous pre-computation overlaps the hash value calculation time with other operations in the main process, effectively reducing the main process's waiting time and improving overall processing efficiency. It should be noted that the specific implementation of asynchronous pre-computation can be through starting a separate background thread, using coroutines, or a task queue, etc. This disclosure does not limit the specific parallel execution mechanism.

[0101] Furthermore, in some embodiments of this disclosure, a linked storage mechanism can be used to manage potential hash collisions in the hash map table. For example, in response to the existence of multiple key-value cache blocks that correspond to the same hash value but have different associated lexical sequences, the identifiers of the multiple key-value cache blocks can be stored in a linked list at the position corresponding to the hash value in the hash map table.

[0102] Specifically, when multiple key-value cache blocks correspond to the same hash value and the lexical sequences associated with these key-value cache blocks are different from each other, the identifiers of these multiple key-value cache blocks can be stored in the hash map table at the positions corresponding to the hash value in the form of a linked list. In the hash map table, the same hash value no longer corresponds to a single key-value cache block identifier, but rather a linked list, where each node stores a key-value cache block identifier.

[0103] During the query process, when it's necessary to search for a corresponding candidate key-value cache block in the hash map table based on the hash value of a query term block (e.g., the first term block), the corresponding linked list in the hash map table can be located based on the hash value. Then, each key-value cache block identifier in the linked list is traversed, and the term sequence associated with each key-value cache block is compared with the term sequence of the query term block. If a key-value cache block with the same term sequence is found, it is determined as a matching key-value cache block; if no matching key-value cache block is found after traversing the entire linked list, it is confirmed that no matching key-value cache block exists.

[0104] Through the aforementioned linked storage mechanism, when a hash collision occurs, the original key-value cache block identifier will not be overwritten by a newly added identifier, thus ensuring that all stored key-value cache blocks can be correctly searched. Furthermore, during a query, by traversing the linked list and performing a precise comparison of the term sequences, key-value cache blocks with truly identical term sequences can be accurately identified, reducing erroneous matches caused by hash collisions. This mechanism, while ensuring efficient lookup of the hash map table, improves the correctness and reliability of hash collision handling, contributing to the overall performance improvement of the key-value cache management method.

[0105] Therefore, according to at least one embodiment of this disclosure, multiple lexical units of the input sequence of a large language model are divided into multiple lexical blocks. In response to a predetermined number of lexical units in a lexical block, a key-value cache block matching that lexical block is determined based on the hash value and hash mapping table of the lexical block. In response to a predetermined number of lexical units in a lexical block, an idle key-value cache block is allocated to that lexical block. By refining the input sequence into multiple lexical blocks, the management granularity of the key-value cache is reduced from the entire input sequence to lexical blocks, thereby providing a basis for prefix sharing across requests. Even if the input sequences of different requests only have some identical prefixes, multiple lexical blocks formed based on the same prefix can share the same key-value cache block. Compared to a storage strategy based on the entire input sequence, the granularity of lexical blocks can significantly improve the utilization of storage resources and sharing flexibility.

[0106] Furthermore, the first word block contains a predetermined number of words and is in a "filled" state. For each first word block, a matching key-value cache block can be determined based on its hash value and hash mapping table. The second word block contains fewer words than the predetermined number and is in an "unfilled" state, allowing for the addition of more words. For second word blocks with fewer words than the predetermined number, no hash calculation is performed, and an empty key-value cache block can be directly allocated. In streaming input or dynamically appended word scenarios, the last word block of the input sequence is usually in an unfilled state, and its content changes frequently as new words are added. If a hash calculation is performed on this second word block and sharing is attempted, the hash value needs to be recalculated each time a word is appended, resulting in significant computational overhead.

[0107] Therefore, the key-value caching management scheme provided by the embodiments of this disclosure, while ensuring fine-grained sharing at the lexical block level and reducing storage resource consumption, adopts a strategy of direct allocation and temporary non-sharing for incomplete lexical blocks, delaying hash calculation and sharing detection until the lexical block is filled, thereby reducing unnecessary computational overhead, helping to improve the overall efficiency in the dynamic decoding stage, reducing computational load and storage fragmentation while maintaining sharing efficiency, improving the throughput of the large language model inference system, and reducing the first lexical latency of the large language model.

[0108] Figure 8 This is a block diagram of a key-value cache management device 2000 provided according to an exemplary embodiment of the present disclosure.

[0109] like Figure 8 As shown, some embodiments of this disclosure provide a key-value cache management device 2000. The key-value cache management device 2000 may include: a partitioning unit 100 and a matching unit 200, wherein the partitioning unit 100 is configured to partition a plurality of lexical units included in a first input sequence of a large language model into a plurality of lexical blocks, wherein a first lexical block among the plurality of lexical blocks includes lexical units that meet a predetermined number of conditions; the matching unit 200 is configured to determine a key-value cache block matching the first lexical block based on the hash value of the first lexical block and a hash mapping table, wherein the hash mapping table is used to store the correspondence between hash values ​​and identifiers of allocated key-value cache blocks.

[0110] The input sequence of a large language model can be segmented into multiple tokens by a tokenizer. Each token has a corresponding identifier, with tokens containing the same content sharing the same identifier. When processing the input sequence and generating the output, attention calculation is required for each token in the input sequence. This calculation involves repeatedly reading the key and value vectors of the tokens. To avoid redundant calculations, the inference system of a large language model can employ a key-value cache technique, storing the calculated key and value vectors in the CPU's memory or the GPU's video memory, which can then be reused in subsequent processing, thus significantly reducing computational overhead.

[0111] In embodiments of this disclosure, the first input sequence of the large language model is divided into multiple word blocks, wherein the first word block among the multiple word blocks includes word blocks that meet a predetermined number of conditions; based on the hash value of the first word block and a hash mapping table, a key-value cache block matching the first word block is determined, wherein the hash mapping table is used to store the correspondence between the hash value and the identifier of the allocated key-value cache block.

[0112] The first word block includes word elements that meet a predetermined quantity condition, which can be understood as a pre-set criterion used to determine whether the word block is in a full state. Specifically, this predetermined quantity condition manifests as whether the number of word elements contained in the word block reaches a certain predetermined standard, and the specific value of this standard (e.g., 256) can be flexibly determined by those skilled in the art based on hardware configuration or application scenario. In other words, this predetermined quantity condition defines the transition boundary of the word block from "not full" to "full," and is not limited to a fixed specific value, but rather indicates the boundary state of the number of word elements contained in the word block relative to the predetermined condition.

[0113] Based on this, the first word block can be understood as a complete word block, which is in a "filled" state. In addition, multiple word blocks may also include a second word block. If the number of words in the second word block does not reach the predetermined number, it can be understood as an incomplete word block, which is in a "not filled" state, and words can be added later.

[0114] By refining the input sequence into multiple token blocks, the management granularity of the key-value cache is reduced from the entire input sequence to token blocks, thus providing a foundation for prefix sharing across requests. Even if the input sequences of different requests only share some prefixes, multiple token blocks formed based on the same prefixes can share the same key-value cache block. Compared to storage strategies based on the entire input sequence, the token block granularity significantly improves the utilization of storage resources and sharing flexibility.

[0115] A key-value cache block is a contiguous storage unit used to store the key vectors and value vectors of all words within a word block. Each key-value cache block corresponds to a word block, storing the key vector and value vector of each word in that word block for direct reuse during subsequent decoding.

[0116] A hash value is a fixed-length numerical value calculated by a hash function on an input of arbitrary length. It can be used to quickly compare whether data might be the same. For example, a hash value can be a fixed-length numerical value calculated by a hash function on a sequence of words within a word block. It can be used to quickly compare whether word sequences within different word blocks might be the same, where a word sequence can be understood as a sequence formed by the sequential arrangement of all words within a word block.

[0117] A hash table is a data structure that uses hash values ​​as keys and storage locations or data identifiers as values, enabling fast lookups. In embodiments of this disclosure, a hash table is used to store the correspondence between hash values ​​and identifiers of allocated key-value cache blocks. It may include multiple pairs of mappings, where each pair may include a hash value and the identifier of the allocated key-value cache block corresponding to that hash value.

[0118] Therefore, in the KV cache management scheme disclosed herein, the identifier of a possible corresponding candidate key-value cache block can be directly located in the hash map table based on the hash value of the first word block. The identifier of the candidate key-value cache block is selected from the identifiers of the allocated key-value cache blocks in the hash map table. This method eliminates the need to compare the word sequence of the first word block with the word sequences associated with each of the allocated key-value cache blocks one by one, thereby significantly reducing the number of content comparisons during the matching process and reducing the computational load required for matching.

[0119] Furthermore, since the time consumed by the lookup operation of the hash map is independent of the total number of allocated key-value cache blocks, even if the number of allocated key-value cache blocks increases, the time consumed by the matching operation will not increase significantly. This improves the efficiency of determining the key-value cache block that matches the first word block while ensuring the feasibility of matching, thereby improving the throughput of the large language model inference system and reducing the TTFT of the large language model.

[0120] In some embodiments of this disclosure, the matching unit 200 may also be configured to query a hash map to obtain the identifier of a candidate key-value cache block corresponding to the hash value of the first lexical block; and in response to the lexical sequence associated with the candidate key-value cache block being the same as the lexical sequence of the first lexical block, to determine that the first lexical block reuses the candidate key-value cache block.

[0121] By first quickly locating candidate key-value cache blocks using hash values ​​and then performing precise comparisons using word sequences, the system leverages the efficient lookup capabilities of hash maps while avoiding erroneous matches that might result from hash collisions, thus improving matching efficiency while ensuring correctness. Upon successful matching, the first word block does not need to recalculate its key and value vectors; it directly reuses the already stored key-value cache block, reducing the computational resources consumed by redundant calculations. Furthermore, first word blocks with identical content in different input sequences can share the same key-value cache block, reducing redundant storage resource usage and contributing to improved overall throughput of large language model inference systems.

[0122] Furthermore, the number of identifiers of candidate key-value cache blocks corresponding to the hash value of the first word block obtained by querying the hash mapping table based on the hash value of the first word block may not be unique. In this case, the word sequence associated with each candidate key-value cache block can be compared with the word sequence of the first word block to determine whether there is a candidate key-value cache block that can be reused by the first word block.

[0123] In some embodiments of this disclosure, the matching unit 200 may also be configured to query a hash map, allocate an idle key-value cache block for the first word block if the hash value of the first word block is not found, and update the hash map based on the hash value of the first word block and the identifier of the key-value cache block allocated to the first word block; or query the hash map, and if the identifier of the candidate key-value cache block corresponding to the hash value of the first word block is obtained, and the word sequence associated with the candidate key-value cache block is different from the word sequence of the first word block, allocate an idle key-value cache block for the first word block, and update the hash map based on the hash value of the first word block and the identifier of the key-value cache block allocated to the first word block.

[0124] When querying the hash map table, if no candidate key-value cache block identifier corresponding to the hash value of the first word block is found, it means that no allocated key-value cache block with the same content as the first word block is currently stored. In this case, an idle key-value cache block can be allocated to the first word block, and the hash map table can be updated with a new mapping relationship based on the hash value of the first word block and the identifier of the allocated key-value cache block. Furthermore, if one or more candidate key-value cache block identifiers corresponding to the hash value of the first word block are found, but after comparing the word sequences, it is found that the word sequences associated with all candidate key-value cache blocks are different from the word sequence of the first word block (i.e., a hash collision occurs), in this case, an idle key-value cache block can also be allocated to the first word block, and the hash map table can be updated with a new mapping relationship based on the hash value of the first word block and the identifier of the allocated key-value cache block.

[0125] Through the above processing, whether the hash value never appears or appears but the content does not match, new storage space can be allocated for the first key-value cache block, and the correspondence between its hash value and the identifier of the allocated key-value cache block can be recorded. This ensures that the first key-value cache block can be correctly hit when it is reused by other requests in the future. This mechanism guarantees the integrity of key-value cache management and avoids processing failures due to mismatches. In addition, by including the identifier of the newly allocated key-value cache block in the hash mapping table, the reusable cache resource pool is continuously expanded, which helps to improve the hit rate of subsequent requests, further reduce redundant calculations, and reduce the overall consumption of storage resources.

[0126] In some embodiments of this disclosure, the key-value cache management device 2000 may further include an allocation unit (not shown) for allocating a free key-value cache block to a second lexical block among a plurality of lexical blocks, wherein the number of lexical units in the second lexical block does not reach a predetermined number condition.

[0127] As described above, the first input sequence of the large language model is divided into multiple word blocks, which may include a first word block and a second word block. The first word block contains a predetermined number of words and can be understood as a complete word block, in a "filled" state. The second word block does not meet the predetermined number of words and can be understood as an incomplete word block, in an "unfilled" state, where words can be added later.

[0128] For a first token block that is in a "filled" state, the matching key-value cache block can be determined based on its hash value and hash mapping table. This helps to quickly locate allocated key-value cache blocks with the same content. When a match is successful, the first token block can directly reuse this key-value cache block, thereby reducing the computational resources consumed by repeatedly calculating the key and value vectors of token blocks with the same content. Furthermore, this mechanism allows first token blocks with the same content in different requests to share the same key-value cache block, reducing storage resource consumption.

[0129] For incomplete word blocks (i.e., the second word block) where the number of words has not reached the predetermined number, they are in an "unfilled" state and can be further supplemented with words. Therefore, hashing is not performed on the second word block where the number of words has not reached the predetermined number; instead, an idle key-value cache block can be allocated to it directly. In streaming input or dynamically appended word scenarios, the last word block of the input sequence is usually in an unfilled state, and its content changes frequently as new words are added. If hashing is performed on this second word block and attempts are made to share it, the hash value needs to be recalculated every time a word is appended, resulting in significant computational overhead.

[0130] Therefore, the key-value caching management scheme provided by the embodiments of this disclosure, while ensuring fine-grained sharing at the lexical block level and reducing storage resource consumption, adopts a strategy of direct allocation and temporary non-sharing for incomplete lexical blocks, delaying hash calculation and sharing detection until the lexical block is filled, thereby reducing unnecessary computational overhead, helping to improve the overall efficiency in the dynamic decoding stage, reducing computational load and storage fragmentation while maintaining sharing efficiency, improving the throughput of the large language model inference system, and reducing the first lexical latency of the large language model.

[0131] Optionally, the partitioning unit 100 is further configured to add at least one lexical element included in the second input sequence of the large language model to the second lexical block; in response to the number of lexical elements in the second lexical block after adding at least one lexical element reaching a predetermined number condition, determine the hash value of the second lexical block; based on the hash value of the second lexical block and the hash mapping table, obtain the identifier of the candidate key-value cache block corresponding to the hash value of the second lexical block from the identifiers of the allocated key-value cache blocks; in response to the lexical sequence associated with the candidate key-value cache block being the same as the lexical sequence of the second lexical block, determine that the second lexical block reuses the candidate key-value cache block, and release the free key-value cache block allocated for the second lexical block.

[0132] Optionally, the partitioning unit 100 is further configured to add at least one lexical element included in the second input sequence of the large language model to the second lexical block; in response to the number of lexical elements in the second lexical block after adding at least one lexical element reaching a predetermined number condition, determine the hash value of the second lexical block; query the hash map table, and if the hash value of the second lexical block is not found, update the hash map table based on the hash value of the second lexical block and the identifier of the free key-value cache block allocated to the second lexical block; or query the hash map table, and if the identifier of the candidate key-value cache block corresponding to the hash value of the second lexical block is obtained, and the lexical sequence associated with the candidate key-value cache block is different from the lexical sequence of the second lexical block, update the hash map table based on the hash value of the second lexical block and the identifier of the free key-value cache block allocated to the second lexical block.

[0133] Specifically, for ease of description, the currently processed input sequence can be referred to as the first input sequence, and the next input sequence processed after the first input sequence can be referred to as the second input sequence. It is understood that the naming of the first and second input sequences is only used to distinguish different input sequences within the same large language model and does not constitute any restriction on the order or number of input sequences.

[0134] The first input sequence is divided into multiple word blocks, including a first word block with a predetermined number of words and a second word block with a less than predetermined number of words (i.e., an incomplete word block). The second input sequence includes at least one word, which can be added to the second word block of the first input sequence. After adding at least one word to the second word block of the first input sequence, the number of words in the second word block reaches the predetermined number, or in other words, the second word block changes from an unfilled state to a filled state. After the second word block becomes filled, the hash value of the second word block can be determined by referring to the processing method of the first word block above. The hash mapping table is queried to obtain the identifier of the candidate key-value cache block corresponding to the hash value, and the word sequence associated with the candidate key-value cache block is compared with the word sequence of the second word block. If the word sequence associated with the candidate key-value cache block is the same as the word sequence of the second word block, it is determined that the filled second word block can reuse the candidate key-value cache block.

[0135] In this way, the initially incomplete second word block in the first input sequence can be filled with the words from the second input sequence, thus transforming it into a complete word block and enabling it to participate in the lookup and sharing of the hash map table. This mechanism allows word blocks from different input sequences to be filled collaboratively, reducing waste caused by incomplete end blocks that cannot be shared, improving the word block filling rate and sharing reuse rate, further reducing storage redundancy, and helping to improve the overall efficiency of large language model inference systems.

[0136] Optionally, the segmentation unit 100 is further configured to: in response to the existence of a stored third word block, wherein the number of words in the third word block has not reached a predetermined number condition, add at least one word from the first input sequence to the third word block until the number of words in the third word block reaches the predetermined number condition or all words in the first input sequence have been added; and in response to the presence of unadded words in the first input sequence, divide the unadded words into multiple word blocks.

[0137] In embodiments of this disclosure, the inference system of the large language model may have processed one or more other input sequences before the currently processing first input sequence. When these previously processed input sequences are segmented into lexical blocks, incomplete lexical blocks (i.e., unfilled lexical blocks) may be generated, where the number of lexical units does not meet a predetermined condition. For ease of description, these are referred to as third lexical blocks in the following text. The third lexical block is similar to the second lexical block in the first input sequence, both being in an unfilled state, but the difference is that the third lexical block already existed and was stored in the system before processing the current first input sequence.

[0138] To fully utilize allocated storage resources, the inference system of the large language model checks for the existence of a stored third word block before dividing the multiple words included in the first input sequence into multiple word blocks. If a third word block exists, at least one word can be taken from the first input sequence and added to it. This addition process can continue until the number of words in the third word block reaches a predetermined condition (i.e., it becomes full), or until all words in the first input sequence have been added.

[0139] Through this filling operation, the initially incomplete third word block has the opportunity to be transformed into a complete word block, thus enabling it to participate in subsequent hash map table queries and shared reuse. After the third word block is transformed into a complete word block, the hash value of the third word block can be determined by referring to the processing method for the first word block described above. The hash map table is then queried to obtain the identifier of the candidate key-value cache block corresponding to the hash value, and the word sequence associated with the candidate key-value cache block is compared with the word sequence of the third word block. If the word sequence associated with the candidate key-value cache block is the same as the word sequence of the third word block, it is determined that the filled third word block can reuse the candidate key-value cache block. The specific processing details are the same as for the first word block and will not be repeated here.

[0140] Alternatively, the matching unit 200 is further configured to: record the number of lexical blocks that use the key-value cache block, and obtain the reference count corresponding to the key-value cache block; wherein, in response to the reuse of the candidate key-value cache block by the first lexical block or the second lexical block, the reference count corresponding to the candidate key-value cache block is increased; and in response to the de-reuse of the candidate key-value cache block by the first lexical block or the second lexical block, the reference count corresponding to the candidate key-value cache block is decreased.

[0141] The reference counting mechanism ensures that key-value cache blocks are released only after all users have finished processing them, preventing other currently used key-value cache blocks from accessing invalid data due to premature release, thus improving the security and reliability of key-value cache management. Based on the reference counting mechanism, as an option, a key-value cache block can be released when its reference count reaches zero. Alternatively, the most recent usage time of a key-value cache block can be recorded; and in response to the number of idle key-value cache blocks falling below a threshold, the key-value cache block with the earliest most recent usage time from among multiple key-value cache blocks with a reference count of zero can be released.

[0142] Furthermore, based on the reference counting mechanism, as an option, a key-value cache block can be released in response to its reference count reaching zero. Alternatively, the most recently used time of a key-value cache block can be recorded; and in response to the number of free key-value cache blocks falling below a threshold, the key-value cache block with the earliest most recently used time can be selected from among multiple key-value cache blocks with a reference count of zero for release.

[0143] Specifically, some embodiments of this disclosure can further optimize the eviction strategy for key-value cache blocks by incorporating the most recent usage time. Specifically, when a candidate key-value cache block is reused in response to a first or second lexical block, and the reference count of the candidate key-value cache block is increased, the most recent usage time of the candidate key-value cache block can be updated to the current time. When the number of idle key-value cache blocks falls below a threshold, the key-value cache block with the earliest most recent usage time can be selected from multiple key-value cache blocks with a reference count of zero for release. Whenever a lexical block reuses a key-value cache block, in addition to increasing the reference count of that key-value cache block, the most recent usage time of that key-value cache block can also be updated to the current time.

[0144] When the number of free key-value cache blocks falls below a threshold, an eviction operation can be triggered: from all key-value cache blocks with a reference count of zero, one or more key-value cache blocks with the earliest recent use time are selected for release. The earliest recent use time means that the key-value cache block has not been reused for a long time and is considered "cold" data. Prioritizing the release of these colder data preserves the recently reused "hot" data, thereby improving cache hit rate within limited storage resources and reducing the computational overhead caused by frequently recalculating key and value vectors.

[0145] The reference counting mechanism and the eviction policy based on recent usage time described above are independent yet complementary. The reference counting mechanism ensures that cache blocks currently in use are not mistakenly released, while the eviction policy proactively cleans up less frequently used, colder data when cache resources are scarce, providing storage space for new lexical blocks. The combination of the two helps improve the utilization efficiency of cache resources and reduce redundant computations while ensuring data security, thereby increasing the overall throughput of large language model inference systems and reducing response latency.

[0146] Alternatively, the matching unit 200 is also configured to: in response to determining that the hash value of the first word block exists in the hash map table through the Bloom filter, perform the step of determining the key-value cache block that matches the first word block based on the hash value of the first word block and the hash map table.

[0147] In addition, in response to the Bloom filter determining that the hash value of the first word block does not exist in the hash map, an idle key-value cache block is allocated for the first word block, and the hash map is updated based on the hash value of the first word block and the identifier of the key-value cache block allocated to the first word block.

[0148] A Bloom filter is a probabilistic data structure used to determine whether an element may exist or is definitely not present in a set. By introducing a Bloom filter as a pre-filter, a large number of non-existent query requests can be quickly eliminated, reducing the number of invalid accesses to the hash table, thereby reducing query latency and improving overall processing efficiency. It should be noted that the Bloom filter is only an exemplary pre-filter implementation, and the embodiments disclosed herein are not limited to using a Bloom filter. Any probabilistic data structure or deterministic filtering mechanism capable of quickly determining whether an element may exist can be applied to the technical solutions of this disclosure.

[0149] Furthermore, the matching unit 200 is also configured to determine the hash value of the first word block. For example, in response to the first word block having an adjacent preceding word block, the hash value of the preceding word block is obtained as a prefix hash, and the hash value of the first word block is determined based on the prefix hash and the word sequence of the first word block. Furthermore, in response to the first word block not having an adjacent preceding word block, the hash value of the first word block is determined based on the word sequence of the first word block. In other words, in the two embodiments described above, the hash value of the first word block can be calculated in different ways depending on whether it has an adjacent preceding word block.

[0150] When the first word block has an adjacent preceding word block, the hash value already calculated for that preceding word block is used as a prefix hash. This prefix hash is then concatenated with the word sequence of the first word block. The concatenated data is then input into a hash function to calculate the hash value of the first word block. Here, the prefix hash refers to the hash value of the preceding word block. Incorporating information from previous word blocks when calculating the hash value of the current word block distinguishes it from cases where the input sequence contains the same word sequence but different contexts. When the first word block does not have an adjacent preceding word block, the word sequence of the first word block can be directly input into the hash function to calculate its hash value.

[0151] Through the aforementioned prefix hashing propagation method, the hash value of the current word block not only depends on its own word sequence but also implicitly contains information from all preceding word blocks. This mechanism guarantees the uniqueness and contextual relevance of hash values, improving the correctness and security of sharing. Furthermore, employing efficient hash functions such as xxhash can generate hash values ​​with low collision rates while maintaining computational speed, reducing the overhead of additional content comparisons caused by hash collisions.

[0152] In some embodiments of this disclosure, an asynchronous pre-computation mechanism can be employed to further reduce the impact of hash value calculation on the main processing flow. Specifically, during the process of querying the hash map table based on the hash value of the m-th first word block among multiple first word blocks, the hash value of the (m+n)-th first word block among the multiple first word blocks is calculated, where m and n are both positive integers. In other words, the process in which the inference system can query the hash map table based on the hash value of one of the multiple first word blocks can be understood as the main flow processing the current first word block. Furthermore, the inference system can also compute the hash value of another first word block among the multiple first word blocks in parallel in a background thread, such as the hash value of the next first word block to be processed or the hash value of a predicted first word block that may be accessed. This computation process is executed in parallel with the query and allocation operations of the main flow and does not block the main flow.

[0153] In this way, when the main process finishes processing the current first word block and moves on to the next first word block, the hash value of that next first word block may have already been pre-calculated. This allows it to be directly used to query the hash mapping table without waiting for the hash value calculation. Asynchronous pre-computation overlaps the hash value calculation time with other operations in the main process, effectively reducing the main process's waiting time and improving overall processing efficiency. It should be noted that the specific implementation of asynchronous pre-computation can be through starting a separate background thread, using coroutines, or a task queue, etc. This disclosure does not limit the specific parallel execution mechanism.

[0154] Furthermore, in some embodiments of this disclosure, a linked storage mechanism can be used to manage potential hash collisions in the hash map table. For example, in response to the existence of multiple key-value cache blocks that correspond to the same hash value but have different associated lexical sequences, the identifiers of the multiple key-value cache blocks can be stored in a linked list at the position corresponding to the hash value in the hash map table.

[0155] Through the aforementioned linked storage mechanism, when a hash collision occurs, the original key-value cache block identifier will not be overwritten by a newly added identifier, thus ensuring that all stored key-value cache blocks can be correctly searched. Furthermore, during a query, by traversing the linked list and performing a precise comparison of the term sequences, key-value cache blocks with truly identical term sequences can be accurately identified, reducing erroneous matches caused by hash collisions. This mechanism, while ensuring efficient lookup of the hash map table, improves the correctness and reliability of hash collision handling, contributing to an overall performance improvement in the key-value cache management scheme.

[0156] According to embodiments of this disclosure, this disclosure also provides an electronic device, a non-volatile computer-readable storage medium, and a computer program product.

[0157] Figure 9 A schematic block diagram of an example electronic device 3000 that can be used to implement embodiments of the present disclosure is shown. The electronic device is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device may also represent various forms of mobile devices, such as personal digital assistants, cellular phones, smartphones, wearable devices, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely illustrative and are not intended to limit the implementation of the present disclosure described and / or claimed herein.

[0158] like Figure 9 As shown, the electronic device 3000 includes a computing unit 301, which can perform various appropriate actions and processes based on a computer program stored in a read-only memory (ROM) 302 or a computer program loaded from a storage unit 308 into a random access memory (RAM) 303. The RAM 303 may also store various programs and data required for the operation of the electronic device 3000. The computing unit 301, ROM 302, and RAM 303 are interconnected via a bus 304. An input / output (I / O) interface 305 is also connected to the bus 304.

[0159] Multiple components in electronic device 3000 are connected to I / O interface 305, including: input unit 306, such as keyboard, mouse, etc.; output unit 307, such as various types of displays, speakers, etc.; storage unit 308, such as disk, optical disk, etc.; and communication unit 309, such as network card, modem, wireless transceiver, etc. Communication unit 309 allows electronic device 3000 to exchange information / data with other devices through computer networks such as the Internet and / or various telecommunications networks.

[0160] The computing unit 301 can be a variety of general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of the computing unit 301 include, but are not limited to, a central processing unit, a graphics processing unit, various special-purpose artificial intelligence computing chips, various computing units running machine learning model algorithms, digital signal processors, and any suitable processor, controller, microcontroller, etc. The computing unit 301 performs the various means and processes described above, such as the key-value cache management method. For example, in some embodiments, the key-value cache management method can be implemented as a computer software program tangibly included in a machine-readable medium, such as storage unit 308. In some embodiments, part or all of the computer program can be loaded and / or installed on the electronic device 3000 via ROM 302 and / or communication unit 309. When the computer program is loaded into RAM 303 and executed by the computing unit 301, one or more steps of the key-value cache management method described above can be performed. Alternatively, in other embodiments, the computing unit 301 can be configured to perform the key-value cache management method by any other suitable means (e.g., by means of firmware).

[0161] Various embodiments of the systems and techniques described above herein can be implemented in digital electronic circuit systems, integrated circuit systems, field-programmable gate arrays, application-specific integrated circuits (ASICs), application-specific standard products (ASICs), systems-on-a-chip (SoCs), complex programmable logic devices, computer hardware, firmware, software, and / or combinations thereof. These embodiments can be implemented as one or more computer programs that can execute on a programmable system including at least one programmable processor, which can be a dedicated or general-purpose programmable processor, capable of receiving data and instructions from a storage system, at least one input device, and at least one output device, and transferring data and instructions to the storage system, the at least one input device, and the at least one output device.

[0162] Program code for implementing the apparatus of this disclosure may be written in any combination of one or more programming languages. This program code may be provided to a processor or controller of a general-purpose computer, a special-purpose computer, or other programmable data processing method, such that when executed by the processor or controller, the program code causes the functions / operations specified in the flowcharts and / or block diagrams to be implemented. The program code may be executed entirely on a machine, partially on a machine, as a standalone software package partially on a machine and partially on a remote machine, or entirely on a remote machine or server.

[0163] In the context of this disclosure, a machine-readable medium can be a tangible medium that may include or store a program for use by or in conjunction with an instruction execution system, method, or apparatus. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can be, but is not limited to, an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, method, or apparatus, or any suitable combination of the foregoing. More specific examples of machine-readable storage media include electrical connections based on one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory, optical fibers, portable compact disk read-only memory, optical storage devices, magnetic storage devices, or any suitable combination of the foregoing.

[0164] To provide interaction with a user, the systems and techniques described herein can be implemented on a computer having: a display method for showing information to the user (e.g., a CRT (Cathode Ray Tube) or an LCD (Liquid Crystal Display) monitor); and a keyboard and pointing method (e.g., a mouse or trackball) through which the user provides input to the computer. Other types of methods can also be used to provide interaction with the user; for example, feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including sound input, voice input, or tactile input).

[0165] The systems and technologies described herein can be implemented in computing systems that include backend components (e.g., as a data server), or computing systems that include middleware components (e.g., an application server), or computing systems that include frontend components (e.g., a user computer with a graphical user interface or web browser through which a user can interact with implementations of the systems and technologies described herein), or any combination of such backend, middleware, or frontend components. The components of the system can be interconnected via digital data communication of any form or medium (e.g., a communication network). Examples of communication networks include local area networks (LANs), wide area networks (WANs), and the Internet.

[0166] Computer systems can include clients and servers. Clients and servers are generally located far apart and typically interact through communication networks. Client-server relationships are created by computer programs running on the respective computers and having a client-server relationship with each other.

[0167] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of systems, apparatuses, and computer program products according to various embodiments of the present disclosure. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of an instruction, which includes one or more executable instructions for implementing a specified logical function. In some alternative implementations, the functions marked in the blocks may occur in a different order than those marked in the drawings. For example, two consecutive blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, can be implemented using a dedicated hardware-based system that performs the specified function or action, or using a combination of dedicated hardware and computer instructions.

[0168] Various embodiments of this disclosure have been described above. These descriptions are exemplary and not exhaustive, nor are they limited to the disclosed embodiments. Many modifications and variations will be apparent to those skilled in the art without departing from the scope and spirit of the described embodiments. The terminology used herein is chosen to best explain the principles, practical applications, or technological improvements to the embodiments in the market, or to enable others skilled in the art to understand the embodiments disclosed herein.

[0169] It should be understood that the various forms of processes shown above can be used to rearrange, add, or delete steps. For example, the steps described in this disclosure can be executed in parallel, sequentially, or in different orders, as long as the desired result of the technical solution of this disclosure can be achieved, and this is not limited herein.

[0170] The specific embodiments described above do not constitute a limitation on the scope of protection of this disclosure. Those skilled in the art should understand that various modifications, combinations, sub-combinations, and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of this disclosure should be included within the scope of protection of this disclosure.

Claims

1. A key-value cache management method, comprising: The first input sequence of the large language model is divided into multiple word blocks, wherein the first word block of the multiple word blocks includes word blocks that meet a predetermined number of conditions; Based on the hash value of the first lexical block and the hash mapping table, a key-value cache block matching the first lexical block is determined, wherein the hash mapping table is used to store the correspondence between the hash value and the identifier of the allocated key-value cache block.

2. The method according to claim 1, wherein, The step of determining the key-value cache block matching the first lexical block based on its hash value and hash mapping table includes: Query the hash mapping table to obtain the identifier of the candidate key-value cache block corresponding to the hash value of the first word block from the identifiers of the allocated key-value cache blocks; In response to the fact that the lexical sequence associated with the candidate key-value cache block is the same as the lexical sequence of the first lexical block, it is determined that the first lexical block reuses the candidate key-value cache block.

3. The method according to claim 1, wherein, The method further includes: The hash mapping table is queried. If the hash value of the first word block is not found, an idle key-value cache block is allocated to the first word block, and the hash mapping table is updated based on the hash value of the first word block and the identifier of the key-value cache block allocated to the first word block; or The hash mapping table is queried. If the identifier of the candidate key-value cache block corresponding to the hash value of the first word block is obtained, and the word sequence associated with the candidate key-value cache block is different from the word sequence of the first word block, an idle key-value cache block is allocated to the first word block, and the hash mapping table is updated based on the hash value of the first word block and the identifier of the key-value cache block allocated to the first word block.

4. The method according to claim 1, wherein, The method further includes: Allocate an idle key-value cache block for the second lexical block among the plurality of lexical blocks, wherein the number of lexical units in the second lexical block does not reach the predetermined number condition.

5. The method according to claim 4, wherein, The method further includes: Add at least one of the lexical units included in the second input sequence of the large language model to the second lexical unit block; In response to the condition that the number of terms in the second term block after adding at least one of the terms reaches the predetermined number, the hash value of the second term block is determined; Based on the hash value of the second word block and the hash mapping table, the identifier of the candidate key-value cache block corresponding to the hash value of the second word block is obtained from the identifiers of the allocated key-value cache blocks; In response to the fact that the lexical sequence associated with the candidate key-value cache block is the same as the lexical sequence of the second lexical block, it is determined that the second lexical block reuses the candidate key-value cache block, and the free key-value cache block allocated to the second lexical block is released.

6. The method according to claim 4, wherein, The method further includes: Add at least one of the lexical units included in the second input sequence of the large language model to the second lexical block; In response to the condition that the number of terms in the second term block after adding at least one of the terms reaches the predetermined number, the hash value of the second term block is determined; If the hash value of the second term block is not found in the hash mapping table, the hash mapping table is updated based on the hash value of the second term block and the identifier of the free key-value cache block allocated to the second term block; or The hash mapping table is queried. If the identifier of the candidate key-value cache block corresponding to the hash value of the second lexical block is obtained, and the lexical sequence associated with the candidate key-value cache block is different from the lexical sequence of the second lexical block, the hash mapping table is updated based on the hash value of the second lexical block and the identifier of the free key-value cache block allocated to the second lexical block.

7. The method according to claim 1, wherein, The process of dividing the multiple lexical units included in the first input sequence of the large language model into multiple lexical blocks includes: In response to the existence of a stored third word block, wherein the number of words in the third word block does not reach the predetermined number condition, at least one word from the first input sequence is added to the third word block until the number of words in the third word block reaches the predetermined number condition or all words in the first input sequence have been added; and In response to the presence of unadded lexical units in the first input sequence, the unadded lexical units are divided into multiple lexical blocks.

8. The method according to claim 5, wherein, The method further includes: Record the number of lexical blocks that use the key-value cache block to obtain the reference count corresponding to the key-value cache block; Wherein, in response to the first lexical block or the second lexical block reusing the candidate key-value cache block, the reference count corresponding to the candidate key-value cache block is increased; and In response to the first lexical block or the second lexical block de-reusing the candidate key-value cache block, the reference count corresponding to the candidate key-value cache block is reduced.

9. The method according to claim 8, wherein, The method further includes: The key-value cache block is released in response to the reference count of the key-value cache block being zero.

10. The method according to claim 8, wherein, The method further includes: Record the most recent usage time of the key-value cache block; and In response to the number of idle key-value cache blocks falling below a threshold, the key-value cache block with the earliest recent use time is selected from the plurality of key-value cache blocks with a reference count of zero and released.

11. The method according to claim 1, wherein, The method further includes: In response to the first lexical block having an adjacent preceding lexical block, the hash value of the preceding lexical block is obtained as a prefix hash, and the hash value of the first lexical block is determined based on the prefix hash and the lexical sequence of the first lexical block.

12. The method according to claim 11, wherein, The method further includes: The following steps are executed in parallel: The hash mapping table is queried based on the hash value of the m-th first word block among multiple first word blocks; and Determine the hash value of the (m+n)th first word block among multiple first word blocks, where m and n are positive integers.

13. The method according to claim 1, wherein, The method further includes: In response to the determination by the Bloom filter that the hash value of the first term block exists in the hash map table, the step of determining the key-value cache block that matches the first term block based on the hash value of the first term block and the hash map table is performed.

14. The method according to claim 1, wherein, The method further includes: In response to the determination by the Bloom filter that the hash value of the first term block does not exist in the hash map table, an idle key-value cache block is allocated for the first term block, and the hash map table is updated based on the hash value of the first term block and the identifier of the key-value cache block allocated to the first term block.

15. The method according to any one of claims 1 to 6, wherein, The method further includes: In response to the existence of multiple key-value cache blocks that correspond to the same hash value but have different associated lexical sequences, the identifiers of the multiple key-value cache blocks are stored in the hash mapping table in the form of a linked list at the position corresponding to the hash value.

16. A key-value cache management device, comprising: The segmentation unit is configured to divide a plurality of lexical units included in a first input sequence of a large language model into a plurality of lexical blocks, wherein a first lexical block among the plurality of lexical blocks includes lexical units that meet a predetermined number of conditions; The matching unit determines the key-value cache block that matches the first lexical block based on the hash value of the first lexical block and the hash mapping table, wherein the hash mapping table is used to store the correspondence between the hash value and the identifier of the allocated key-value cache block.

17. An electronic device comprising: A processor and a memory, wherein the memory stores at least one instruction, at least one program, a code set, or an instruction set, wherein the at least one instruction, the at least one program, the code set, or the instruction set is loaded and executed by the processor to implement the key-value cache management method as described in any one of claims 1 to 15.

18. A non-volatile computer-readable storage medium having stored thereon computer program instructions, which, when executed by a processor, implement the key-value cache management method according to any one of claims 1 to 15.

19. A computer program product comprising a computer program that, when executed by a processor, implements the steps of the key-value cache management method according to any one of claims 1 to 15.