Method, device, equipment and product for storing key value pairs
By determining the lexical storage location based on the cumulative attention value in a large language model, the problem of information loss caused by deleting or merging lexical units in existing technologies is solved, achieving a balance between efficient storage and accurate reasoning.
Patent Information
- Application Number
- CN202511014332.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-22
- Publication Date
- 2025-11-07
AI Technical Summary
Existing technologies, when reducing key-value cache memory usage in large language models by deleting or merging lexical units, result in irreversible information loss, affecting the model's information integrity and semantic expression.
The storage location of historical lexical units is determined by the cumulative attention value. Lexical units with a cumulative attention value greater than a preset value are stored in the first storage space of the key-value cache, while lexical units with a cumulative attention value less than or equal to the preset value are stored in the second storage space of the multi-bucket hash table, ensuring the integrity and efficient storage of lexical information.
This approach achieves reduced memory usage while maintaining model inference quality, avoids information loss due to deletion or merging, and ensures the model's accurate understanding and response to context.
Smart Images

Figure CN120910183A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present disclosure relates to the field of computers, and more specifically, to methods, apparatuses, devices, and products for storing key-value pairs. BACKGROUND
[0002] In large language models (LLMs), tokens are the smallest semantic units of text after being segmented, which can be words, subwords, characters, or symbols, used to convert original text into numerical forms that models can process. Tokens are the basic units of model input and output. During model inference, the history dialogue context is usually stored in the form of key-value cache. By retaining the embedding vectors of the keys and values of the generated tokens, the repeated calculation of the self-attention layer is avoided, thereby improving the inference efficiency. With the increase of the length of the context window, the memory occupancy of the key-value cache increases linearly.
[0003] To alleviate the memory pressure, there are two compression strategies in the related technology. One is selective pruning, which permanently removes part of the low importance tokens based on attention scores and other indicators. The other is dynamic merging, which compresses the embedding vectors of multiple tokens into a small number of new token representations through weighted fusion. Both of the above two methods reduce the number of stored tokens to reduce memory overhead. SUMMARY
[0004] In a first aspect of embodiments of the present disclosure, a method for storing key-value pairs is provided. The method includes determining a plurality of accumulated attention values of a plurality of historical tokens in response to receiving a new token of the key-value pair. The method also includes determining, for the plurality of historical tokens, whether the accumulated attention value of each historical token is greater than a preset value. The method further includes storing, in response to the accumulated attention value of the historical token being greater than the preset value, the historical token to a first storage space of a key-value cache. In addition, the method also includes storing, in response to the accumulated attention value of the historical token being less than or equal to the preset value, the historical token to a second storage space of the key-value cache, the second storage space including a plurality of fixed-size multi-bucket hash tables.
[0005] In a second aspect of embodiments of the present disclosure, an apparatus for storing key-value pairs is provided. The apparatus includes a cumulative attention value determination module configured to determine, in response to receiving a new token of a key-value pair, a plurality of cumulative attention values of a plurality of historical tokens. The apparatus further includes a first determination module configured to determine, for the plurality of historical tokens, whether a cumulative attention value of each historical token is greater than a preset value. The apparatus further includes a first storage module configured to store, in response to the cumulative attention value of a historical token being greater than the preset value, the historical token to a first storage space of a key-value cache. In addition, the apparatus further includes a second storage module configured to store, in response to the cumulative attention value of the historical token being less than or equal to the preset value, the historical token to a second storage space of the key-value cache, the second storage space comprising a plurality of fixed-size multi-bucket hash tables.
[0006] In a third aspect of embodiments of the present disclosure, an electronic device is provided. The electronic device includes one or more processors; and a storage storing one or more programs, when executed by the one or more processors, cause the one or more processors to implement a method for storing key-value pairs. The method includes determining, in response to receiving a new token of a key-value pair, a plurality of cumulative attention values of a plurality of historical tokens. The method further includes determining, for the plurality of historical tokens, whether a cumulative attention value of each historical token is greater than a preset value. The method further includes storing, in response to the cumulative attention value of a historical token being greater than the preset value, the historical token to a first storage space of a key-value cache. In addition, the method further includes storing, in response to the cumulative attention value of the historical token being less than or equal to the preset value, the historical token to a second storage space of the key-value cache, the second storage space comprising a plurality of fixed-size multi-bucket hash tables.
[0007] In a fourth aspect of embodiments of the present disclosure, a computer program product is provided. The computer program product is tangibly stored on a non-transitory computer-readable medium and includes machine executable instructions that, when executed, cause a machine to implement a method for storing key-value pairs. The method includes determining, in response to receiving a new token of a key-value pair, a plurality of cumulative attention values of a plurality of historical tokens. The method further includes determining, for the plurality of historical tokens, whether a cumulative attention value of each historical token is greater than a preset value. The method further includes storing, in response to the cumulative attention value of a historical token being greater than the preset value, the historical token to a first storage space of a key-value cache. In addition, the method further includes storing, in response to the cumulative attention value of the historical token being less than or equal to the preset value, the historical token to a second storage space of the key-value cache, the second storage space comprising a plurality of fixed-size multi-bucket hash tables.
[0008] The following presents a summary of the disclosure in order to provide a basic understanding of some aspects of the application. This summary is not an extensive overview of the application. It is not intended to identify key or critical elements of the application or to delineate the scope of the application. Its sole purpose is to present some concepts of the application in a simplified form as a prelude to the more detailed description that is presented later. BRIEF DESCRIPTION OF DRAWINGS
[0009] The above and other features, aspects, and advantages of various embodiments of the present disclosure will become more apparent from the following detailed description when taken in conjunction with the accompanying drawings. In the drawings:
[0010] Figure 1 A schematic diagram of a key-value pair storage system is shown in accordance with some embodiments of the present disclosure;
[0011] Figure 2 A flowchart of a method for storing key-value pairs is shown in accordance with some embodiments of the present disclosure;
[0012] Figure 3 A schematic diagram of a key-value pair storage system is shown in accordance with some embodiments of the present disclosure;
[0013] Figure 4A A flowchart of a method for managing storage space for key-value caches is shown in accordance with some embodiments of the present disclosure;
[0014] Figure 4B A flowchart of another method for managing storage space for key-value caches is shown in accordance with some embodiments of the present disclosure;
[0015] Figure 5 A flowchart of a method for restoring a key-value matrix from a key-value cache is shown in accordance with some embodiments of the present disclosure;
[0016] Figure 6 A flowchart of a method for storing historical tokens to a plurality of multi-bucket hash tables is shown in accordance with some embodiments of the present disclosure;
[0017] Figure 7 A flowchart of a method for restoring historical tokens from a plurality of multi-bucket hash tables is shown in accordance with some embodiments of the present disclosure;
[0018] Figure 8 A block diagram of an apparatus for storing key-value pairs is shown in accordance with some embodiments of the present disclosure; and
[0019] Figure 9 A block diagram of an apparatus capable of implementing various embodiments of the present disclosure is shown. DETAILED DESCRIPTION
[0020] It can be understood that all user-related data involved in the technical solution should be acquired and used after the user's authorization. This means that in the technical solution, if the user's personal information needs to be used, the user's explicit consent and authorization are required before the data is acquired, otherwise the relevant data collection and use will not be performed. It should also be understood that in the implementation of the technical solution, relevant laws and regulations should be strictly followed in the collection, use and storage of data, and necessary technical and measures should be taken to protect the user's data security and ensure the safe use of data.
[0021] It can be understood that before using the technical solution disclosed in each embodiment of the present disclosure, the type of personal information involved in the present disclosure, the use range, the use scenario, etc. should be informed to the user and the authorization of the user should be obtained through appropriate means according to relevant laws and regulations.
[0022] For example, when receiving the user's active request, prompt information is sent to the user to explicitly prompt the user that the operation requested to be performed will require the acquisition and use of the user's personal information. Thus, the user can voluntarily choose whether to provide personal information to the electronic device, application program, server or storage medium, etc. software or hardware that performs the operation of the technical solution of the present disclosure according to the prompt information.
[0023] As an optional but non-limiting implementation manner, in response to receiving the user's active request, the way of sending prompt information to the user may be, for example, the way of pop-up window, in which the prompt information can be presented in the form of text. In addition, the pop-up window can also carry selection controls for the user to select "agree" or "disagree" to provide personal information to the electronic device.
[0024] It can be understood that the above notification and user authorization process is only illustrative and does not limit the implementation of the present disclosure. Other ways that meet relevant laws and regulations can also be applied to the implementation of the present disclosure.
[0025] Embodiments of the present disclosure will be described in more detail below with reference to the accompanying drawings. Although certain embodiments of the present disclosure are shown in the drawings, it should be understood that the present disclosure can be implemented in various forms, and should not be interpreted as being limited to the embodiments set forth herein, but rather these embodiments are provided to more thoroughly and completely understand the present disclosure. It should be understood that the drawings and embodiments of the present disclosure are only for exemplary purposes, and are not intended to limit the scope of protection of the present disclosure.
[0026] In the description of embodiments of the disclosure, the term "includes" and its similar terms are understood to be open-ended, i.e., "includes but is not limited to". The term "based on" is understood to be "based, at least in part, on". The term "one embodiment" or "the embodiment" is understood to be "at least one embodiment". The terms "first", "second", and the like can refer to different or identical objects unless otherwise specified. The following can also include other explicit and implicit definitions.
[0027] As described above, the related art reduces the memory occupancy of the key-value cache by deleting key-value pairs or merging key-value pairs. Although the memory occupancy can be reduced by pruning, the original information of the removed tokens will be permanently lost and cannot be recovered. This irreversible pruning causes the model to be unable to access the information of the discarded tokens during subsequent decoding, especially in tasks involving long-term dependencies, such as multi-round question answering and document analysis, which can cause context amnesia due to the lack of key details, thereby compromising the information integrity of the model.
[0028] On the other hand, although the merging strategy preserves more token information, the fused token generated by linear weighting destroys the distribution characteristics of the original embedding space. Such methods cause distortion of the semantic expression of the key vector during attention calculation, and the merged token cannot be decoupled and restored to the original independent token, also causing irreversible loss of information, reducing the model's ability to capture fine-grained semantics.
[0029] To this end, the present disclosure provides a method for storing key-value pairs. First, after receiving a new token of a key-value pair, a plurality of accumulated attention values of a plurality of historical tokens are determined, then it is determined whether the accumulated attention value of each historical token is greater than a preset value, and finally the historical tokens with accumulated attention values greater than the preset value are stored in a first storage space of the key-value cache, and the historical tokens with accumulated attention values less than or equal to the preset value are stored in a second storage space of the key-value cache, the second storage space including a plurality of fixed-size multi-bucket hash tables. In this way, the tokens stored in the multi-bucket hash table can be reconstructed with high precision, so that all token information can be recovered, and the fixed-size storage structure can significantly reduce the memory occupancy of the key-value cache, thereby achieving reduced memory occupancy while improving model inference quality.
[0030] Figure 1 A schematic diagram of an example environment 100 in which embodiments of the present disclosure can be implemented is shown. As shown in FIG. 1, the environment 100 includes a client device 102, a server 104, and a network 106. Figure 1As shown, the example environment 100 can include a context, which refers to historical interaction content generated by the user and the model during the conversation process, including past questions, answers, and intermediate generated text sequences. The context can include multiple historical tokens, such as historical token 105-1, historical token 105-2, historical token 105-3, historical token 105-4, and historical token 105-N. A historical token is a basic unit that constitutes the context and can be defined as a word, symbol, or subword unit after tokenization processing. Each historical token can be stored in the key-value cache 107 in the form of a key-value pair inside the model. Specifically, the key is usually a unique identifier or position encoding of the token, and the value is a semantic embedding vector (such as a word embedding or a position embedding) corresponding to the token, which is used to represent the semantic information and context association of the token. Taking a dialogue scenario as an example, if the context is “he loves to eat apples”, after tokenization processing, historical token 105-1 can correspond to the key-value pair of “he”, historical token 105-2 can correspond to the key-value pair of “loves”, historical token 105-3 can correspond to the key-value pair of “to eat”, and historical token 105-4 can correspond to the key-value pair of “apples”.
[0031] In some embodiments, the model can be a large language model capable of generating natural language text or understanding the meaning of language text. The large language model can store all tokens (such as user questions and model answers) generated during the user and model interaction process in the key-value cache 107 in the form of key-value pairs to form a set of historical tokens. When a new dialogue turn arrives, the model can understand the context semantics by retrieving historical tokens in the key-value cache 107, thereby generating answers that conform to the context, and realizing the ability of continuous dialogue based on the context. When a new token 101 appears, the cumulative attention value of each historical token can be determined by the cumulative attention value determination module 103. The cumulative attention value determination module 103 can be a computing component integrated inside the large language model or an external control module independent of the model. As a model component, it can be linked with the self-attention mechanism. At each time a new token 101 is generated, the attention value of the new token 101 to each historical token is calculated, and the attention value is added to the cumulative attention value of the corresponding historical token. For example, if the attention value of the new token 101 to the historical token 105-1 is 0.12, and the cumulative attention value of the historical token 105-1 before the model receives the new token 101 is 0.85, then the updated cumulative attention value is 0.97. As an external module, the cumulative attention value determination module 103 can interact with the model through a pre-defined API interface, real-time access to the attention value calculation result and perform accumulation operation. After determining the cumulative attention value of each historical token, the position of the stored historical token can be determined according to the multiple cumulative attention values.
[0032] In the embodiments of the present disclosure, the key-value cache 107 can include a first storage space 109 and a second storage space 111. The first storage space 109 can be used to store the complete history tokens, which means that the original key-value pair data of the tokens is retained without compression processing, and the key-value vector can be directly read when needed. The second storage space 111 can include a plurality of fixed-size multi-bucket hash tables, such as multi-bucket hash table 113, multi-bucket hash table 117, and multi-bucket hash table 119. The fixed-size multi-bucket hash table can control the storage overhead within a preset memory budget, and by limiting the number and capacity of the buckets, high-density storage can be achieved in a limited space. Each multi-bucket hash table can include a plurality of bucket positions, such as bucket position 115, which is used to store the key-value vector of the history token by means of hash function mapping. In some embodiments, when different history tokens are mapped to the same bucket position, the key-value vector can be stored cumulatively, so as to achieve high-density storage in a limited space.
[0033] In some embodiments, the storage space of the history tokens and the new token 101 can be determined by the model according to a preset strategy, or can be performed by other modules according to a preset strategy. For example, the model can store the history token 105-4 with a cumulative attention value exceeding a preset value to the first storage space 109 to completely retain its original key-value pair data, and ensure that it can be directly and quickly obtained subsequently; and store the history token 105-N with a cumulative attention value lower than the preset value to the second storage space 111. In the above manner, the high-frequency activated and possibly frequently used history tokens are completely stored in the first storage space 109, so that they can be directly called during inference; and the history tokens with less possibility of use are stored in the second storage space 111 by compression, so as to reduce the memory occupation while ensuring that all tokens can be accurately reconstructed during inference, avoiding information loss caused by permanent deletion or merging.
[0034] In this way, the high-frequency used history tokens can be stored in a complete form to ensure fast and accurate calling during inference, and the low-frequency tokens can be stored by compression to save memory, while the token information can be reconstructed from the compressed data when needed, achieving a balance between storage efficiency and information integrity, avoiding context semantic discontinuity caused by simple deletion or merging of tokens, and providing a guarantee for the model to accurately understand the context based on the history interaction and generate a response.
[0035] Figure 2 A flowchart of a method 200 for storing key-value pairs is shown, according to some embodiments of the present disclosure. The method 200 can be performed by a model capable of generating natural language text or understanding language text. Figure 1 The method 200 includes block 202, block 204, block 206, and block 208.
[0036] As Figure 2As shown, at block 202, in response to receiving the new token of the key-value pair, a plurality of accumulated attention values of a plurality of historical tokens is determined. Referring to Figure 1 The accumulated attention value of each historical token can be determined by the accumulated attention value determination module 103, which can be a computing component integrated inside the large language model or an external control module independent of the model. In some embodiments, the method of determining the accumulated attention value can include calculating the attention value of the new token 101 to each historical token and accumulating the attention value into the accumulated attention value of the corresponding historical token. For example, if the attention value of the new token 101 to the historical token 105-1 is 0.12 and the accumulated attention value of the historical token 105-1 before the model receives the new token 101 is 0.85, the updated accumulated attention value is 0.97. Of course, the accumulated attention values of the plurality of historical tokens can also be determined by other methods, which are not limited in the present disclosure.
[0037] At block 204, for the plurality of historical tokens, it is determined whether the accumulated attention value of each historical token is greater than a preset value. Referring to Figure 1 In response to receiving the new token 101 of the key-value pair, the model can calculate the attention value of the new token 101 to each historical token based on the self-attention mechanism and accumulate it into the accumulated attention value of the corresponding historical token, and then compare the accumulated attention value of each historical token with the preset value. The preset value can be dynamically adjusted according to the memory budget and task requirements in the model, for example, in the experiment of 2k length context, the preset value can be set to a relatively high attention weight threshold, while in the 32k length context, considering the memory limit and the change of information density, the preset value can be adjusted accordingly to balance the storage efficiency and information integrity.
[0038] At block 206, in response to the accumulated attention value of the historical token being greater than the preset value, the historical token is stored in the first storage space of the key-value cache. Referring to Figure 1 When the accumulated attention value of the historical token exceeds the preset value, it indicates that the token has a high importance in the current context and needs to be stored in the first storage space 109 of the key-value cache 107. The first storage space 109 uses a direct storage method, which completely retains the original key-value pair data of the historical token, including the key (such as the unique identifier or position encoding of the token) and the value (semantic embedding vector), without any compression processing. This storage method ensures the integrity of the token information, allowing the model to directly and quickly read the key-value vector during inference, improving response speed.
[0039] At block 208, in response to the accumulated attention value of the historical token being less than or equal to the preset value, the historical token is stored in the second storage space of the key-value cache, and the second storage space includes a plurality of fixed-size multi-bucket hash tables. Referring to Figure 1If the cumulative attention value of a historical word is less than or equal to a preset value, it is stored in the second storage space 111 of the key-value cache 107. The second storage space 111 consists of multiple fixed-size multi-bucket hash tables, each containing several buckets. In some embodiments, historical words can be mapped to corresponding buckets for storage using a hash function. Fixed-size multi-bucket hash tables can control storage overhead within a preset memory budget, achieving high-density storage within a limited space by limiting the number and capacity of buckets.
[0040] In this way, the tokens stored in the multi-bucket hash table can be reconstructed with high accuracy, so that all token information can be recovered. Moreover, the fixed-size storage structure can significantly reduce the memory usage of the key-value cache, thereby improving the model inference quality while reducing memory usage.
[0041] Figure 3 A schematic diagram of a process 300 for storing key-value pairs according to some embodiments of the present disclosure is shown. Figure 3 As shown, the new lexical unit 301, the cumulative attention value determination module 303, the historical lexical units 305-1, 305-2, 305-3, 305-4, 305-N, the key-value cache 307, the first storage space 309, the second storage space 311, the multi-bucket hash table 313, the buckets 315, the multi-bucket hash table 317, and the multi-bucket hash table 319 are respectively located in... Figure 1 The cumulative attention value determination module 103, historical lexicon 105-1, historical lexicon 105-2, historical lexicon 105-3, historical lexicon 105-4, historical lexicon 105-N, key-value cache 107, first storage space 109, second storage space 111, multi-bucket hash table 113, bucket 115, multi-bucket hash table 117, and multi-bucket hash table 119 are consistent and will not be described in detail here.
[0042] In some embodiments, the accumulated attention value determination module 303 can include a normalization exponential function 303-1. The accumulated attention value determination module 303 can first calculate the degree of association between the new token 101 and each historical token based on the calculation logic of the attention mechanism to obtain an original association score, and then normalize these scores by the normalization exponential function 303-1 to convert them into a plurality of attention values 303-2 between 0 and 1, the sum of which is 1, so as to accurately measure the degree of attention of the new token 301 to each historical token. At the same time, the accumulated attention value determination module 303 can also include an accumulation function 303-3, which can correspond to the accumulation of the plurality of attention values 303-2 determined by the normalization exponential function 303-1 into the accumulated attention value of each historical token. For example, if the attention value 303-2 of the new token 301 to the historical token 305-1 is 0.12, and the accumulated attention value of the historical token 305-1 before the model receives the new token 301 is 0.85, then the updated accumulated attention value is 0.97. Through such a process, the model can dynamically track the accumulation of attention between tokens, and provide accurate and real-time quantitative basis for judging the importance of tokens in subsequent storage strategies.
[0043] As shown in Figure 3 The key-value cache 307 can also include a third storage space 321. The third storage space 321 and the first storage space 309 are both used to store the tokens of the context completely. The model can store the latest received preset number of historical tokens and the new token 301 in the third storage space 321 according to the receiving order. For example, the new token 301, the historical token 305-1 and the historical token 305-2 can be stored in the third storage space 321, regardless of whether the accumulated attention values of the historical token 305-1 and the historical token 305-2 are greater than the preset value. For the remaining historical tokens, the preset value is divided, and the smaller one is stored in the second storage space 311, and the larger one is stored in the first storage space 309. Unlike the strategy of determining the storage location only according to the accumulated attention value, the third storage space 321 and the first storage space 309 can both undertake the function of storing tokens completely, and can follow independent storage strategies. For example, the latest received preset number of historical tokens and the new token 301 can be stored in the third storage space 321 according to the receiving order. Figure 3As shown, the new word item 301, the historical word item 305-1, and the historical word item 305-2 can be stored in the third storage space 321, and the original key-value pair data of the two historical word items can be completely retained regardless of whether the cumulative attention value of the two historical word items exceeds the preset value. After this part of storage is completed, for the remaining historical word items that are not included in the third storage space 321, the historical word items are further divided according to the relationship between the cumulative attention value and the preset value: the cumulative attention value is less than or equal to the preset value, and the historical word item is stored in the second storage space 311; and the cumulative attention value is greater than the preset value, and the historical word item is stored in the first storage space 309.
[0044] In this way, the third storage space 321 can quickly obtain recent key interaction information to assist in understanding the current context when processing continuous dialog by virtue of complete retention of the latest interaction word item; the first storage space 309 focuses on high-frequency activation and long-term important word items to ensure efficient calling of core semantics; and the second storage space 311 compressively stores low-frequency word items to save memory. The three spaces cooperate to realize hierarchical management of word items of different heat and different time sequences, which not only meets the dependence of real-time dialog on the latest information, but also takes into account the retention and memory optimization of important information in the long-term context, thereby providing a basis for the model to accurately understand complex dialog scenarios and generate coherent responses.
[0045] In some embodiments, the historical word items stored in the second storage space 311 can be reconstructed. The second storage space 311 includes a plurality of multi-bucket hash tables. For example, when the historical word item 323 is to be reconstructed, the model can first find the corresponding bucket position of the historical word item 323 in each multi-bucket hash table according to the feature information of the historical word item 323 by using a hash function. In some embodiments, since the key-value vector of the historical word item is mapped by a hash function and stored in the bucket position in an accumulated form, the model can extract the accumulated key-value vector related data from the corresponding bucket position, integrate and calculate the information scattered in different bucket positions of the plurality of multi-bucket hash tables according to a preset reconstruction algorithm, and gradually restore the original key-value pair data of the historical word item 323. Thus, the compressed historical word item 323 can be reconstructed in a complete and usable form to provide semantic information support for the model inference process, ensuring that even after compression storage, the context association relationship and other information contained in the word item can be effectively called to maintain the coherence of the model's understanding of the context.
[0046] Figure 4A A flowchart of a method 400A of managing storage spaces of a key-value cache is shown according to some embodiments of the present disclosure. At block 402, it is determined whether the third storage space reaches a first preset capacity. For example, as shown in FIG. 3, the first preset capacity can be a storage upper limit previously set by the model for the third storage space 321 to regulate the storage size of the latest interaction word item. Figure 3 Figure 3 The third storage space 321 is used to store the latest preset number of word elements in the order they are received, such as new word element 301, historical word element 305-1, and historical word element 305-2. When storing a new word element 301, it is necessary to check whether the number of word elements currently stored in the third storage space 321 and the memory occupied have reached the first preset capacity. If they have, it means that the third storage space 321 cannot directly accommodate the new word element 301.
[0047] At box 404, determine whether the first storage space has reached the second preset capacity. For example, as... Figure 3 As shown, after determining that the third storage space 321 has reached the first preset capacity, it is then determined whether the first storage space 309 has reached the second preset capacity. The second preset capacity of the first storage space 309 is also a storage threshold set by the model based on memory resources, high-frequency terminology call requirements, etc. Figure 3 It can be seen that the first storage space 309 is used to store historical words with high cumulative attention values and long-term importance. At this time, the storage status of the first storage space 309 is checked. If it has not reached the second preset capacity, it means that there is still space remaining, which can accept words transferred from the third storage space 321. If it has reached the second preset capacity, it means that the first storage space 309 cannot directly receive newly transferred words. It is necessary to first filter and compress the words stored inside it (such as the process in the subsequent box 406) to free up space, so as to ensure the orderly storage management of words with different popularity and balance the memory occupation and semantic information retention requirements.
[0048] At box 406, historical lexical units whose cumulative attention values are below a threshold in the first storage space are stored in a multi-bucket hash table in the second storage space. For example, as shown... Figure 3 As shown, if it is determined in box 404 that the first storage space 309 has reached the second preset capacity, the operation in box 406 is executed, that is, the historical words in the first storage space 309 with a cumulative attention value lower than the threshold are stored in the multi-bucket hash table of the second storage space 311 (such as multi-bucket hash table 313, multi-bucket hash table 317, multi-bucket hash table 319, etc.). The first storage space 309 originally stores important words with high frequency activation and high cumulative attention value, but as new word storage needs arise, it is necessary to transfer the relatively "less important" words (with a cumulative attention value lower than the threshold). These selected historical words can be mapped to the corresponding buckets in the multi-bucket hash table of the second storage space 311 according to the hash function, and compressed by accumulating the key-value vector. In this way, part of the memory of the first storage space 309 can be released to prepare for receiving new important words, and the semantic information of these words can be preserved while saving memory through the compression storage mechanism of the second storage space 311. They can also be reconstructed when needed later, avoiding the semantic gap caused by simple deletion.
[0049] At block 408, the target historical tokens in the third storage space are stored to the first storage space. For example, as shown in FIG. 3, the latest tokens stored in the third storage space 321 in the order of receiving, as the conversation progresses, some of the tokens will change from the latest interaction to tokens that are important to the long-term context, these tokens are the “target historical tokens” to be transferred. Transferring them from the third storage space 321 to the first storage space 309 can allow the first storage space 309 to continuously collect important tokens at different stages, ensuring that the model can efficiently call these long-term important semantic information when processing subsequent conversations, maintaining the understanding of the core context in complex conversation scenarios, and realizing the transition of token management from the latest interaction to long-term importance. Figure 3
[0050] At block 410, the new token is stored to the third storage space. For example, as shown in FIG. 3, after completing the previous space check, token transfer, and other operations, the new token 301 is stored to the third storage space 321 at block 410. The third storage space 321 is responsible for storing the latest interaction tokens. Regardless of the cumulative attention value of the historical tokens associated with the new token 301, as long as it is the latest received token, it will be stored here in order. This ensures that the model can obtain the latest interaction content in the first time when processing continuous conversations, which helps to understand the ongoing conversation context and provides a basis for generating responses that conform to real-time communication logic. By continuously storing new tokens 301 into the third storage space 321 and cooperating with subsequent space management processes (such as blocks 402-408), dynamic hierarchical management of tokens with different time sequences and different importance by the three storage spaces is realized, balancing memory utilization and semantic information integrity. Figure 3
[0051] In some embodiments, determining whether the first storage space reaches the second preset capacity can also not be a prerequisite for determining that the third storage space reaches the first preset capacity, that is, the model can determine whether the first storage space reaches the second preset capacity every preset time period or at a preset time point. When the first storage space reaches the second preset capacity, the historical tokens in the first storage space with a cumulative attention value lower than the threshold value are stored to the multi-bucket hash table of the second storage space. In this way, the capacity limit of the third storage space is eliminated, and the first storage space is actively checked by a preset time strategy, which flexibly adapts to various interaction scenarios and avoids storage overload.
[0052] Figure 4B A flowchart illustrating another method 400B of managing the storage space of a key-value cache according to some embodiments of the present disclosure is shown. At block 412, a first historical token with the lowest accumulated attention value in the first storage space and a second historical token with the highest accumulated attention value in the second storage space are determined. In the process of managing the storage space of the key-value cache, it can be determined whether there is a possibility of optimizing the storage structure by comparing the tokens with extreme attention values in the two spaces. For example, as shown in Figure 3 the first storage space 309 originally stores high-frequency activated tokens with high accumulated attention values. If there is a token with extremely low attention value, it may mean that its importance has decreased. The second storage space 311 stores low-frequency tokens. If there is a token with higher attention value, it may mean that its importance has increased, and the storage location needs to be reevaluated.
[0053] When it is determined that the attention value of the first historical token is higher than that of the second historical token, it means that the existing hierarchical storage strategy is still effective. At this time, there is no need to adjust the token location, which can maintain the fast access capability of the first storage space 309 to high-frequency tokens and ensure the compression storage efficiency of the second storage space 311 to low-frequency tokens, avoiding the consumption of computing resources due to meaningless location exchange. At this time, block 416 is executed to keep the storage structure unchanged.
[0054] When the attention value of the first historical token is lower than that of the second historical token, it means that there is a low-value token in the first storage space 309, and there is a high-value token in the second storage space 311. At this time, block 418 is executed to transfer the first historical token from the first storage space 309 to the multi-bucket hash of the second storage space 311, releasing the location of the first storage space 309 for more important tokens, while retaining the low-value token in a compressed form to avoid information loss.
[0055] At block 420, the second historical token is reconstructed from the second storage space and stored in the first storage space. After the transfer of the low-value token in block 418 is completed, the second historical token with the highest accumulated attention value in the second storage space 311 needs to be reconstructed and transferred to the first storage space 309 to be included in the fast access region again, ensuring that the model can directly call high-frequency activated tokens during inference.
[0056] Figure 5 A flowchart illustrating a method 500 of restoring a key-value matrix from a key-value cache according to some embodiments of the present disclosure is shown. At block 502, the storage location of each historical token is determined for a plurality of stored historical tokens. For example, as shown in Figure 3 when the key-value matrix restoration process is started in the model inference stage, the historical token storage location can be determined by querying the historical token index table or the metadata label, whether the historical token is stored in the first storage space 309, the second storage space 311, or the third storage space 321.
[0057] At block 504, when the historical token is stored in the first storage space or the third storage space, the key-value vector of the historical token is directly obtained. For example, as shown in FIG. 6, if it is determined that the historical token is located in the first storage space 309 or the third storage space 321, since the first storage space 309 stores the original key-value pair of the high-frequency active token in a dynamic array structure, and the third storage space 321 retains the latest token in the form of a queue in the order of reception, neither of them has compression processing on the data, the key-value vector can be directly obtained. Figure 3
[0058] At block 506, when the historical token is stored in the second storage space, the key-value vector of the historical token is reconstructed by the multi-bucket hash table. For example, as shown in FIG. 6, if the historical token is stored in the second storage space 311, the model can locate the corresponding bucket position in the multi-bucket hash table by using a hash function according to the index of the historical token, and then extract data from the multiple bucket positions to restore the original vector. Figure 3
[0059] At block 508, based on the obtained key-value vector and the reconstructed key-value vector, a key-value matrix is generated. The model can splice the directly obtained key-value vector of the first storage space or the third storage space and the reconstructed key-value vector of the second storage space by column to form a complete key-value matrix. For example, each column of the key matrix corresponds to the unique identifier or position encoding vector of a token, and the value matrix stores the semantic embedding vector thereof. Finally, the model completes full-scale reasoning under limited memory, achieving a balance between compression rate and accuracy.
[0060] Figure 6 A flowchart of a method 600 of storing historical tokens into multiple multi-bucket hash tables according to some embodiments of the present disclosure is shown. At block 602, for each hash function in a plurality of preset hash functions, the key mapping position and the word mapping position of the historical token in each multi-bucket hash table are determined. The model can preset a plurality of different hash functions (such as h1, h2, …, h r ), each of which corresponds to a multi-bucket hash table. For the historical token, the index thereof is calculated by each hash function to determine the bucket position in the corresponding multi-bucket hash table, i.e., the key mapping position and the word mapping position. For example, the hash function h i maps the token index to an integer in the range of [1, b], pointing to the h i bucket position of the i-th multi-bucket hash table, where b represents the number of bucket positions, i.e., the length of the multi-bucket hash table. In this way, a distributed storage structure of one token and multiple bucket positions is achieved.
[0061] At block 604, the key vector of the historical token is stored to the key mapping location. For each key mapping location, i.e. bucket position, determined by a hash function, the model can directly accumulate the key vector of the historical token to the existing key vector in that bucket position. For example, as shown in FIG. 3B, if the historical token 305-N is mapped to the bucket position 115 of the multi-bucket hash table 313 by the hash function h1, and the bucket position 115 has stored key vectors of other tokens, the key vector of the historical token 305-N will be added to the existing vectors in the bucket position 115 to form a new accumulated key vector. This accumulated storage manner allows multiple tokens to share the same bucket position, and realizes high-density storage in limited space. Figure 3
[0062] At block 606, the value vector of the historical token is stored to the value mapping location after being multiplied by a random sign. The model can preset a random sign generation function (e.g. g1, g2, …, gr) for each hash function to generate random signs of ±1, e.g. +1 or -1. For each value mapping location determined by a hash function, the value vector of the historical token is multiplied by the corresponding random sign and then accumulated to the existing value vector in that bucket position. For example, the value vector of the historical token 305-N is multiplied by the random sign generated by g2, e.g. +1, after being mapped to the bucket position of the multi-bucket hash table 317 by the hash function h2, and then accumulated to the existing value vector in that bucket position. The introduction of the random sign ensures the unbiasedness of the value vector accumulation, and provides mathematical guarantee for subsequent reconstruction.
[0063] Figure 7 A flowchart of a method 700 of reconstructing a historical token from multiple multi-bucket hash tables according to some embodiments of the present disclosure is shown. At block 702, based on the index of the historical token stored in the multi-bucket hash table and the multiple hash functions, multiple hash values are determined, which indicate the multiple mapping positions of the historical token in the multiple multi-bucket hash tables. When the model needs to reconstruct the historical token in the second storage space during the inference process of the model, the model can calculate r hash values (h1(J), h2(J), …, hr(J)) according to the unique index (e.g. position encoding J) of the historical token and the r preset hash functions, each of which corresponds to a bucket position of a multi-bucket hash table. For example, the index J is calculated by h1 to obtain h1(J) = 5, which points to the 5th bucket position of the multi-bucket hash table, and so on, to determine the specific mapping positions of the historical token in the r multi-bucket hash tables. r
[0064] At 704, based on the multiple hash values, the multiple key vectors and the multiple value vectors of the historical token are obtained from the multiple multi-bucket hash tables. According to the bucket positions determined in block 702, the accumulated key vectors and value vectors are extracted from the corresponding positions of each multi-bucket hash table. For the value vectors, the accumulated values stored in the bucket positions need to be divided by the random signs generated by the corresponding hash functions, i.e. the ±1 multiplied when storing. For example, as shown in FIG. 3B, the value vector of the historical token 305-N is multiplied by the random sign generated by g2, e.g. +1, after being mapped to the bucket position of the multi-bucket hash table 317 by the hash function h2, and then accumulated to the existing value vector in that bucket position. Figure 3 As shown, the key vector K1 and value vector V1 are obtained from bucket 5 of the multi-bucket hash table 313. V1 needs to be divided by g1(J) to offset the influence of random symbols during storage and obtain an unbiased estimate.
[0065] At box 706, the medians of multiple key vectors and multiple value vectors are taken to reconstruct historical lexical units. The medians of the key vectors and value vectors obtained from the r hash functions are taken respectively. For example, given the r key vectors K1, K2, ..., K... r The median is used as the key vector for reconstruction, and r value vectors V1, V2, ..., V r The median is used as the reconstructed value vector. Since the mapping results of multiple hash functions contain independent noise, taking the median can effectively reduce the variance, ensure that the reconstructed key-value vector is close to the original value, and achieve low-error lexical information restoration.
[0066] Figure 8 A block diagram of an apparatus 800 for storing key-value pairs according to some embodiments of the present disclosure is shown. Figure 8 As shown, device 800 includes a cumulative attention value determination module 802, configured to determine multiple cumulative attention values for multiple historical words in response to receiving a new word word with a key-value pair. Device 800 includes a first determination module 804, configured to determine whether the cumulative attention value of each historical word word is greater than a preset value. Furthermore, device 800 includes a first storage module 806, configured to store the historical word word in a first storage space of a key-value cache in response to the historical word word's cumulative attention value being greater than the preset value. Additionally, device 800 includes a second storage module 808, configured to store the historical word word in a second storage space of a key-value cache in response to the historical word word's cumulative attention value being less than or equal to the preset value. The second storage space includes multiple fixed-size multi-bucket hash tables.
[0067] Figure 9 A block diagram of a device 900 capable of implementing various embodiments of the present disclosure is shown. (See diagram for example.) Figure 9 As shown, device 900 includes a central processing unit (CPU) 901, which can perform various appropriate actions and processes according to computer program instructions stored in read-only memory (ROM) 902 or loaded from storage unit 908 into random access memory (RAM) 903. The RAM 903 can also store various programs and data required for the operation of device 900. The CPU / GPU 901, ROM 902, and RAM 903 are interconnected via bus 904. Input / output (I / O) interface 905 is also connected to bus 904. Although not shown in... Figure 9 As shown, device 900 may also include a coprocessor.
[0068] A number of the components in device 900 are connected to the I / O interface 905, including an input unit 906, such as a keyboard, mouse, etc.; an output unit 907, such as various types of displays, speakers, etc.; a storage unit 908, such as a disk, a CD, etc.; and a communication unit 909, such as a network card, a modem, a wireless communication transceiver, etc. The communication unit 909 allows device 900 to exchange information / data with other devices over a computer network, such as the Internet, and / or various telecommunication networks.
[0069] The various methods or processes described above can be performed by the CPU / GPU 901. For example, in some embodiments, the methods can be implemented as a computer software program tangibly embodied in a machine readable medium, such as the storage unit 908. In some embodiments, portions or all of the computer program can be loaded and / or installed onto device 900 via the ROM 902 and / or the communication unit 909. When the computer program is loaded onto the RAM 903 and executed by the CPU / GPU 901, one or more steps or actions of the methods or processes described above can be performed.
[0070] In some embodiments, the methods and processes described above can be implemented as a computer program product. The computer program product can include a computer readable storage medium having computer readable program instructions embodied therewith.
[0071] The computer readable storage medium can be a tangible device that can retain and store instructions for use by an instruction execution device. The computer readable storage medium can be, for example, but is not limited to, an electronic storage device, a magnetic storage device, an optical storage device, an electromagnetic storage device, a semiconductor storage device, or any suitable combination of the foregoing. More specific examples (a non-exhaustive list) of the computer readable storage medium include the following: a portable computer diskette, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or Flash memory), a static random access memory (SRAM), a portable compact disc read-only memory (CD-ROM), a digital versatile disk (DVD), a memory stick, a floppy disk, a mechanically encoded device such as punch-cards or punched tape, a
[0072] Computer readable program instructions described herein can be downloaded to respective computing / processing devices from a computer readable storage medium or to an external computer or external storage device via a network, for example, the Internet, a local area network, a wide area network and / or a wireless network. The network can comprise copper transmission cables, optical transmission fibers, wireless transmission, routers, firewalls, switches, gateway computers and / or edge servers. A network adapter card or network interface in each computing / processing device receives computer readable program instructions from the network and forwards the computer readable program instructions for storage in a computer readable storage medium within the respective computing / processing device.
[0073] Computer readable program instructions for carrying out operations of the present disclosure can be assembler instructions, instruction-set-architecture (ISA) instructions, machine instructions, machine dependent instructions, microcode, firmware instructions, state-setting data, or either source code or object code written in any combination of one or more programming languages, including an object oriented programming language and conventional procedural programming languages. The computer readable program instructions can execute entirely on the user's computer, partly on the user's computer, as a stand-alone software package, partly on the user's computer and partly on a remote computer or entirely on the remote computer or server. In the latter scenario, the remote computer can be connected to the user's computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or the connection can be made to an external computer (for example, through the Internet using an Internet Service Provider). In some embodiments, electronic circuitry including, for example, programmable logic circuitry, field-programmable gate array (FPGA), or programmable logic array (PLA) can execute the computer readable program instructions by utilizing state information of the computer readable program instructions to personalize the electronic circuitry, in order to perform aspects of the present disclosure.
[0074] These computer readable program instructions can be provided to a processor of a general purpose computer, special purpose computer, or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing apparatus, create means for implementing the functions / acts specified in the flowchart and / or block diagram block or blocks. These computer readable program instructions can also be stored in a computer readable storage medium that can include a non-transitory computer readable storage medium that can direct a computer, a programmable data processing apparatus, and / or other devices to function in a particular manner, such that the computer readable storage medium having instructions for causing an apparatus to implement various aspects of the functions / acts specified in the flowchart and / or block diagram block or blocks is produced.
[0075] The computer readable program instructions can also be loaded onto a computer, other programmable data processing apparatus, or other device to cause a series of operational steps to be performed on the computer, other programmable data processing apparatus or other device to produce a computer implemented process such that the instructions which execute on the computer, other programmable data processing apparatus, or other device implement the functions / acts specified in the flowchart and / or block diagram block or blocks.
[0076] The flow diagrams and the block diagrams in the drawings are presented to illustrate the architecture, functionality, and operation of possible implementations of devices, methods and computer program products according to various embodiments of the present disclosure. In this regard, each block in the flow diagrams and the block diagrams can represent a module, segment, or portion of instructions, which comprises one or more executable instructions for implementing the specified logic functions. In some alternative implementations, the functions noted in the blocks can occur out of the order noted in the figures. For example, two blocks shown in succession may, in fact, be executed substantially concurrently, or the blocks may sometimes be executed in the reverse order, depending upon the functionality involved. It will also be noted that each block of the block diagrams and / or flow diagrams, and combinations thereof, can be implemented by special purpose hardware-based systems that perform the specified functions or acts, or combinations of special purpose hardware and
[0077] Embodiments of the present disclosure have been described above, with the understanding that these embodiments are exemplary and not exhaustive, and are not limited to the embodiments disclosed. Many modifications and variations will be apparent to those of ordinary skill in the art without departing from the scope and spirit of the described embodiments. The selection of terms used herein is intended to best explain the principles of the embodiments, practical application, or technical improvements in the art made by the embodiments disclosed herein, or to enable other ordinary skilled persons in the art to understand the embodiments disclosed herein.
[0078] Some example implementations of the present disclosure are listed below.
[0079] Example 1. A method for storing key-value pairs, comprising:
[0080] determining, in response to receiving a new token of the key-value pair, a plurality of accumulated attention values of a plurality of historical tokens;
[0081] determining, for the plurality of historical tokens, whether an accumulated attention value of each historical token is greater than a preset value;
[0082] storing, in response to the accumulated attention value of the historical token being greater than the preset value, the historical token to a first storage space of a key-value cache; and
[0083] in response to the accumulated attention value of the historical token being less than or equal to the preset value, storing the historical token to a second storage space of the key-value cache, the second storage space comprising a plurality of fixed-size multi-bucket hash tables.
[0084] Example 2. The method of example 1, wherein the key-value cache further comprises a third storage space, the method further comprising:
[0085] storing, in the order of reception, a preset number of most recently received historical tokens and the new token to the third storage space.
[0086] Example 3. The method of any one of examples 1-2, the method further comprising:
[0087] in response to receiving the new token, determining whether the third storage space reaches a first preset capacity;
[0088] in response to the third storage space reaching the first preset capacity, storing a target historical token in the third storage space to the first storage space; and
[0089] storing the new token to the third storage space.
[0090] Example 4. The method of any one of examples 1-3, the method further comprising:
[0091] determining whether the first storage space reaches a second preset capacity; and
[0092] in response to the first storage space reaching the second preset capacity, storing a historical token in the first storage space with an accumulated attention value below a threshold to a multi-bucket hash table of the second storage space.
[0093] Example 5. The method of any one of examples 1-4, the method further comprising:
[0094] determining a first historical token in the first storage space with a lowest accumulated attention value and a second historical token in the second storage space with a highest accumulated attention value;
[0095] determining whether the accumulated attention value of the first historical token is lower than the accumulated attention value of the second historical token;
[0096] in response to the accumulated attention value of the first historical token being lower than the accumulated attention value of the second historical token, storing the first historical token to the multi-bucket hash table of the second storage space; and
[0097] reconstructing the second historical token from the second storage space and storing to the first storage space.
[0098] Example 6. The method of any one of examples 1-5, further comprising:
[0099] determining, for a plurality of stored historical tokens, a storage location of each historical token;
[0100] directly obtaining a key-value vector of the historical token when the historical token is stored in the first storage space or the third storage space;
[0101] reconstructing the key-value vector of the historical token through the multi-bucket hash table when the historical token is stored in the second storage space;
[0102] generating a key-value matrix based on the obtained key-value vectors and the reconstructed key-value vector.
[0103] Example 7. The method of any one of examples 1-6, wherein determining, in response to receiving a new token of the key-value pair, a plurality of accumulated attention values of a plurality of historical tokens comprises:
[0104] determining, for the plurality of historical tokens, an attention value of the new token to each historical token; and
[0105] accumulating the attention value into a historical accumulated attention value of the corresponding historical token to determine the plurality of accumulated attention values.
[0106] Example 8. The method of any one of examples 1-7, wherein storing the historical token to the second storage space of the key-value cache comprises:
[0107] determining, for each hash function of a plurality of preset hash functions, a key mapping location and a token mapping location of the historical token in each multi-bucket hash table;
[0108] storing a key vector of the historical token to the key mapping location; and
[0109] storing a value vector of the historical token to the token mapping location after multiplying the value vector by a random sign.
[0110] Example 9. The method of any one of examples 1-8, wherein the historical token comprises a first historical token and a second historical token, the first historical token and the second historical token corresponding to a same key mapping location and a same value mapping location, the method further comprising:
[0111] accumulatively storing a key vector of the first historical token and a key vector of the second historical token; and
[0112] The value vector of the first historical token and the value vector of the second historical token are multiplied by the random symbol and then accumulated and stored.
[0113] Example 10. The method of any one of examples 1-9, the apparatus further comprising:
[0114] determining, based on the indices of the historical tokens stored in the multi -bucket hash table and the plurality of hash functions, a plurality of hash values indicating a plurality of mapping locations of the historical tokens in a plurality of the multi -bucket hash tables;
[0115] obtaining, based on the plurality of hash values, a plurality of key vectors and a plurality of value vectors of the historical tokens from a plurality of the multi -bucket hash tables; and
[0116] respectively taking a median of the plurality of key vectors and the plurality of value vectors to reconstruct the historical tokens.
[0117] Example 11. An apparatus for storing key-value pairs, comprising:
[0118] a cumulative attention value determination module configured to determine, in response to receiving a new token of the key-value pair, a plurality of cumulative attention values of a plurality of historical tokens;
[0119] a first determination module configured to determine, for the plurality of historical tokens, whether a cumulative attention value of each historical token is greater than a preset value;
[0120] a first storage module configured to, in response to the cumulative attention value of the historical token being greater than the preset value, store the historical token to a first storage space of a key-value cache; and
[0121] a second storage module configured to, in response to the cumulative attention value of the historical token being less than or equal to the preset value, store the historical token to a second storage space of the key-value cache, the second storage space comprising a plurality of fixed-size multi -bucket hash tables.
[0122] Example 12. The apparatus of example 11, wherein the key-value cache further comprises a third storage space, the apparatus further comprising:
[0123] a third storage module configured to store, in a receiving order, a preset number of the historical tokens most recently received and the new token to the third storage space.
[0124] Example 13. The apparatus of any one of examples 11-12, the apparatus further comprising:
[0125] a second determination module configured to determine, in response to receiving the new token, whether the third storage space reaches a first preset capacity;
[0126] a fourth storage module configured to store a target historical token in the third storage space to the first storage space in response to the third storage space reaching the first preset capacity; and
[0127] a fifth storage module configured to store the new token to the third storage space.
[0128] Example 14. The apparatus according to any one of examples 11-13, further comprising:
[0129] a third determination module configured to determine whether the first storage space reaches a second preset capacity; and a sixth storage module configured to store a historical token in the first storage space with a cumulative attention value lower than a threshold value to a multi-bucket hash table of the second storage space in response to the first storage space reaching the second preset capacity.
[0130] Example 15. The apparatus according to any one of examples 11-14, further comprising:
[0131] a fourth determination module configured to determine a first historical token with a lowest cumulative attention value in the first storage space and a second historical token with a highest cumulative attention value in the second storage space;
[0132] a fifth determination module configured to determine whether the cumulative attention value of the first historical token is lower than the cumulative attention value of the second historical token;
[0133] a seventh storage module configured to store the first historical token to the multi-bucket hash table of the second storage space in response to the cumulative attention value of the first historical token being lower than the cumulative attention value of the second historical token; and
[0134] a first reconstruction module configured to reconstruct the second historical token from the second storage space and store to the first storage space.
[0135] Example 16. The apparatus according to any one of examples 11-15, further comprising:
[0136] a sixth determination module configured to determine a storage location of each historical token for a plurality of stored historical tokens;
[0137] a direct acquisition module configured to directly acquire a key-value vector of the historical token when the historical token is stored in the first storage space or the third storage space;
[0138] a second reconstruction module configured to reconstruct a key vector of the history token by the multi-bucket hash table when the history token is stored in the second storage space;
[0139] a matrix generation module configured to generate a key-value matrix based on the obtained key vector and the reconstructed key vector.
[0140] Example 17. The apparatus of any one of examples 11-16, wherein the cumulative attention value determination module comprises:
[0141] a seventh determination module configured to determine, for the plurality of history tokens, an attention value of the new token to each history token; and
[0142] an accumulation module configured to accumulate the attention value into a history cumulative attention value of the corresponding history token to determine the plurality of cumulative attention values.
[0143] Example 18. The apparatus of any one of examples 11-17, wherein the second storage module comprises:
[0144] an eighth determination module configured to determine, for each hash function in a plurality of preset hash functions, a key mapping position and a token mapping position of the history token in each multi-bucket hash table;
[0145] an eighth storage module configured to store a key vector of the history token to the key mapping position; and
[0146] a ninth storage module configured to store a value vector of the history token multiplied by a random sign to the token mapping position.
[0147] Example 19. The apparatus of any one of examples 11-18, wherein the history token comprises a first history token and a second history token, the first history token and the second history token corresponding to a same key mapping position and a same value mapping position, the apparatus further comprising:
[0148] a tenth storage module configured to accumulate store a key vector of the first history token and a key vector of the second history token; and
[0149] an eleventh storage module configured to accumulate store a value vector of the first history token and a value vector of the second history token multiplied by the random sign.
[0150] Example 20. The apparatus of any one of examples 11-19, the apparatus further comprising:
[0151] a ninth determining module configured to determine a plurality of hash values based on indexes of historical tokens stored in the plurality of bucketed hash tables and the plurality of hash functions, the plurality of hash values indicating a plurality of mapping locations of the historical tokens in the plurality of bucketed hash tables;
[0152] a vector obtaining module configured to obtain a plurality of key vectors and a plurality of value vectors of the historical tokens from the plurality of bucketed hash tables based on the plurality of hash values; and
[0153] a token reconstructing module configured to take a median of the plurality of key vectors and the plurality of value vectors respectively to reconstruct the historical tokens.
[0154] Example 21. An electronic device, comprising:
[0155] a processor; and
[0156] a memory coupled with the processor, the memory having stored therein instructions that, when executed by the processor, cause the electronic device to perform acts comprising:
[0157] in response to receiving a new token of the key-value pair, determining a plurality of accumulated attention values of a plurality of historical tokens;
[0158] for the plurality of historical tokens, determining whether an accumulated attention value of each historical token is greater than a preset value;
[0159] in response to the accumulated attention value of the historical token being greater than the preset value, storing the historical token to a first storage space of a key-value cache; and
[0160] in response to the accumulated attention value of the historical token being less than or equal to the preset value, storing the historical token to a second storage space of the key-value cache, the second storage space comprising a plurality of bucketed hash tables of fixed sizes.
[0161] Example 22. The electronic device of example 21, wherein the key-value cache further comprises a third storage space, the acts further comprising:
[0162] storing a preset number of most recently received historical tokens and the new token in the third storage space in a receiving order.
[0163] Example 23. The electronic device of any one of examples 21-22, the acts further comprising:
[0164] in response to receiving the new token, determining whether the third storage space reaches a first preset capacity;
[0165] in response to the third storage space reaching the first preset capacity, storing target historical tokens in the third storage space to the first storage space; and
[0166] storing the new token to the third storage space.
[0167] Example 24. The electronic device of any of examples 21-23, the actions further comprising:
[0168] determining whether the first storage space reaches a second preset capacity; and
[0169] in response to the first storage space reaching the second preset capacity, storing historical tokens in the first storage space with accumulated attention values below a threshold to a multi-bucket hash table of the second storage space.
[0170] Example 25. The electronic device of any of examples 21-24, the actions further comprising:
[0171] determining a first historical token with a lowest accumulated attention value in the first storage space and a second historical token with a highest accumulated attention value in the second storage space;
[0172] determining whether the accumulated attention value of the first historical token is lower than the accumulated attention value of the second historical token;
[0173] in response to the accumulated attention value of the first historical token being lower than the accumulated attention value of the second historical token, storing the first historical token to the multi-bucket hash table of the second storage space; and
[0174] reconstructing the second historical token from the second storage space and storing to the first storage space.
[0175] Example 26. The electronic device of any of examples 21-25, the actions further comprising:
[0176] for a plurality of stored historical tokens, determining a storage location of each historical token;
[0177] when the historical token is stored in the first storage space or the third storage space, directly obtaining a key-value vector of the historical token;
[0178] when the historical token is stored in the second storage space, reconstructing a key-value vector of the historical token through the multi-bucket hash table;
[0179] based on the obtained key-value vector and the reconstructed key-value vector, generating a key-value matrix.
[0180] Example 27. The electronic device of any of examples 21-25, wherein responsive to receiving the new token of the key-value pair, determining a plurality of accumulated attention values for a plurality of historical tokens comprises:
[0181] determining, for the plurality of historical tokens, an attention value of the new token for each historical token; and
[0182] accumulating the attention values into historical accumulated attention values for corresponding historical tokens to determine the plurality of accumulated attention values.
[0183] Example 28. The electronic device of any of examples 21-27, wherein storing the historical token to the second storage space of the key-value cache comprises:
[0184] determining, for each hash function of a plurality of preset hash functions, a key mapping location and a word mapping location of the historical token in each multi-bucket hash table;
[0185] storing a key vector of the historical token to the key mapping location; and
[0186] storing a value vector of the historical token to the word mapping location after multiplying the value vector by a random sign.
[0187] Example 29. The electronic device of any of examples 21-28, wherein the historical token comprises a first historical token and a second historical token, the first historical token and the second historical token corresponding to a same key mapping location and a same value mapping location, the actions further comprising:
[0188] accumulatively storing a key vector of the first historical token and a key vector of the second historical token; and
[0189] accumulatively storing a value vector of the first historical token and a value vector of the second historical token after multiplying the value vector by the random sign.
[0190] Example 30. The electronic device of any of examples 21-29, the actions further comprising:
[0191] determining, based on an index of the historical token stored in the multi-bucket hash table and the plurality of hash functions, a plurality of hash values, the plurality of hash values indicating a plurality of mapping locations of the historical token in a plurality of the multi-bucket hash tables;
[0192] obtaining, based on the plurality of hash values, a plurality of key vectors and a plurality of value vectors of the historical token from a plurality of the multi-bucket hash tables; and
[0193] respectively taking a median of the plurality of key vectors and the plurality of value vectors to reconstruct the historical token.
[0194] Example 31. A computer-readable storage medium having stored thereon computer- executable instructions, wherein the computer-executable instructions are executed by a processor to implement the method of any of examples 1-10.
[0195] Example 32. A computer program product tangibly stored in a computer-readable medium and comprising computer-executable instructions that, when executed by a device, cause the device to perform the method of any of examples 1-10.
[0196] Although the present disclosure has been described in some detail with specific reference to structure features and / or method logical actions, it is understood that the subject matter defined in the appended claims is not necessarily limited to the particular features or acts described above. Rather, the particular features and acts described above are merely illustrative examples of implementing the claims.
Claims
1.A method for storing key-value pairs, comprising: determining a plurality of accumulated attention values of a plurality of historical tokens in response to receiving a new token of a key-value pair; determining whether an accumulated attention value of each historical token is greater than a preset value for the plurality of historical tokens; storing the historical token to a first storage space of a key-value cache in response to the accumulated attention value of the historical token being greater than the preset value; and storing the historical token to a second storage space of the key-value cache in response to the accumulated attention value of the historical token being less than or equal to the preset value, the second storage space comprising a plurality of fixed-size multi-bucket hash tables. 2.The method of claim 1, wherein the key-value cache further comprises a third storage space, the method further comprising: storing a preset number of latest received historical tokens and the new token in the third storage space in a receiving order. 3.The method of claim 2, the method further comprising: determining whether the third storage space reaches a first preset capacity in response to receiving the new token; storing a target historical token in the third storage space to the first storage space in response to the third storage space reaching the first preset capacity; and storing the new token to the third storage space. 4.The method of claim 3, the method further comprising: determining whether the first storage space reaches a second preset capacity; and storing historical tokens in the first storage space with accumulated attention values below a threshold to the multi-bucket hash tables of the second storage space in response to the first storage space reaching the second preset capacity. 5.The method of claim 4, the method further comprising: determining a first historical token with a lowest accumulated attention value in the first storage space and a second historical token with a highest accumulated attention value in the second storage space; determining whether the accumulated attention value of the first historical token is lower than the accumulated attention value of the second historical token; storing the first historical token to the multi-bucket hash tables of the second storage space in response to the accumulated attention value of the first historical token being lower than the accumulated attention value of the second historical token; and reconstructing the second historical token from the second storage space and storing to the first storage space. 6.The method of claim 5, the method further comprising: determining a storage location of each historical token for a plurality of stored historical tokens; directly obtaining a key-value vector of the historical token when the historical token is stored in the first storage space or the third storage space; reconstructing a key-value vector of the historical token through the multi-bucket hash tables when the historical token is stored in the second storage space; and generating a key-value matrix based on the obtained key-value vectors and the reconstructed key-value vector. 7.The method of claim 1, wherein determining a plurality of accumulated attention values of a plurality of historical tokens in response to receiving a new token of a key-value pair comprises: determining an attention value of the new token to each historical token for the plurality of historical tokens; and accumulate the attention value into a historical accumulated attention value of a corresponding historical token to determine a current one of the plurality of accumulated attention values. 8.The method of claim 1, wherein storing the historical token to the second storage space of the key-value cache comprises: determining, for each of a plurality of preset hash functions, a key mapping location and a word mapping location of the historical token in each of the plurality of bucketed hash tables; storing a key vector of the historical token to the key mapping location; and storing a value vector of the historical token to the word mapping location after being multiplied by a random sign. 9.The method of claim 8, wherein the historical token comprises a first historical token and a second historical token, the first historical token and the second historical token corresponding to a same key mapping location and a same value mapping location, the method further comprising: accumulatively storing a key vector of the first historical token and a key vector of the second historical token; and accumulatively storing a value vector of the first historical token and a value vector of the second historical token after being multiplied by the random sign. 10.The method of claim 9, the method further comprising: determining, based on an index of the historical token stored in the plurality of bucketed hash tables and the plurality of hash functions, a plurality of hash values, the plurality of hash values indicating a plurality of mapping locations of the historical token in a plurality of the plurality of bucketed hash tables; obtaining, based on the plurality of hash values, a plurality of key vectors and a plurality of value vectors of the historical token from a plurality of the plurality of bucketed hash tables; and respectively taking a median of the plurality of key vectors and the plurality of value vectors to reconstruct the historical token. 11.An apparatus for storing key-value pairs, comprising: an accumulated attention value determining module configured to determine, in response to receiving a new token of the key-value pair, a plurality of accumulated attention values of a plurality of historical tokens; a first determining module configured to determine, for the plurality of historical tokens, whether an accumulated attention value of each historical token is greater than a preset value; a first storing module configured to, in response to the accumulated attention value of the historical token being greater than the preset value, store the historical token to a first storage space of a key-value cache; and a second storing module configured to, in response to the accumulated attention value of the historical token being less than or equal to the preset value, store the historical token to a second storage space of the key-value cache, the second storage space comprising a plurality of bucketed hash tables of a fixed size. 12.An electronic device, comprising: a processor; and a memory coupled with the processor, the memory having stored therein instructions that, when executed by the processor, cause the electronic device to perform the method of any one of claims 1-10. 13.A computer-readable storage medium having computer-executable instructions stored therein, wherein the computer-executable instructions are executed by a processor to implement the method of any one of claims 1-10.