KV cache processing method and apparatus for large language model, and data processing method and apparatus
By updating the window length and attention map update KV cache, the memory overflow problem of large language models when processing ultra-long text is solved, and the processing efficiency and text processing capabilities are improved.
Patent Information
- Application Number
- PCT/CN2024/130182
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2023-12-30
- Filing Date
- 2024-11-06
- Publication Date
- 2025-07-03
AI Technical Summary
Large language models can easily lead to an increase in the KV cache length when processing ultra-long text, resulting in memory overflow, affecting processing speed and efficiency.
When decoding the current semantic unit, if the current KV cache length is greater than or equal to the first window length, the update window length is the second window length, and the KV cache is updated according to the attention diagram so that its length does not exceed the preset window maximum length.
It avoids memory overflow, improves the processing efficiency of large language models, and expands the text processing capabilities of the model.
Smart Images

Figure CN2024130182_03072025_PF_FP_ABST
Abstract
Description
KV cache processing method, data processing method and device for large language model Technical Field
[0001] The present application belongs to the field of artificial intelligence technology, and in particular relates to a KV cache processing method, a data processing method and a device for a large language model.
[0002] This application claims priority to the Chinese patent application filed with the China Patent Office on December 30, 2023, with application number 202311865545.4 and invention name “KV cache processing method, data processing method and device for large language model”, the entire contents of which are incorporated by reference into this application. Background Art
[0003] Large Language Models (LLMs), also known as large language models, are artificial intelligence models designed to understand, generate, and process human language. LLMs are trained on large amounts of text data and can perform a wide range of tasks, such as text summarization, text translation, and sentiment analysis.
[0004] LLM models are typically based on deep learning architectures, are large in scale, and have numerous parameters. During training and application, they generate a large amount of intermediate vector cache (i.e., KVCache). The length of the KVCache increases with the length of the generated text. When processing extremely long text, this can cause the LLM model to slow down and even cause memory overflow. Technical issues
[0005] The embodiments of the present application provide a KV cache processing method, a data processing method, and an apparatus for a large language model, which avoid memory overflow and improve the processing efficiency of the large language model.
[0006] In a first aspect, an embodiment of the present application provides a KV cache processing method for a large language model, including:
[0007] When decoding the current semantic unit, if the length of the current KV cache is greater than or equal to the first window length, obtaining the attention map of each processing layer in the large language model after decoding the previous semantic unit before the current semantic unit;
[0008] Updating the first window length to a second window length according to the attention maps of all processing layers; wherein the first window length and the second window length are both less than or equal to a preset maximum window length;
[0009] According to the first attention map and the second window length, the current KV cache is updated to the target KV cache; wherein the first attention map and the current KV cache correspond to the same processing layer.
[0010] In a second aspect, an embodiment of the present application provides a data processing method for a large language model, including:
[0011] Receive input text;
[0012] Generate a semantic unit and a current KV cache corresponding to the semantic unit according to the input text;
[0013] Processing the current KV cache using the method described in the first aspect above, and updating the current KV cache to the target KV cache;
[0014] The semantic unit is decoded according to the target KV cache to generate output text.
[0015] In a third aspect, an embodiment of the present application provides a large language model processing device, including:
[0016] An acquisition module is configured to, when decoding a current semantic unit, obtain an attention map of each processing layer in the large language model after decoding a previous semantic unit before the current semantic unit if the length of the current KV cache is greater than or equal to the first window length;
[0017] A window length updating module, configured to update the first window length to a second window length according to the attention maps of all processing layers; wherein the first window length and the second window length are both less than or equal to a preset maximum window length;
[0018] A KV cache update module is used to update the current KV cache to a target KV cache according to the first attention map and the second window length; wherein the first attention map and the current KV cache correspond to the same processing layer.
[0019] In a fourth aspect, an embodiment of the present application provides a large language model processing device, including:
[0020] A receiving module, used for receiving input text;
[0021] A first generating module is configured to generate a semantic unit and a current KV cache corresponding to the semantic unit according to the input text;
[0022] A KV cache processing module, configured to process the current KV cache using the method described in the first aspect above, and update the current KV cache to a target KV cache;
[0023] The second generating module is used to perform decoding processing on the semantic unit according to the target KV cache to generate output text.
[0024] In the fifth aspect, an embodiment of the present application provides a processing device for a large language model, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor implements the method described in the first aspect above when executing the computer program.
[0025] In a sixth aspect, an embodiment of the present application provides a processing device for a large language model, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor implements the method described in the second aspect above when executing the computer program.
[0026] In the seventh aspect, an embodiment of the present application provides a large language model, including multiple processing layers, each processing layer corresponds to a KV cache, and the processing layer is used to execute the method described in the first aspect above, or the large language model is used to execute the method described in the first or second aspect above.
[0027] In an eighth aspect, an embodiment of the present application provides a computer-readable storage medium, wherein the computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the method described in the first aspect above is implemented.
[0028] In a ninth aspect, an embodiment of the present application provides a computer-readable storage medium, wherein the computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the method described in the second aspect above is implemented.
[0029] The KV cache processing method for a large language model provided in this application is as follows: when the large language model decodes the current semantic unit, if the current KV cache length is greater than or equal to the first window length determined when decoding the previous semantic unit, the window length is updated to the second window length, and the current KV cache is updated based on the updated second window length and the attention map, so that the KV cache length does not exceed the preset maximum window length. This avoids memory overflow, reduces the computational complexity of the large language model, and improves the processing efficiency of the large language model. BRIEF DESCRIPTION OF THE DRAWINGS
[0030] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the following briefly introduces the drawings required for use in the embodiments or descriptions of the prior art. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without any creative work.
[0031] FIG1 is a schematic diagram of an application scenario of the LLM model provided in an embodiment of the present application;
[0032] FIG2 is a schematic diagram of a working principle of the Transformer Block in FIG1 ;
[0033] 3A and 3B are a set of schematic diagrams of a KV cache provided in an embodiment of the present application;
[0034] FIG4 is a flow chart of a KV cache processing method for a large language model provided in an embodiment of the present application;
[0035] FIG5 is a schematic diagram of a current KV cache and a target KV cache provided in an embodiment of the present application;
[0036] FIG6 is another schematic diagram of the current KV cache and the target KV cache provided in an embodiment of the present application;
[0037] FIG7 is a schematic diagram of KV vector re-encoding according to an embodiment of the present application;
[0038] FIG8 is a flow chart of a data processing method for a large language model provided in an embodiment of the present application;
[0039] FIG9 is a schematic structural diagram of a large language model processing device provided in an embodiment of the present application;
[0040] FIG10 is another structural diagram of a large language model processing device provided in an embodiment of the present application;
[0041] FIG11 is another structural diagram of a large language model processing device provided in an embodiment of the present application. Modes for Carrying Out the Invention
[0042] In the following description, specific details such as specific system structures and techniques are provided for purposes of illustration rather than limitation to facilitate a thorough understanding of the embodiments of the present application. However, it will be apparent to those skilled in the art that the present application may be implemented in other embodiments without these specific details. In other cases, detailed descriptions of well-known systems, devices, circuits, and methods are omitted to avoid obscuring the description of the present application with unnecessary detail.
[0043] It should be understood that when used in the present specification and the appended claims, the term "comprising" indicates the presence of described features, integers, steps, operations, elements and / or components, but does not preclude the presence or addition of one or more other features, integers, steps, operations, elements, components and / or collections thereof.
[0044] It will also be understood that the term "and / or" used in this specification and the appended claims refers to and includes any and all possible combinations of one or more of the associated listed items.
[0045] As used in this specification and the appended claims, the term "if" can be interpreted as "when" or "upon" or "in response to determining" or "in response to detecting," depending on the context. Similarly, the phrase "if it is determined" or "if [described condition or event] is detected" can be interpreted as meaning "upon determination" or "in response to determining" or "upon detection of [described condition or event]" or "in response to detecting [described condition or event]," depending on the context.
[0046] In addition, in the description of the present application specification and the appended claims, the terms "first", "second", "third", etc. are only used to distinguish the descriptions and cannot be understood as indicating or implying relative importance.
[0047] References to "one embodiment" or "some embodiments" in this specification mean that a particular feature, structure, or characteristic described in conjunction with that embodiment is included in one or more embodiments of the present application. Thus, phrases such as "in one embodiment," "in some embodiments," "in other embodiments," and "in other embodiments" appearing in various places in this specification do not necessarily refer to the same embodiment, but rather mean "one or more but not all embodiments," unless otherwise specifically emphasized. The terms "including," "comprising," "having," and variations thereof all mean "including but not limited to," unless otherwise specifically emphasized.
[0048] The KV cache processing method and data processing method for a large language model provided in this application are applied to large language models. This application does not limit the implementation of large language models, for example, the Transformer model. The Transformer model is a neural network model based on an attention mechanism, which is used to process sequence data. It captures contextual information from the entire sequence through the attention mechanism to obtain prediction results, and has better parallel performance and shorter training time. For ease of explanation, this application uses the LLM model as an example of the Transformer model.
[0049] For example, FIG1 is a schematic diagram of an application scenario of the LLM model provided in an embodiment of the present application. As shown in FIG1 , the LLM model includes multiple transformer blocks or transformer layers (Transformer Block), marked as Transformer Block 0 to Transformer Block N. The input text is "I love Beijing Tiananmen", [CLS] is classification, which indicates the beginning of the sentence. The output text is generated based on the previous text sequence. After processing by the LLM model, the text "The sun rises over Tiananmen" will be output to achieve text continuation. In FIG1 , the Transformer Block can cache the intermediate results generated by the previous input to save computing time. Such cached results are usually called KV cache (kv cache or KV Cache or KVCache). As the length of the generated text increases, the KV cache will become larger and larger. In the processing scenario of ultra-long texts, the amount of model calculation will become larger and larger, resulting in a decrease in the model processing speed and even memory overflow.
[0050] Currently, it is possible to modify the position encoding scheme to encode more position information in the KV cache, thereby achieving the effect of being able to process longer contexts. However, this requires retraining the model and cannot process infinitely long text due to the memory limitations of hardware processing devices.
[0051] The present application provides a KV cache processing method and data processing method for a large language model. When decoding a current semantic unit, if the length of the KV cache is greater than or equal to the window length determined when decoding the previous semantic unit, the window length is updated, and the KV cache is updated based on the updated window length, so that the length of the KV cache does not exceed the preset maximum window length. Thus, regardless of the length of the generated text, the length of the KV cache is limited to a certain range, avoiding memory overflow, and improving the processing efficiency of the large language model. The text length processed by the large language model is not limited, thus expanding the model's text processing capabilities.
[0052] The concepts involved in this application are explained below.
[0053] 1. Semantic unit (token)
[0054] In natural language processing (NLP), the smallest semantic unit in a text can be called a subword or token. Tokenization can be used to segment the text to obtain tokens. This embodiment of the application does not limit the segmentation method. For example, a sentence can be segmented into several words, each of which is a token. For example, the sentence "I love Tiananmen, Beijing" can be segmented into three tokens: I, love, and Tiananmen, Beijing.
[0055] As shown in Figure 1, each token needs to be processed in sequence by Transformer Block 0 to Transformer Block N to complete encoding and decoding.
[0056] 2. Encoding and decoding
[0057] During text processing, the process of converting subwords or tokens into numerical vectors is called encoding. The goal of encoding is to map discrete, unordered tokens into a continuous, ordered vector space, facilitating computation and learning for large language models. For example, "Beijing Tiananmen" can be encoded as the numerical vector [1, 2, 3, 4, 5, 6, 7, 8, 4]. This numerical vector can be used to generate a set of QKV vectors, where Q is the query vector, K is the key vector, and V is the value vector.
[0058] Correspondingly, each numerical code is replaced with its corresponding subword, and then adjacent subwords are merged into the longest matching word to obtain a text, which is called decoding. For example, the numerical vector [1, 2, 3, 4, 5, 6, 7, 8, 4] can be decoded into the word "Beijing Tiananmen".
[0059] 3. KV cache, KV cache length, KV vector, vector location information
[0060] In Figure 1, the LLM model includes multiple Transformer Blocks, and the processing of each Transformer Block is similar. For the convenience of explanation, the embodiment of the present application uses the processing of one Transformer Block as an example for explanation.
[0061] Figure 2 illustrates the operating principle of the Transformer Block in Figure 1. As shown in Figure 2, the Transformer Block includes a position module (POS) and an attention module (atten). The position module (POS) encodes the position of the key-value vector, while the attention module (atten) calculates the attention value, which is stored as an attention map. In Figure 2, the Transformer Block is currently processing the nth input token. The n-1 tokens preceding it have already been decoded. The current KV cache of the Transformer Block is obtained after the n-1th token is decoded. As shown in Figure 2, the current KV cache includes the following KV vectors: (k1, v1), (k2, v2), (k3, v3), etc. After the nth token is decoded, the KV vector (kn, vn) corresponding to the nth token is added to the KV cache to facilitate decoding of the next token (i.e., the n+1th token). As you can see, the KV cache typically grows larger as the length of the generated text increases.
[0062] It should be noted that each Transformer Block in the LLM model has a corresponding KV cache. The KV cache contains KV vectors and their vector position information, which indicates the position of the KV vector in the KV cache. The KV cache length refers to the number of KV vectors included in the KV cache. The KV cache length of each Transformer Block is the same.
[0063] For example, Figures 3A and 3B are schematic diagrams of a KV cache provided in an embodiment of the present application. The large language model includes three Transformer Blocks, represented as Transformer Block 1 to Transformer Block 3, or as layer 1 to layer 3. Each Transformer Block corresponds to a KV cache.
[0064] As shown in Figure 3A, the large language model is currently decoding the fourth token. The KV cache for Transformer Block 1 is {(k1,v1), (k2,v2), (k3,v3)}, the KV cache for Transformer Block 2 is {(k1,v1), (k2,v2), (k3,v3)}, and the KV cache for Transformer Block 3 is {(k1,v1), (k2,v2), (k3,v3)}. The KV cache length is 3. In each KV cache, the vector position information for the KV vectors (k1,v1), (k2,v2), and (k3,v3) is 1, 2, and 3, indicating the position of the KV vector in the KV cache.
[0065] Similarly, as shown in Figure 3B, the large language model is currently decoding the fifth token. The KV cache for Transformer Block 1 is {(k1, v1), (k2, v2), (k3, v3), (k4, v4)}, the KV cache for Transformer Block 2 is {(k1, v1), (k2, v2), (k3, v3), (k4, v4)}, and the KV cache for Transformer Block 3 is {(k1, v1), (k2, v2), (k3, v3), (k4, v4)}. The KV cache length is 4. In each KV cache, the vector positions of the KV vectors (k1, v1), (k2, v2), (k3, v3), and (k4, v4) are 1, 2, 3, and 4, respectively.
[0066] 4. Attention map, attention value, and attention location information
[0067] As shown in Figure 2, after the Transformer Block decodes the current input token, it can save the attention value of the current decoded token into an attention map. The attention map includes multiple attention values and the attention position information of the attention values. The attention position information is used to indicate the position of the attention value in the attention map. For example, the attention map is (0.2, 0.06, 0.6, 0.1, 0.01, 0.03), that is, the attention map includes 6 attention values, and the corresponding attention position information is 1, 2, 3, 4, 5, and 6.
[0068] It should be noted that after decoding the current token, the number of attention values in the attention map is equal to the number of KV vectors in the KV cache.
[0069] 5. LLM pre-training maximum length (max_seq_len), preset window minimum length (min_len), preset window maximum length (max_len)
[0070] The preset window minimum length (min_len) defines the minimum length of the KV cache, and the preset window maximum length (max_len) defines the maximum length of the KV cache. Three parameters need to be met:
[0071] 0 < min_len <= max_len <=max_seq_len
[0072] The selection of max_len and max_seq_len is usually related to the video memory size of the processing device used by the LLM model. Setting max_len can ensure that the LLM model does not overflow memory during inference. Setting min_len can minimize attenuation and ensure the inference effect of the LLM model.
[0073] Optional, min_len=0.5×max_len.
[0074] The KV cache processing method, data processing method and device for the large language model provided in this application are described in detail below with reference to the accompanying drawings.
[0075] FIG4 is a flow chart of a KV cache processing method for a large language model provided in an embodiment of the present application. The KV cache processing method for a large language model provided in this embodiment can be executed by a processing device for a large language model. As shown in FIG4 , the KV cache processing method for a large language model may include:
[0076] S401. When decoding the current semantic unit, if the length of the current KV cache is greater than or equal to the first window length, obtain the attention map of each processing layer in the large language model after decoding the previous semantic unit before the current semantic unit.
[0077] For a large language model, multiple semantic units are processed sequentially. Each semantic unit is processed sequentially by different processing layers (Transformer Blocks) within the large language model. For ease of distinction, the semantic unit to be decoded by the large language model at the current moment is called the current semantic unit, and the semantic unit immediately preceding the current semantic unit that has already been decoded is called the previous semantic unit.
[0078] The current KV cache can be understood as the KV cache of any processing layer in the large language model, or the KV cache of each processing layer in the large language model. It is understood that the KV cache length of each processing layer is the same, and the KV cache of each processing layer needs to be similarly processed according to the method provided in this embodiment. This embodiment uses one processing layer as an example for explanation.
[0079] When decoding the current semantic unit, the current KV cache is the KV cache obtained after decoding the previous semantic unit. If the length of the current KV cache is greater than or equal to the first window length, the KV cache may be too long. Therefore, it is necessary to obtain the attention map of each processing layer obtained after decoding the previous semantic unit to re-determine the optimal length of the KV cache.
[0080] S402: Update the first window length to a second window length according to the attention maps of all processing layers, wherein both the first window length and the second window length are less than or equal to a preset maximum window length.
[0081] The attention map includes attention values, the size of which can reflect the importance of different KV vectors in the KV cache. Based on the attention maps of all processing layers in the large language model, the first window length is updated to the second window length, resulting in the optimal length of the KV cache, namely the second window length. This ensures that the KV cache is not too large while also taking into account the importance of the KV vectors. Moreover, both the first window length and the second window length are less than or equal to the preset maximum window length, so that the number of KV vectors in the KV cache does not exceed a certain range, avoiding an unlimited increase in the model's computational complexity.
[0082] The first window length refers to the window length of the KV cache determined when decoding the previous semantic unit, and the second window length refers to the window length of the KV cache determined again when decoding the current semantic unit if the length of the current KV cache is greater than or equal to the first window length. This embodiment does not limit the names of the first window length and the second window length.
[0083] S403: Update the current KV cache to the target KV cache according to the first attention map and the second window length, wherein the first attention map and the current KV cache correspond to the same processing layer.
[0084] Since the second window length of the KV cache has been redefined, and the size of the attention value can reflect the importance of different KV vectors, for each processing layer in the large language model, the current KV cache of that processing layer can be updated to the target KV cache based on the attention map and the second window length of the processing layer, taking into account both the size of the KV cache and the data quality. For ease of distinction and description, the attention map corresponding to the same processing layer as the current KV cache is referred to as the first attention map.
[0085] It can be seen that the KV cache processing method for the large language model provided in this embodiment is applied to the large language model. When decoding the current semantic unit, if the length of the current KV cache is greater than or equal to the first window length determined when the previous semantic unit was decoded, the window length is updated to the second window length, and the current KV cache is updated according to the updated second window length and the attention map. The KV cache of each processing layer needs to be updated, and the length of the KV cache does not exceed the preset maximum window length. Therefore, no matter how long the generated text is, the length of the KV cache is limited to a certain range, avoiding memory overflow, reducing the computational complexity of the large language model, improving the processing efficiency of the large language model, not limiting the length of the text processed by the large language model, and expanding the text processing capability of the model.
[0086] Optionally, the KV cache processing method for a large language model also includes:
[0087] If the length of the current KV cache is less than the first window length, the current KV cache is determined as the target KV cache.
[0088] Specifically, when decoding the current semantic unit, if the length of the current KV cache is less than the first window length, then the first window length may not be updated, the window length remains unchanged, and the current KV cache is not updated, or it is understood that the current KV cache and the target KV cache are the same. When decoding the current semantic unit, the current KV cache or the target KV cache is used. After the decoding of the current semantic unit is completed, the KV vector corresponding to the current semantic unit is directly added to the current KV cache or the target KV cache for calculation when decoding the next semantic unit after the current semantic unit.
[0089] Optionally, the attention map includes an attention value. In S402, updating the first window length to a second window length according to the attention map of all processing layers includes:
[0090] For each processing layer, a first value is determined based on a plurality of attention values included in an attention map of the processing layer, wherein the first value is a minimum value of at least one attention value when a sum of at least one attention value among the plurality of attention values is greater than a preset threshold.
[0091] The maximum value between the first values corresponding to all processing layers and the preset minimum window length is determined as the second window length.
[0092] This process can be simply understood as follows: when the large language model decodes the current semantic unit, if the KV Cache length is less than window_len (the first window length), the window length remains unchanged. When the KV Cache length is greater than or equal to window_len (the first window length), a new window_len (the second window length) is calculated based on the attention map after decoding the previous semantic unit.
[0093] Specifically, define the following parameters: window length (window_len), minimum preset window length (min_len), and maximum preset window length (max_len), where min_len ≤ window_len ≤ max_len. Initially, window_len = max_len.
[0094] When decoding the current semantic unit, if the length of the current KV cache is less than the first window length (window_len), the first value of each processing layer is obtained by calculating the attention map of each processing layer using formula (1).
[0095] SUM(TOP (Ki,atten))>Preset threshold formula (1)
[0096] Among them, atten represents the attention map of the i-th processing layer, the TOP() function represents returning a fixed amount of data, and SUM() represents the sum operation. . SUM(TOP (Ki,atten))>preset threshold means that Ki attention values are added together in the attention map of the i-th processing layer, and the sum obtained is greater than the preset threshold. This embodiment does not limit the value of the preset threshold. For example, the preset threshold is 0.95. Ki is required to be as small as possible. Optionally, the minimum value of Ki is called the first numerical value, and the first numerical value is represented by Qi.
[0097] After obtaining the first numerical value corresponding to each processing layer in the large language model, the updated value of the first window length (i.e., the second window length) is determined according to formula (2), and window_len is reset.
[0098] window_len=max(min_len, Q1, Q2,…, Q(layer_num)) Formula (2)
[0099] Among them, layer_num represents the number of processing layers in the large language model.
[0100] Optionally, if the model uses a multi-head attention mechanism, the shape of the attention map is (layer_num, head_num, window_len), where head_num represents the number of heads. In this case, the average value is first taken over the dimension of head_num, and the shape of the attention map is converted to (layer_num, window_len), and then the calculation is performed using formula (1).
[0101] The following explains with examples.
[0102] Assume that the large language model consists of three processing layers, denoted as layer 1 through layer 3, with layer_num = 3. The first window length is 6, with a preset minimum window length of min_len = 3 and a preset maximum window length of max_len = 10. The attention map for layer 1 is (0.2, 0.06, 0.6, 0.1, 0.01, 0.03), with the attention positions 1, 2, 3, 4, 5, 6. The preset threshold is 0.95.
[0103] The current KV cache length is 6, equal to the first window length. The first window length needs to be updated to the second window length. For processing layer 1, the sum of the attention values 0.6+0.2+0.1+0.06 is greater than 0.95, and the first value Q1 corresponding to layer 1 is 4. Assume that the first value Q2 corresponding to layer 2 is 5, and the first value Q3 corresponding to layer 3 is 2. Then, with min_len=3, Q1=4, Q2=5, and Q3=2, the second window length is determined to be 5.
[0104] Optionally, the KV cache includes the KV vector and the vector position information of the KV vector, and the attention map also includes the attention position information of the attention value. In S403, according to the first attention map and the second window length, the current KV cache is updated to the target KV cache, including:
[0105] According to the order of attention values from large to small in the first attention map, attention position information of attention values of the second window length is obtained in the first attention map.
[0106] In the current KV cache, determine the second window length vector position information that corresponds one-to-one to the attention position information of the second window length attention values.
[0107] In the current KV cache, the KV vectors corresponding to the second window length vector position information are determined as the target KV cache.
[0108] This process can be simply understood as follows: when decoding the current token, if the length of the current KV cache is greater than or equal to the first window length, after updating the first window length to the second window length, the attention map obtained after decoding the previous token is used to reselect more important KV vectors, thereby updating the current KV cache to the target KV cache for decoding the current token. This KV vector update prevents the KV cache from continuously increasing in length when processing infinitely long sequences, which could eventually lead to memory overflow. It also retains important KV vectors while discarding less important ones, ensuring accurate model inference and balancing both efficiency and accuracy.
[0109] It should be noted that the current KV cache of each processing layer in the large language model needs to be updated, and the execution principle is similar.
[0110] Specifically, after the previous token is decoded, looking at one of the processing layers, we will obtain an attention map of (1, window_len) and a KV cache of length window_len. After calculating the second window length, we select the KV vector to retain or discard according to formula (3).
[0111] TOP_INDEX(attention map, window_len) formula (3)
[0112] Among them, attention map represents the attention map, window len represents the length of the second window, and TOP_INDEX() is used to retrieve the index (attention position information) of the top N (N=window_len) attention map values. After obtaining the index, the KV vector cached at the corresponding index position is retrieved from KVCache.
[0113] The following explains with examples.
[0114] Figure 5 is a schematic diagram of the current KV cache and the target KV cache provided by an embodiment of the present application. As shown in Figure 5, the length of the current KV cache is 6, including 6 KV vectors: (k1, v1), (k2, v2), (k3, v3), (k4, v4), (k5, v5), (k6, v6), and the vector position information is 1, 2, 3, 4, 5, and 6 respectively. Assume that the second window length is 3. According to the order of attention values from large to small in the attention map, the attention position information of the three attention values is obtained in the attention map. Assume that the attention map is (0.2, 0.1, 0.06, 0.6, 0.01, 0.03), and according to the order of attention values from large to small, the attention position information of the three attention values is determined to be 1, 2, and 4. Then select the important KV vectors (k1, v1), (k2, v2), (k4, v4) at positions 1, 2, and 4 in the current KV cache, and discard the relatively unimportant KV vectors (k3, v3), (k5, v5), and (k6, v6) in the current KV cache. The updated target KV cache includes 3 KV vectors: (k1, v1), (k2, v2), and (k4, v4).
[0115] Optionally, the KV cache processing method for a large language model provided in this embodiment may further include:
[0116] Determine other KV vectors in the KV vectors included in the current KV cache excluding the KV vectors included in the target KV cache.
[0117] A second value of KV vectors is selected from other KV vectors, where the second value is smaller than a difference between the first window length and the second window length.
[0118] The second value KV vector is added to the target KV cache to update the target KV cache.
[0119] This process can be simply understood as follows: when decoding the current token, if the length of the current KV cache is greater than or equal to the first window length, after discarding some KV vectors in the current KV cache to form the target KV cache, additional KV vectors can be selected from the discarded KV vectors for compensation. Typically, when decoding a token, the attention map distribution of the previous token and the attention map distribution of the current token are similar, but occasionally different distributions may occur. By selecting additional KV vectors from the discarded KV vectors for compensation, the loss of KV cache that requires attention due to different distributions is alleviated, thereby improving the inference accuracy of the model.
[0120] The number of compensated KV vectors is a second value, which is less than the difference between the first window length and the second window length. This embodiment does not limit the value of the second value. For example, the second value is 0.5×detla_len rounded up or rounded down. Detla_len represents the difference between the first window length and the second window length, that is, the number of KV vectors originally discarded in the current KV cache.
[0121] Optionally, a second value of KV vectors is selected from other KV vectors, and the second value of KV vectors can be selected based on the size of the attention value in the first attention map. The larger the attention value, the more important the KV vector at the corresponding position.
[0122] Optionally, a second numerical KV vector is selected from other KV vectors, and the selection may be performed randomly.
[0123] The following explains with examples.
[0124] Figure 6 is another schematic diagram of the current KV cache and the target KV cache provided by an embodiment of the present application. As shown in Figure 6, the length of the current KV cache is 6, including 6 KV vectors: (k1, v1), (k2, v2), (k3, v3), (k4, v4), (k5, v5), (k6, v6). Assume that the second window length is 3. According to the order of attention values from large to small in the attention graph, the attention position information of the three attention values is determined to be 1, 2, and 4. The important KV vectors (k1, v1), (k2, v2), and (k4, v4) at positions 1, 2, and 4 are selected in the current KV cache to form the target KV cache. Assume that the second value is 2. From the discarded KV vectors (k3, v3), (k5, v5), and (k6, v6), additional KV vectors (k3, v3) and (k6, v6) are selected and added to the target KV cache to form the updated target KV cache. Finally, when decoding the current semantic unit, the target KV cache is determined to include five KV vectors: (k1, v1), (k2, v2), (k3, v3), (k4, v4), and (k6, v6). The target KV cache will be used to decode the current semantic unit.
[0125] Optionally, the KV cache processing method for a large language model provided in this embodiment may further include:
[0126] The positions of the multiple KV vectors included in the target KV cache are re-encoded.
[0127] Specifically, each KV vector in the current KV cache has the original vector position information, which is used to indicate the position of the KV vector in the current KV cache. After the current KV cache is updated to the target KV cache u, some KV vectors are retained and some KV vectors are discarded, resulting in uneven positions of the KV vectors, which is not conducive to storage and calculation. Therefore, it is necessary to re-position-encode the multiple KV vectors included in the target KV cache. For the target KV cache of each processing layer in the large language model, the KV vectors need to be position-encoded again. Through the re-position encoding, the positions of the KV vectors in the target KV cache of each processing layer in the large language model are uniformly aligned, which is conducive to storage and calculation.
[0128] An exemplary description is given below with reference to FIG7 .
[0129] Assume that the large language model consists of three processing layers, denoted as layer 1 through layer 3. When decoding the current semantic unit, the current KV cache is 6 in length and contains six KV vectors: (k1, v1), (k2, v2), (k3, v3), (k4, v4), (k5, v5), and (k6, v6). The original vector positions are 1, 2, 3, 4, 5, and 6, respectively. After updating the current KV cache to the target KV cache, the target KV cache for layer 1 is {(k2, v2), (k4, v4), (k6, v6)}, the target KV cache for layer 2 is {(k4, v4), (k5, v5), (k6, v6)}, and the target KV cache for layer 3 is {(k1, v1), (k2, v2), (k4, v4)}. The KV vectors in the target KV cache of each processing layer are repositioned, as shown on the right side of Figure 7. For example, for the target KV cache of layer 1, the re-encoded vector position information corresponding to the KV vectors (k2, v2), (k4, v4), and (k6, v6) are 1, 2, and 3 respectively.
[0130] Optionally, the KV cache processing method for a large language model provided in this embodiment may further include:
[0131] Decode the current semantic unit according to the target KV cache.
[0132] Specifically, when decoding the current semantic unit, if the length of the current KV cache is less than the first window length, then the first window length may not be updated, the window length remains unchanged, the current KV cache is not updated, the current KV cache is the same as the target KV cache, and the current semantic unit is decoded according to the target KV cache. If the length of the current KV cache is greater than or equal to the first window length, then the first window length is updated to the second window length, and the current KV cache is updated according to the second window length and the attention map. By selecting important KV vectors in the current KV cache, or also selecting additional KV vectors, the target KV cache is obtained, and then the current semantic unit is decoded according to the target KV cache, the length of the KV cache is controlled, and the model processing efficiency is ensured.
[0133] FIG8 is a flow chart of a data processing method for a large language model provided in an embodiment of the present application. The data processing method for a large language model provided in this embodiment may be executed by a processing device for the large language model. As shown in FIG8 , the data processing method for a large language model may include:
[0134] S801: Receive input text.
[0135] For example, in the scenario shown in Figure 1, the input text could be "I love Tiananmen Square in Beijing." The large language model performs inference processing on the input text and ultimately generates the output text.
[0136] S802: Generate a semantic unit and a current KV cache corresponding to the semantic unit according to the input text.
[0137] Among them, the semantic unit and the current KV cache can be found in the relevant description in the previous part of this application and will not be repeated here.
[0138] S803: Process the current KV cache using the KV cache processing method of the large language model, and update the current KV cache to the target KV cache.
[0139] Among them, the KV cache processing method of the large language model can be referred to the embodiment shown in Figure 4, which will not be repeated here.
[0140] S804: Decode the semantic unit according to the target KV cache to generate output text.
[0141] The data processing method of the large language model provided in this embodiment is applied to the large language model, and the input text is processed by the large language model to obtain the output text. A semantic unit is obtained based on the input text. When decoding the current semantic unit, if the length of the current KV cache is greater than or equal to the first window length determined when the previous semantic unit is decoded, the window length is updated to the second window length, and the current KV cache is updated according to the updated second window length and the attention map to obtain the target KV cache, and the subsequent semantic unit decoding processing is performed based on the target KV cache to obtain the output text. During the data processing process, no matter how long the generated text is, the length of the KV cache is limited to a certain range, which avoids memory overflow, reduces the computational complexity of the large language model, improves the processing efficiency of the large language model, does not limit the length of the text processed by the large language model, and expands the text processing capability of the model.
[0142] It should be understood that the size of the serial numbers of the steps in the above embodiments does not mean the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of this application.
[0143] Figure 9 is a structural diagram of a processing device for a large language model provided in an embodiment of the present application. The modules included are used to execute the steps in the embodiment corresponding to Figure 4. For ease of explanation, only the parts related to the embodiment of the present application are shown.
[0144] 9 , the apparatus for processing a large language model provided in this embodiment may include:
[0145] An acquisition module 91 is configured to, when decoding a current semantic unit, obtain an attention map of each processing layer in the large language model after decoding a previous semantic unit before the current semantic unit if the length of the current KV cache is greater than or equal to the first window length;
[0146] a window length updating module 92, configured to update the first window length to a second window length according to the attention maps of all processing layers; wherein the first window length and the second window length are both less than or equal to a preset maximum window length;
[0147] The KV cache update module 93 is used to update the current KV cache to the target KV cache according to the first attention map and the second window length; wherein the first attention map and the current KV cache correspond to the same processing layer.
[0148] Optionally, the attention map includes attention values; the window length updating module 92 is used to:
[0149] For each processing layer, determining a first value based on a plurality of attention values included in an attention map of the processing layer; wherein the first value is a minimum value of the number of at least one attention value in the plurality of attention values when a sum of the at least one attention value is greater than a preset threshold;
[0150] The maximum value between the first values corresponding to all processing layers and the preset minimum window length is determined as the second window length.
[0151] Optionally, the KV cache includes the KV vector and the vector position information of the KV vector, and the attention map also includes the attention position information of the attention value;
[0152] The KV cache update module 93 is used to:
[0153] Obtaining attention position information of the second window length attention values in the first attention map in descending order of the attention values in the first attention map;
[0154] Determine, in the current KV cache, vector position information of the second window length that corresponds one-to-one to attention position information of attention values of the second window length;
[0155] In the current KV cache, the KV vectors corresponding to the second window length vector position information are determined as the target KV cache.
[0156] Optionally, the KV cache update module 93 is further configured to:
[0157] Determine other KV vectors included in the current KV cache excluding the KV vector included in the target KV cache;
[0158] Selecting a second value of KV vectors from the other KV vectors; the second value is smaller than the difference between the first window length and the second window length;
[0159] The second value of KV vectors is added to the target KV cache to update the target KV cache.
[0160] Optionally, the KV cache update module 93 is further configured to:
[0161] If the length of the current KV cache is less than the first window length, the current KV cache is determined as the target KV cache.
[0162] Optionally, a position encoding module is also included for:
[0163] Position encoding is performed again on a plurality of KV vectors included in the target KV cache.
[0164] Optionally, a decoding module is also included, for:
[0165] The current semantic unit is decoded according to the target KV cache.
[0166] Figure 10 is another structural diagram of the processing device of the large language model provided in an embodiment of the present application. The modules included are used to execute the steps in the embodiment corresponding to Figure 8. For the sake of convenience, only the parts related to the embodiment of the present application are shown.
[0167] 10 , the apparatus for processing a large language model provided in this embodiment may include:
[0168] Receiving module 101, for receiving input text;
[0169] A first generating module 102 is configured to generate a semantic unit and a current KV cache corresponding to the semantic unit according to the input text;
[0170] A KV cache processing module 103 is configured to process the current KV cache using the method described in the first aspect above, and update the current KV cache to a target KV cache;
[0171] The second generating module 104 is configured to decode the semantic unit according to the target KV cache to generate an output text.
[0172] It should be noted that the large language model processing device shown in FIG9 and the large language model processing device shown in FIG10 can be integrated into the same device.
[0173] It should be noted that the information interaction, execution process, etc. between the above-mentioned devices / units are based on the same concept as the method embodiment of this application. Their specific functions and technical effects can be found in the method embodiment section and will not be repeated here.
[0174] Those skilled in the art can clearly understand that, for the convenience and brevity of description, only the division of the above-mentioned functional units and modules is used as an example for illustration. In actual applications, the above-mentioned functions can be distributed and completed by different functional units and modules as needed, that is, the internal structure of the device can be divided into different functional units or modules to complete all or part of the functions described above. The functional units and modules in the embodiment can be integrated into one processing unit, or each unit can exist physically alone, or two or more units can be integrated into one unit. The above-mentioned integrated unit can be implemented in the form of hardware or in the form of software functional units. In addition, the specific names of the functional units and modules are only for the convenience of distinguishing each other, and are not used to limit the scope of protection of this application. The specific working process of the units and modules in the above-mentioned system can refer to the corresponding process in the aforementioned method embodiment, and will not be repeated here.
[0175] Figure 11 is another schematic diagram of the structure of a large language model processing device provided in an embodiment of the present application. As shown in Figure 11, the large language model processing device includes: at least one processor 110, a memory 111, and a computer program 112 stored in the memory 111 and executable on the at least one processor 110. When the processor 110 executes the computer program 112, it implements the steps of any of the above-mentioned method embodiments.
[0176] The processor may be a central processing unit (CPU), other general-purpose processors, digital signal processors (DSP), application-specific integrated circuits (ASIC), field-programmable gate arrays (FPGA), other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. A general-purpose processor may be a microprocessor or any conventional processor.
[0177] The so-called memory may include non-volatile and / or volatile memory. Non-volatile memory may include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM), or flash memory. Volatile memory may include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM is available in various forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (DDR SDRAM), enhanced SDRAM (ESDRAM), synchronous link DRAM (SLDRAM), RAMbus direct RAM (RDRAM), direct RAMbus dynamic RAM (DRDRAM), and RAMbus dynamic RAM (RDRAM).
[0178] An embodiment of the present application also provides a large language model, including multiple processing layers, each processing layer corresponding to a KV cache, and the processing layer is used to execute the KV cache processing method of the large language model provided in the present application, or the large language model is used to execute the KV cache processing method of the large language model or the data processing method of the large language model provided in the present application.
[0179] An embodiment of the present application further provides a computer-readable storage medium, wherein the computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the steps in the above-mentioned various method embodiments can be implemented.
[0180] An embodiment of the present application provides a computer program product. When the computer program product is run on an electrode discharge machining device, the electrode discharge machining device can implement the steps in the above-mentioned various method embodiments when executing the computer program product.
[0181] Those skilled in the art will appreciate that if the above-mentioned integrated unit is implemented as a software functional unit and sold or used as a standalone product, it can be stored in a computer-readable storage medium. Based on this understanding, the present application can implement all or part of the process steps in the above-mentioned method embodiments by using a computer program to instruct the relevant hardware. The computer program can be stored in a computer-readable storage medium. When executed by a processor, the computer program can implement the steps of each of the above-mentioned method embodiments. The computer program includes computer program code, which can be in source code form, object code form, executable file, or some intermediate form. The computer-readable medium can include at least: any entity or device capable of carrying computer program code to a camera / terminal device, recording medium, computer memory, read-only memory (ROM), random access memory (RAM), electric carrier signals, telecommunication signals, and software distribution media. Examples include USB flash drives, removable hard drives, magnetic disks, or optical disks. In some jurisdictions, based on legislation and patent practice, computer-readable media cannot be electric carrier signals or telecommunication signals.
[0182] Those skilled in the art will appreciate that in the above embodiments, the description of each embodiment has its own focus, and for parts that are not described or recorded in detail in a certain embodiment, reference can be made to the relevant descriptions of other embodiments.
[0183] It will be understood by those skilled in the art that the units and algorithm steps of each example described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Professional and technical personnel can use different methods to implement the described functions for each specific application, but such implementation should not be considered to be beyond the scope of this application.
[0184] The above-described embodiments are only used to illustrate the technical solutions of the present application, rather than to limit them. Although the present application has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or make equivalent replacements for some of the technical features therein. These modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the various embodiments of the present application, and should all be included in the scope of protection of the present application.
Claims
1. A method for processing KV caches of a large language model, characterized in that, Including: When decoding the current semantic unit, if the length of the current KV cache is greater than or equal to the first window length, obtain the attention maps of each processing layer in the large language model after decoding the previous semantic unit before the current semantic unit; Update the first window length to the second window length according to the attention maps of all processing layers; wherein, both the first window length and the second window length are less than or equal to a preset maximum window length; Update the current KV cache to the target KV cache according to the first attention map and the second window length; wherein, the first attention map and the current KV cache correspond to the same processing layer.
2. The method according to claim 1, characterized in that, The attention map includes attention values; the updating the first window length to the second window length according to the attention maps of all processing layers includes: For each processing layer, determine a first value according to the multiple attention values included in the attention map of the processing layer; wherein, the first value is the minimum value of the number of at least one attention value when the sum of at least one attention value in the multiple attention values is greater than a preset threshold; Determine the maximum value among the first values corresponding to all processing layers and a preset minimum window length as the second window length.
3. The method according to claim 2, wherein The KV cache includes KV vectors and vector position information of the KV vectors, and the attention map further includes attention position information of the attention values; The updating the current KV cache to the target KV cache according to the first attention map and the second window length includes: According to the order of the attention values in the first attention map from large to small, obtain the attention position information of the second window length of attention values in the first attention map; In the current KV cache, determine the second window length of vector position information corresponding one-to-one to the attention position information of the second window length of attention values respectively; In the current KV cache, determine the KV vectors corresponding to the second window length of vector position information respectively as the target KV cache.
4. The method according to claim 1, wherein The method further includes: Determine other KV vectors in the KV vectors included in the current KV cache except for the KV vectors included in the target KV cache; Select the second number of KV vectors among the other KV vectors; the second number is less than the difference between the first window length and the second window length; Add the second number of KV vectors to the target KV cache to update the target KV cache.
5. The method according to any one of claims 1 to 4, characterized in that, The method further includes: If the length of the current KV cache is less than the first window length, determine the current KV cache as the target KV cache.
6. The method according to any one of claims 1-4, characterized in that, The method further includes: Re-perform position encoding on the multiple KV vectors included in the target KV cache.
7. The method according to any one of claims 1-4, characterized in that, The method further includes: Decode the current semantic unit according to the target KV cache.
8. A data processing method for a large language model, characterized in that, Including: Receive the input text; Generate a semantic unit and the corresponding current KV cache according to the input text; Process the current KV cache by using the method according to any one of claims 1-7, and update the current KV cache to the target KV cache; Perform decoding processing of the semantic unit according to the target KV cache to generate an output text.
9. A processing device for a large language model, characterized in that, Including: An acquisition module, configured to, when decoding the current semantic unit, if the length of the current KV cache is greater than or equal to the first window length, acquire the attention maps of each processing layer in the large language model after decoding the previous semantic unit before the current semantic unit; A window length update module, configured to update the first window length to a second window length according to the attention maps of all processing layers; wherein, both the first window length and the second window length are less than or equal to a preset maximum window length; A KV cache update module, configured to update the current KV cache to a target KV cache according to the first attention map and the second window length; wherein, the first attention map and the current KV cache correspond to the same processing layer.
10. A processing device for a large language model, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, When the processor executes the computer program, it implements the method described in any one of claims 1-7, or implements the method described in claim 8.
11. A computer-readable storage medium storing a computer program, characterized in that, When the computer program is executed by the processor, it implements the method described in any one of claims 1-7, or implements the method described in claim 8.
Citation Information
Patent Citations
Weight-based key value storage cache elimination method and system and storage medium
CN114153760A
Cited By
Model output security detection method and device, and storage medium
CN121435223A
Text2SQL semantic caching method based on context and mode matching
CN121681573A
Fixed capacity KV cache management method, device and equipment based on multi-dimensional dynamic importance evaluation and storage medium
CN122450392A