A method and device for constructing a large model with a pre-set KV cache capacity

By using preset key-value vectors and key-vector sequences to update the cache in the Transformer structure of the large model, the memory constraint problem caused by the growth of KV cache capacity is solved, the system throughput is improved and the inference effect is maintained.

CN119377133BActive Publication Date: 2025-07-08BEIJING DEEPAI TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411445684.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-10-16
Publication Date
2025-07-08
Estimated Expiration
2044-10-16

AI Technical Summary

Technical Problem

In the prior art, the sequence length of the KV cache continues to grow with the progress of large-model inference, resulting in memory constraint problems, limiting system throughput, and traditional methods discard KV information, resulting in a decrease in the inference effect.

Method used

Using the method of preset KV cache capacity, the key value vectors in the cache are updated and replaced by the preset key value vector sequence and key vector sequence through the attention layer in the Transformer structure of the large model, the key value vectors in the cache are updated and replaced, the integrity of KV information is maintained, the steps are avoided, and the KV cache with limited capacity is realized.

Benefits of technology

It effectively alleviates memory constraint problems, improves system throughput, and maintains the inference effect of the large model, without relying on abandoning KV information.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119377133B_ABST
    Figure CN119377133B_ABST
Patent Text Reader

Abstract

The present application provides a method and device for constructing a large model with a preset KV cache capacity, which are applied to the attention layer in the Transformer structure of the large model. The attention layer includes a key-value vector sequence MV composed of M key-value vectors and a key vector sequence MK composed of M key vectors; where M is equal to a preset value. The method includes, for the Nth input vector, mapping it to a write query vector wq and a first write key-value vector wv; calculating the write weight vector ww by using the write query vector wq and the M key vectors; using the write weight vector ww and the first write key-value vector wv to update the M key-value vectors in the historical key-value vector sequence MV', and writing the updated key-value vector sequence MV into the cache. In this way, a KV cache capacity scheme with a preset length can be implemented to replace the KV cache capacity scheme that grows infinitely with the context length.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of artificial intelligence technology, and in particular, to a large model construction method and device with a preset KV cache capacity. Background Art

[0002] The KV cache technology is an efficient caching method. Based on the KV cache technology, the model can store information in the form of a KV cache sequence during inference. In this way, during subsequent inference processes, the large model can quickly access historical information, thereby improving the overall inference efficiency and speed. However, the sequence length of the KV cache gradually increases with the inference of the large model. As the sequence length increases, the cache requirements also continuously grow, which makes the large model inference become a memory constraint problem, greatly limiting the system throughput.

[0003] Using a KV cache with a limited capacity to handle infinite context inference has become the key to solving this problem. Streaming LLM (StreamingLLM) and Low-rank Embedding Sidekick with Sparse policy (LESS) are two common means of implementing a KV cache with a limited capacity.

[0004] However, both StreamingLLM and LESS adopt the method of discarding some KV information without caching, which easily leads to a decline in the inference effect of the large model. Summary of the Invention

[0005] The embodiments of this application provide a large model construction method and device with a preset KV cache capacity to solve the problem of poor inference effect caused by discarding KV information in traditional caching methods.

[0006] In a first aspect, the embodiments of this application provide a large model construction method with a preset KV cache capacity, which is applied to the attention layer in the Transformer structure of the large model. The attention layer includes a key-value vector sequence MV composed of M key-value vectors and a key vector sequence MK composed of M key vectors; where M is equal to a preset value, and the key-value vector sequence MV is composed of initial values.

[0007] The method includes: for the Nth input vector of the input attention layer, mapping it to a write query vector wq using a first matrix Wwq, and mapping it to a first write key-value vector wv using a second matrix Wwv; where N≥1, and the input vector is determined based on the embedding of the text tokens of the input large model or the output of the previous Transformer block; calculating the write weight vector ww by computing the write query vector wq with M key vectors; when N = 1, obtaining a key-value vector sequence MV composed of initial values from the cache and determining it as the historical key-value vector sequence MV'; when N > 1, obtaining the key-value vector sequence MV corresponding to the (N - 1)th input vector from the cache and determining it as the historical key-value vector sequence MV'; using the write weight vector ww and the first write key-value vector wv to update M key-value vectors in the historical key-value vector sequence MV' to obtain an updated key-value vector sequence MV; writing the updated key-value vector sequence MV into the cache to replace the historical key-value vector sequence MV', and the updated key-value vector sequence MV corresponds to the Nth input vector, i.e., MV(N). MV(0) is the initial form of MV, MV(0) is composed of initial values, and is used as the historical key-value vector sequence when N = 1, and MV(N - 1) is the historical key-value vector sequence when N > 1.

[0008] Generally, MK, MV(0), Wwq, and Wwv are obtained by training the large model, where MV(0) can also be preset to all 0 values or random values and does not participate in training updates.

[0009] In one implementable way, the dimension of the key-value vector is D, and the dimension of the key vector is E.

[0010] In one implementable way, the dimension of the first matrix Wwq is [D, E], so that the dimension of the write query vector wq is the same as that of the key vector, both being E; the dimension of the second matrix Wwv is [D, D], so that the dimension of the first write key-value vector wv is the same as that of the key-value vector, both being D.

[0011] In one implementable way, the step of calculating the write weight vector ww by computing the write query vector wq with M key vectors includes: calculating the similarities between the write query vector wq and the M key vectors respectively to obtain M weight scores; normalizing the M weight scores to obtain the write weight vector ww, and the dimension of the write weight vector ww is M.

[0012] In an implementable manner, the steps of updating M key-value vectors in the historical key-value vector sequence MV' by using the write weight vector ww and the first write key-value vector wv to obtain the updated key-value vector sequence MV include: transposing the write weight vector ww to convert the write weight vector ww into a write weight matrix ww' with a dimension of [M, 1], so that the number of columns of the write weight matrix ww' matches the number of rows of the first write key-value vector wv; performing matrix multiplication on the write weight matrix ww' and the first write key-value vector wv to obtain M second write key-value vectors wMV, and the dimension of the second write key-value vector wMV is D; performing an addition calculation on the M second write key-value vectors wMV and the M key-value vectors in the historical key-value vector sequence MV' to obtain the updated key-value vector sequence MV, that is, MV(N).

[0013] In an implementable manner, M≥1; and the historical key-value vector sequence MV(0) corresponding to the first input vector is obtained by training or preset.

[0014] In an implementable manner, the process of using Wwq or Wwv for mapping can be replaced by a single-layer or multi-layer perceptron (MLP), and the embodiments of the present application do not make specific limitations on this.

[0015] In an implementable manner, during the inference process, when processing tokens one by one, the same memory block / video memory block can be used for MV for different Ns; during the training process, N independent memory blocks / video memory blocks can be used for parallel training.

[0016] In an implementable manner, the updated MV is used as a KV cache for subsequent calculation processes, and the subsequent calculation processes can be the same as those of traditional large model methods, and the embodiments of the present application do not make specific limitations on this.

[0017] In a second aspect, an embodiment of the present application further provides a large model construction method with a preset KV cache capacity, which is applied to the attention layer in the Transformer structure of a large model. The attention layer includes a key-value vector sequence MV composed of M key-value vectors and a key vector sequence MK composed of M key vectors; where M is equal to a preset value, and the M key-value vectors are updated based on the large model construction method provided in the first aspect and its various implementation manners. Before the update, the key-value vector sequence MV is composed of initial values;

[0018] The method includes: for the Nth input vector of the input attention layer, using the third matrix Wq to map the input vector into the current query vector q; where N≥1, the input vector is the embedding vector corresponding to the text tokenized by the input large model or the process vector output by the previous Transformer block of the current Transformer block; calculating the current weight vector w by using the current query vector q and M key vectors; when N = 1, obtaining the key-value vector sequence MV composed of initial values from the cache, and when N>1, obtaining the key-value vector sequence MV corresponding to the (N-1)th input vector from the cache; using the current weight vector w to perform weighted summation on the obtained key-value vector sequence MV to obtain the output vector.

[0019] In a feasible manner, the dimension of the third matrix Wq is [F, E], F is the dimension of the input vector, and the dimension of the current query vector q is E; the dimension of the current weight vector w is M; the attention layer further includes a fourth matrix Wk with a dimension of [F, E], and the fourth matrix Wk is used to map the input vector into the current key vector k with a dimension of E; and, the dimension of the key vector is the same as the dimension of the current key vector k; the attention layer further includes a fifth matrix Wv with a dimension of [F, D], and the fifth matrix Wv is used to map the input vector into the current key-value vector v with a dimension of D; and, the dimension of the key-value vector is the same as the dimension of the current key-value vector v.

[0020] In a third aspect, an embodiment of the present application further provides a large model construction device with a preset KV cache capacity, which is applied to the attention layer in the Transformer structure of the large model. The attention layer includes a key-value vector sequence MV composed of M key-value vectors and a key vector sequence MK composed of M key vectors; where M is equal to a preset value, and the key-value vector sequence MV is composed of initial values;

[0021] The apparatus includes: a first mapping module, configured to map the Nth input vector of the input attention layer to a write query vector wq by using a first matrix Wwq, and map it to a first write key-value vector wv by using a second matrix Wwv, where N≥1, and the input vector is an embedding vector corresponding to the text tokenized by the input large model or a process vector output by the previous Transformer block of the current Transformer block; a first calculation module, configured to calculate the write weight vector ww by using the write query vector wq and M key vectors; a first acquisition module, when N = 1, configured to acquire a key-value vector sequence MV composed of initial values from the cache and determine it as the historical key-value vector sequence MV', when N>1, configured to acquire the key-value vector sequence MV corresponding to the N-1th input vector from the cache and determine it as the historical key-value vector sequence MV'; an update module, configured to update the M key-value vectors in the historical key-value vector sequence MV' by using the write weight vector ww and the first write key-value vector wv to obtain an updated key-value vector sequence MV; a cache module, configured to write the updated key-value vector sequence MV into the cache to replace the historical key-value vector sequence MV', and the updated key-value vector sequence MV corresponds to the Nth input vector.

[0022] In a fourth aspect, an embodiment of the present application further provides a large model construction apparatus with a preset KV cache capacity, which is applied to the attention layer in the Transformer structure of the large model. The attention layer includes a key-value vector sequence MV composed of M key-value vectors and a key vector sequence MK composed of M key vectors; where M is equal to a preset value, and the M key-value vectors are updated based on the large model construction method provided in the first aspect and its various implementation manners. Before the update, the key-value vector sequence MV is composed of initial values;

[0023] The apparatus includes: a second mapping module, configured to map the Nth input vector of the input attention layer to a current query vector q by using a third matrix Wq, where N≥1, and the input vector is an embedding vector corresponding to the text tokenized by the input large model or a process vector output by the previous Transformer block of the current Transformer block; a second calculation module, configured to calculate the current weight vector w by using the current query vector q and M key vectors; a second acquisition module, when N = 1, configured to acquire the key-value vector sequence MV composed of initial values from the cache, when N>1, configured to acquire the key-value vector sequence MV corresponding to the N-1th input vector from the cache; an output module, configured to perform weighted summation on the acquired key-value vector sequence MV by using the current weight vector w to obtain an output vector.

[0024] As can be seen from the above, the embodiments of the present application provide a method and apparatus for constructing a large model with a preset KV cache capacity, which are applied to the attention layer in the Transformer structure of the large model. The attention layer includes a key-value vector sequence MV composed of M key-value vectors and a key vector sequence MK composed of M key vectors; where M is equal to a preset value, and the key-value vector sequence MV is composed of initial values. The method includes: for the Nth input vector input to the attention layer, mapping it to a write query vector wq using the first matrix Wwq, and mapping it to a first write key-value vector wv using the second matrix Wwv; where N ≥ 1, and the input vector is the embedding vector corresponding to the text tokenized by the large model or the process vector output by the previous Transformer block of the current Transformer block; calculating the write weight vector ww using the write query vector wq and the M key vectors; when N = 1, obtaining the key-value vector sequence MV composed of initial values from the cache and determining it as the historical key-value vector sequence MV'; when N > 1, obtaining the key-value vector sequence MV corresponding to the (N - 1)th input vector from the cache and determining it as the historical key-value vector sequence MV'; updating the M key-value vectors in the historical key-value vector sequence MV' using the write weight vector ww and the first write key-value vector wv to obtain an updated key-value vector sequence MV; writing the updated key-value vector sequence MV into the cache to replace the historical key-value vector sequence MV', and the updated key-value vector sequence MV corresponds to the Nth input vector. The embodiments of the present application can use the MKV cache to replace the KV cache, and by continuously updating the key-value vector sequence MV, a finite KV cache capacity can be used to handle an infinite context length. A KV (i.e., MKV) cache capacity scheme with a preset length can be implemented to replace the KV cache capacity scheme that grows infinitely with the context length, relieve the cache pressure, solve the memory constraint problem, and does not involve steps of discarding KV information. BRIEF DESCRIPTION OF THE DRAWINGS

[0025] Figure 1 is a schematic flowchart of the method for constructing a large model provided by the embodiment of the present application;

[0026] Figure 2 is a schematic flowchart of the first method for constructing a large model with a preset KV cache capacity provided by the embodiment of the present application;

[0027] Figure 3 is a schematic diagram of the process of constructing a large model provided by the embodiment of the present application;

[0028] Figure 4 is a schematic flowchart of the second method for constructing a large model with a preset KV cache capacity provided by the embodiment of the present application;

[0029] Figure 5Schematic diagram of the first large model construction device with a pre - settable KV cache capacity provided by the embodiments of the present application;

[0030] Figure 6 Schematic diagram of the second large model construction device with a pre - settable KV cache capacity provided by the embodiments of the present application. Detailed implementation manners

[0031] In order to enable those skilled in the art to better understand the technical solutions in the present application, the following will clearly and completely describe the technical solutions in the embodiments of the present application in conjunction with the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the scope of protection of the present application.

[0032] Before introducing the technical solutions of the embodiments of the present application, an exemplary introduction to the terms involved in the embodiments of the present application will be made first.

[0033] KV caching (Key - Value Caching) technology: A data storage technology. Specifically, KV caching refers to storing data in the form of key - value pairs (Key - Value Pair) in memory or video memory. The key is the unique identifier of the data, and the key value (Value) is the data associated with the key. Since the access speed of memory or video memory is much faster than that of the hard disk, storing frequently accessed data in the cache can significantly improve the speed of data retrieval.

[0034] The construction of a large model includes two main stages: training and inference.

[0035] Training stage: In this stage, the large model learns through a large amount of data. The training process usually involves processes such as data preparation and parameter tuning. Once the model is trained, it can be used for inference in actual applications. Inference process of the large model: It refers to the process of generating a predicted output using the trained model (such as a language model) given data (such as input text).

[0036] The following takes the Large Language Model (LLM) as an example to make an exemplary introduction to the model inference process.

[0037] The inference process includes at least the following steps:

[0038] 1. Input processing:

[0039] (1) Text preprocessing: Convert the input text into a format that the model can understand. This process usually includes tokenization of the text data, which is specifically the process of dividing the text data into smaller units (such as words, sub-words, or characters), and the resulting units can be called tokens. Tokenization is a key step in text preprocessing, which helps the model better understand the underlying structure of the text and process it more effectively.

[0040] (2) Embedding representation: Convert the tokenized text into embedding vectors (or word vectors).

[0041] 2. Model inference:

[0042] (1) Input embedding: Input the processed input text (embedding vectors) into the neural network of the language model.

[0043] (2) Forward propagation: The language model processes the input embedding through multiple neural network structures (such as the Transformer structure). Forward propagation mainly consists of two important components: the self-attention mechanism and the feed-forward neural network.

[0044] Self-attention mechanism: In the attention layer, the model uses the self-attention mechanism to calculate the attention weights of each token to other tokens, so as to capture the dependencies and context information in the input sequence. It includes at least the following three steps:

[0045] ① Key-value pair conversion: Convert the embedding vector corresponding to each token into three different vectors: query vector, key vector, and value vector. These vectors are generated from the embedding through matrix multiplication.

[0046] ② Calculate attention weights: Calculate the dot product of the query vector and all key vectors, and then calculate the attention weights of each token through the Softmax function.

[0047] ③ Weighted average: Use the attention weights to perform a weighted average on all value vectors to generate the context representation of the token.

[0048] (3) Generate output: After processing through multiple layers, the model generates the prediction distribution of each token, and these distributions represent the probability estimates of the model for each possible output token.

[0049] 3. Output generation: According to the generated probability distribution and predefined strategies, select the output text generated by the model.

[0050] It is understandable that the self-attention mechanism involves generating key and value vectors for each input word, that is, generating key-value pairs. As the vectors are continuously generated, the number of key-value pairs keeps increasing. Moreover, the self-attention mechanism also involves obtaining previously generated key-value pairs for subsequent reasoning. Therefore, an efficient caching method is needed to store the key-value pairs to ensure that the previously computed results can be reused during the reasoning process.

[0051] The KV cache technology is an efficient caching means. Based on the KV cache technology, the model can quickly access historical information when generating new content, thereby improving the overall reasoning efficiency and speed. In the actual application process, using the KV cache technology as a caching method has become a common means to accelerate the reasoning generation speed of large models (including large language models LLM).

[0052] When applying the KV cache to the large model reasoning process, it specifically involves the following two terms.

[0053] Key / Value (KV): A collective term for two paired vector data (key and value) used in the attention layer of the neural network model transformer. In each attention layer of the large model network structure, one token generally corresponds to one KV.

[0054] Key-Value Cache (KV Cache): A sequence composed of one or more KV data, generally referring to the KV sequence of all tokens before (and can also include the KV of the current token, without limitation) the token currently being processed.

[0055] KV Cache Example: The sentence "I love you very much" can be tokenized as [I, very, love, you]. When processing [love], the KV cache refers to the sequence of [I, very]; when processing [you], the KV cache refers to the sequence of [I, very, love]. It can be seen that the sequence length of the KV cache of the transformer gradually increases as the token processing progresses.

[0056] However, as the sequence length increases, the cache requirement also keeps growing, which makes the large model reasoning become a memory constraint problem, greatly limiting the system throughput. Using a KV cache with a limited capacity to handle infinite context reasoning has become the key to solving this problem. Streaming LLM and Low-rank Embedding Sidekick with Sparse policy (LESS) are two common means to implement a KV cache with a limited capacity.

[0057] The key idea of StreamingLLM is to use window attention, which modifies global attention to local attention, enabling the KV cache to be maintained at a fixed size during inference, i.e., the window length. However, this solution leads to incomplete utilization of historical KV information. When the context length is long, most KV information is discarded and cannot provide all the information for subsequent inferences, easily resulting in a decline in the performance of large models.

[0058] LESS refers to learning the residual between the original attention output and the attention output approximated by the sparse policy. This is achieved by accumulating the information discarded by the sparse policy into a low-rank cache or state of a constant size. The sparse policy means retaining the KV information considered crucial and discarding other unimportant KV information. However, this solution increases the additional overhead of the low-rank cache and the complexity of implementation. Moreover, the weak importance of KV information does not mean that this KV information is completely useless. Directly discarding unimportant KV information will lead to a decline in the inference performance of large models.

[0059] To solve the above problems, the embodiments of this application first provide a large model construction method, which can be applied to the attention layer in the Transformer structure of a large model. The attention layer may include a third matrix Wq, a fourth matrix Wk, and a fifth matrix Wv.

[0060] Figure 1 It is a schematic flowchart of the large model construction method provided by the embodiments of this application.

[0061] As Figure 1 shown, the large model construction method provided by the embodiments of this application includes the following steps S101 - S103.

[0062] S101: For the Nth input vector input to the attention layer, use the fourth matrix Wk to map the input vector to the current key vector k.

[0063] In a large model, first, the input text can be converted into embedding vectors, and each vector represents a word or sub-word unit in the text, i.e., a token. Then, the embedding vectors can be used as input vectors and input into the attention layer in the Transformer structure for the model construction process. It can be understood that a piece of text can be tokenized into multiple embedding vectors (input vectors) to form a vector sequence, and then input into the attention layer. That is to say, in step S100, the input vector is determined based on the text input to the large model. The Nth input vector is one of the vector sequences, and N ≥ 1.

[0064] It should also be noted that the Transformer structure is generally a multi-layer structure, which may include multiple attention layers. The attention layer in step S100 can be any layer in the multiple attention layers, not limited to the first layer. Correspondingly, the input vector can be the embedding vector directly from the input text (the input of the first attention layer), or the intermediate vector processed by a certain previous attention layer.

[0065] The fourth matrix Wk is a weight matrix, which is used to calculate the key key of the input vector, that is, to map the input vector to the current key vector k. The dimension of the fourth matrix Wk can be [F, E], and F is the dimension of the input vector. Then, the dimension of the current key vector k is E. The input vector is mapped from the original F-dimensional space to a new E-dimensional space. The specific values of F and E can be designed based on the actual situation, and the embodiments of the present application do not make specific limitations on this.

[0066] For the input vector x, the mapping formula is k = X·Wk; where, Wk represents the fourth matrix, and k represents the current key vector.

[0067] S102: For the Nth input vector of the input attention layer, use the fifth matrix Wv to map the input vector to the current key-value vector v.

[0068] The fifth matrix Wv is also a weight matrix, which is used to calculate the key-value value of the input vector, that is, to map the input vector to the current key-value vector v. The dimension of the fifth matrix can be [F, D]. Then, the dimension of the current key-value vector v is D. The input vector is mapped from the original F-dimensional space to a new D-dimensional space. The specific value of D can be designed based on the actual situation, and the embodiments of the present application do not make specific limitations on this.

[0069] For the input vector x, the mapping formula is V = X·Wv; where, Wv represents the fifth matrix, and V represents the current key-value vector v.

[0070] S103: For the Nth input vector of the input attention layer, use the third matrix Wq to map the input vector to the current query vector q.

[0071] The third matrix Wq is also a weight matrix, which is used to calculate the query query of the input vector, that is, to map the input vector to the current query vector q. The dimension of the third matrix Wq can be [F, E], then the dimension of the current query vector q is E. The input vector is mapped from the original F-dimensional space to a new E-dimensional space.

[0072] It should be noted that for the current query vector q, the current key vector k, and the current key-value vector v, operations such as calculating similarity and normalization can also be performed. The specific steps will be described in detail below and will not be elaborated here.

[0073] In some implementations, steps S101, S102, and S103 can be executed in other sequences, and the embodiments of the present application do not make specific limitations on this.

[0074] The embodiments of the present application can also provide a large model construction method with a preset KV cache capacity. This method can use a KV cache with a preset length (i.e., MKV) to cache KV information. As the construction progresses, every time a new token is generated, the newly added KV information can be written (memorized) into MKV in a target manner, and this target manner will not change the length of MKV. It can be understood that in the traditional KV cache step, the KV information corresponding to the current token is appended to the corresponding position of the KV cache, and the sequence length of the KV cache will increase by 1 accordingly. The method provided by the embodiments of the present application can write the KV information corresponding to the current token into the KV cache with a preset length, i.e., MKV, without discarding KV information, can maintain the integrity of KV information, is not limited by the memory capacity, and can increase the system throughput.

[0075] Specifically, the large model construction method with a preset KV cache capacity provided by the embodiments of the present application can be applied to the attention layer in the Transformer structure of the large model. The attention layer includes a key-value vector sequence MV composed of M key-value vectors and a key vector sequence MK composed of M key vectors; where M is equal to a preset value, M≥1. Exemplarily, M can be equal to 1, 2, 3, or others, and can also be dynamically expanded during actual use. The embodiments of the present application do not make specific limitations on this. The key-value vector sequence MV is composed of initial values, and the initial values can be obtained by training the large model, or can be all 0 values or random values. The embodiments of the present application do not make specific limitations on this.

[0076] In the embodiments of the present application, both the key-value vector sequence MV and the key vector sequence MK can be stored in the cache, and each element of the vectors in the key-value vector sequence MV and the key vector sequence MK can have a preset value. The preset value can be obtained through vector initialization. The preset value is, for example, 0 or 1, and can also be other numerical values. The embodiments of the present application do not make specific limitations on this.

[0077] It should be noted that the dimension of the key-value vector can be the same as the dimension of the aforementioned current key-value vector v, i.e., both are D. Then, the key-value vector sequence MV is equivalent to a matrix with dimensions [M, D]. The dimension of the key vector can be the same as the dimension of the aforementioned current key vector k, i.e., both are E. Then the key vector sequence MK is equivalent to a matrix with dimensions [M, E] to ensure the feasibility of matrix calculations.

[0078] Figure 2Schematic flowchart of the first large model construction method with a pre - settable KV cache capacity provided by an embodiment of the present application.

[0079] Figure 3 Schematic diagram of the large model construction process provided by an embodiment of the present application.

[0080] As Figure 2 and Figure 3 shown, the large model construction method with a pre - settable KV cache capacity provided by an embodiment of the present application may include the following steps S201 - S205.

[0081] S201: For the Nth input vector input to the attention layer, use the first matrix Wwq to map it to a write query vector wq, and use the second matrix Wwv to map it to a first write key - value vector wv; where N≥1, and the input vector is determined based on the embedding of the text tokens input to the large model or the output of the previous Transformer block.

[0082] Among them, the Nth input vector of the input attention layer corresponds to the current token. The first matrix Wwq is a weight matrix used to calculate the query of the input vector, that is, to map the input vector to the write query vector wq (wquery). The dimension of the first matrix Wwq can be [D, E], so that the dimension of the write query vector wq is the same as that of the key vector, and both are E.

[0083] The second matrix Wwv is also a weight matrix used to calculate the key - value of the input vector, that is, to map the input vector to the write key - value vector wv (wvalue). The dimension of the second matrix Wwv can be [D, D], so that the dimension of the first write key - value vector wv is the same as that of the key - value vector, and both are D.

[0084] S202: Use the write query vector wq to calculate with M key vectors to obtain a write weight vector ww.

[0085] Step S202 is specifically an attention calculation step, which at least includes the processes of calculating similarity and normalization. Specifically, step S202 includes the following steps S2021 - S2022.

[0086] S2021: Calculate the similarity between the write query vector wq and M key vectors respectively to obtain M weight scores.

[0087] In the embodiments of the present application, the similarity can be calculated specifically by the method of dot product or cosine similarity, and it can also be the method of calculating the Euclidean distance, etc. The embodiments of the present application do not make specific limitations on this.

[0088] Exemplarily, M can be equal to 3. The M key vectors mkey are, for example, mkey1, mkey2, and mkey3. Then the M weight scores can be score1, score2, and score3.

[0089] S2022: Normalize the M weight scores to obtain the write weight vector ww, and the dimension of the write weight vector ww is M.

[0090] In the embodiments of the present application, the Softmax function can be specifically used for normalization. The sum of the normalized weight scores is equal to 1, and the normalized weight scores can form the write weight vector ww. Since the number of weight scores is M, the dimension of the write weight vector ww is M.

[0091] Exemplarily, when M is equal to 3, the normalized weight scores can be w1, w2, and w3.

[0092] It should be noted that the above steps S2021 - S2022 are exemplary introductions to step S202. In practical applications, the write query vector wq and the M key vectors can also be calculated through other calculation steps to obtain the write weight vector ww.

[0093] S203: When N = 1, obtain the key - value vector sequence MV composed of initial values from the cache and determine it as the historical key - value vector sequence MV', that is, MV(0). When N > 1, obtain the key - value vector sequence MV corresponding to the (N - 1)th input vector from the cache and determine it as the historical key - value vector sequence MV', that is, MV(N - 1).

[0094] Among them, the (N - 1)th input vector corresponds to the previous token of the current token. For example, the sentence "I love you very much" can be tokenized into [I, very, love, you]. When processing [love], [love] is the current token, and [very] is the previous token.

[0095] It should be noted that MV(0) is the initial form of MV and serves as the historical key - value vector sequence when N = 1. MV(N - 1) is the historical key - value vector sequence when N > 1. MK, MV(0), Wwq, Wwv are obtained by training the large - model. Among them, MV(0) can also be preset as all 0 values or random values and does not participate in training updates.

[0096] As the construction progresses, the key-value vector sequence MV can be continuously updated, and each input vector corresponds to a different key-value vector sequence MV. Specifically as follows:

[0097] S204: Use the write weight vector ww and the first write key-value vector wv to update M key-value vectors in the historical key-value vector sequence MV', and obtain the updated key-value vector sequence MV, that is, MV(N).

[0098] The update steps may include the following steps S2041 - S2043.

[0099] S2041: Transpose the write weight vector ww so that the write weight vector ww is converted into a write weight matrix ww' with a dimension of [M, 1], so that the number of columns of the write weight matrix ww' matches the number of rows of the first write key-value vector wv.

[0100] It can be understood that the write weight vector ww is composed of M normalized weight scores. Therefore, by transposing, the write weight vector ww can be converted into a write weight matrix ww' with a dimension of [M, 1].

[0101] S2042: Perform matrix multiplication on the write weight matrix ww' and the first write key-value vector wv to obtain M second write key-value vectors wMV, and the dimension of the second write key-value vector wMV is D.

[0102] According to the rules of matrix multiplication, the number of columns of the write weight matrix ww' matches the number of rows of the first write key-value vector wv, both equal to 1. Therefore, the two can perform matrix multiplication. Moreover, since the dimension of the write weight matrix ww' is [M, 1] and the dimension of the first write key-value vector wv is [1, D], the dimension of the second write key-value vector wMV is D, and M second write key-value vectors wMV can form a matrix with a dimension of [M, D]. As Figure 3 shown, when M is equal to 3, the M second write key-value vectors wMV can be wvalue1, wvalue2, and wvalue3 respectively.

[0103] S2043: Perform an addition calculation on the M second write key-value vectors wMV and the M key-value vectors in the historical key-value vector sequence MV' to obtain the updated key-value vector sequence MV, that is, MV(N).

[0104] It can be understood that adding and calculating the M second write key-value vectors wMV with the M key-value vectors in the historical key-value vector sequence MV' to obtain the updated key-value vector sequence MV (new MV) is equivalent to completing the memory process of the KV information corresponding to the current token. Exemplarily, when M equals 3, the updated key-value vector sequence MV includes three key-value vectors, namely new mvalue1, new mvalue2, and new mvalue3.

[0105] It is worth noting that when N equals 2, the key-value vector sequence MV corresponding to its previous input vector, i.e., the first input vector, can be preset. The first input vector corresponds to the first token of the text input to the large model.

[0106] S205: Write the updated key-value vector sequence MV into the cache to replace the historical key-value vector sequence MV'. The updated key-value vector sequence MV corresponds to the Nth input vector.

[0107] In this way, caching of KV information can be achieved. Further, when processing the next token of the current token, the key-value vector sequence MV of the Nth input vector can be obtained, and the key-value vector sequence MV can be updated using the (N + 1)th input vector, which will not be elaborated here.

[0108] It should be supplemented that steps S201 - S205 can also be referred to as the process of writing, memorizing, or updating MKV.

[0109] It should be supplemented that, like traditional transformers, MKV can be fused with positional encoding, multi-head attention, group convolution, or any other form, and the embodiments of the present application do not make any limitations in this regard.

[0110] It should be supplemented that the construction method provided by the embodiments of the present application only modifies the KV cache formation mechanism of the attention layer of the large model and does not impose limitations on other parts of the large model.

[0111] It can be understood that for an inference process with an expected context length (including the prompt and the generated content) of L, using a KV cache with a preset length of M instead of a KV cache with a maximum of L can achieve the purpose of compressing the KV cache when M is less than L, thereby further reducing the video memory occupancy of large model inference and improving the inference speed of the large model.

[0112] It should be noted that the process of using Wwq or Wwv for mapping can be replaced by a single-layer or multi-layer perceptron (MLP), and the embodiments of the present application do not make specific limitations in this regard.

[0113] It should also be noted that during the inference process, when processing tokens one by one, MV can use the same memory block / video memory block for different Ns; during the training process, it can be N independent memory blocks / video memory blocks for parallel training.

[0114] It should further be noted that the updated MV is used as the KV cache for subsequent calculation processes, and the subsequent calculation processes can be the same as those of traditional large model methods, and the embodiments of the present application do not make specific limitations in this regard.

[0115] From the above, the embodiments of the present application provide a method for constructing a large model with a preset KV cache capacity, which is applied to the attention layer in the Transformer structure of the large model. The attention layer includes a key-value vector sequence MV composed of M key-value vectors and a key vector sequence MK composed of M key vectors; where M is equal to a preset value, and the key-value vector sequence MV is composed of initial values; the method includes: for the Nth input vector input to the attention layer, using the first matrix Wwq to map it to a write query vector wq, and using the second matrix Wwv to map it to a first write key-value vector wv; where N≥1, the input vector is the embedding vector corresponding to the text tokenized input to the large model or the process vector output by the previous Transformer block of the current Transformer block; calculating the write weight vector ww by using the write query vector wq and the M key vectors; when N = 1, obtaining the key-value vector sequence MV composed of initial values from the cache and determining it as the historical key-value vector sequence MV'; when N>1, obtaining the key-value vector sequence MV corresponding to the (N-1)th input vector from the cache and determining it as the historical key-value vector sequence MV'; using the write weight vector ww and the first write key-value vector wv to update the M key-value vectors in the historical key-value vector sequence MV' to obtain an updated key-value vector sequence MV; writing the updated key-value vector sequence MV into the cache to replace the historical key-value vector sequence MV', and the updated key-value vector sequence MV corresponds to the Nth input vector. The embodiments of the present application can use the MKV cache to replace the KV cache, and by continuously updating the key-value vector sequence MV, a finite KV cache capacity can be used to handle an infinite context length. A KV (i.e., MKV) cache capacity scheme with a preset length can be implemented to replace the KV cache capacity scheme that grows infinitely with the context length, relieve the cache pressure, solve the memory constraint problem, and does not involve steps of discarding KV information.

[0116] Figure 4Schematic flowchart of the second large model construction method with a preset KV cache capacity provided by the embodiments of the present application.

[0117] As Figure 4 shown, the embodiments of the present application also provide a large model construction method with a preset KV cache capacity. This method can be combined with the two large model construction methods in the foregoing embodiments and applied to the attention layer in the Transformer structure of the large model to implement model construction. The attention layer may include a key-value vector sequence MV composed of M key-value vectors, and a key vector sequence MK composed of M key vectors; where M is equal to a preset value, and the M key-value vectors are updated based on the foregoing first large model construction method with a preset KV cache capacity. Before the update, the key-value vector sequence MV is composed of initial values. For more introductions to the key-value vector sequence MV and the key vector sequence MK, reference may be made to the foregoing content, and details are not described here.

[0118] Continue to refer to Figure 4 , the large model construction method with a preset KV cache capacity provided by the embodiments of the present application may include the following steps S301-S304. Steps S301-S304 are steps for performing an attention operation on MKV.

[0119] S301: For the Nth input vector input to the attention layer, use the third matrix Wq to map the input vector to the current query vector q; where N≥1, and the input vector is the embedding vector corresponding to the text tokenized by the input large model or the process vector output by the previous Transformer block of the current Transformer block.

[0120] For the introduction to the third matrix Wq, reference may be made to S103 above, and details are not described here. The dimension of the third matrix Wq may be [F, E], where F is the dimension of the input vector, and then the dimension of the current query vector q is E. The embodiments of the present application do not specifically limit the values of F and E.

[0121] In some implementation manners, step S301 and step S103 may be alternatively executed.

[0122] S302: Use the current query vector q to calculate with the M key vectors to obtain the current weight vector w.

[0123] In this step, specifically, the current query vector q may be used to perform a dot product calculation with each key vector to obtain an attention score. These scores represent the correlation between the current query vector q and each key vector. Then, perform a normalization SoftMax operation on the attention scores to obtain attention weights, which represent the degree of attention of the current query vector q to each key vector.

[0124] It can be understood that the current weight vector w is composed of M normalized attention scores, and the dimension of the current weight vector w is M.

[0125] S303: When N = 1, obtain the key-value vector sequence MV composed of initial values from the cache; when N > 1, obtain the key-value vector sequence MV corresponding to the (N - 1)-th input vector from the cache.

[0126] The introduction of the key-value vector sequence MV composed of initial values and the key-value vector sequence MV corresponding to the (N - 1)-th input vector can refer to the aforementioned step S203, which will not be elaborated here.

[0127] S304: Use the current weight vector w to perform weighted summation on the obtained key-value vector sequence MV to obtain the output vector.

[0128] Step S304 means using the attention weights to perform weighted averaging on the M key-value vectors in the key-value vector sequence MV to generate the final attention-weighted output. The contribution of each key-value vector is determined by its corresponding attention weight. It can be understood that the key-value vectors in the key-value vector sequence MV refer to the key-value vectors of MKV, rather than the current key-value vector v in step S102.

[0129] The weighted sum value calculation formula can be:

[0130]

[0131] where Output represents the output vector, V i represents the i-th key-value vector, 0 < i ≤ M, Attention represents the attention weight of the i-th key-value vector.

[0132] It should be added that steps S201 - S205 can also be referred to as the process of recalling or reading MKV.

[0133] In some implementation manners, the attention layer further includes a fourth matrix Wk with dimensions [F, E], and the fourth matrix Wk is used to map the input vector into a current key vector k with dimensions E; moreover, the dimension of the key vector is the same as the dimension of the current key vector k.

[0134] In some implementation manners, the attention layer further includes a fifth matrix Wv with dimensions [F, D], and the fifth matrix Wv is used to map the input vector into a current key-value vector v with dimensions D; moreover, the dimension of the key-value vector is the same as the dimension of the current key-value vector v.

[0135] Figure 5 This is the structural schematic diagram of the first large model construction device with a pre-settable KV cache capacity provided by the embodiments of the present application.

[0136] As Figure 5 shown, an embodiment of the present application provides a large model construction device with a preset KV cache capacity, which is applied to the attention layer in the Transformer structure of the large model. The attention layer includes a key-value vector sequence MV composed of M key-value vectors, and a key vector sequence MK composed of M key vectors; where M is equal to a preset value, and the key-value vector sequence MV is composed of initial values;

[0137] The device includes:

[0138] A first mapping module 1001, which is used to map the Nth input vector input to the attention layer to a write query vector wq by using a first matrix Wwq, and map it to a first write key-value vector wv by using a second matrix Wwv; where N≥1, and the input vector is the embedding vector corresponding to the text tokenized input to the large model or the process vector output by the previous Transformer block of the current Transformer block;

[0139] A first calculation module 1002, which is used to calculate the write weight vector ww by using the write query vector wq and M key vectors;

[0140] A first acquisition module 1003, when N = 1, is used to obtain the key-value vector sequence MV composed of initial values from the cache and determine it as the historical key-value vector sequence MV', when N>1, is used to obtain the key-value vector sequence MV corresponding to the N-1th input vector from the cache and determine it as the historical key-value vector sequence MV';

[0141] An update module 1004, which is used to update the M key-value vectors in the historical key-value vector sequence MV' by using the write weight vector ww and the first write key-value vector wv to obtain the updated key-value vector sequence MV;

[0142] A cache module 1005, which is used to write the updated key-value vector sequence MV into the cache to replace the historical key-value vector sequence MV', and the updated key-value vector sequence MV corresponds to the Nth input vector.

[0143] In some implementation manners, the dimension of the key-value vector is D, and the dimension of the key vector is E.

[0144] In some implementation manners, the dimension of the first matrix Wwq is [D, E], so that the dimension of the write query vector wq is the same as the dimension of the key vector, and both are E; the dimension of the second matrix Wwv is [D, D], so that the dimension of the first write key-value vector wv is the same as the dimension of the key-value vector, and both are D.

[0145] In some implementations, the first calculation module 1002 is specifically configured to: calculate the similarities between the written query vector wq and the M key vectors respectively to obtain M weight scores; perform normalization processing on the M weight scores to obtain a written weight vector ww, and the dimension of the written weight vector ww is M.

[0146] In some implementations, the update module 1004 is specifically configured to: transpose the written weight vector ww so that the written weight vector ww is converted into a written weight matrix ww' with a dimension of [M, 1], thereby making the number of columns of the written weight matrix ww' match the number of rows of the first written key value vector wv; perform matrix multiplication calculation on the written weight matrix ww' and the first written key value vector wv to obtain M second written key value vectors wMV, and the dimension of the second written key value vector wMV is D; perform addition calculation on the M second written key value vectors wMV and the M key value vectors in the historical key value vector sequence MV' to obtain an updated key value vector sequence MV.

[0147] In some implementations, M ≥ 1; and, the key vector sequence MK, the first matrix Wwq, and the second matrix Wwv are obtained by training a large model; the key value vector sequence MV is obtained by training a large model, or the initial value of the key value vector sequence MV is all 0 values or random values.

[0148] Figure 6 FIG. 10 is a schematic structural diagram of a second large model construction device with a preset KV cache capacity provided by an embodiment of the present application.

[0149] As Figure 6 shown, an embodiment of the present application further provides a large model construction device with a preset KV cache capacity, which is applied to the attention layer in the Transformer structure of a large model. The attention layer includes a key value vector sequence MV composed of M key value vectors and a key vector sequence MK composed of M key vectors; where M is equal to a preset value, and the M key value vectors are updated based on the foregoing first large model construction method with a preset KV cache capacity. Before the update, the key value vector sequence MV is composed of initial values;

[0150] The device includes:

[0151] A second mapping module 2001, configured to map the input vector for the Nth input vector input to the attention layer into a current query vector q by using a third matrix Wq; where N ≥ 1, and the input vector is an embedding vector corresponding to the text tokenized by the input large model or a process vector output by the previous Transformer block of the current Transformer block;

[0152] The second calculation module 2002 is used to calculate the current weight vector w by using the current query vector q and M key vectors.

[0153] The second acquisition module 2003 is used to acquire the key-value vector sequence MV composed of initial values from the cache when N = 1, and to acquire the key-value vector sequence MV corresponding to the (N - 1)-th input vector from the cache when N > 1.

[0154] The output module 2004 is used to perform weighted summation on the acquired key-value vector sequence MV by using the current weight vector w to obtain the output vector.

[0155] In some implementation manners, the dimension of the third matrix Wq is [F, E], where F is the dimension of the input vector and the dimension of the current query vector q is E; the dimension of the current weight vector w is M; the attention layer further includes a fourth matrix Wk with the dimension of [F, E], and the fourth matrix Wk is used to map the input vector to the current key vector k with the dimension of E; and the dimension of the key vector is the same as the dimension of the current key vector k; the attention layer further includes a fifth matrix Wv with the dimension of [F, D], and the fifth matrix Wv is used to map the input vector to the current key-value vector v with the dimension of D; and the dimension of the key-value vector is the same as the dimension of the current key-value vector v.

[0156] In specific implementation, the present invention further provides a computer storage medium. The computer storage medium can store a program, and when the program is executed, it can include some or all of the steps in the embodiments of the large model construction method with a preset KV cache capacity provided by the present invention. The storage medium can be a magnetic disk, an optical disk, a read-only memory (ROM for short), a random access memory (RAM for short), etc.

[0157] It is easy to understand that those skilled in the art can combine, split, recombine, etc. the embodiments of the present application based on several embodiments provided by the present application to obtain other embodiments, and these embodiments do not exceed the protection scope of the present application.

[0158] The above specific implementation manners further elaborate the purpose, technical solutions and beneficial effects of the embodiments of the present application. It should be understood that the above is only the specific implementation manners of the embodiments of the present application, and is not used to limit the protection scope of the embodiments of the present application. Any modifications, equivalent replacements, improvements, etc. made on the basis of the technical solutions of the embodiments of the present application shall be included in the protection scope of the embodiments of the present application.

Claims

1. A method for constructing a large model with a pre - set KV cache capacity, characterized in that, The attention layer in the Transformer structure applied to the large model, the attention layer includes a key-value vector sequence MV composed of M key-value vectors, and a key vector sequence MK composed of M key vectors; where M is equal to a preset value, and the key-value vector sequence MV is composed of initial values; The method includes: For the Nth input vector input to the attention layer, use the first matrix Wwq to map it to a write query vector wq, and use the second matrix Wwv to map it to a first write key-value vector wv; where N≥1, and the input vector is the embedding vector corresponding to the text tokenized input to the large model or the process vector output by the previous Transformer block of the current Transformer block; Calculate the write weight vector ww by using the write query vector wq and the M key vectors; When N = 1, obtain the key-value vector sequence MV composed of the initial values from the cache and determine it as the historical key-value vector sequence MV', when N>1, obtain the key-value vector sequence MV corresponding to the N-1th input vector from the cache and determine it as the historical key-value vector sequence MV'; Use the write weight vector ww and the first write key-value vector wv to update the M key-value vectors in the historical key-value vector sequence MV' to obtain the updated key-value vector sequence MV; The step of using the write weight vector ww and the first write key-value vector wv to update the M key-value vectors in the historical key-value vector sequence MV' to obtain the updated key-value vector sequence MV includes: Transpose the write weight vector ww so that the write weight vector ww is converted into a write weight matrix ww' with a dimension of [M, 1], so that the number of columns of the write weight matrix ww' matches the number of rows of the first write key-value vector wv; Perform matrix multiplication on the write weight matrix ww' and the first write key-value vector wv to obtain M second write key-value vectors wMV, and the dimension of the second write key-value vector wMV is D; Perform an addition calculation on the M second write key-value vectors wMV and the M key-value vectors in the historical key-value vector sequence MV' to obtain the updated key-value vector sequence MV; Write the updated key-value vector sequence MV into the cache to replace the historical key-value vector sequence MV', and the updated key-value vector sequence MV corresponds to the Nth input vector.

2. The large model construction method with a pre - settable KV cache capacity according to claim 1, wherein, The dimension of the key-value vector is D, and the dimension of the key vector is E.

3. The large model construction method with a pre - settable KV cache capacity according to claim 2, wherein, The dimension of the first matrix Wwq is [D, E] so that the dimension of the write query vector wq is the same as the dimension of the key vector, both are E; the dimension of the second matrix Wwv is [D, D] so that the dimension of the first write key-value vector wv is the same as the dimension of the key-value vector, both are D.

4. The method for constructing a large model with a pre - set KV cache capacity according to claim 3, wherein, The steps of calculating the write weight vector ww by using the write query vector wq and the M key vectors include: Calculating the similarities between the write query vector wq and the M key vectors respectively to obtain M weight scores; Performing normalization processing on the M weight scores to obtain the write weight vector ww, and the dimension of the write weight vector ww is M.

5. The large model construction method with a pre - settable KV cache capacity according to claim 1, wherein, M≥1; Moreover, the key vector sequence MK, the first matrix Wwq, and the second matrix Wwv are obtained by training the large model; The key-value vector sequence MV is obtained by training the large model, or the initial value of the key-value vector sequence MV is all 0 or a random value.

6. A large model construction method with a preset KV cache capacity, characterized in that Applied to the attention layer in the Transformer structure of the large model, the attention layer includes a key-value vector sequence MV composed of M key-value vectors and a key vector sequence MK composed of M key vectors; where M is equal to a preset value, and the M key-value vectors are updated based on the large model construction method according to any one of claims 1-5. Before the update, the key-value vector sequence MV is composed of initial values; The method includes: For the Nth input vector input into the attention layer, using the third matrix Wq to map the input vector into the current query vector q; where N≥1, and the input vector is the embedding vector corresponding to the text tokenized input into the large model or the process vector output by the previous Transformer block of the current Transformer block; Calculating the current weight vector w by using the current query vector q and the M key vectors; When N = 1, obtaining the key-value vector sequence MV composed of the initial values from the cache, and when N>1, obtaining the key-value vector sequence MV corresponding to the (N-1)th input vector from the cache; Performing weighted summation on the obtained key-value vector sequence MV by using the current weight vector w to obtain the output vector.

7. The large model construction method with a pre - settable KV cache capacity according to claim 6, wherein, The dimension of the third matrix Wq is [F, E], F is the dimension of the input vector, and the dimension of the current query vector q is E; The dimension of the current weight vector w is M; The attention layer further includes a fourth matrix Wk with a dimension of [F, E], and the fourth matrix Wk is used to map the input vector into the current key vector k with a dimension of E; and the dimension of the key vector is the same as the dimension of the current key vector k; The attention layer further includes a fifth matrix Wv with a dimension of [F, D], and the fifth matrix Wv is used to map the input vector into the current key-value vector v with a dimension of D; and the dimension of the key-value vector is the same as the dimension of the current key-value vector v.

8. A large model construction device with a pre - settable KV cache capacity, characterized in that, Applied to the attention layer in the Transformer structure of the large model, the attention layer includes a key-value vector sequence MV composed of M key-value vectors and a key vector sequence MK composed of M key vectors; where M is equal to a preset value, and the key-value vector sequence MV is composed of initial values; The device includes: The first mapping module is used to map the Nth input vector input to the attention layer to a write query vector wq by using a first matrix Wwq, and to map it to a first write key-value vector wv by using a second matrix Wwv; where N≥1, and the input vector is an embedding vector corresponding to the text tokenized input to the large model or a process vector output by the previous Transformer block of the current Transformer block. The first calculation module is used to calculate the write weight vector ww by using the write query vector wq and the M key vectors. The first acquisition module, when N = 1, is used to acquire the key-value vector sequence MV composed of the initial values from the cache and determine it as the historical key-value vector sequence MV'; when N>1, it is used to acquire the key-value vector sequence MV corresponding to the N-1th input vector from the cache and determine it as the historical key-value vector sequence MV'. The update module is used to update the M key-value vectors in the historical key-value vector sequence MV' by using the write weight vector ww and the first write key-value vector wv to obtain the updated key-value vector sequence MV. Specifically, the update module is used to update the M key-value vectors in the historical key-value vector sequence MV' by using the write weight vector ww and the first write key-value vector wv to obtain the updated key-value vector sequence MV. The steps include: Transpose the write weight vector ww to convert the write weight vector ww into a write weight matrix ww' with a dimension of [M, 1], so that the number of columns of the write weight matrix ww' matches the number of rows of the first write key-value vector wv. Perform matrix multiplication on the write weight matrix ww' and the first write key-value vector wv to obtain M second write key-value vectors wMV, and the dimension of the second write key-value vector wMV is D. Perform an addition calculation on the M second write key-value vectors wMV and the M key-value vectors in the historical key-value vector sequence MV' to obtain the updated key-value vector sequence MV. The cache module is used to write the updated key-value vector sequence MV into the cache to replace the historical key-value vector sequence MV', and the updated key-value vector sequence MV corresponds to the Nth input vector.

9. A large model construction device with a pre - settable KV cache capacity, characterized in that, Applied to the attention layer in the Transformer structure of the large model, the attention layer includes a key-value vector sequence MV composed of M key-value vectors and a key vector sequence MK composed of M key vectors; where M is equal to a preset value, and the M key-value vectors are updated based on the large model construction method according to any one of claims 1-5. Before the update, the key-value vector sequence MV is composed of initial values. The device includes: The second mapping module is used to map the input vector to the current query vector q by using the third matrix Wq for the Nth input vector input to the attention layer; where N≥1, and the input vector is the embedding vector corresponding to the text tokenized input to the large model or the process vector output by the previous Transformer block of the current Transformer block; The second calculation module is used to calculate the current weight vector w by using the current query vector q and the M key vectors; The second acquisition module is used to acquire the key-value vector sequence MV composed of the initial values from the cache when N = 1, and is used to acquire the key-value vector sequence MV corresponding to the N-1th input vector from the cache when N>1; The output module is used to perform weighted summation on the obtained key-value vector sequence MV by using the current weight vector w to obtain an output vector.

Citation Information

Patent Citations

  • A method and device for named entity recognition

    CN111291565A

  • Text translation method and device, kernel function combination method, server and medium

    CN114580443A