Key value compression method in self-attention mechanism, large language model and electronic equipment

By performing multiple residual decomposition and cluster compression on the key-value matrix, and combining the query matrix for attention calculation, the limitations of the traditional self-attention mechanism in long-sequence tasks are solved, and efficient long-sequence processing and flexible sequence representation are achieved.

CN120106150AActive Publication Date: 2025-06-06BEIJING ACAD OF ARTIFICIAL INTELLLIGENCE
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
CN202510344978.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-21
Publication Date
2025-06-06
Estimated Expiration
2045-03-21

AI Technical Summary

Technical Problem

The traditional self-attention mechanism based on softmax has limitations in memory capacity and computing efficiency when processing long-sequence tasks, making it difficult to effectively process long-sequence data.

Method used

By performing multiple residual decompositions on the key matrix and the value matrix, clustering and compression of the key residual vector and the value residual vector after each decomposition, attention calculation is performed in combination with the query matrix, and all attention calculation results are accumulated to achieve approximate attention output.

Benefits of technology

Improves the efficiency of Transformer in long-sequence tasks, solves the problem that Linear Transformer cannot use standard Softmax Transformer parameters, and implements a more flexible and robust sequence representation to adapt to various attention mechanisms.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120106150A_ABST
    Figure CN120106150A_ABST
Patent Text Reader

Abstract

The invention discloses a key value compression method in a self-attention mechanism, a large language model and electronic equipment, and relates to the technical field of computers. The compression method comprises the following steps: respectively carrying out multiple times of residual decomposition on a key matrix and a value matrix to obtain a key residual vector and a value residual vector after each time of decomposition; respectively carrying out clustering compression on the key residual vector and the value residual vector after each decomposition, and carrying out attention calculation on the query matrix, the compressed key residual vector and the compressed value residual vector; and accumulating all attention calculation results. According to the method and the device, the problems that standard Softmax Transform parameters cannot be used by the Linear Transform, and the difference between the standard Softmax Transform parameters and the standard Softmax Transform parameters is large are solved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of computer technology, and in particular to a key-value compression method in a self-attention mechanism, a large language model, and an electronic device. Background Art

[0002] Attention mechanisms have been the cornerstone of recent advances in natural language processing, computer vision, and related fields. However, traditional softmax-based self-attention has quadratic complexity, which greatly limits memory capacity and computational efficiency, especially in tasks involving long sequences. To address this bottleneck, a variety of linearized attention mechanisms have been proposed. Linearized attention mechanisms aim to reduce the complexity of self-attention from quadratic to linear, thereby achieving more efficient reasoning with fewer time and space resources. This improvement is particularly beneficial for long sequence processing tasks, where the computational overhead of traditional attention methods can be prohibitive.

[0003] Linear Attention is one of the pioneering methods in this field, which replaces the softmax operation with kernel functions such as Exponential Linear Unit (ELU) activation and rearranges the order of calculation to achieve linear complexity. Linear Attention is compatible with recurrent neural network structures, thus bridging the gap between Transformer models and recurrent neural networks (RNNs). The currently introduced models such as GLA (Gated Linear Attention), GSA (Gated Slot Attention), and MetaLA (Meta Linear Attention) further improve the efficiency and accuracy of linear attention. For example, GLA adopts a dynamic gating mechanism to more efficiently manage and update storage information, further unifying previous linear methods including structured state space models (SSM) into this broader framework and becoming a special form of GLA. At the same time, GSA introduces an extended ABC storage mechanism to compress the key matrix and value matrix into fixed-size slots, making it closer to softmax attention than the earlier linearized attention models. In addition to linear Attention and SSM models, many recursive models also use the Neural Turing Mechanism (NTM) method to write and read historical information, thereby improving the recursive model's ability to process historical information. Summary of the invention

[0004] The present application provides a key-value compression method, a large language model and an electronic device in a self-attention mechanism to achieve a more flexible and robust sequence representation, which can adapt to various attention mechanisms and achieve a high level of performance in long sequence tasks.

[0005] In a first aspect, the present application provides a key value compression method in a self-attention mechanism, which is applied to a large language model, and the compression method includes:

[0006] Perform multiple residual decompositions on the key matrix and the value matrix respectively, to obtain a key residual vector and a value residual vector after each decomposition; wherein the key residual vector after each decomposition includes the key residual vector after the 0th decomposition, and the value residual vector after each decomposition includes the value residual vector after the 0th decomposition;

[0007] The key residual vector and value residual vector after each decomposition are clustered and compressed respectively, and attention calculation is performed on the query matrix, the compressed key residual vector and the value residual vector;

[0008] All attention calculation results are accumulated.

[0009] Furthermore, the specific formula for the residual decomposition of the bond matrix is:

[0010] K i =K i-1 -gather(C i-1 [0], index1);

[0011] index1 = argmax(matmul(K i-1 , C i-1 [0].transpose));

[0012] Among them, K i represents the key residual vector after the ith residual decomposition, K i-1 represents the key residual vector after the i-1th residual decomposition. When i=1, K 0 represents the key matrix; gather represents selecting C according to index position index1 i-1 [0] and splice; C i-1 represents the trainable matrix, C i-1 [0] represents the trainable matrix C i-1 argmax means returning the index of the maximum value of the matrix in the specified dimension, matmul means matrix multiplication, and transpose means matrix transposition;

[0013] The specific formula for the residual decomposition of the value matrix is:

[0014] V i =V i-1 -gather(C i-1 [1], index2);

[0015] index2 = argmax(matmul(V i-1 , C i-1 [1].transpose));

[0016] Among them, V i Represents the residual vector after the i-th residual decomposition, V i-1 Represents the residual vector after the i-1th residual decomposition. When i=1, V 0 Represents a value matrix; gather represents selecting C according to index position index2 i-1 [1] and splice; C i-1 [1] represents the trainable matrix C i-1 The first dimension of .

[0017] Furthermore, clustering and compression are performed on the key residual vector and the value residual vector after each decomposition, including:

[0018] Calculating a correlation matrix between the key residual vector and the trainable matrix;

[0019] Performing mapping transformation on the correlation matrix to obtain a weight matrix;

[0020] Clustering the key residual vector and the weight matrix to obtain a compressed key residual vector;

[0021] The value residual vector and the weight matrix are clustered to obtain a compressed value residual vector.

[0022] Further, mapping transformation is performed on the correlation matrix, including:

[0023] The correlation matrix is ​​mapped and transformed using a feedforward neural network or a sofimax function.

[0024] In a second aspect, the present application provides a large language model, including an embedding layer, a position encoding layer, an output layer, and a plurality of Transformer Blocks; each Transformer Block includes an attention layer and an MLP layer;

[0025] The embedding layer is used to convert each word in the text data into a feature vector of fixed length; the position encoding layer is used to map the position of each word in the text data to the feature space and superimpose it on the feature vector corresponding to the word to obtain an input embedding matrix;

[0026] The attention layer includes a first regularization module, a self-attention transformation module and a first superposition module; the first regularization module is used to regularize the input embedding matrix; the self-attention transformation module is used to perform QKV transformation on the regularized matrix to obtain a query matrix, a key matrix and a value matrix, and perform multiple residual decompositions on the key matrix and the value matrix respectively to obtain a key residual vector and a value residual vector after each decomposition; cluster and compress the key residual vector and the value residual vector after each decomposition respectively, and perform attention calculation on the query matrix, the compressed key residual vector and the value residual vector; accumulate all attention calculation results to obtain an attention output matrix; the first superposition module is used to superimpose the input embedding matrix and the attention output matrix to obtain an attention layer output matrix;

[0027] The MLP layer includes a second regularization module, an MLP transformation module and a second superposition module; the second regularization module is used to regularize the attention layer output matrix; the MLP transformation module is used to extract features from the regularized matrix to obtain a feature matrix; the second superposition module is used to superimpose the attention layer output matrix and the feature matrix.

[0028] In a third aspect, the present application also provides an electronic device, including a memory, a processor, and a computer program / instructions stored on the memory, wherein the processor executes the computer program / instructions to implement the steps of the key-value compression method in the self-attention mechanism as described above in the present application.

[0029] In a fourth aspect, the present application also provides a computer-readable storage medium having a computer program / instruction stored thereon, which, when executed by a processor, implements the steps of the key-value compression method in the self-attention mechanism as described above in the present application.

[0030] In a fifth aspect, the present application also provides a computer program product, including a computer program / instruction, which, when executed by a processor, implements the steps of the key-value compression method in the self-attention mechanism as described above in the present application.

[0031] The beneficial effects of this application are:

[0032] This application performs multiple residual decompositions on the key matrix and the value matrix, and then clusters and compresses the key matrix, the value matrix, the key residual vector and the value residual vector after each decomposition, respectively, and performs attention calculation and accumulation to obtain approximate attention output, thereby improving the efficiency of Transformer in processing long sequences; at the same time, it solves the problem that the Linear Transformer cannot use the standard Softmax Transformer parameters and has large differences from the standard Softmax Transformer.

[0033] This application implements limited space key-value storage of long sequences, reducing the training storage overhead and computational complexity from N 2 The inference storage overhead is reduced from N to 1, and the inference computation is reduced from N to 1. 2 Compared with LinearAttention, the method of this application has the advantage of not requiring re-pre-training. In addition, the model weights fine-tuned by the method of this application can be used for reasoning of the large language model of this application (i.e., the approximate Transformer model) and the standard Transformer model, so it has better model application flexibility than Linear Attention. In addition, since the method of this application still uses softmax for weight rearrangement, the model effect is closer to the standard transformer than various Linear Attention models. BRIEF DESCRIPTION OF THE DRAWINGS

[0034] In order to more clearly illustrate the technical solution of the present application, the drawings required for use in the description of the embodiments will be briefly introduced below. Obviously, the drawings described below are only one embodiment of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.

[0035] Figure 1 This is a flow chart of the key value compression method in the self-attention mechanism in the embodiment of the present application;

[0036] Figure 2 This is a detailed process diagram of key value compression in an embodiment of the present application;

[0037] Figure 3 This is the loss curve in the embodiment of this application;

[0038] Figure 4 This is a diagram of the large language model architecture in the embodiment of this application;

[0039] Figure 5 This is a schematic diagram of the structure of an electronic device in an embodiment of the present application. DETAILED DESCRIPTION

[0040] The following is a clear and complete description of the technical solutions in this application in conjunction with the drawings in the embodiments of this application. Obviously, the described embodiments are only part of the embodiments of this application, not all of them. Based on the embodiments in this application, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of this application.

[0041] The technical solution of the present application is described in detail with specific embodiments below. The following specific embodiments can be combined with each other, and the same or similar concepts or processes may not be described in detail in some embodiments.

[0042] Embodiment 1

[0043] Figure 1 The flowchart of the key value compression method in the self-attention mechanism provided by this application is shown. Figure 1 As shown, the compression method comprises the following steps:

[0044] Step S1: performing multiple residual decompositions on the key matrix and the value matrix respectively, to obtain a key residual vector and a value residual vector after each decomposition; wherein the key residual vector after each decomposition includes the key residual vector (i.e., the key matrix) after the 0th decomposition, and the value residual vector after each decomposition includes the value residual vector (i.e., the value matrix) after the 0th decomposition;

[0045] Step S2: cluster and compress the key residual vector and the value residual vector after each decomposition, and perform attention calculation on the query matrix, the compressed key residual vector and the value residual vector;

[0046] Step S3: Accumulate all attention calculation results.

[0047] The query, key, and value in the attention mechanism are represented by Q, K, and V, respectively. Therefore, the query matrix, key matrix, and value matrix are represented by Q, K, and V, respectively. 0 , K 0 、V 0 Indicates that K 0 、V 0 They also represent the key residual vector and value residual vector after the 0th residual decomposition. Performing QKV decomposition on the attention input matrix X, we can get the query matrix, key matrix, and value matrix. The specific QKV decomposition formula is:

[0048] Q 0 =XW q , K 0 =XW k , V 0 =XW v (1)

[0049] Among them, W q , W k , W v All of them are weight matrices. Note that the size of the input matrix X is N×d, where N represents the sequence length and d represents the embedding dimension size. Then the size of the query matrix, key matrix, and value matrix are all N×d, and the weight matrix W q , W k , Wv The size of the weight matrix W is d×d. q , W k , W v The value of is gradually adjusted during training as the loss function gradient descents. It should be noted that the embedding dimension d of the query matrix can be further divided into multiple heads, and the number of heads of the embedding dimensions of the query matrix, key matrix, and value matrix may not be the same. For ease of introduction, but without affecting generality, it is assumed that the embedding dimensions of the query matrix, key matrix, and value matrix are the same, and are d.

[0050] The idea of ​​key-value compression is to train a fixed-length multi-head Codebook token sequence, where each head is used for the residual clustering of the token vectors of the key matrix and the value matrix, and after clustering, the key residual vector and the value residual vector of each decomposition are compressed to finally obtain the compressed key residual vector and value residual vector. This embodiment takes three-head KV compression as an example to illustrate the specific compression process of the key matrix and the value matrix. Figure 2 As shown, the specific compression process is:

[0051] Step B1: Perform a residual decomposition on the key matrix and the value matrix respectively to obtain the key residual vector and the value residual vector after the decomposition. The specific decomposition formula is:

[0052] K 1 =K 0 -gather(C 0 [0],index1) (2)

[0053] index1 = argmax(matmul(K 0 , C 0 [0].transpose)) (3)

[0054] V 1 =V 0 -gather(C 0 [1],index2) (4)

[0055] index2 = argmax(matmul(V 0 , C 0 [1].transpose)) (5)

[0056] Among them, K 1 represents the key residual vector after the first residual decomposition (i.e. the key residual vector after the first decomposition), K 0 represents the key residual vector (that is, the original key matrix) after the 0th residual decomposition; gather represents the selection of C according to the index position indexl 0 [0] and splice; C 0Represents a trainable matrix and has a size of C×d, that is, C 0 is a Codebook token sequence of length C, C 0 [0] represents the trainable matrix C 0 The 0th dimension, C 0 [1] represents the trainable matrix C 0 argmax means returning the index of the maximum value of the matrix in the specified dimension, matmul means matrix multiplication, and transpose means matrix transposition; V 1 Represents the residual vector after the first residual decomposition (i.e., the residual vector after the first decomposition), V 0 Represents the residual vector (i.e., the original value matrix) after the 0th residual decomposition; gather represents the selection of C according to the index position index2 0 [1] and splice.

[0057] Step B2: Perform secondary residual decomposition on the key residual vector and value residual vector after the primary decomposition, respectively, to obtain the key residual vector and value residual vector after the secondary decomposition. The specific decomposition formula is:

[0058] K 2 =K 1 -gather(C 1 [0],index1) (6)

[0059] index1 = argmax(matmul(K 1 , C 1 [0].transpose)) (7)

[0060] V 2 =V 1 -gather(C 1 [1],index2) (8)

[0061] index2 = argmax(matmul(V 1 , C 1 [1].transpose)) (9)

[0062] Among them, K 2 represents the key residual vector after the second residual decomposition (i.e. the key residual vector after the secondary decomposition); C 1 Represents a trainable matrix and has a size of C×d, that is, C 1 is a Codebook token sequence of length C, C 1 [0] represents the trainable matrix C 1 The 0th dimension, C 1 [1] represents the trainable matrix C1 The first dimension of 2 Represents the residual vector of the value after the second residual decomposition (i.e., the residual vector of the value after the secondary decomposition).

[0063] Step B3: Key matrix K 0 , the key residual vector K after the first decomposition 1 , the key residual vector K after the second decomposition 2 , value matrix V 0 , the residual vector V after the first decomposition 1 , the residual vector V after the secondary decomposition 2 Cluster compression is performed separately.

[0064] Key matrix K 0 Sum value matrix V 0 Corresponding to a compression head, the key residual vector K after one decomposition 1 And the residual vector V after the first decomposition 1 Corresponding to a compression head, the key residual vector K after secondary decomposition 2 And the residual vector V after the second decomposition 2 Corresponding to one compression head, the clustering compression process of each compression head is the same. 0 Sum value matrix V 0 Taking cluster compression as an example, the specific cluster compression process is as follows:

[0065] Step B3.1: Calculate the bond matrix K 0 With the trainable matrix C 0 The correlation matrix of is calculated as follows:

[0066] A 0 =matmul(C 0 [0],K 0 ) (10)

[0067] Among them, A 0 Denotes the key matrix K 0 With the trainable matrix C 0 The correlation matrix of is of size C×N. Figure 2 A in 1 Represents the key residual vector K after a decomposition 1 With the trainable matrix C 1 The correlation matrix, A 2 represents the key residual vector K after secondary decomposition 2 With the trainable matrix C 2 The correlation matrix of .

[0068] Step B3.2: For the correlation matrix A 0 Perform mapping transformation to obtain the weight matrix F0 , the specific formula is:

[0069] F 0 =F(A 0 ) (11)

[0070] Wherein, F represents the mapping transformation function. The mapping transformation function of this embodiment is a feedforward neural network or a softmax function, that is, the correlation matrix A 0 After a feedforward neural network or softmax function transformation, the weight matrix F is obtained. 0 The weight matrix F 0 The nonlinear characteristics of are more prominent, and its size is still C×N. Figure 2 F 1 For A 1 The corresponding weight matrix, F 2 For A 2 The corresponding weight matrix.

[0071] Step B3.3: Key matrix K 0 With the weight matrix F 0 Perform clustering to obtain the compressed key residual vector K 0_z , the specific formula is:

[0072] K 0_z =matmul(F 0 , K 0 ) (12)

[0073] The compressed key residual vector K 0_z Can be regarded as the key matrix K 0 Cluster compression with the Codebook tokens sequence as the cluster center, the compressed key residual vector K 0_z The size is C×d. Figure 2 The key residual vector K in 1_z is the key residual vector K after a decomposition 1 The clustering compression result, key residual vector K 2_z is the key residual vector K after secondary decomposition 2 Clustering compression results.

[0074] Step B3.4: Pair value matrix V 0 With the weight matrix F 0 Perform clustering to obtain the compressed value residual vector V 0_z , the specific formula is:

[0075] V 0_z =matmul(V 0 , K 0 ) (13)

[0076] The compressed value residual vector V 0_z Can be viewed as a value matrix V 0 Cluster compression with the Codebook tokens sequence as the cluster center, the compressed value residual vector V 0_z The size is C×d. Figure 2 The residual vector V 1_z is the residual vector V after a decomposition 1 The clustering compression result, the residual vector V 2_z is the residual vector V after the secondary decomposition 2 Clustering compression results.

[0077] The clustering compression process of the three compression heads is exactly the same, so in the implementation, the compression head can be used as the batch dimension for parallelization, and the outputs of multiple heads can be merged into K _z ={K 0_z , K 1_z , K 2_z}、V _z = {V 0_z , V 1_z , V 2_z}.

[0078] Step B4: Perform attention calculation on the query matrix, the compressed key residual vector and the value residual vector.

[0079] The clustering compression result of each compression head is calculated with the query matrix for attention. The specific calculation formula is:

[0080] P[i]=softmax(matmul(Q 0 , K _z [i].transpose)) (14)

[0081] Out[i] = matmul(P[i], V _z [i]) (15)

[0082] Where P[i] represents the intermediate matrix, i = 0, 1, 2, and the size of P[i] is N×C; K _z [i] indicates K _z The i-th element in K i_z ; Out[i] represents the clustering compression result of the i-th compression head and the attention calculation result of the query matrix. The size of Out[i] is N×d.

[0083] Step B5: Accumulate all attention calculation results. The specific accumulation formula is:

[0084] Out=Out[0]+Out[1]+Out[2] (16)

[0085] Among them, Out represents the attention output matrix, that is, the final approximate attention output.

[0086] In order to illustrate the effectiveness of the method of this application, the weight parameters of the Mistral-7B model are inherited, and then its standard Attention is replaced with the key value compression method of this application. The loss curve of the Mistral-7B model fine-tuned is as follows: Figure 3 When fine-tuning, only the weight matrix W is trained q , W k , W v As well as the parameters of the embedding layer and the regularization module, other parameters are frozen to increase the training speed. The training uses two A100 computing cards, the sequence length is 2K, and the gradient accumulation is 64. Figure 3 It can be seen that in less than 200 steps, the linear attention of this application converges to the loss of the original Mistral-7B model.

[0087] After the Mistral-7B model is trained for 1000 steps, the perplexity results tested on the PG-19 dataset and the Proof-Pile dataset are shown in Table 1. The test lengths are 512, 1K, and 2K, respectively.

[0088] Table 1. The perplexity results of Mistral-7B model on different datasets

[0089] Test length PG-19 Dataset Proof-Pile Dataset 512 7.39 3.58 1024 7.76 3.69 2048 8.17 3.83

[0090] As can be seen from Table 1, this application achieves considerable perplexity on multiple data sets.

[0091] The effectiveness of multi-head KV compression comes from the residual decomposition and cluster compression of the key matrix and value matrix. Because when compressing KV using the Codebook tokens sequence, the feature-related KV is compressed into cluster centers and stored in K i_z 、V i_z. However, due to the limitation of storage size, the shared features stored in a single cluster center are very rough, so the present application adopts a multi-level residual decomposition and cluster compression method to solve this problem, that is, a multi-head compression method. For example, the first-level KV maintains the coarsest but stored information-rich features as cluster centers, and the second-level performs a finer decomposition on the rough cluster centers, that is, decomposing the large categories of the first level into small categories. The second-level KV stores cluster features relative to the first-level categories, and its effectiveness lies in the shared feature differences between the first and second levels. The more shared, the better the performance. The third level is similar to the above process, with finer clustering within the second-level categories. The optimal situation is that the feature differences can be fully shared, thereby achieving a storage space with a power exponential of the parameter size. The upper-level application can adjust the compression rate and restoration degree to a suitable ratio by adjusting the number of compression heads.

[0092] This application uses a multi-level codebook to effectively cluster and compress the KV matrix. The compressed approximate KV matrix (i.e., K i_z 、V i_z ) has a fixed length, which makes the attention calculation between the query matrix and the approximate KV matrix linear with the length of the original key and value matrices, and only requires constant-level storage, with the same computational and storage complexity as linear attention. Compared with traditional linear attention, this application still supports the use of softmax to adjust the weights of tokens similarity, making the result closer to the standard softmax attention.

[0093] The method of this application relies on a set of fine-tuned trained Codebook token sequences to map long sequence tokens of size N×d to a fixed size C×d token sequence; the method of this application improves the standard SoftmaxAttention method, calculates the approximate attention output, and achieves a better approximation to the existing Softmax Attention. This application inherits the weights of the current large language model based on Softmax Attention and fine-tunes it to convert SoftmaxAttention into a linear recursive method: sliding window combined with the storage mechanism of multi-head KV compression of this application. Experiments show that this application can achieve convergence with very few training resources and achieve considerable perplexity on multiple data sets.

[0094] The method of this application realizes the limited space key-value storage of long sequences, reducing the training storage overhead and computational complexity from N 2 The inference storage overhead is reduced from N to 1, and the inference computation is reduced from N to 1. 2Compared with Linear Attention, the method of this application has the advantage of not requiring re-pre-training. In addition, the model weights fine-tuned by the method of this application can be used for reasoning of the large language model of this application (i.e., the approximate Transformer model) and the standard Transformer model, so it has better model application flexibility than Linear Attention. In addition, since the method of this application still uses softmax for weight rearrangement, the model effect is closer to the standard transformer than various Linear Attention models.

[0095] Embodiment 2

[0096] Figure 4 The large language model architecture diagram provided by this application is shown, and the square brackets indicate the parameters that the model needs to train, mainly including the weight parameters of the embedding layer, the weight matrix W q , W k , W v , trainable matrix C 0 ~C 2 , MLP weight parameters. Figure 4 As shown, the large language model (i.e., the approximate Transformer model) includes an embedding layer, a position encoding layer, an output layer, and multiple Transformer Blocks; each Transformer Block includes an attention layer and an MLP layer, the attention layer includes a first regularization module, a self-attention transformation module, and a first superposition module connected in sequence, and the MLP layer includes a second regularization module, an MLP transformation module, and a second superposition module connected in sequence.

[0097] The embedding layer is used to convert each word in the text data into a feature vector of fixed length: the position encoding layer is used to map the position of each word in the text data to the feature space and superimpose it on the feature vector corresponding to the word to obtain the input embedding matrix.

[0098] The first regularization module is used to regularize the input embedding matrix; the self-attention transformation module is used to perform QKV transformation on the regularized matrix to obtain a query matrix, a key matrix and a value matrix, and perform multiple residual decompositions on the key matrix and the value matrix respectively to obtain a key residual vector and a value residual vector after each decomposition; cluster and compress the key residual vector and the value residual vector after each decomposition respectively, and perform attention calculation on the query matrix, the compressed key residual vector and the value residual vector; accumulate all attention calculation results to obtain an attention output matrix Out, for the specific process, refer to the method in Example 1 of the present application; the first superposition module is used to superimpose the input embedding matrix with the attention output matrix Out to obtain an attention layer output matrix.

[0099] The second regularization module is used to regularize the output matrix of the attention layer; the MLP transformation module is used to extract features from the regularized matrix to obtain a feature matrix; the second superposition module is used to superimpose the output matrix of the attention layer with the feature matrix to obtain the output matrix of the MLP layer.

[0100] The entire large language model can call Transformer Block multiple times in a loop to implement multiple TransformerBlocks. The input of the first call of Transformer Block is the input embedding matrix obtained by superimposing the input layer and the position encoding layer, and the input of subsequent calls is the output of the previous call, that is, the output matrix of the previous MLP layer.

[0101] The trainable matrix C of the large language model 0 ~C 2 Introduced in the fine-tuning stage, other parameters that need to be trained have the same parameter structure as the model pre-training stage, and the model inference stage can reuse the parameter results obtained in the pre-training stage.

[0102] The compression mechanism provided by this application theoretically covers and expands the most advanced key-value (KV) compression technology used in current linear attention models, including the gating mechanism of GLA, the ABC gating architecture of GSA, and the neural Turing machine (NTM) method. By separating storage from the current hidden state, the method of this application achieves a more flexible and robust sequence representation, which can adapt to various attention mechanisms, including the standard Softmax Attention and its linearized counterpart. The approximate Transformer model constructed using this compression mechanism of this application has better adaptability and is suitable for achieving high levels of performance in a wide range of long sequence tasks.

[0103] Embodiment 3

[0104] The embodiment of the present invention further provides an electronic device, such as Figure 5 As shown, the electronic device includes: a memory, a processor, and a computer program / instructions stored in the memory, and the processor executes the computer program / instructions to implement the key-value compression method in the self-attention mechanism in Example 1 of the present application.

[0105] Although not shown, the electronic device includes a processor, which can perform various appropriate operations and processes according to the program and / or data stored in the read-only memory (ROM) or the program and / or data loaded from the storage portion into the random access memory (RAM). The processor can be a multi-core processor, or it can include multiple processors. In some embodiments, the processor can include a general main processor and one or more special coprocessors, such as a central processing unit, a graphics processing unit (GPU), a neural network processor (NPU), a digital signal processor (DSP), etc. In RAM, various programs and data required for device operation are also stored. The processor, ROM and RAM are connected to each other via a bus. The input / output (I / O) interface is also connected to the bus.

[0106] The processor and memory are used together to execute the program / instructions stored in the memory. When the program / instructions are executed by the computer, the methods, steps or functions described in the above embodiments can be implemented.

[0107] Although not shown, an embodiment of the present invention further provides a computer-readable storage medium having a computer program / instruction stored thereon, which, when executed by a processor, implements the key-value compression method in the self-attention mechanism in the first embodiment of the present application.

[0108] Readable storage media include permanent and non-permanent, removable and non-removable media that can be implemented by any method or technology to store information. Information can be computer-readable instructions, data structures, program modules or other data. Examples of computer storage media include, but are not limited to, phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technology, compact disk read-only memory (CD-ROM), digital versatile disk (DVD) or other optical storage, magnetic cassettes, disk storage or other magnetic storage devices or any other non-transmission media that can be used to store information that can be accessed by a computing device. As defined herein, computer-readable media does not include temporary computer-readable media (transitory media), such as modulated data signals and carrier waves.

[0109] Although not shown, an embodiment of the present invention also provides a computer program product, including: a computer program / instruction, which, when executed by a processor, implements the key-value compression method in the self-attention mechanism in the first embodiment of the present application.

[0110] Although the preferred embodiments of the present application have been described, those skilled in the art may make additional changes and modifications to these embodiments once they are aware of the basic creative concepts. Therefore, the appended claims are intended to be interpreted as including the preferred embodiments and all changes and modifications that fall within the scope of the present application. Obviously, those skilled in the art may make various changes and modifications to the present application without departing from the spirit and scope of the present application. Thus, if these modifications and variations of the present application fall within the scope of the claims of the present application and their equivalents, the present application is also intended to include these modifications and variations.

Claims

1. A key value compression method in a self-attention mechanism, applied to a large language model, characterized in that: The compression method comprises: Perform multiple residual decompositions on the key matrix and the value matrix respectively, to obtain a key residual vector and a value residual vector after each decomposition; wherein the key residual vector after each decomposition includes the key residual vector after the 0th decomposition, and the value residual vector after each decomposition includes the value residual vector after the 0th decomposition; The key residual vector and value residual vector after each decomposition are clustered and compressed respectively, and attention calculation is performed on the query matrix, the compressed key residual vector and the value residual vector; All attention calculation results are accumulated.

2. The key value compression method in the self-attention mechanism according to claim 1, characterized in that: The specific formula for the residual decomposition of the bond matrix is: K i =K i-1 -gather(C i-1 [0],index1); index1=argmax(matmul(K i-1 ,C i-1 [0].transpose)); Among them, K i represents the key residual vector after the ith residual decomposition, K i-1 represents the key residual vector after the i-1th residual decomposition. When i=1, K0 represents the key matrix; gather represents selecting C according to index position index1 i-1 [0] and splice; C i-1 represents the trainable matrix, C i-1 [0] represents the trainable matrix C i-1 argmax means returning the index of the maximum value of the matrix in the specified dimension, matmul means matrix multiplication, and transpose means matrix transposition; The specific formula for the residual decomposition of the value matrix is: V i =V i-1 -gather(C i-1 [1],index2); index2=argmax(matmul(V i-1 ,C i-1 [1].transpose)); Among them, V i Represents the residual vector after the i-th residual decomposition, V i-1 Represents the residual vector after the i-1th residual decomposition. When i=1, V0 represents the value matrix; gather represents selecting C according to index position index2 i-1 [1] and splice; C i-1 [1] represents the trainable matrix C i-1 The first dimension of .

3. The key value compression method in the self-attention mechanism according to claim 1 or 2, characterized in that: The key residual vector and value residual vector after each decomposition are clustered and compressed respectively, including: Calculating a correlation matrix between the key residual vector and the trainable matrix; Performing mapping transformation on the correlation matrix to obtain a weight matrix; Clustering the key residual vector and the weight matrix to obtain a compressed key residual vector; The value residual vector and the weight matrix are clustered to obtain a compressed value residual vector.

4. The key value compression method in the self-attention mechanism according to claim 3, characterized in that: Performing a mapping transformation on the correlation matrix includes: The correlation matrix is ​​mapped and transformed using a feedforward neural network or a softmax function.

5. A large language model, comprising an embedding layer, a position encoding layer, an output layer, and a plurality of Transformer Blocks; each Transformer Block comprises an attention layer and an MLP layer; characterized in that, The embedding layer is used to convert each word in the text data into a feature vector of fixed length; the position encoding layer is used to map the position of each word in the text data to the feature space and superimpose it on the feature vector corresponding to the word to obtain an input embedding matrix; The attention layer includes a first regularization module, a self-attention transformation module and a first superposition module; the first regularization module is used to perform regularization processing on the input embedding matrix; the self-attention transformation module is used to perform QKV transformation on the regularized matrix to obtain a query matrix, a key matrix and a value matrix, and perform multiple residual decompositions on the key matrix and the value matrix respectively to obtain a key residual vector and a value residual vector after each decomposition; clustering and compressing the key residual vector and the value residual vector after each decomposition respectively, and performing attention calculation on the query matrix, the compressed key residual vector and the value residual vector; All attention calculation results are accumulated to obtain an attention output matrix; the first superposition module is used to superimpose the input embedding matrix and the attention output matrix to obtain an attention layer output matrix; The MLP layer includes a second regularization module, an MLP transformation module and a second superposition module; the second regularization module is used to regularize the attention layer output matrix; the MLP transformation module is used to extract features from the regularized matrix to obtain a feature matrix; the second superposition module is used to superimpose the attention layer output matrix and the feature matrix.

6. The large language model according to claim 5, characterized in that: The specific formula for the self-attention transformation module to perform multiple residual decompositions on the key matrix and the value matrix is: K i =K i-1 -gather(C i-1 [0],index1); index1=argmax(matmul(K i-1 ,C i-1 [0].transpose)); V i =V i-1 -gather(C i-1 [1],index2); index2=argmax(matmul(V i-1 ,C i-1 [1].transpose)); Among them, K i represents the key residual vector after the ith residual decomposition, K i-1 represents the key residual vector after the i-1th residual decomposition. When i=1, K0 represents the key matrix; gather represents selecting C according to index position index1 i-1 [0] and splice; Ci -1 represents the trainable matrix, C i-1 [0] represents the trainable matrix C i-1 argmax means returning the index of the maximum value of the matrix in the specified dimension, matmul means matrix multiplication, and transpose means matrix transposition; V i Represents the residual vector after the i-th residual decomposition, V i-1 Represents the residual vector after the i-1th residual decomposition. When i=1, V0 represents the value matrix; gather represents selecting C according to index position index2 i-1 [1] and splice; C i-1 [1] represents the trainable matrix C i-1 The first dimension of .

7. The large language model according to claim 5, characterized in that: The self-attention transformation module is used to perform clustering compression on the key residual vector and the value residual vector after each decomposition, including: Calculating a correlation matrix between the key residual vector and the trainable matrix; Performing mapping transformation on the correlation matrix to obtain a weight matrix; Clustering the key residual vector and the weight matrix to obtain a compressed key residual vector; The value residual vector and the weight matrix are clustered to obtain a compressed value residual vector.

8. An electronic device comprising a memory, a processor, and a computer program / instructions stored on the memory, wherein the processor executes the computer program / instructions to implement the steps in the key-value compression method in the self-attention mechanism as described in any one of claims 1-4.

9. A computer-readable storage medium having a computer program / instruction stored thereon, which, when executed by a processor, implements the steps in the key-value compression method in the self-attention mechanism as described in any one of claims 1 to 4.

10. A computer program product comprising a computer program / instructions, characterized in that When the computer program / instructions are executed by a processor, the steps in the key-value compression method in the self-attention mechanism as described in any one of claims 1 to 4 are implemented.

Citation Information

Patent Citations

  • Point cloud attribute compression method based on inter-block prediction and graph Fourier transform

    CN115802048A

  • DETR improved model-based sparse attention target detection method

    CN117152416A

  • Large language model training method and device

    CN118445379A

  • Text prediction model training method and device, and text prediction method and device

    WO2024228666A1

  • Data processing method and related device

    WO2025030840A1