Transform low-rank decomposition method based on Query-Key joint decomposition

By using the low-rank decomposition method of Query-Key joint decomposition in Transformer, the problem of the inability to remove redundant information in the Attention module in the prior art is solved, and efficient model compression and calculation efficiency improvement are achieved.

CN120087424APending Publication Date: 2025-06-03TIANJIN UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510245919.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-04
Publication Date
2025-06-03

AI Technical Summary

Technical Problem

The existing Transformer lightweight method based on low-rank decomposition cannot remove the redundant information introduced after multiplication of the weight matrix of Query and Key when compressing the linear layer in the Attention module, resulting in limited model compression effect.

Method used

The Transformer low-rank decomposition method based on Query-Key joint decomposition is used to decompose the weight matrix of the attention module in Transformer through the weight joint decomposition strategy, and the decomposed bias vector is initialized using the bias alignment strategy to achieve efficient compression of the model.

Benefits of technology

With almost no loss of model accuracy, extremely high compression effect is achieved, reducing the calculation amount and parameter amount of the model, and significantly improving the calculation efficiency and deployment ability of the model.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120087424A_ABST
    Figure CN120087424A_ABST
Patent Text Reader

Abstract

The invention relates to a Transform low-rank decomposition method based on Query-Key joint decomposition. The method comprises the following steps: S1, decomposing a weight matrix of an attention module in Transform by using a weight joint decomposition strategy; s2, using an offset alignment strategy to initialize an offset vector after decomposition of an attention module in the Transform; s3, evaluating a calculation amount and a parameter amount after the decomposition of the Transform model; and S4, performing fine tuning on the decomposed Transform model. According to the method, the compression effect and the expression ability of the Transform model are effectively improved, and the performance close to that of an original model can be kept under the condition that the parameter quantity and the calculation quantity are remarkably reduced.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of deep learning, and specifically relates to a Transformer low-rank decomposition method based on Query-Key joint decomposition. Background Art

[0002] Transformer has developed rapidly in recent years and demonstrated advanced performance in various vision tasks such as image classification, object detection, and semantic segmentation. Due to its flexibility in processing various input formats, Transformer is also widely used in self-supervised learning and other modalities. However, since Transformer usually stacks multiple identical blocks, the attention module in each Transformer block involves the calculation of multiple key, query, and value matrices, and the sizes of these matrices are often very large. This design makes the Transformer model have a large number of parameters, high demands on memory and computing resources, resulting in high latency and high energy consumption problems during the inference stage, which limits the deployment of the model on resource-constrained platforms and poses challenges to practical applications. Therefore, compressing Transformer to improve computational efficiency is crucial for promoting its wider application.

[0003] Existing Transformer lightweight methods based on low-rank decomposition usually decompose each linear layer (including the Attention module and the FFN module) in Transformer independently. However, this independent decomposition strategy has deficiencies in compressing the linear layer in the Attention module, that is, it cannot remove the redundant information introduced in the product matrix after the weight matrices of Query and Key in the linear layer are multiplied. Specifically, the direct multiplication of the outputs of Query and Key in the linear layer of the Attention module is equivalent to the multiplication of the weight matrices of these two linear layers and forms a new weight matrix. Matrix multiplication introduces additional redundant information to the product matrix. Separately decomposing the weight matrices of Query and Key in the linear layer only removes their own redundant information and cannot remove the additional redundant information in the product matrix. This leads to the inability to fully exploit the compression potential of the Attention module, and the model compression effect is limited.

[0004] To solve the above problems, the present invention proposes a Transformer low-rank decomposition method based on Query-Key joint decomposition. After retrieval, no literature of prior art identical or similar to the present invention has been found. Summary of the Invention

[0005] The purpose of the present invention is to overcome the deficiencies of the prior art and provide a Transformer low-rank decomposition method based on Query-Key joint decomposition, achieving an extremely high compression effect while hardly sacrificing the model accuracy.

[0006] The present invention solves its technical problems through the following technical solutions:

[0007] A Transformer low-rank decomposition method based on Query-Key joint decomposition, and the steps of the method are as follows:

[0008] S1. Use the weight joint decomposition strategy to decompose the weight matrix of the attention module in the Transformer;

[0009] S2. Use the bias alignment strategy to initialize the bias vector after decomposition of the attention module in the Transformer;

[0010] S3. Evaluate the computational complexity and the number of parameters after decomposition of the Transformer model;

[0011] S4. Fine-tune the decomposed Transformer model through knowledge distillation.

[0012] Moreover, the specific content of S1 is as follows:

[0013] 1) Collect the weight matrices W Q and W K of the linear layer Query and the linear layer Key included in the attention module of the Transformer;

[0014] 2) Calculate the product of the weight matrices W Q and W K of the linear layer Query and Key to obtain the joint matrix W QK :

[0015]

[0016] 3) Use singular value decomposition to decompose the joint matrix W QK , and after decomposition, obtain three matrices U QK , S QK and V QK :

[0017] U QK , S QK , V QK = f svd (W QK )

[0018] Where: U QK and V QKThey are the left and right singular value matrices respectively;

[0019] S QK is the singular value matrix:

[0020] 4) Combine the diagonal values of the diagonal matrix S QK into the left and right orthogonal matrices S QK and V QK to form matrix U and matrix V. The combination relationship satisfies:

[0021]

[0022] Where: represents the diagonal matrix formed by taking the square root of the singular values of the diagonal matrix S QK ;

[0023] The combined matrix W QK is represented by the product of matrix U and matrix V:

[0024] W QK = UV T

[0025] 5) Retain the first r row vectors of matrix U and the first r column vectors of matrix V to obtain the approximate matrices W U and W V of matrix U and V, and approximate the combined matrix W U and W V with matrix W QK :

[0026]

[0027] 6) Matrix W U and W V serve as the weight matrices of the linear layer formed after decomposition, and their matrix sizes are smaller than the original weight matrices W Q and W K . The mathematical formula of the attention mechanism is:

[0028]

[0029] Moreover, the specific S2 is:

[0030] 1) Expand the original mathematical formula of the attention score to:

[0031]

[0032] Where: b q and b k are the bias vectors corresponding to the weight matrices W Q and W K of the linear layer Query and Key respectively;

[0033] 2) The mathematical formula of the attention score expands after decomposition as follows:

[0034]

[0035] Where: b u and b v are the bias vectors corresponding to the weight matrices W U and W V respectively, and b u and b v are variables to be solved;

[0036] 3) Align the first-order term coefficients of the mathematical formulas of the original attention score and the decomposed attention score to obtain the following equation:

[0037]

[0038] 4) Solve the equation in 3), and the calculation formulas for the bias vectors b u and b v are:

[0039]

[0040] The bias vectors b u and b v satisfy the consistency of the constant term coefficients of the mathematical formulas of the original attention score and the decomposed attention score;

[0041] 5) Use the weight matrices W U and W V and the corresponding bias vectors b u and b v to construct a new linear layer, and replace the original linear layers Query and Key with the new linear layer to obtain the finally decomposed Transformer model.

[0042] The positive effects that can be produced by the present invention are:

[0043] 1. Starting from the product characteristics of the linear layer in the attention module of the Transformer, the present invention designs a weight joint decomposition strategy, effectively avoiding the redundancy problem that may be introduced by multiplying the independently decomposed weight matrices, and significantly improving the efficiency and performance of model compression.

[0044] 2. The present invention proposes a bias alignment strategy, effectively ensuring the mathematical consistency of the attention score before and after decomposition, solving the problem of performance degradation in the initial stage of fine-tuning caused by initializing the bias vector to zero in the past, and effectively shortening the convergence time in the fine-tuning stage. Brief Description of the Drawings

[0045] Figure 1 This is the overall framework diagram of the present invention;

[0046] Figure 2 This is the schematic diagram of the weight joint decomposition strategy of the present invention;

[0047] Figure 3 This is the schematic diagram of the bias alignment strategy of the present invention. Detailed implementation manners

[0048] The present invention will be further described in detail below through specific embodiments. The following embodiments are only descriptive and not restrictive, and the protection scope of the present invention cannot be limited thereby.

[0049] A Transformer low-rank decomposition method based on Query-Key joint decomposition, and its innovation lies in that the steps of the method are as follows:

[0050] Step 1: Use the weight joint decomposition strategy to decompose the weight matrix of the attention module in the Transformer; the specific steps of Step 1 include:

[0051] 1.1. Collect the weight matrices W Q and W K of the linear layer Query and the linear layer Key included in the attention module of the Transformer;

[0052] 1.2. Convert the calculation formula of the attention mechanism in the Transformer:

[0053]

[0054]

[0055] 1.3. Calculate the product of the weight matrices W Q and W K of the linear layer Query and Key to obtain the joint matrix W QK :

[0056]

[0057] 1.4. Use singular value decomposition to decompose the joint matrix W QK , and after decomposition, obtain three matrices U QK , S QK and V QK :

[0058] U QK , S QK , V QK = f svd (W QK )

[0059] In the formula: where U QK and V QK are orthogonal matrices, S QK is a diagonal matrix, and the product of the three matrices after decomposition is equal to the joint matrix W QK :

[0060]

[0061] 1.5. Merge the diagonal values of the diagonal matrix S QK into the left and right orthogonal matrices S QK and V QK to form the matrix U and the matrix V. The merging relationship satisfies:

[0062]

[0063] In the formula: represents the diagonal matrix formed by taking the square root of the singular values of the diagonal matrix S QK . The joint matrix W QK can be expressed as the product of the matrix U and the matrix V:

[0064] W QK = UV T

[0065] 1.6. Retain the first r row vectors of the matrix U and the first r column vectors of the matrix V to obtain the approximate matrices W U and W V of the matrices U and V, and approximate the joint matrix W U and W V with the matrices W QK :

[0066]

[0067] 1.7. The matrices W U and W V serve as the weight matrices of the linear layer formed after decomposition, and their matrix sizes are smaller than the original weight matrices W Q and W K . The mathematical formula of the attention mechanism is:

[0068]

[0069] Step 2. Initialize the bias vector after decomposition in the attention module of the Transformer using the bias alignment strategy; the specific steps of the said Step 2 include:

[0070] 2.1. Expand the original mathematical formula of the attention score into:

[0071]

[0072] In the formula: b q and b k are the bias vectors corresponding to the weight matrices W Q and W K of the linear layer Query and Key respectively;

[0073] 2.2. The mathematical formula of the attention score is expanded after decomposition as:

[0074]

[0075] In the formula: b u and b v are the bias vectors corresponding to the weight matrices W U and W V respectively, and b u and b v are variables to be solved;

[0076] 2.3. Align the first-order term coefficients of the mathematical formulas of the original attention score and the decomposed attention score to obtain the following equation:

[0077]

[0078] 2.4. Solve the equation in step 2.3, and the calculation formulas for the bias vectors b u and b v are:

[0079]

[0080] The bias vectors b u and b v satisfy that the constant term coefficients of the mathematical formulas of the original attention score and the decomposed attention score are the same;

[0081] 2.5. Use the weight matrices W U and W V and the corresponding bias vectors b u and b v to construct a new linear layer, and replace the original linear layer Query and Key with the new linear layer to obtain the finally decomposed Transformer model.

[0082] Step 3: Evaluate the computational complexity and the number of parameters of the decomposed Transformer model; the calculation methods of the model's computational complexity and the number of parameters in step 3 are:

[0083] Use FLOPs and Params to represent the computational complexity and the number of parameters of the Transformer. In the structure of the Transformer, since the calculations of each layer are based on linear transformations, both the FLOPs and Params of the Transformer can be measured by the calculation formulas for the computational complexity and the number of parameters of the linear layer;

[0084] FLOPs: The computational complexity of the model, representing the amount of floating-point operations required for a single forward pass. The calculation formula for the computational complexity of the linear layer is as follows:

[0085] FLOPs = 2 × N in × N out

[0086] Where: N in is the input feature dimension, and N out is the output feature dimension.

[0087] Params: The number of parameters of the model, representing the total number of trainable parameters (weights and biases) in the neural network. The calculation formula for the number of parameters of the linear layer is as follows:

[0088] Params = N in × N out + N out

[0089] Where: N in is the input feature dimension, N out is the output feature dimension, and the additional N out is the number of bias terms.

[0090] Step 4, fine-tune the decomposed Transformer model; the loss function in the fine-tuning process of Step 4 is:

[0091] Fine-tune the model by distillation. The loss function of distillation includes soft label loss, hard label loss, and intermediate feature loss. The distillation loss L is expressed as:

[0092] L = L soft + L hard + L intermediate

[0093] The soft label loss is L soft :

[0094]

[0095] Where: p i is the output of the teacher model, that is, the probability distribution after temperature scaling, and q i is the output probability of the student model;

[0096] The hard label loss is L hard :

[0097]

[0098] Where: y i is the output of the teacher model, and q i is the output probability of the student model;

[0099] The intermediate feature loss is L intermediate :

[0100]

[0101] Where: and represent the feature representations of the student and teacher models at a certain layer, l represents the number of layers of the Transformer, and the loss metric used is the Euclidean distance.

[0102] In this embodiment, the overall framework diagram of the Transformer low-rank decomposition method based on Query-Key joint decomposition is as shown in Figure 1 where the attention module is the core part of the Transformer network, and the low-rank decomposition method of the attention module includes a weight joint decomposition strategy and a bias alignment strategy.

[0103] Figure 2 As shown in the weight joint decomposition strategy of the present invention, this strategy is applied to decompose the weight matrices of the linear layers Query and Key in the attention module of the Transformer, W Q and W K of the product matrix W QK .

[0104] Figure 3 As shown in the bias alignment strategy of the present invention, this strategy is applied to initialize the bias vectors b of the linear layers Query and Key in the attention module of the Transformer u and b v .

[0105] The present invention compares various advanced Transformer compression methods on a classic Transformer structure - the DeiT model.

[0106] All methods are evaluated on the DeiT model, as shown in Table 1, where Baseline represents the original uncompressed DeiT model, and the remaining 7 methods are comparative methods. Compared with the original model, the model compressed by the method of the present invention has a 45% reduction in FLOPs and the number of parameters compared to the uncompressed original DeiT model, but the accuracy only loses 0.3% compared to the original model. In addition, compared with the other 7 advanced Transformer compression methods, the compression effect and accuracy of this method are the best. Specifically, this method has the largest reduction in FLOPs and the number of parameters while having the highest accuracy.

[0107] Table 1

[0108] Method FLOPs FLOPssaving Params Paramssaving Acc Baseline 17.6G 0 86.4M 0 81.8% WDPruning 11.0G 37% 60.6M 30% 81.1% VTP 9.9G 43% 49.2M 43% 80.7% ToMe 11.5G 35% - - 80.5% EViT 11.5G 35% - - 80.3% LRKD 10.5G 40% 51.8M 40% 79.1% FWSVD 10.5G 40% 51.8M 40% 80.3% AFM 9.9G 43% 49.2M 43% 81.1% Ours 9.7G 45% 47.5M 45% 81.5%

[0109] Although embodiments and drawings of the present invention are disclosed for illustrative purposes, those skilled in the art can understand that: without departing from the spirit and scope of the present invention and the appended claims, various substitutions, changes, and modifications are possible. Therefore, the scope of the present invention is not limited to the content disclosed in the embodiments and drawings.

Claims

1. A Transformer low-rank decomposition method based on Query-Key joint decomposition, characterized by: The steps of the method are: S1. Decompose the weight matrix of the attention module in Transformer using the weight joint decomposition strategy; S2, use the bias alignment strategy to initialize the bias vector after the decomposition of the attention module in Transformer; S3, evaluate the computational effort and parameter count of the Transformer model after decomposition; S4. Fine-tune the decomposed Transformer model through knowledge distillation.

2. The Transformer low-rank decomposition method based on Query-Key joint decomposition according to claim 1, characterized in that: The S1 is specifically: 1) Collect the weight matrix W of the linear layer Query and linear layer Key contained in the Transformer's attention module Q and W K ; 2) Calculate the weight matrix W of the linear layer Query and Key Q and W K The product of QK : 3) Decompose the joint matrix W using singular value decomposition QK , after decomposition, we get 3 matrices U QK , S QK and V QK : U QK ,S QK ,V QK =f svd (W QK ) Among them: U QK and V QK are the left and right singular value matrices respectively; S QK is the singular value matrix: 4) The diagonal matrix S QK The diagonal values ​​of are merged into the left and right orthogonal matrices S QK and V QK In the above example, we form the matrix U and the matrix V, and the merging relationship satisfies: in: Denotes the diagonal matrix S QK The diagonal matrix formed by the square roots of the singular values ​​of ; Joint matrix W QK It can be expressed as the product of matrix U and matrix V: W QK =UV T 5) Keep the first r row vectors of matrix U and the first r column vectors of matrix V to get the approximate matrix W of matrix U and V U and W V , using the matrix W U and W V Approximate joint matrix W QK : 6) Matrix W U and W V As the weight matrix of the linear layer formed after decomposition, its matrix size is smaller than the original weight matrix W Q and W K , the mathematical formula of the attention mechanism is:

3. The Transformer low-rank decomposition method based on Query-Key joint decomposition according to claim 1, characterized in that: The S2 is specifically: 1) Expand the original mathematical formula of the attention score into: Where: b q and b k are the weight matrices W of the linear layer Query and Key respectively. Q and W K The corresponding bias vector; 2) The mathematical formula of the attention score is expanded after decomposition: Where: b u and b v are respectively the weight matrix W U and W V The corresponding bias vector, b u and b v is the variable to be solved; 3) Align the first-order coefficients of the mathematical formula of the original attention score and the decomposed attention score to obtain the following equation: 4) Solve equation 3) and find the bias vector b u and b v The calculation formula is: Bias vector b u and b v The constant coefficients of the mathematical formulas satisfying the original attention score and the decomposed attention score are consistent; 5) Use the weight matrix W U and W V and the corresponding bias vector b u and b v Construct a new linear layer and use the new linear layer to replace the original linear layer Query and Key to obtain the final decomposed Transformer model.