Generalized approximation-based adaptive LoRA low-rank method

Through the adaptive LoRA low-rank method of generalized approximation, the low-rank structure of the parameter change matrix of each module and level of the large language model is dynamically adjusted, which solves the problems of large calculation volume and information loss of existing LoRA methods, and achieves flexible adaptation and saving of computing resources.

CN119940415APending Publication Date: 2025-05-06CHINESE PEOPLES LIBERATION ARMY UNIT 91776
View PDF 0 Cites 2 Cited by

Patent Information

Application Number
CN202510090158.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-21
Publication Date
2025-05-06

AI Technical Summary

Technical Problem

In the specialized field knowledge fine-tuning training of large language models, the existing LoRA method fixes low rank values ​​and ignores the differences in importance of different modules and hierarchical parameters for fine-tuning, resulting in large amounts of calculation and information loss.

Method used

The adaptive LoRA low-rank method of generalized approximation is used to dynamically adjust the low-rank structure of the parameter change matrix at each module and level through singular value decomposition and iterative calculation to avoid manually setting the low-rank value.

Benefits of technology

It realizes flexible adaptation to different downstream tasks, reduces the consumption of computing resources, retains information about the changes in the original parameters, and reduces the computational complexity of the training loss function.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119940415A_ABST
    Figure CN119940415A_ABST
Patent Text Reader

Abstract

The invention provides a generalized approximation-based adaptive LoRA low-rank method, which is characterized by comprising the following steps of: directly carrying out generalized low-rank approximation solution on a plurality of parameter variation matrixes of a transformer attention layer, and carrying out alternative iterative calculation solution to obtain left and right projection transformation matrixes of the parameter variation matrixes; and performing rank reduction on the matrix by a bilateral dimension reduction iteration method according to a convergence condition of an optimization target, and finally obtaining a low-rank structure of each parameter variation matrix. On the basis of LoRA efficient fine tuning idea logic used in a professional training process in the field of large language models, a generalized low-rank approximation method of a matrix is adopted to solve a low-rank structure of parameter variation, and compared with traditional LoRA efficient fine tuning, low-rank structures of different matrixes can be automatically calculated; compared with an AdaLoRA method based on SVD decomposition, the method has the advantage that low-rank structure calculation of different matrixes can be realized without adding a complex penalty term into a loss function trained by a large language model. The method has the advantages of being better in flexibility, small in calculation amount and high in robustness.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of professional knowledge training in the field of large language models, and in particular to an adaptive LoRA low-rank method of generalized approximation. Background Art

[0002] At present, in the application of large language models, various vertical fields are conducting various fine-tuning training studies in order to better train the specialized capabilities of large language models. Since the number of model parameters is tens of billions, it is not realistic to retrain all parameters from scratch for different downstream tasks. The efficient fine-tuning method of some parameters represented by LoRA (Low-Rank Adaptation) solves these problems well. The core idea of ​​LoRA is that after the language model is fine-tuned for a specific task, the weight matrix usually has a very low intrinsic rank, so even if the parameter update is projected into a smaller subspace, it will not affect the effectiveness of learning. For specific downstream tasks, the LoRA method fixes the pre-trained model parameters unchanged, adds the product of the low-rank matrix as a trainable parameter in the weight matrix bypass of the transformer architecture to simulate the change in the parameters, and finally merges the original weights and bypass weights when the large language model is used for inference, thereby training the specialized capabilities of the large language model with a smaller computing cost.

[0003] However, the LoRA method used in domain training of large language models assigns a unique rank to all low-rank matrices, thus ignoring the differences in the importance of parameters at different levels of different modules for fine-tuning specific downstream tasks, and the selection of low-rank values ​​relies entirely on manual experimental settings, which is time-consuming and labor-intensive. If the setting is too low, a large amount of information may be lost. Therefore, with the development of the diversity of downstream tasks, although this method saves resources, it can no longer meet the effectiveness of training. AdaLoRA approximates the parameter change matrix through singular value decomposition, and dynamically adjusts the parameter weights for different tasks by adding regularization terms to the loss function, which increases the amount of training computation. At the same time, the calculation of singular values ​​is a difficult problem. When the matrix dimension is very high and the scale is large, the complexity grows exponentially. Summary of the invention

[0004] In order to solve the problem of the low-rank structure of the parameter change matrix in the existing LoRA efficient fine-tuning solution during the fine-tuning training of specialized domain knowledge in large language models, the present invention provides a generalized approximation adaptive LoRA low-rank method with good robustness, small computational complexity and high flexibility, which can adaptively and dynamically calculate the low-rank structure of the parameter change matrix of each module and each level for different downstream tasks. And compared with AdaLoRA singular value decomposition, it can handle different low-rank structures more flexibly, without affecting the calculation of the large language model training loss function itself, and save more computing resources. The technical solution adopted by the present invention includes the following steps:

[0005] First, freeze the pre-trained parameter matrix of the large language model and only train the newly added network layer parameters. For any input matrix of the downstream task calculate The forward propagation loss is then back-propagated based on the loss to obtain the weight gradient update matrix ΔW for each query layer Q, key layer K, and value layer V in the attention layer i ,i=q,k,v. Next, we need to i ,i=q,k,v to solve the reduced rank structure M i , specifically the minimum value optimization problem of the following formula:

[0006]

[0007] This formula can be decomposed into:

[0008]

[0009] Since the first term on the right side above is a constant, minimizing is equivalent to minimizing:

[0010]

[0011] So only when M i =L i T ΔW i R i When the minimization is achieved, that is, M i L i and R i The only decision is to iterate the simulation to calculate L i and R i To approximate ΔW i low-rank structure.

[0012] L i t=0 Initialize to the d-dimensional identity matrix, time step t is initialized to 1, and calculate the right projection matrix at this time step Calculate the matrix composed of unit orthogonal eigenvectors corresponding to all eigenvalues ​​of the right projection matrix in accordance with Alternate iteratively obtain the left projection matrix Similarly, the matrix composed of unit orthogonal eigenvectors corresponding to all eigenvalues ​​of the left projection matrix is ​​calculated The possible low-rank matrix M is obtained by calculating the left and right projection matrices i =(L i t ) T ΔW i R i t as well as Observe whether the convergence condition (ε t-1 -ε t ) / ε t-1 <η,η=10 -6 (ε t=0 =1). If it does not meet the requirement, set t=t+1 to continue iterating and calculating the left and right projection matrices. If it meets the requirement, calculate and output ΔW i The low-rank matrix M i .

[0013] The advantages of the invention are:

[0014] (1) It can directly compress the parameter change matrices of several levels at a time to obtain low rank, without adding penalty terms to the loss function of large language model training to adapt to the low rank requirements of different downstream tasks;

[0015] (2) This method is more flexible in dealing with low-rank structures and can better adapt to the "quasi-low-rank" structure and sparse matrix of data;

[0016] (3) This method has a small overall computational complexity and can better preserve the information of the original parameter changes. BRIEF DESCRIPTION OF THE DRAWINGS

[0017] In order to more clearly illustrate the technical solution of the embodiment of the present invention, the following is a brief introduction to the drawings required for use in the embodiment of the present invention. Obviously, the drawings described below are only some examples of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.

[0018] Figure 1 This is a calculation flow chart of the generalized approximation adaptive LoRA low-rank method of the present invention. DETAILED DESCRIPTION

[0019] In order to make the purpose, technical solution and advantages of the present application clearer, the technical solution of the present application will be clearly and completely described below in combination with the specific embodiments of the present application and the corresponding drawings. Obviously, the described embodiments are only part of the embodiments of the present application, not all of the embodiments. Based on the embodiments in the present application, all other embodiments obtained by ordinary technicians in this field without making creative work are within the scope of protection of the present application.

[0020] The specific steps of this embodiment are as follows: Figure 1 As shown, the technical solution of the present invention will be further described below by taking an example.

[0021] Taking a query layer Q of the transformer self-attention layer as an example, the initial weight matrix of the query layer Q is According to the new input matrix The forward propagation calculates the loss value based on the difference between the predicted result and the actual value, and then in the backward propagation process, the gradient is calculated based on the loss value. Finally, the parameter change matrix of the query layer Q is obtained according to the gradient descent method. Initialize L q t=0 is the d-dimensional identity matrix, initialized to 1 at time step t.

[0022] When executing the alternating iterative method, the calculation steps at each time step are as follows:

[0023] (1) Based on L q t-1 and Calculate the right projection matrix (at the initial time step, L q t-1 =L q t=0 );

[0024] (2) Calculate the matrix composed of all unit orthogonal eigenvectors of the right projection matrix

[0025] (3) Basis and Calculate the left projection matrix;

[0026] (4) Calculate the matrix composed of all unit orthogonal eigenvectors of the left projection matrix

[0027] (5) Calculate whether the low-rank matrix of level Q obtained at this time step meets the convergence condition of minimum optimization;

[0028] (6) If the condition is satisfied, the time step loop iteration is exited. If the condition is not satisfied, the time step is set to t = t + 1 and the solution is continued at the new time step. The left and right projection matrices.

[0029] Although the above describes the specific implementation mode of the present invention in conjunction with the accompanying drawings, it is not intended to limit the scope of protection of the invention. Those skilled in the art should understand that various modifications or variations that can be made by those skilled in the art on the basis of the technical solution of the present invention without creative work are still within the scope of protection of the present invention.

Claims

1. A generalized approximation adaptive LoRA low-rank method, characterized in that The following steps are involved: S1. Determine the weight matrix ΔW to be updated when training downstream tasks for each query layer Q, key layer K, and value layer V of the attention layer under the transformer architecture i ,i=q,k,v, taking the query layer Q as an example, the specific example is: The initial weight matrix of the query layer Q is According to the new input matrix The forward propagation calculates the loss value according to the difference between the predicted result and the actual value, and then in the back propagation process, the gradient is calculated according to the loss value. Finally, the parameter change matrix of the query layer Q is obtained according to the gradient descent method. S2. For all weight matrices ΔW to be updated i The alternating iterative low-rank approximation is adopted, specifically: ΔW i The matrix L that needs to be iterated alternately in the process of finding low rank i t=0 Initialize to 0 matrix, time step t is initialized to 1; S3. Calculate the time step by L i t Generate calculated ΔW i Right projection transformation matrix S4. Calculate the right projection transformation matrix M under the time step i,R The matrix consisting of the unit orthogonal eigenvectors corresponding to all the largest eigenvalues S5. The calculation time step is Generate calculated ΔW i Left projection transformation matrix S6. Calculate the left projection transformation matrix M under the time step i,L The matrix L consisting of the unit orthogonal eigenvectors corresponding to all the largest eigenvalues t i ; S7. Calculate the low-rank structure approximation matrix M i =(L i t ) T ΔW i R i t ; S8. Calculate whether the convergence condition (ε t-1 -ε t ) / ε t-1 <η,η=10 -6 , In particular, ε t=0 =1, if the method converges, proceed to the following steps; otherwise, increase the time step value (t=t+1), continue to calculate the right projection transformation matrix, and return to step S3; S9. Output ΔW i The low-rank structure approximation matrix M i .

Citation Information

Cited By

  • Attention approximation method and system based on low-rank adaptation and partial binarization

    CN120745699A

  • Multifidelity bayesian low-rank approximation modeling method with low-fidelity trend embedded

    CN122692454A