Fine adjustment method and device for large language model
By decomposing the weight matrix of the large language model into amplitude vectors and direction matrix, and performing low-rank decomposition of the direction matrix, the problems of high fine-tuning costs and high computing resources requirements of the large language model are solved, and more efficient fine-tuning effects are achieved.
Patent Information
- Application Number
- CN202510199077.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-21
- Publication Date
- 2025-05-13
AI Technical Summary
When large language models need further fine-tuning in specific downstream tasks, conventional fine-tuning methods are costly and require a high demand for computing resources.
By decomposing the weight matrix of the pre-trained large language model into amplitude vectors and direction matrix, and performing low-rank decomposition on the direction matrix with large parameters, fitting the incremental matrix of the direction matrix, the model can independently adjust the size and direction of the weight during the fine-tuning process, reducing the demand for computing resources.
It significantly improves the learning ability and fine-tuning effect of the model, while reducing the cost of fine-tuning and reducing the demand for computing resources.
Smart Images

Figure CN119990183A_ABST
Abstract
Description
Technical Field
[0001] One or more embodiments of the present specification relate to the field of machine learning, and more particularly, to a method and apparatus for fine-tuning a large language model using text data. Background Art
[0002] Large language models (LLMs) pre-trained using a large number of general domain datasets have demonstrated strong reasoning and generalization capabilities in various tasks, but they often require further fine-tuning when directly applied to specific downstream tasks. Large language models have a large number of parameters, and the cost of fine-tuning large language models using conventional methods is high. Therefore, it is hoped that there will be an improved solution to reduce the cost of fine-tuning large language models. Summary of the invention
[0003] One or more embodiments of the present specification describe a fine-tuning scheme for a large language model, which decomposes the weight matrix of a pre-trained large language model into an amplitude vector and a first direction matrix, and performs low-rank decomposition on the first direction matrix with a large number of parameters to fit the incremental matrix of the first direction matrix, so that the learning ability and fine-tuning effect of the model can be significantly improved while reducing the demand for computing resources.
[0004] According to a first aspect, a fine-tuning method for a large language model is provided, comprising: obtaining a weight matrix of a pre-trained large language model, decomposing the weight matrix into an amplitude vector and a first direction matrix, and initializing a first low-rank matrix and a second low-rank matrix, wherein the product of the first low-rank matrix and the second low-rank matrix is used to fit an incremental matrix of the first direction matrix; keeping the first direction matrix unchanged, performing multiple rounds of fine-tuning, wherein each round of fine-tuning comprises: inputting a training text into the large language model, calculating a loss function based on the difference between a predicted text output by the large language model and a marked text of the training text; updating the amplitude vector based on the loss function to obtain an updated amplitude vector; determining a first gradient matrix of the loss function relative to a current direction matrix; obtaining an updated first low-rank matrix and a second low-rank matrix according to the first gradient matrix; determining an incremental matrix of this round based on the product of the updated first low-rank matrix and the second low-rank matrix, and superimposing the incremental matrix of this round on the first direction matrix as the updated direction matrix of this round.
[0005] According to one embodiment, after the multiple rounds of fine-tuning, the method further includes: merging the updated direction matrix and the updated amplitude vector into an updated weight matrix, and obtaining a fine-tuned large language model based on the updated weight matrix.
[0006] In one embodiment, initializing a first low-rank matrix and a second low-rank matrix includes: initializing a left singular matrix and a right singular matrix; obtaining an updated first low-rank matrix and a second low-rank matrix according to the first gradient matrix, including: obtaining an updated left singular matrix and a right singular matrix according to the first gradient matrix; selecting r left singular vectors corresponding to the first r largest singular values from the updated left singular matrix, thereby forming an updated first low-rank matrix; and selecting r right singular vectors corresponding to the first r largest singular values from the updated right singular matrix, thereby forming an updated second low-rank matrix.
[0007] In a further embodiment, initializing the first low-rank matrix and the second low-rank matrix also includes initializing a diagonal matrix whose diagonal element values represent singular values; obtaining an updated first low-rank matrix and a second low-rank matrix according to the first gradient matrix, and also includes: obtaining an updated diagonal matrix according to the first gradient; determining r diagonal elements with larger values from the updated diagonal matrix as the r maximum singular values; wherein the r left singular vectors are r column vectors corresponding to the r diagonal elements in the updated left singular matrix; and the r right singular vectors are r row vectors corresponding to the r diagonal elements in the updated right singular matrix.
[0008] In a further embodiment, the updated first low-rank matrix is a matrix that retains the r left singular vectors to mask the remaining vectors, and the updated second low-rank matrix is a matrix that retains the r right singular vectors to mask the remaining vectors; the incremental matrix of this round is determined based on the product of the updated first low-rank matrix and the second low-rank matrix, including: taking the product of multiplying the updated first low-rank matrix, the updated diagonal matrix, and the updated second low-rank matrix in sequence as the incremental matrix of this round.
[0009] In a further embodiment, the updated first low-rank matrix and the second low-rank matrix are obtained according to the first gradient matrix, and the following further includes: before determining the r diagonal elements with larger values, calculating the importance scores of each triplet according to the first gradient matrix, wherein a single triplet includes the diagonal elements of the target position in the diagonal matrix, the column vector corresponding to the target position in the left singular matrix, and the row vector corresponding to the target position in the right singular matrix; setting the values of the diagonal elements corresponding to the first number of triplets with lower importance scores in the updated diagonal matrix to 0.
[0010] In a further embodiment, the first number first increases with the increase of the number of fine-tuning rounds, and then decreases with the increase of the number of fine-tuning rounds.
[0011] In a further embodiment, a loss function is calculated based on the difference between the predicted text output by the large language model and the marked text of the training text, including: calculating a first loss function based on the difference between the predicted text output by the large language model and the marked text of the training text; calculating a second loss function based on the degree to which the left singular matrix deviates from the orthogonal matrix and the degree to which the right singular matrix deviates from the orthogonal matrix, the sum of the first loss function and the second loss function is used as the loss function of this round, and each round of fine-tuning also includes: determining the gradient of the loss function on each parameter in the left singular matrix, the diagonal matrix, and the right singular matrix; determining the importance of each parameter based on each parameter and its gradient; and calculating the importance score of each triple according to the importance of each parameter.
[0012] According to one embodiment, the amplitude vector is updated based on the loss function to obtain an updated amplitude vector, including: calculating the product of the gradient matrix of the loss function relative to the weight matrix and a first scaling factor, and determining the gradient vector of the loss function relative to the amplitude vector according to the product result; wherein the first scaling factor is equal to the direction matrix divided by the vector norm of each column vector in the current direction matrix; and updating the amplitude vector based on the gradient vector.
[0013] According to one embodiment, determining a first gradient matrix of the loss function relative to a current direction matrix includes: determining the first gradient matrix based on the product of a gradient matrix of the loss function relative to a weight matrix and a second scaling factor, wherein the second scaling factor is equal to a magnitude vector divided by a target constant used to approximate the norm of each column vector in the direction matrix.
[0014] According to a second aspect, a fine-tuning device for a large language model is provided, comprising: an acquisition unit, configured to acquire a weight matrix of a pre-trained large language model, decompose the weight matrix into an amplitude vector and a direction matrix, and initialize a first low-rank matrix and a second low-rank matrix, wherein the product of the first low-rank matrix and the second low-rank matrix is used to fit an incremental matrix of the first direction matrix; a fine-tuning unit, configured to keep the first direction matrix unchanged, and perform multiple rounds of fine-tuning, wherein each round of fine-tuning comprises: inputting a training text into the large language model, and calculating a loss function based on the difference between a predicted text output by the large language model and a marked text of the training text; updating the amplitude vector based on the loss function to obtain an updated amplitude vector; determining a first gradient matrix of the loss function relative to an incremental matrix; obtaining an updated first low-rank matrix and a second low-rank matrix according to the first gradient matrix; determining an incremental matrix of this round based on the product of the updated first low-rank matrix and the second low-rank matrix, and superimposing the incremental matrix of this round on the first direction matrix as the updated direction matrix of this round.
[0015] According to a third aspect, a computer program product is provided, comprising a computer program / instruction, which, when executed by a processor, implements the steps of the method described in the first aspect.
[0016] According to a fourth aspect, a computer-readable storage medium is provided, on which a computer program is stored. When the computer program is executed in a computer, the computer is caused to execute the method described in the first aspect.
[0017] According to a fifth aspect, a computing device is provided, comprising a memory and a processor, wherein an executable code is stored in the memory, and when the processor executes the executable code, the method of the first aspect is implemented.
[0018] In the embodiments of the present specification, a novel weight decomposition mechanism is introduced to divide the weight matrix of the pre-trained large language model into two parts: a vector representing the magnitude and a matrix representing the direction. This decomposition allows the model to independently adjust the size and direction of the weights during the fine-tuning process, provides greater flexibility for model fine-tuning, significantly improves the learning ability and fine-tuning effect of the model, and can more flexibly adapt to different task requirements. At the same time, a low-rank decomposition is also implemented for the direction matrix (i.e., the first direction matrix) with a large number of parameters to fit the incremental matrix of the direction matrix, and the direction matrix is kept unchanged and the low-rank matrix of the incremental matrix used to fit the direction matrix is updated instead, so that a large number of parameter updates required for the comprehensive fine-tuning of the direction matrix can be avoided, and the demand for computing resources is reduced. BRIEF DESCRIPTION OF THE DRAWINGS
[0019] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the accompanying drawings required for use in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other accompanying drawings can be obtained based on these accompanying drawings without paying creative work.
[0020] Figure 1 A schematic diagram of the technical concept of fine-tuning a large language model based on a weight decomposition mechanism is shown.
[0021] Figure 2 A flow chart of a method for fine-tuning a large language model according to one embodiment is shown.
[0022] Figure 3 A schematic structural diagram of a device for fine-tuning a large language model according to an embodiment is shown. DETAILED DESCRIPTION
[0023] The solution provided in this specification is described below in conjunction with the accompanying drawings.
[0024] In the description of this specification, words such as "exemplary", "for example" or "for example" are used to indicate examples, illustrations or explanations. Any embodiment or design described as "exemplary", "for example" or "for example" in this specification should not be interpreted as being more preferred or more advantageous than other embodiments or designs. Specifically, the use of words such as "exemplary", "for example" or "for example" is intended to present related concepts in a specific way.
[0025] In the description of this specification, the term "and / or" is only a description of the association relationship of associated objects, indicating that there can be three relationships. For example, A and / or B can mean: A exists alone, B exists alone, and A and B exist at the same time. In addition, unless otherwise specified, the term "plurality" means two or more.
[0026] In addition, the terms "first" and "second" are used for descriptive purposes only and should not be understood as indicating or implying relative importance or implicitly indicating the indicated technical features. Therefore, the features defined as "first" and "second" may explicitly or implicitly include one or more of the features. The terms "include", "comprises", "has" and their variations all mean "including but not limited to", unless otherwise specifically emphasized.
[0027] The Big Language Model is a deep neural network model with a very large number of parameters. It is trained based on massive amounts of text data. It can understand and generate coherent and grammatically correct natural language text, and contains rich knowledge, bringing revolutionary progress in tasks such as text generation, machine translation, and knowledge question and answer.
[0028] Although large language models demonstrate powerful reasoning and generalization capabilities in various tasks, direct application to specific downstream tasks often requires further fine-tuning of the large language models.
[0029] Full fine-tuning (FT) is a common fine-tuning method that involves retraining all parameters of a pre-trained model to adapt to a new task. This approach can better adapt the model to specific task requirements, but it also comes with significant computational costs. Due to the need to update a large number of parameters, FT not only takes a long time to train, but may also cause high latency during inference. These factors limit the feasibility of FT in resource-constrained or real-time application scenarios.
[0030] In view of this, the embodiment of this specification proposes a fine-tuning method for a large language model, which introduces a novel weight decomposition mechanism, dividing the weight matrix of the pre-trained large language model into two parts: a vector representing the magnitude and a matrix representing the direction. This decomposition allows the model to independently adjust the size and direction of the weights during the fine-tuning process, providing greater flexibility for model fine-tuning, significantly improving the learning ability and fine-tuning effect of the model, and can more flexibly adapt to different task requirements. At the same time, a low-rank decomposition is also implemented for the direction matrix with a large number of parameters to fit the incremental matrix of the direction matrix, so that the direction matrix can be kept unchanged and the low-rank matrix of the incremental matrix used to fit the direction matrix is updated instead, thereby avoiding the large number of parameter updates required for the comprehensive fine-tuning of the direction matrix and reducing the demand for computing resources.
[0031] Figure 1 A schematic diagram of the technical concept of fine-tuning a large language model based on a weight decomposition mechanism is shown.
[0032] like Figure 1 As shown, first, you can select a pre-trained large language model and obtain its weight matrix W0.
[0033] Then, the weight matrix W0 can be decomposed into a magnitude vector m (also called a size vector or a magnitude vector) and a direction matrix V using a weight decomposition module. The decomposition form of the weight matrix W0 can be recorded as W0=m·V.
[0034] Among them, W0∈R d×k , m=‖W0‖ c ∈R 1×k , V=∈R d×k And ‖V‖ c =‖W0‖ c , ‖‖ c is the vector norm calculated by the columns of the matrix. m represents the scale or length of each column vector in the weight matrix, and V represents the direction of each weight in the weight matrix. The vector elements (i.e., scalars) in m correspond one-to-one to the column vectors in V and are used to define the magnitude of the column vectors in V. m·V can be understood as multiplying each vector element in m with each column vector element in the corresponding column vector in V.
[0035] Then, the fine-tuning module can be used to update m and V. During the fine-tuning process, the direction matrix V is frozen (that is, the direction matrix V always remains fixed), and the amplitude vector m is regarded as a trainable parameter and m is fine-tuned. For the direction matrix V, the incremental matrix of the direction matrix V can be fitted by low-rank decomposition, so that the direction matrix V can be kept unchanged and the incremental matrix (that is, the low-rank matrix) can be updated.
[0036] By independently adjusting the two properties of amplitude and direction, the effect of full fine-tuning (FT) can be simulated more accurately, thereby improving the model's learning ability on specific tasks; and when adjusting the direction, instead of directly updating the direction matrix V with a large number of parameters, the low-rank matrix of the incremental matrix used to fit the direction matrix V obtained by implementing low-rank decomposition (such as the first low-rank matrix and the second low-rank matrix mentioned below) is updated. This can greatly reduce the amount of fine-tuning parameters, save computing resources, and improve the efficiency of fine-tuning.
[0037] Finally, the fine-tuned amplitude vector m can be ′ and the direction matrix V ′ Merge to generate the final weight matrix W ′ ,W ′ =m ′ ·V ′ The updated weight matrix W ′ It can be applied to downstream tasks for reasoning or further training.
[0038] The specific implementation process of the above technical concept is described below.
[0039] Figure 2 A flow chart of a fine-tuning method for a large language model according to an embodiment is shown. It can be understood that the method can be executed by any device, equipment, platform, or device cluster with computing and processing capabilities.
[0040] See also Figure 2 In step S21, a weight matrix of a pre-trained large language model is obtained, the weight matrix is decomposed into an amplitude vector and a first direction matrix, and a first low-rank matrix and a second low-rank matrix are initialized. The product of the first low-rank matrix and the second low-rank matrix is used to fit an incremental matrix of the first direction matrix.
[0041] A pre-trained large language model refers to a large language model that is pre-trained on a large-scale data set. The weight matrix obtained refers to the weight matrix that needs to be fine-tuned. A large language model may include multiple weight matrices. In some exemplary embodiments, all weight matrices of the pre-trained large language model may be obtained, and all weight matrices of the large language model may be fine-tuned. In other exemplary embodiments, some weight matrices may also be selectively selected, such as selecting a weight matrix that is more closely related to the downstream task to be processed, and only fine-tuning the selected weight matrix.
[0042] The first direction matrix is the above-mentioned direction matrix V. For how to decompose the weight matrix into the amplitude vector and the first direction matrix, refer to the above-mentioned related description.
[0043] The product of the first low-rank matrix and the second low-rank matrix is used to fit the incremental matrix of the first direction matrix, which means that the product of the first low-rank matrix and the second low-rank matrix can be regarded as the incremental matrix of the first direction matrix.
[0044] The relationship between the incremental matrix and the first low-rank matrix and the second low-rank matrix can be expressed as, ΔV = B × A. Where ΔV represents the incremental matrix, B represents the first low-rank matrix, and A represents the second low-rank matrix. d×k For example, B∈R d×r , A∈R r×k , the rank r is usually much smaller than the value of d and k. The rank r can be a preset fixed value or a value that changes with the fine-tuning rounds.
[0045] The purpose of initializing the first low-rank matrix and the second low-rank matrix is to ensure that the product of the first low-rank matrix and the second low-rank matrix is zero in the initial case (i.e., before fine-tuning). In some exemplary embodiments, the initial value of the first low-rank matrix is zero, that is, the first low-rank matrix is an all-zero matrix, and the second low-rank matrix can be initialized using a normal distribution.
[0046] After step S21, the first direction matrix may be kept unchanged and multiple rounds of fine-tuning may be performed.
[0047] That is, during the fine-tuning process, the first direction matrix can be regarded as a non-adjustable part, and the amplitude vector and the first and second low-rank matrices of the incremental matrix used to fit the first direction matrix can be regarded as adjustable parts.
[0048] Figure 2 Steps S22 to S25 in FIG. 2 are used to characterize each round of fine-tuning process.
[0049] In step S22, the training text is input into the large language model, and the loss function is calculated based on the difference between the predicted text output by the large language model and the marked text of the training text.
[0050] The training text is the text data constructed according to the downstream task. The labeled text can be regarded as the standard answer expected to be output by the large language model. The greater the difference between the predicted text output by the large language model and the labeled text, the less accurate the model prediction is. The value of the loss function is positively correlated with the difference between the predicted text and the labeled text. That is, the greater the difference between the predicted text and the labeled text, the greater the value of the loss function.
[0051] After the loss function is obtained, the amplitude vector and the direction matrix can be updated based on the loss function. Updating the direction matrix refers to updating the increment matrix, that is, updating the first low-rank matrix and the second low-rank matrix.
[0052] Figure 2It is shown that the update of the amplitude vector (i.e., step S23) is performed first, and the update of the first low-rank matrix and the second low-rank matrix (i.e., step S24 and step S25) is performed later. It should be known that the update of the amplitude vector and the update of the first low-rank matrix and the second low-rank matrix can also be performed in an interchangeable order or in parallel without particular order.
[0053] In step S23, the amplitude vector is updated based on the loss function to obtain an updated amplitude vector.
[0054] The back-propagation algorithm can be used to learn and optimize (i.e., update) the amplitude vector based on the loss function of the downstream task using the gradient descent algorithm to optimize the performance of the model on a specific task.
[0055] In some exemplary embodiments, the product of the gradient matrix of the loss function relative to the weight matrix and the first scaling factor may be calculated, and the gradient vector of the loss function relative to the magnitude vector may be determined according to the product result; wherein the first scaling factor is equal to the direction matrix (i.e., the current direction matrix) divided by the vector norm of each column vector in the current direction matrix. Then, the magnitude vector is updated based on the gradient vector. Updating the magnitude vector based on the gradient vector means updating the magnitude vector along the direction of gradient descent.
[0056] The gradient vector of the magnitude vector can be exemplarily expressed as formula (1),
[0057]
[0058] Among them, L is the loss function. V represents the gradient of the loss function relative to the magnitude vector m (i.e., the gradient vector). ′ Represents the current direction matrix (the direction matrix after the last update). ′ =V+ΔV, V is the first direction matrix, ΔV is the increment matrix (such as the increment matrix of the previous round). represents the gradient matrix of the loss function with respect to the weight matrix, Indicates the first scaling factor.
[0059] In step S24, a first gradient matrix of the loss function relative to the current direction matrix is determined.
[0060] As mentioned above, the first direction matrix remains unchanged during the fine-tuning process, and the incremental matrix is updated during the fine-tuning process. Therefore, it is necessary to determine the first gradient matrix of the loss function relative to the incremental matrix so as to update the incremental matrix according to the first gradient matrix. To update the incremental matrix, the first low-rank matrix and the second low-rank matrix are updated.
[0061] It should be known that since the first direction matrix remains unchanged during the fine-tuning process, the current direction matrix V′ = V + ΔV. Therefore, the gradient matrix of the loss function relative to the current direction matrix can be regarded as the first gradient matrix of the loss function relative to the increment matrix. The first gradient matrix can be exemplarily expressed as formula (2):
[0062]
[0063] in, Represents the loss function relative to the current direction matrix V ′ The gradient of V ′ =V+ΔV and V remains unchanged, so, Equivalent to I is the unit matrix. The above formula can be understood as the gradient matrix of the loss function relative to the weight matrix Scaling factor Adjust and project out from the current weight matrix.
[0064] In some exemplary embodiments, ‖V ′ ‖ c (i.e. ‖V+ΔV‖ c ) is considered a constant in the back propagation. Thus, It can be redefined as: the first gradient matrix is determined according to the product of the gradient matrix of the loss function relative to the weight matrix and the second scaling factor; wherein the second scaling factor is equal to the magnitude vector divided by the target constant used to approximate the norm of each column vector in the direction matrix. That is, the first gradient matrix can be calculated according to formula (3):
[0065]
[0066] Where C is the target constant used to approximate the norm of each column vector in the direction matrix. That is, C = ‖V ′ ‖ c By optimizing the gradient calculation, the optimization stability and efficiency during fine-tuning are improved; during backpropagation, some terms (such as ‖V ′ ‖ c ) is considered as a constant, which can reduce memory consumption and training costs, making it possible to effectively fine-tune the model even in resource-constrained environments.
[0067] In step S25, an updated first low-rank matrix and a second low-rank matrix are obtained according to the first gradient matrix; the incremental matrix of this round is determined based on the product of the updated first low-rank matrix and the second low-rank matrix, and the incremental matrix of this round is superimposed on the first direction matrix as the updated direction matrix of this round.
[0068] The updated first low-rank matrix and the second low-rank matrix are obtained according to the first gradient matrix, which means that the first low-rank matrix and the second low-rank matrix can be updated along the direction of gradient descent using a gradient descent algorithm based on the first gradient matrix.
[0069] Determining the incremental matrix of this round based on the product of the updated first low-rank matrix and the second low-rank matrix means that the product of the first low-rank matrix and the second low-rank matrix after the update of this round can be used as the incremental matrix of this round.
[0070] The incremental matrix of this round can be exemplarily represented as ΔV(t)=B(t)×A(t). Where t is the fine-tuning round, ΔV(t) represents the incremental matrix of the tth round, B(t) represents the first low-rank matrix after the tth round update, and A(t) represents the second low-rank matrix after the tth round update. The direction matrix after this round of update can be exemplarily represented as V ′ (t) = V + ΔV (t). Where V ′ (t) represents the direction matrix after the tth round of update.
[0071] In some exemplary embodiments, after multiple rounds of fine-tuning (i.e., after the fine-tuning process is completed), the updated direction matrix and the updated amplitude vector may be combined into an updated weight matrix, and a fine-tuned large language model may be obtained based on the updated weight matrix. The updated weight matrix may be exemplarily represented as: ′ =m″·V″. Where, W ′ represents the updated weight matrix, m″ represents the updated amplitude vector, and V″ represents the updated direction matrix.
[0072] The fine-tuned large language model can be applied to downstream tasks, for inference or further training.
[0073] Next Figure 2 Some of the details involved are further illustrated.
[0074] In some exemplary embodiments, the “initializing the first low-rank matrix and the second low-rank matrix” in step S21 includes: initializing a left singular matrix and a right singular matrix, wherein the left singular matrix can be regarded as the first low-rank matrix, and the right singular matrix can be regarded as the second low-rank matrix.
[0075] The step of "obtaining an updated first low-rank matrix and a second low-rank matrix according to the first gradient matrix" in step S25 includes: obtaining an updated left singular matrix and a right singular matrix according to the first gradient matrix; it is generally considered that the left singular vectors and the right singular vectors corresponding to the larger singular values are more important, so after obtaining the updated left singular matrix and the right singular matrix, the r left singular vectors corresponding to the first r largest singular values can be selected from the updated left singular matrix to form an updated first low-rank matrix; and the r right singular vectors corresponding to the first r largest singular values can be selected from the updated right singular matrix to form an updated second low-rank matrix. In this way, the ranks of the first low-rank matrix and the second low-rank matrix can be dynamically adjusted during the fine-tuning process.
[0076] In some further exemplary embodiments, the "initializing the first low-rank matrix and the second low-rank matrix" in step S21 also includes: initializing a diagonal matrix, the values of the diagonal elements of which represent singular values. The diagonal matrix can be initialized to 0, and the left singular matrix and the right singular matrix can be initialized to Gaussian random matrices. The product of the left singular matrix, the diagonal matrix, and the right singular matrix multiplied in sequence is used to fit the incremental matrix of the first direction matrix. It can be understood that the incremental matrix of the first direction matrix is subjected to singular value decomposition to obtain the left singular matrix, the diagonal matrix, and the right singular matrix.
[0077] The "obtaining an updated first low-rank matrix and a second low-rank matrix according to the first gradient matrix" in step S25 also includes: obtaining an updated diagonal matrix according to the first gradient; determining r diagonal elements with larger values from the updated diagonal matrix as r maximum singular values; wherein the r left singular vectors are the r column vectors corresponding to the r diagonal elements in the updated left singular matrix; and the r right singular vectors are the r row vectors corresponding to the r diagonal elements in the updated right singular matrix.
[0078] The updated first low-rank matrix is a matrix that retains r left singular vectors to mask the remaining vectors, and the updated second low-rank matrix is a matrix that retains r right singular vectors to mask the remaining vectors. The "determining the incremental matrix of this round based on the product of the updated first low-rank matrix and the second low-rank matrix" in step S25 includes: taking the product of the updated first low-rank matrix, the updated diagonal matrix, and the updated second low-rank matrix multiplied in sequence as the incremental matrix of this round.
[0079] In some further exemplary embodiments, "obtaining an updated first low-rank matrix and a second low-rank matrix according to the first gradient matrix" in step S25 also includes: before determining the r diagonal elements with larger values, calculating the importance scores of each triplet according to the first gradient matrix, wherein a single triplet includes the diagonal elements of the target position in the diagonal matrix, the column vector corresponding to the target position in the left singular matrix, and the row vector corresponding to the target position in the right singular matrix; setting the values of the diagonal elements corresponding to the first number of triplets with the lowest importance scores in the updated diagonal matrix to 0.
[0080] The values of the diagonal elements corresponding to the first number of triplets with the lowest importance scores are set to 0, so that when the forward propagation is performed in the next round of fine-tuning, the corresponding column vector and row vector in this triplet are equivalent to being masked and have no contribution to the loss function of the next round, thus achieving the effect of changing the rank.
[0081] That is to say, by setting the values of the diagonal elements corresponding to the first number of triplets with the lowest importance scores to 0, not only the r diagonal elements with larger values are determined to be the diagonal elements with higher importance scores, but also the remaining vectors in the left singular matrix and the right-left singular matrix can be masked.
[0082] The first number can be a pre-set fixed value or a value that changes with the number of fine-tuning rounds. Exemplarily, the first number first increases with the number of fine-tuning rounds, and then decreases with the number of fine-tuning rounds. For example, in the early stage of fine-tuning, the first number can be gradually increased (i.e., the rank is gradually increased) to achieve more exploration; in the middle or late stage of fine-tuning, the first number can be gradually reduced until finally training is performed with a stable and smaller first number to leave training resources for the most important parameters.
[0083] In order to calculate the importance score of each triple, the "calculation of the loss function based on the difference between the predicted text output by the large language model and the marked text of the training text" in step S22 may include: calculating the first loss function based on the difference between the predicted text output by the large language model and the marked text of the training text; calculating the second loss function based on the degree to which the left singular matrix deviates from the orthogonal matrix and the degree to which the right singular matrix deviates from the orthogonal matrix, and the sum of the first loss function and the second loss function is used as the loss function of this round. Each round of fine-tuning may also include: determining the gradient of the loss function on each parameter in the left singular matrix, the diagonal matrix, and the right singular matrix; determining the importance of each parameter based on each parameter and its gradient; and calculating the importance score of each triple according to the importance of each parameter.
[0084] According to an embodiment of another aspect, a fine-tuning apparatus for a large language model is provided. Figure 3A schematic diagram of the structure of a fine-tuning device for a large language model according to an embodiment is shown. The fine-tuning device can be deployed in any device, platform or device cluster with data storage, computing and processing capabilities.
[0085] like Figure 3 As shown, the fine-tuning device 300 includes:
[0086] An acquisition unit 31 is configured to acquire a weight matrix of a pre-trained large language model, decompose the weight matrix into an amplitude vector and a direction matrix, and initialize a first low-rank matrix and a second low-rank matrix, wherein the product of the first low-rank matrix and the second low-rank matrix is used to fit an incremental matrix of the first direction matrix;
[0087] The fine-tuning unit 32 is configured to keep the first direction matrix unchanged and perform multiple rounds of fine-tuning, each round of fine-tuning comprising:
[0088] Inputting a training text into the large language model, and calculating a loss function based on the difference between a predicted text output by the large language model and a marked text of the training text;
[0089] Updating the amplitude vector based on the loss function to obtain an updated amplitude vector;
[0090] Determine a first gradient matrix of the loss function relative to the current direction matrix;
[0091] An updated first low-rank matrix and a second low-rank matrix are obtained according to the first gradient matrix; an incremental matrix of this round is determined based on the product of the updated first low-rank matrix and the second low-rank matrix, and the incremental matrix of this round is superimposed on the first direction matrix as the updated direction matrix of this round.
[0092] In some embodiments, the fine-tuning device 300 further includes: a merging module configured to merge the updated direction matrix and the updated amplitude vector into an updated weight matrix after the multiple rounds of fine-tuning, and obtain a fine-tuned large language model based on the updated weight matrix.
[0093] In some embodiments, the acquisition unit 31 is specifically configured as follows: initialize the left singular matrix and the right singular matrix; obtain the updated left singular matrix and the right singular matrix according to the first gradient matrix; select r left singular vectors corresponding to the first r largest singular values from the updated left singular matrix, thereby forming an updated first low-rank matrix; and select r right singular vectors corresponding to the first r largest singular values from the updated right singular matrix, thereby forming an updated second low-rank matrix.
[0094] In a further embodiment, the acquisition unit 31 is specifically configured as follows: initialize a diagonal matrix whose diagonal element values represent singular values; obtain an updated diagonal matrix according to the first gradient; determine r diagonal elements with larger values from the updated diagonal matrix as the r maximum singular values; wherein the r left singular vectors are r column vectors corresponding to the r diagonal elements in the updated left singular matrix; and the r right singular vectors are r row vectors corresponding to the r diagonal elements in the updated right singular matrix.
[0095] In a further embodiment, the updated first low-rank matrix is a matrix that retains the r left singular vectors to mask the remaining vectors, and the updated second low-rank matrix is a matrix that retains the r right singular vectors to mask the remaining vectors; the fine-tuning unit 32 is specifically configured to: use the product of the updated first low-rank matrix, the updated diagonal matrix, and the updated second low-rank matrix in sequence as the incremental matrix of this round.
[0096] In a further embodiment, the fine-tuning unit 32 is specifically configured as follows: before determining the r diagonal elements with larger values, according to the first gradient matrix, calculate the importance score of each triple, wherein a single triple includes the diagonal element of the target position in the diagonal matrix, the column vector corresponding to the target position in the left singular matrix, and the row vector corresponding to the target position in the right singular matrix; and set the values of the diagonal elements corresponding to the first number of triplets with lower importance scores in the updated diagonal matrix to 0. Exemplarily, the first number may first increase with the increase of the number of fine-tuning rounds, and then decrease with the increase of the number of fine-tuning rounds.
[0097] In some embodiments, the fine-tuning unit 32 is specifically configured to: calculate the product of the gradient matrix of the loss function relative to the weight matrix and the first scaling factor, and determine the gradient vector of the loss function relative to the amplitude vector based on the product result; wherein the first scaling factor is equal to the direction matrix divided by the vector norm of each column vector in the current direction matrix; and update the amplitude vector based on the gradient vector.
[0098] In some embodiments, the fine-tuning unit 32 is specifically configured to determine the first gradient matrix based on the product of the gradient matrix of the loss function relative to the weight matrix and a second scaling factor, wherein the second scaling factor is equal to the amplitude vector divided by a target constant used to approximate the norm of each column vector in the direction matrix.
[0099] For specific implementation examples of each unit in the above device, please refer to the previous Figure 2 Description.
[0100] According to another embodiment, there is also provided a computer-readable storage medium having a computer program stored thereon, which, when executed in a computer, causes the computer to execute a combination of Figure 2 The method described.
[0101] According to another embodiment, a computer program product is also provided, including a computer program / instruction, which is executed by a processor to implement the aforementioned combination Figure 2 The method steps described.
[0102] According to another embodiment of the present invention, a computing device is provided, including a memory and a processor, wherein the memory stores executable code, and when the processor executes the executable code, the Figure 2 The method described.
[0103] Those skilled in the art should be aware that in one or more of the above examples, the functions described in the present invention can be implemented by hardware, software, firmware or any combination thereof. When implemented by software, these functions can be stored in a computer-readable medium or transmitted as one or more instructions or codes on a computer-readable medium.
[0104] The specific implementation methods described above further illustrate the objectives, technical solutions and beneficial effects of the present invention in detail. It should be understood that the above description is only a specific implementation method of the present invention and is not intended to limit the scope of protection of the present invention. Any modifications, equivalent substitutions, improvements, etc. made on the basis of the technical solution of the present invention should be included in the scope of protection of the present invention.
Claims
1. A fine-tuning method for a large language model, comprising: Obtain a weight matrix of a pre-trained large language model, decompose the weight matrix into an amplitude vector and a first direction matrix, and initialize a first low-rank matrix and a second low-rank matrix, wherein the product of the first low-rank matrix and the second low-rank matrix is used to fit an incremental matrix of the first direction matrix; Keeping the first direction matrix unchanged, multiple rounds of fine-tuning are performed, each round of fine-tuning includes: Inputting a training text into the large language model, and calculating a loss function based on the difference between a predicted text output by the large language model and a marked text of the training text; Updating the amplitude vector based on the loss function to obtain an updated amplitude vector; Determine a first gradient matrix of the loss function relative to the current direction matrix; An updated first low-rank matrix and a second low-rank matrix are obtained according to the first gradient matrix; an incremental matrix of this round is determined based on the product of the updated first low-rank matrix and the second low-rank matrix, and the incremental matrix of this round is superimposed on the first direction matrix as the updated direction matrix of this round.
2. The method according to claim 1, wherein: After the multiple rounds of fine-tuning, the following steps are also included: The updated direction matrix and the updated amplitude vector are combined into an updated weight matrix, and a fine-tuned large language model is obtained based on the updated weight matrix.
3. The method according to claim 1, wherein: Initializing the first low-rank matrix and the second low-rank matrix, including: initializing a left singular matrix and a right singular matrix; Obtaining an updated first low-rank matrix and a second low-rank matrix according to the first gradient matrix includes: Obtaining updated left singular matrix and right singular matrix according to the first gradient matrix; Select r left singular vectors corresponding to the first r largest singular values from the updated left singular matrix to form an updated first low-rank matrix; and select r right singular vectors corresponding to the first r largest singular values from the updated right singular matrix to form an updated second low-rank matrix.
4. The method according to claim 3, wherein: Initializing the first low-rank matrix and the second low-rank matrix also includes initializing a diagonal matrix whose diagonal element values represent singular values; The method further comprises: obtaining an updated first low-rank matrix and a second low-rank matrix according to the first gradient matrix; Obtaining an updated diagonal matrix according to the first gradient; Determine r diagonal elements with larger values from the updated diagonal matrix as the r maximum singular values; Among them, the r left singular vectors are the r column vectors corresponding to the r diagonal elements in the updated left singular matrix; the r right singular vectors are the r row vectors corresponding to the r diagonal elements in the updated right singular matrix.
5. The method according to claim 4, wherein: The updated first low-rank matrix is a matrix that retains the r left singular vectors to shield the remaining vectors, and the updated second low-rank matrix is a matrix that retains the r right singular vectors to shield the remaining vectors; Determining the current round increment matrix based on the product of the updated first low-rank matrix and the second low-rank matrix includes: The product of the updated first low-rank matrix, the updated diagonal matrix, and the updated second low-rank matrix multiplied in sequence is used as the incremental matrix of this round.
6. The method according to claim 4, wherein: The method further comprises: obtaining an updated first low-rank matrix and a second low-rank matrix according to the first gradient matrix; Before determining the r diagonal elements with larger values, calculating the importance scores of each triplet according to the first gradient matrix, wherein a single triplet includes the diagonal elements of the target position in the diagonal matrix, the column vector corresponding to the target position in the left singular matrix, and the row vector corresponding to the target position in the right singular matrix; The values of the diagonal elements corresponding to the first number of triplets ranked last in importance scores in the updated diagonal matrix are set to 0.
7. The method according to claim 6, wherein: The first number first increases as the number of fine-tuning rounds increases, and then decreases as the number of fine-tuning rounds increases.
8. The method according to claim 6, wherein: Calculating a loss function based on the difference between the predicted text output by the large language model and the marked text of the training text includes: A first loss function is calculated based on the difference between the predicted text output by the large language model and the marked text of the training text; a second loss function is calculated based on the degree to which the left singular matrix deviates from the orthogonal matrix and the degree to which the right singular matrix deviates from the orthogonal matrix, and the sum of the first loss function and the second loss function is used as the loss function of this round, Each round of fine-tuning also includes: determining the gradient of the loss function on each parameter in the left singular matrix, the diagonal matrix, and the right singular matrix; determining the importance of each parameter based on each parameter and its gradient; and calculating the importance score of each triplet based on the importance of each parameter.
9. The method according to claim 1, wherein: Updating the amplitude vector based on the loss function to obtain an updated amplitude vector includes: Calculate the product of the gradient matrix of the loss function relative to the weight matrix and the first scaling factor, and determine the gradient vector of the loss function relative to the amplitude vector according to the product result; wherein the first scaling factor is equal to the direction matrix divided by the vector norm of each column vector in the current direction matrix; The magnitude vector is updated based on the gradient vector.
10. The method according to claim 1, wherein: Determining a first gradient matrix of the loss function relative to the current direction matrix includes: The first gradient matrix is determined according to the product of the gradient matrix of the loss function relative to the weight matrix and a second scaling factor, wherein the second scaling factor is equal to the magnitude vector divided by a target constant for approximating the norm of each column vector in the direction matrix.
11. A fine-tuning device for a large language model, comprising: an acquisition unit configured to acquire a weight matrix of a pre-trained large language model, decompose the weight matrix into an amplitude vector and a direction matrix, and initialize a first low-rank matrix and a second low-rank matrix, wherein the product of the first low-rank matrix and the second low-rank matrix is used to fit an incremental matrix of the first direction matrix; The fine-tuning unit is configured to keep the first direction matrix unchanged and perform multiple rounds of fine-tuning, each round of fine-tuning comprising: Inputting a training text into the large language model, and calculating a loss function based on the difference between a predicted text output by the large language model and a marked text of the training text; Updating the amplitude vector based on the loss function to obtain an updated amplitude vector; Determine a first gradient matrix of the loss function relative to the current direction matrix; An updated first low-rank matrix and a second low-rank matrix are obtained according to the first gradient matrix; an incremental matrix of this round is determined based on the product of the updated first low-rank matrix and the second low-rank matrix, and the incremental matrix of this round is superimposed on the first direction matrix as the updated direction matrix of this round.
12. A computer program product, comprising a computer program / instruction, which, when executed by a processor, implements the steps of the method according to any one of claims 1 to 10.
13. A computing device comprising a memory and a processor, characterized in that: The memory stores executable codes, and when the processor executes the executable codes, the method according to any one of claims 1 to 10 is implemented.
Citation Information
Cited By
Distributed power model updating method and device based on model increment training and electronic equipment
CN120408010A
Large language model zero-order fine tuning method and system based on low-rank projection matrix learning
CN122045829A