Large model fine tuning method and device, equipment, storage medium and program product

By determining the target network layer and freezing some weights in LoRA fine-tuning, and using a low-rank matrix to adjust model parameters, the problems of high computational complexity and large memory usage are solved, achieving faster training speed and stronger cross-domain adaptability.

CN121009946APending Publication Date: 2025-11-25MOORE THREAD INTELLIGENT TECHNOLOGY (HANGZHOU) CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202511087626.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-08-04
Publication Date
2025-11-25

AI Technical Summary

Technical Problem

Existing LoRA fine-tuning methods have high computational complexity and large memory usage during training, and are not adaptable enough to cross-domain transfer training.

Method used

The matching relationship between the first index and the index threshold is determined by the low-rank matrix based on the pre-trained LoRA model. The model parameters are adjusted only for the target network layer. Fine-tuning training is performed using the low-rank matrix, some weights are frozen, and the training process is optimized.

Benefits of technology

It significantly reduces computational complexity and memory usage during training, accelerates model training, and improves adaptability in cross-domain transfer training.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121009946A_ABST
    Figure CN121009946A_ABST
Patent Text Reader

Abstract

The embodiment of the invention discloses a large model fine tuning method and device, equipment, a storage medium and a program product, and the method comprises the steps: determining a first index based on a low-rank matrix of a pre-trained low-rank adaptive LoRA model; the first index represents the matrix size of the pre-trained LoRA model; the pre-trained LoRA model is obtained by training on the basis of partial data selected from first training data of a large model; based on a matching relationship between the first index and an index threshold value, determining a target network layer which needs to adopt the low-rank matrix to carry out model parameter adjustment from a plurality of network layers of the large model; and performing fine tuning training on the model parameters of the target network layer based on the low-rank matrix by adopting the first training data.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This disclosure relates to, but is not limited to, the field of machine learning technology, and in particular to a method, apparatus, device, storage medium, and program product for fine-tuning large models. Background Technology

[0002] Low-Rank Adaptation of Large Language Models (LoRA) is a low-rank adaptation technique for fine-tuning large language models, aiming to efficiently adjust pre-trained models with a small number of trainable parameters. However, LoRA fine-tuning still requires further optimization. Summary of the Invention

[0003] In view of this, the present disclosure provides at least one method, apparatus, device, storage medium, and program product for fine-tuning large models.

[0004] The technical solution of this disclosure embodiment is implemented as follows:

[0005] On one hand, embodiments of this disclosure provide a method for fine-tuning a large model, the method comprising:

[0006] Based on the low-rank matrix of the pre-trained low-rank adaptive LoRA model, a first metric is determined; the first metric characterizes the matrix size of the pre-trained LoRA model; the pre-trained LoRA model is obtained by training on a subset of data selected from the first training data of the large model.

[0007] Based on the matching relationship between the first indicator and the indicator threshold, the target network layer that needs to be adjusted by using a low-rank matrix is ​​determined from multiple network layers of the large model.

[0008] Using the first training data, the model parameters of the target network layer are fine-tuned based on the low-rank matrix.

[0009] On the other hand, embodiments of this disclosure provide a large model fine-tuning device, which includes:

[0010] The processing module is configured to determine a first metric based on the low-rank matrix of the pre-trained low-rank adaptive LoRA model; the first metric characterizes the matrix size of the pre-trained LoRA model; the pre-trained LoRA model is obtained by training on a subset of data selected from the first training data of the large model.

[0011] The processing module is also configured to determine the target network layer from multiple network layers of the large model that requires model parameter adjustment using a low-rank matrix, based on the matching relationship between the first indicator and the indicator threshold.

[0012] The training module is configured to fine-tune the model parameters of the target network layer based on the first training data and the low-rank matrix.

[0013] In another aspect, embodiments of this disclosure provide a computer device, including a memory and a processor. The memory stores a computer program that can run on the processor, and the processor executes the program to implement some or all of the steps in the above-described method.

[0014] In another aspect, embodiments of this disclosure provide a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements some or all of the steps in the above-described method.

[0015] In another aspect, embodiments of this disclosure provide a computer program including computer-readable code, which, when executed in a computer device, causes a processor in the computer device to perform some or all of the steps in the above-described method.

[0016] In another aspect, embodiments of this disclosure provide a computer program product, which includes a non-transitory computer-readable storage medium storing a computer program. When the computer program is read and executed by a computer, it implements some or all of the steps in the above-described method.

[0017] In this embodiment, based on the matching relationship between the first index determined by the low-rank matrix of the pre-trained LoRA model and the index threshold, the target network layer that needs to be adjusted using the low-rank matrix can be determined. Then, the model parameters of the target network layer are fine-tuned only based on the low-rank matrix. In this way, compared with LoRA fine-tuning training of all network layers, the number of parameters that need to be trained can be greatly reduced, the computational complexity and memory usage during training can be significantly reduced, the training speed of the model can be accelerated, and in cross-domain transfer training, it can quickly adapt to new domains and has strong applicability.

[0018] It should be understood that the above general description and the following detailed description are merely exemplary and explanatory, and are not intended to limit the technical solutions of this disclosure. Attached Figure Description

[0019] The accompanying drawings, which are incorporated in and form part of this specification, illustrate embodiments consistent with this disclosure and, together with the specification, serve to illustrate the technical solutions of this disclosure.

[0020] Figure 1 A schematic diagram of the implementation process of a large model fine-tuning method provided in this embodiment of the disclosure. Figure 1 ;

[0021] Figure 2 A schematic diagram of the implementation process of a large model fine-tuning method provided in this embodiment of the disclosure. Figure 2 ;

[0022] Figure 3 This is a schematic diagram of the composition structure of a large model fine-tuning device provided in an embodiment of the present disclosure;

[0023] Figure 4 This is a schematic diagram of the hardware entity of a computer device provided in an embodiment of this disclosure. Detailed Implementation

[0024] To make the objectives, technical solutions, and advantages of this disclosure clearer, the technical solutions of this disclosure are further described in detail below with reference to the accompanying drawings and embodiments. The described embodiments should not be regarded as limitations on this disclosure. All other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this disclosure.

[0025] In the following description, references are made to “some embodiments,” which describe a subset of all possible embodiments. However, it is understood that “some embodiments” may be the same subset or different subsets of all possible embodiments and may be combined with each other without conflict.

[0026] The terms “first / second / third” are used merely to distinguish similar objects and do not represent a specific ordering of objects. It is understood that “first / second / third” may be interchanged in a specific order or sequence where permitted, so that the embodiments of this disclosure described herein can be implemented in an order other than that illustrated or described herein.

[0027] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this disclosure pertains. The terminology used herein is for descriptive purposes only and is not intended to limit the scope of this disclosure.

[0028] This disclosure provides a method for fine-tuning a large model, which can be executed by a processor of a computer device. The computer device refers to a device with data processing capabilities, such as a server, laptop, tablet, desktop computer, smart TV, set-top box, or mobile device (e.g., mobile phone, portable video player, personal digital assistant, dedicated messaging device, portable gaming device). Figure 1 As shown, the method includes the following steps 101 to 103:

[0029] Step 101: Determine a first index based on the low-rank matrix of the pre-trained low-rank adaptive LoRA model; the first index characterizes the matrix size of the pre-trained LoRA model; the pre-trained LoRA model is obtained by training based on a portion of the data selected from the first training data of the large model.

[0030] A pre-trained LoRA model refers to a LoRA model trained using a small batch of data. The first metric reflects the matrix size of the pre-trained LoRA model. The low-rank matrix refers to the low-rank matrix derived from the weight decomposition of the large model. The large model refers to any large model to be processed. A large model can be used for image segmentation, speech recognition, traffic person capture, etc. The first training data refers to the training data of the large model.

[0031] In some implementations, the initial determination of the low-rank matrix can be achieved by performing a low-rank decomposition on the weights of the large model to obtain a first low-rank matrix A and a second low-rank matrix B; the first low-rank matrix A and the second low-rank matrix B are then used as the low-rank matrices of the LoRA model.

[0032] In some implementations, the pre-trained LoRA model can be obtained by selecting a portion of data from the first training data and using this portion of data to train the initial LoRA model, thus obtaining the pre-trained LoRA model. Training the initial LoRA model can refer to fine-tuning the first low-rank matrix A and the second low-rank matrix B in the initial LoRA model. Furthermore, the portion of data used to train the initial LoRA model can also be data of the same data type as the first training data but not included in the first training data.

[0033] Step 102: Based on the matching relationship between the first indicator and the indicator threshold, determine the target network layer from the multiple network layers of the large model that needs to be adjusted using the low-rank matrix.

[0034] A large model can include multiple network layers, such as an input layer, multiple Transformer layers, and an output layer. Each network layer can incorporate a low-rank matrix to fine-tune its weights. A threshold metric is used to determine the network layers requiring LoRA fine-tuning. The target network layer is the one undergoing LoRA fine-tuning.

[0035] In some implementations, the indicator threshold may include a first indicator threshold and a second indicator threshold. Step 102 may be implemented by determining the target network layer based on the matching relationship between the first indicator and the first indicator threshold and the second indicator threshold.

[0036] Step 103: Using the first training data, fine-tune the model parameters of the target network layer based on the low-rank matrix.

[0037] Fine-tuning the model parameters of the target network layer based on the low-rank matrix means performing LoRA fine-tuning only on the network layers that require LoRA fine-tuning. Compared to performing LoRA fine-tuning on all network layers, this can significantly reduce the number of parameters that need to be trained, significantly reduce the computational complexity during the training process, speed up the training of the model, and enable it to quickly adapt to new domains in cross-domain transfer training, making it highly applicable.

[0038] In some implementations, step 103 can be specifically implemented as follows: using the first training data, fine-tuning the model parameters of the target network layer based on the low-rank matrix and the fine-tuning direction until a well-trained large model is obtained; the fine-tuning direction represents the correlation between the weights of the large model, the target matrix, and the low-rank matrix, and the target matrix is ​​used to control the fine-tuning amplitude.

[0039] In this embodiment, based on the matching relationship between the first index determined by the low-rank matrix of the pre-trained LoRA model and the index threshold, the target network layer that needs to be adjusted using the low-rank matrix can be determined. Then, the model parameters of the target network layer are fine-tuned only based on the low-rank matrix. In this way, compared with LoRA fine-tuning training of all network layers, the number of parameters that need to be trained can be greatly reduced, the computational complexity and memory usage during training can be significantly reduced, the training speed of the model can be accelerated, and in cross-domain transfer training, it can quickly adapt to new domains and has strong applicability.

[0040] This disclosure provides a method for fine-tuning a large model, which can be executed by the processor of a computer device. For example... Figure 2 As shown, the method includes the following steps 201 to 206:

[0041] Step 201: Select a portion of the data from the first training data to obtain the second training data.

[0042] The second training data refers to a subset of data selected from the first training data.

[0043] In some implementations, step 201 can be specifically implemented by: using a random selection method to select a portion of data from the first training data to obtain the second training data. Alternatively, a balanced selection method based on data acquisition channels can be used to select a portion of data from the first training data to obtain the second training data.

[0044] Step 202: Train the initial LoRA model using the second training data to obtain the pre-trained LoRA model.

[0045] The initial LoRA model refers to the untrained LoRA model. Since the second training data is collected from the first training data, using the second training data to train the initial LoRA model makes the pre-trained LoRA model more suitable for fine-tuning large models.

[0046] In some implementations, the first low-rank matrix A in the initial LoRA model can be initialized using a normal distribution, and the second low-rank matrix B can be initialized to 0.

[0047] Step 203: Based on the low-rank matrix of the pre-trained LoRA model, determine the first index threshold and the second index threshold.

[0048] The first and second threshold indicators are used together to determine the network layers that need LoRA fine-tuning training.

[0049] In some implementations, step 203 can be specifically implemented by: determining a first index threshold and a second index threshold based on the low-rank matrix of the pre-trained LoRA model and the training objective of the large model.

[0050] In some implementations, step 203 can be specifically implemented by: manually tuning parameters to determine the first and second index thresholds based on the low-rank matrix of the pre-trained LoRA model.

[0051] Step 204: Determine the first index based on the low-rank matrix of the pre-trained low-rank adaptive LoRA model.

[0052] In some implementations, step 204 can be specifically implemented as follows: determining the first low-rank matrix and the second low-rank matrix in the low-rank matrix; and determining the first index based on the norms of the first low-rank matrix and the second low-rank matrix.

[0053] In some implementations, the first index can be determined by using the first norm of the first low-rank matrix and the second low-rank matrix as the first index; or by using the second norm of the first low-rank matrix and the second low-rank matrix as the first index.

[0054] Step 205: Based on the matching relationship between the first indicator and the indicator threshold, determine the target network layer from the multiple network layers of the large model that needs to be adjusted using the low-rank matrix.

[0055] In some embodiments, the specific implementation of step 205 can be: determining the first metric threshold and the second metric threshold among the metric thresholds; for any one of the candidate network layers in the large model, when the first metric of the candidate network layer is greater than the first metric threshold and less than the second metric threshold, determining the candidate network layer as the target network layer.

[0056] In some embodiments, the first metric threshold can be denoted as t1, and the second metric threshold can be denoted as t2. At this time, if t1 < first metric < t2, it means that this network layer needs to perform LoRA fine-tuning training; if the first metric < t1 or the first metric > t2, it means that this network layer does not need to perform LoRA fine-tuning training and can be discarded.

[0057] Step 206: Using the first training data, based on the low-rank matrix, fine-tune the model parameters of the target network layer.

[0058] In some embodiments, the specific implementation of step 206 can be: inputting the first training data into the large model to obtain the model output of the large model; determining the gradient of the target network layer based on the model output of the large model and the loss function; fine-tuning the low-rank matrix of the target network layer based on the gradient of the target network layer to obtain a fine-tuned low-rank matrix; fine-tuning the weights of the target network layer based on the adjusted low-rank matrix to obtain the fine-tuned weights of the target network layer; updating the fine-tuned weights of the target network layer to the large model, and continuing to input the first training data into the updated large model, and so on, until a trained large model is obtained.

[0059] In some embodiments, during the process of fine-tuning the model parameters of the target network layer based on the first training data and the low-rank matrix, the second low-rank matrix is updated, and the first low-rank matrix is not updated.

[0060] The weights W0 of the large model and the first low-rank matrix A are frozen and not trained, and only the second low-rank matrix B is trained. In this way, the video memory can be reduced. After a large number of experiments, it is proved that freezing A and only training B will not affect the performance and accuracy of the model.

[0061] In some embodiments, the specific implementation of "updating the second low-rank matrix" can be: determining the first parameter and the learning rate; the first parameter is used to control the rate of gradient descent; updating the second low-rank matrix based on the first parameter, the learning rate, and the gradient of the second low-rank matrix.

[0062] The first parameter can be the hyperparameter λ. The learning rate can be represented as η. The update formula for the second low-rank matrix can be: B(t+1)=B(t)-ληΔB(t). B(t) represents the unupdated B(t), and B(t+1) represents the updated second low-rank matrix.

[0063] During implementation, it was found that initializing the second low-rank matrix B to 0 required a larger update step. Experiments showed that setting λ to 16 could speed up the training time of models such as llama-7b by a factor of 2.

[0064] In some embodiments, the present disclosure may also determine a weight fine-tuning expression through steps A to C, so as to update the weights of the target network layer based on the weight fine-tuning expression.

[0065] Step A: Initialize the second parameter based on the weights of the large model; the second parameter is used to control the fine-tuning amplitude.

[0066] The second parameter can be an m-matrix, used to control the fine-tuning amplitude.

[0067] In some implementations, step A can be specifically implemented as follows: initialize the second parameter based on the norm of the weights of the large model.

[0068] In some implementations, step A can be specifically implemented as follows: initialize the second parameter of each network layer based on the weights of each network layer in the large model.

[0069] For example, the relationship between the weights of the large model and the second matrix can be expressed as: ||m|| n =||w0|| n Where m is the second parameter and w0 is the weight of the large model (or network layer).

[0070] It should be noted that since the large model consists of multiple network layers, a second parameter can be set for each network layer, and the second parameter of each network layer can be set based on the weights of the corresponding network layer. Alternatively, the second parameter of each network layer can be initialized based on the total weights of the large model.

[0071] Step B: Determine the fine-tuning direction based on the weights of the large model and the low-rank matrix.

[0072] In some implementations, the fine-tuning direction can be an expression determined based on the weights and low-rank matrix of the large model, which can be expressed as:

[0073] Step C: Determine the weight fine-tuning expression based on the second parameter and the fine-tuning direction.

[0074] In some implementations, the weight fine-tuning expression can be represented as: Among them, the first low-rank matrix A, the second low-rank matrix B, and the second parameter m are all trainable, while W0 is frozen (not trained).

[0075] Based on the weight fine-tuning expression, step 206 above can be achieved through steps 2061 to 2065 as follows:

[0076] Step 2061: Input the first training data into the large model to obtain the model output of the large model.

[0077] In some implementations, step 2061 can be specifically implemented as follows: the first training data x is transformed by the weights W0 of the large model to obtain W0x; at the same time, W0x is also transformed by the low-rank matrices A and B to obtain BAx; based on the weighted sum of W0x and BAx, the final model output is obtained.

[0078] Step 2062: Based on the model output and loss function of the large model, determine the gradient of the target network layer.

[0079] In some implementations, the loss function can be the cross-entropy function, or mean squared error, mean absolute error, etc.

[0080] In some implementations, step 2062 can be specifically implemented as follows: the model output is processed using a loss function to obtain the total gradient; the total gradient is processed using the chain rule to obtain the gradient of the target network layer.

[0081] Step 2063: Fine-tune the low-rank matrix and the second parameter of the target network layer based on the gradient of the target network layer to obtain the fine-tuned low-rank matrix and the fine-tuned second parameter.

[0082] In some implementations, step 2063 can be specifically implemented by using an optimizer to fine-tune the low-rank matrix and second parameter of the target network layer based on the gradient of the target network layer. The optimizer can be stochastic gradient descent (SGD) or adaptive moment estimation (Adam), etc.

[0083] Step 2064: Using the weight fine-tuning expression, fine-tune the weights of the target network layer based on the adjusted low-rank matrix and the adjusted second parameter to obtain the fine-tuned weights of the target network layer.

[0084] Step 2065: Update the weights of the fine-tuned target network layer to the large model, and continue to input the first training data into the updated large model until the trained large model is obtained.

[0085] In some implementations, step 2065 can be specifically implemented as follows: performing a first operation on the adjusted first low-rank matrix and the adjusted second low-rank matrix to obtain a first operation result; the adjusted first low-rank matrix and the adjusted second low-rank matrix are included in the adjusted low-rank matrix; performing a second operation on the weights of the target network layer and the first operation result to obtain a second operation result; performing a third operation on the norm of the second operation result and the second operation result to obtain a third operation result; performing a fourth operation on the third operation result and the adjusted second parameter to obtain the fine-tuned weights of the target network layer; the weight fine-tuning expression characterizes the operational relationship between the low-rank matrix, the target matrix, and the weights of the target network layer.

[0086] For example, the first operation refers to matrix multiplication, the second operation refers to addition, the third operation refers to division, and the fourth operation refers to multiplication.

[0087] In some embodiments, based on the foregoing embodiments, the large model fine-tuning method provided in this disclosure further includes the following steps 207 to 208:

[0088] Step 207: Initialize the first low-rank matrix in the low-rank matrix using a normal distribution; the initial value of the first low-rank matrix is ​​one order of magnitude smaller than the weights of the large model.

[0089] For example, if the weights of the large model are 1, then the value of the first low-rank matrix is ​​0.1 or 0.08. The initial value of A should be as small as possible, an order of magnitude smaller than the weights of the large model, to ensure that the values ​​of AB are not too large and affect the performance of the original large model.

[0090] Step 208: Initialize the second low-rank matrix in the low-rank matrix to the target threshold.

[0091] The target threshold can be 0. A is initialized randomly using a normal distribution, and B is initialized to 0, thus ensuring that the initial values ​​of AB are 0.

[0092] In this embodiment, the matching relationship between the first index and the index threshold determined by the low-rank matrix of the pre-trained LoRA model can identify the target network layer for which the model parameters need to be adjusted using the low-rank matrix. Then, the model parameters of the target network layer are fine-tuned only based on the low-rank matrix. In this way, compared with LoRA fine-tuning training of all network layers, the number of parameters that need to be trained can be significantly reduced, the computational complexity and memory usage during training can be significantly reduced, the training speed of the model can be accelerated, and in cross-domain transfer training, it can quickly adapt to new domains and has strong applicability.

[0093] The following describes the application of the large model fine-tuning method provided in the embodiments of this disclosure in real-world scenarios.

[0094] In existing solutions, full fine-tuning consumes too much GPU resources. LoRA fine-tuning still has many shortcomings, such as: 1) updating some parameters does not have a significant impact on model accuracy, but it will occupy a lot of GPU memory during training; 2) the fine-tuning effect of LoRA is not ideal.

[0095] This disclosure provides three optimization methods for LoRA fine-tuning.

[0096] Optimization Method 1:

[0097] (1) Initialize A (corresponding to the first low-rank matrix above) and B (corresponding to the second low-rank matrix above). A is randomly initialized using a normal distribution, and B is initialized to 0, ensuring that the initial values ​​of AB are 0.

[0098] (2) The initial value of A should be as small as possible, one order of magnitude smaller than the weight W of the large model, so as to ensure that the value of AB is not too large and will affect the performance of the original large model.

[0099] (3) Fix the weights of A and train only B. The weight update formula is as follows:

[0100] W = W0 + AB;

[0101] Where W0 is the large model weight, and A and B are LoRA weights. If W0 and A are frozen and not trained, and only B is trained, the GPU memory will be reduced. However, after a lot of experiments, it has been proven that freezing A and only training B will not affect the model performance and accuracy.

[0102] (4) The parameters of the second low-rank matrix B are updated using SGD. The update formula is as follows:

[0103] B(t+1)=B(t)-ληΔB(t);

[0104] Where ΔB is the gradient of B, η is the learning rate, and λ is a hyperparameter, an additional parameter we add to speed up training.

[0105] The initialization of the second low-rank matrix B is 0, so larger update steps are required. Proven by experiments, setting λ to 16 speeds up the training time of models such as llama-7b by 2 times.

[0106] Optimization method two:

[0107] (5) Sample the training data and select a small batch of data (corresponding to the above-mentioned second training data).

[0108] (6) Use this small batch of data for LoRA fine-tuning and save the fine-tuned LoRA model.

[0109] (7) Calculate the norm of the AB matrix in the fine-tuned LoRA model. The second norm, the first norm, or other norms can be used.

[0110] (8) Manually adjust the lower threshold t1 (corresponding to the above-mentioned first index threshold) and the upper threshold t2 (corresponding to the above-mentioned second index threshold). When the following condition is met:

[0111] t1 < AB norm < t2 indicates that LoRA needs to be applied to this network layer; AB norm < t1 or AB norm > t2 means LoRA is not required and can be dropped.

[0112] (9) Use all the training data and only perform LoRA fine-tuning training on the layers that require LoRA fine-tuning. Save the model after training.

[0113] Optimization method three:

[0114] (10) Initialize A and B. A is initialized with a normal distribution, and B is initialized to 0.

[0115] (11) Initialize m (corresponding to the above-mentioned second parameter). m is a matrix, and the values of the m matrix are randomly initialized, but the following conditions must be met: ||m|| n = ||w0|| n ;

[0116] where W0 is the large model weight, and ||m|| n represents the n-norm of m.

[0117] (12) The weight fine-tuning expression can be:

[0118] where the m matrix controls the fine-tuning amplitude, controls the fine-tuning direction.

[0119] (13) A, B, and m are trainable, and the w0 parameter is frozen.

[0120] (14) The loss function is cross-entropy or other, the optimizer is SGD or Adam, etc., and A, B and m are updated.

[0121] (15) After training is complete, save matrices A, B and m.

[0122] (16) The above three optimization methods can be used in combination or individually. Combining them will have a better effect, improving performance and reducing video memory usage.

[0123] Based on the foregoing embodiments, this disclosure provides a large model fine-tuning device, which includes the included units and the modules included in each unit, and can be implemented by a processor in a computer device; of course, it can also be implemented by specific logic circuits; in the implementation process, the processor can be a central processing unit (CPU), a microprocessor unit (MPU), a digital signal processor (DSP), or a field programmable gate array (FPGA), etc.

[0124] Figure 3 This is a schematic diagram of the composition structure of a large model fine-tuning device provided in an embodiment of this disclosure, as shown below. Figure 3 As shown, the large model fine-tuning device 300 includes: a processing module 310 and a training module 320, wherein:

[0125] The processing module 310 is configured to determine a first index based on the low-rank matrix of a pre-trained low-rank adaptive LoRA model; the first index characterizes the matrix size of the pre-trained LoRA model; the pre-trained LoRA model is obtained by training based on a portion of the data selected from the first training data of the large model.

[0126] The processing module 310 is further configured to determine, based on the matching relationship between the first indicator and the indicator threshold, the target network layer from multiple network layers of the large model that needs to be adjusted using the low-rank matrix;

[0127] The training module 320 is configured to use the first training data to fine-tune the model parameters of the target network layer based on the low-rank matrix.

[0128] In some embodiments, the processing module 310 is further configured to: determine a first low-rank matrix and a second low-rank matrix in the low-rank matrix; and determine the first index based on the norms of the first low-rank matrix and the second low-rank matrix.

[0129] In some embodiments, the processing module 310 is further configured to: determine a first indicator threshold and a second indicator threshold among the indicator thresholds; and for any candidate network layer in the large model, if the first indicator of the candidate network layer is greater than the first indicator threshold and less than the second indicator threshold, determine the candidate network layer as the target network layer.

[0130] In some embodiments, the processing module 310 is further configured to: select a portion of data from the first training data to obtain second training data; train an initial LoRA model using the second training data to obtain the pre-trained LoRA model; determine a first indicator threshold and a second indicator threshold based on the low-rank matrix of the pre-trained LoRA model; and determine the indicator threshold based on the first indicator threshold and the second indicator threshold.

[0131] In some embodiments, the processing module 310 is further configured to: initialize a first low-rank matrix in the low-rank matrix using a normal distribution; the initial value of the first low-rank matrix is ​​an order of magnitude smaller than the weights of the large model; and initialize a second low-rank matrix in the low-rank matrix as a target threshold.

[0132] In some embodiments, the training module 320 is further configured to: update the second low-rank matrix but not update the first low-rank matrix during the process of fine-tuning the model parameters of the target network layer based on the low-rank matrix using the first training data.

[0133] In some embodiments, the processing module 310 is further configured to: determine a first parameter and a learning rate; the first parameter is used to control the descent speed of the gradient; and update the second low-rank matrix based on the first parameter, the learning rate, and the gradient of the second low-rank matrix.

[0134] In some embodiments, the processing module 310 is further configured to: initialize a second parameter based on the weights of the large model; the second parameter is used to control the fine-tuning amplitude; determine the fine-tuning direction based on the weights of the large model and the low-rank matrix; and determine a weight fine-tuning expression based on the second parameter and the fine-tuning direction.

[0135] In some embodiments, the training module 320 is further configured to: input the first training data into the large model to obtain the model output of the large model; determine the gradient of the target network layer based on the model output and loss function of the large model; fine-tune the low-rank matrix and the second parameter of the target network layer based on the gradient of the target network layer to obtain the fine-tuned low-rank matrix and the fine-tuned second parameter; fine-tune the weights of the target network layer based on the adjusted low-rank matrix and the adjusted second parameter using the weight fine-tuning expression to obtain the fine-tuned weights of the target network layer; update the weights of the fine-tuned target network layer to the large model, and continue to input the first training data into the updated large model until the trained large model is obtained.

[0136] In some embodiments, the training module 320 is further configured to: perform a first operation on the adjusted first low-rank matrix and the adjusted second low-rank matrix to obtain a first operation result; the adjusted low-rank matrix includes the adjusted first low-rank matrix and the adjusted second low-rank matrix; perform a second operation on the weights of the target network layer and the first operation result to obtain a second operation result; perform a third operation on the second operation result and the norm of the second operation result to obtain a third operation result; perform a fourth operation on the third operation result and the adjusted second parameter to obtain the fine-tuned weights of the target network layer; the weight fine-tuning expression characterizes the operational relationship between the low-rank matrix, the target matrix, and the weights of the target network layer.

[0137] The descriptions of the apparatus embodiments above are similar to those of the method embodiments above, and have similar beneficial effects. In some embodiments, the functions or modules included in the apparatus provided in this disclosure can be used to perform the methods described in the method embodiments above. For technical details not disclosed in the apparatus embodiments of this disclosure, please refer to the descriptions of the method embodiments of this disclosure for understanding.

[0138] It should be noted that, in the embodiments of this disclosure, if the above-described large model fine-tuning method is implemented as a software functional module and sold or used as an independent product, it can also be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the embodiments of this disclosure, or the part that contributes to related technologies, can be embodied in the form of a software product. This software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the methods described in the various embodiments of this disclosure. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), magnetic disks, or optical disks. Thus, the embodiments of this disclosure are not limited to any specific hardware, software, or firmware, or any combination of hardware, software, and firmware.

[0139] This disclosure provides a computer device including a memory and a processor. The memory stores a computer program that can run on the processor. When the processor executes the program, it implements some or all of the steps in the above-described method.

[0140] This disclosure provides a computer-readable storage medium storing a computer program thereon, which, when executed by a processor, implements some or all of the steps in the above-described method. The computer-readable storage medium may be transient or non-transient.

[0141] This disclosure provides a computer program including computer-readable code, wherein when the computer-readable code is executed in a computer device, a processor in the computer device performs some or all of the steps in the above-described method.

[0142] This disclosure provides a computer program product, which includes a non-transitory computer-readable storage medium storing a computer program. When the computer program is read and executed by a computer, it implements some or all of the steps in the above-described method. This computer program product can be implemented specifically through hardware, software, or a combination thereof. In some embodiments, the computer program product is specifically embodied as a computer storage medium; in other embodiments, the computer program product is specifically embodied as a software product, such as a software development kit (SDK), etc.

[0143] It should be noted that the descriptions of the various embodiments above tend to emphasize the differences between them, while their similarities or commonalities can be referenced interchangeably. The descriptions of the above embodiments of the device, storage medium, computer program, and computer program product are similar to the descriptions of the above method embodiments and have similar beneficial effects. For technical details not disclosed in the embodiments of the device, storage medium, computer program, and computer program product of this disclosure, please refer to the descriptions of the method embodiments of this disclosure for understanding.

[0144] It should be noted that, Figure 4 This is a schematic diagram of a hardware entity of a computer device in an embodiment of this disclosure, such as... Figure 4 As shown, the hardware entity of the computer device 400 includes: a processor 401, a communication interface 402, and a memory 403, wherein:

[0145] Processor 401 typically controls the overall operation of computer device 400.

[0146] Communication interface 402 enables computer devices to communicate with other terminals or servers via a network.

[0147] The memory 403 is configured to store instructions and applications executable by the processor 401, and can also cache data to be processed or already processed (e.g., image data, audio data, voice communication data, and video communication data) of the processor 401 and various modules in the computer device 400. It can be implemented using flash memory or random access memory (RAM). Data transfer between the processor 401, the communication interface 402, and the memory 403 can be performed via bus 404.

[0148] It should be understood that the phrase "an embodiment" or "one embodiment" throughout the specification means that a specific feature, structure, or characteristic related to the embodiment is included in at least one embodiment of this disclosure. Therefore, "in one embodiment" or "one embodiment" appearing throughout the specification does not necessarily refer to the same embodiment. Furthermore, these specific features, structures, or characteristics can be combined in any suitable manner in one or more embodiments. It should be understood that in the various embodiments of this disclosure, the sequence numbers of the above steps / processes do not imply a sequential order of execution; the execution order of each step / process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of this disclosure. The sequence numbers of the above embodiments of this disclosure are merely descriptive and do not represent the superiority or inferiority of the embodiments.

[0149] It should be noted that, in this document, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Unless otherwise specified, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes that element.

[0150] In the several embodiments provided in this disclosure, it should be understood that the disclosed devices and methods can be implemented in other ways. The device embodiments described above are merely illustrative. For example, the division of units is only a logical functional division, and in actual implementation, there may be other division methods, such as: multiple units or components may be combined, or integrated into another system, or some features may be ignored or not executed. In addition, the coupling, direct coupling, or communication connection between the various components shown or discussed may be through some interfaces, and the indirect coupling or communication connection between devices or units may be electrical, mechanical, or other forms.

[0151] The units described above as separate components may or may not be physically separate. The components shown as units may or may not be physical units. They may be located in one place or distributed across multiple network units. Some or all of the units may be selected to achieve the purpose of this embodiment according to actual needs.

[0152] In addition, each functional unit in the various embodiments of this disclosure can be integrated into one processing unit, or each unit can be a separate unit, or two or more units can be integrated into one unit; the integrated unit can be implemented in hardware or in the form of hardware plus software functional units.

[0153] Those skilled in the art will understand that all or part of the steps of the above method embodiments can be implemented by hardware related to program instructions. The aforementioned program can be stored in a computer-readable storage medium. When the program is executed, it performs the steps of the above method embodiments. The aforementioned storage medium includes various media that can store program code, such as mobile storage devices, read-only memory (ROM), magnetic disks, or optical disks.

[0154] Alternatively, if the integrated units described above are implemented as software functional modules and sold or used as independent products, they can also be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this disclosure, or the part that contributes to related technologies, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the methods described in the various embodiments of this disclosure. The aforementioned storage medium includes various media capable of storing program code, such as mobile storage devices, ROM, magnetic disks, or optical disks.

[0155] The above description is merely an embodiment of this disclosure, but the scope of protection of this disclosure is not limited thereto. Any changes or substitutions that can be easily conceived by those skilled in the art within the scope of the technology disclosed in this disclosure should be included within the scope of protection of this disclosure.

Claims

1. A method for fine-tuning a large model, characterized in that, The large model fine-tuning method includes: Based on the low-rank matrix of the pre-trained low-rank adaptive LoRA model, a first index is determined; the first index characterizes the matrix size of the pre-trained LoRA model; the pre-trained LoRA model is obtained by training on a subset of data selected from the first training data of a large model. Based on the matching relationship between the first indicator and the indicator threshold, the target network layer that needs to be adjusted by the low-rank matrix is ​​determined from the multiple network layers of the large model. Using the first training data, the model parameters of the target network layer are fine-tuned based on the low-rank matrix.

2. The large model fine-tuning method according to claim 1, characterized in that, The low-rank matrix of the pre-trained low-rank adaptive LoRA model is used to determine the first metric, including: Determine the first low-rank matrix and the second low-rank matrix in the low-rank matrix; The first index is determined based on the norms of the first low-rank matrix and the second low-rank matrix.

3. The large model fine-tuning method according to claim 1, characterized in that, The step of determining the target network layer from multiple network layers of the large model that requires model parameter adjustment using the low-rank matrix, based on the matching relationship between the first indicator and the indicator threshold, includes: Determine the first and second indicator thresholds among the indicator thresholds; For any candidate network layer in the large model, if the first index of the candidate network layer is greater than the first index threshold and less than the second index threshold, the candidate network layer is determined to be the target network layer.

4. The method for fine-tuning a large model according to any one of claims 1 to 3, characterized in that, The large model fine-tuning method also includes: A portion of the data is selected from the first training data to obtain the second training data; The initial LoRA model is trained using the second training data to obtain the pre-trained LoRA model; Based on the low-rank matrix of the pre-trained LoRA model, the first index threshold and the second index threshold are determined. The indicator threshold is determined based on the first indicator threshold and the second indicator threshold.

5. The method for fine-tuning a large model according to any one of claims 1 to 3, characterized in that, The large model fine-tuning method also includes: The first low-rank matrix in the low-rank matrix is ​​initialized using a normal distribution; the initial value of the first low-rank matrix is ​​an order of magnitude smaller than the weights of the large model. The second low-rank matrix in the low-rank matrix is ​​initialized to the target threshold.

6. The method for fine-tuning a large model according to any one of claims 1 to 3, characterized in that, The step of fine-tuning the model parameters of the target network layer based on the low-rank matrix using the first training data includes: Using the first training data, the model parameters of the target network layer are fine-tuned based on the low-rank matrix. During the fine-tuning process of the model parameters of the target network layer using the first training data and the low-rank matrix, the second low-rank matrix is ​​updated, but the first low-rank matrix is ​​not updated.

7. The large model fine-tuning method according to claim 6, characterized in that, The updating of the second low-rank matrix includes: Determine the first parameter and the learning rate; the first parameter is used to control the descent speed of the gradient. The second low-rank matrix is ​​updated based on the first parameter, the learning rate, and the gradient of the second low-rank matrix.

8. The method for fine-tuning a large model according to any one of claims 1 to 3, characterized in that, The large model fine-tuning method also includes: Based on the weights of the large model, the second parameter is initialized; the second parameter is used to control the fine-tuning amplitude. Based on the weights of the large model and the low-rank matrix, the fine-tuning direction is determined; Based on the second parameter and the fine-tuning direction, the weight fine-tuning expression is determined.

9. The large model fine-tuning method according to claim 8, characterized in that, The step of fine-tuning the model parameters of the target network layer based on the low-rank matrix using the first training data includes: The first training data is input into the large model to obtain the model output of the large model; Based on the model output and loss function of the large model, the gradient of the target network layer is determined; Based on the gradient of the target network layer, the low-rank matrix and the second parameter of the target network layer are fine-tuned to obtain the fine-tuned low-rank matrix and the fine-tuned second parameter. Using the weight fine-tuning expression, the weights of the target network layer are fine-tuned based on the adjusted low-rank matrix and the adjusted second parameter to obtain the fine-tuned weights of the target network layer. The weights of the fine-tuned target network layer are updated to the large model, and the first training data is continued to be input into the updated large model until the trained large model is obtained.

10. The large model fine-tuning method according to claim 9, characterized in that, The step of fine-tuning the weights of the target network layer using the weight fine-tuning expression, based on the adjusted low-rank matrix and the adjusted second parameter, to obtain the fine-tuned weights of the target network layer includes: A first operation is performed on the adjusted first low-rank matrix and the adjusted second low-rank matrix to obtain a first operation result; the adjusted low-rank matrix includes the adjusted first low-rank matrix and the adjusted second low-rank matrix; A second operation is performed on the weights of the target network layer and the first operation result to obtain a second operation result; Perform a third operation on the second operation result and the norm of the second operation result to obtain the third operation result; A fourth operation is performed on the third operation result and the adjusted second parameter to obtain the weights of the fine-tuned target network layer; the weight fine-tuning expression represents the operational relationship between the low-rank matrix, the target matrix, and the weights of the target network layer.

11. A large-scale model fine-tuning device, characterized in that, The large model fine-tuning device includes: The processing module is configured to determine a first metric based on the low-rank matrix of a pre-trained low-rank adaptive LoRA model; the first metric characterizes the matrix size of the pre-trained LoRA model; the pre-trained LoRA model is trained based on a subset of data selected from the first training data of a large model. The processing module is further configured to determine, based on the matching relationship between the first indicator and the indicator threshold, the target network layer from the multiple network layers of the large model that needs to be adjusted using the low-rank matrix; The training module is configured to use the first training data to fine-tune the model parameters of the target network layer based on the low-rank matrix.

12. A computer device comprising a memory and a processor, the memory storing a computer program executable on the processor, characterized in that, When the processor executes the program, it implements the steps of the method according to any one of claims 1 to 10.

13. A computer-readable storage medium having a computer program stored thereon, characterized in that, When executed by a processor, the computer program implements the steps of the method according to any one of claims 1 to 10.

14. A computer program product, characterized in that, The computer program product includes a non-transitory computer-readable storage medium storing a computer program that, when read and executed by a computer, implements the steps of the method according to any one of claims 1 to 10.