Method and system for performing zero-order optimization fine tuning on large model in random subspace
By adopting a hierarchical low-rank perturbation method and lazy update strategy in the fine-tuning of large language models, the problem of large gradient estimation variance in high-dimensional fine-tuning is solved, and the memory and computing cost is reduced, and the efficiency and accuracy of model fine-tuning is improved.
Patent Information
- Application Number
- CN202510217768.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-26
- Publication Date
- 2025-06-10
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
During the high-dimensional fine-tuning process of large language models (LLMs), the gradient estimation variance of zero-order optimizers is large, resulting in performance degradation. The traditional random subspace method requires storing large projection matrices, and the memory and computational costs are too high.
A hierarchical low-rank perturbation method is adopted to generate a low-rank perturbation matrix by combining column orthogonal matrix and Gaussian random matrix for gradient estimation, and a lazy update strategy is introduced to regularly update the perturbation matrix to reduce overhead.
The variance and angle error of gradient estimation is significantly reduced, memory and computing costs are reduced, and the efficiency and accuracy of model fine-tuning is improved.
Smart Images

Figure CN120123710A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of large language model fine-tuning, and specifically relates to a method and system for zero-order optimization fine-tuning of a large model on a random subspace. Background Art
[0002] Large language models (LLMs), such as the GPT and LLaMA series, have demonstrated remarkable capabilities in natural language processing tasks and other fields in recent years. These models learn complex patterns in language data through deep learning, especially based on the Transformer architecture. However, for professional tasks that require domain-specific knowledge, LLMs may perform poorly. Fine-tuning provides an effective solution for the model to more effectively adapt to specific tasks by moderately adjusting the pre-trained LLMs using domain data.
[0003] In fine-tuning, first-order (FO) optimizers (such as SGD or Adam) are usually used to achieve better performance on domain datasets. However, as the scale of the LLMs model grows, due to the gradient calculation required for backpropagation (BP), the memory consumption of the first-order optimizer also increases significantly. To improve memory efficiency, MeZO first introduced a zero-order (ZO) optimizer into LLM fine-tuning, completely eliminating the need for BP. It only requires forward propagation and calculates gradient estimates through finite differences of the training loss value. However, the variance of the ZO gradient estimate linearly depends on the perturbation dimension, which corresponds to the number of model parameters. In LLMs, this dimension can be extremely large, resulting in a significant performance degradation compared to first-order optimizers.
[0004] To address the problem of high variance in ZO gradient estimates, there are mainly two solutions. The first solution is to increase the batch size as the training progresses, which can reduce the noise and variance in ZO gradient estimates. However, this method incurs significant runtime and memory overhead due to the large amount of data in the later training stages. The second solution is to reduce the number of perturbed parameters through sparse parameter perturbations, such as random and sparse pruning masks and block coordinate perturbations, or to reduce the number of trainable parameters through techniques such as parameter-efficient fine-tuning (PEFT) and tensorized adapters. Recent theoretical progress has proposed reducing the dependence on dimension in the ZO optimizer by applying low-dimensional perturbations in a random subspace using random projection. However, a major drawback of this method is the need to store a large projection matrix proportional to the dimension of the model parameters, which is impractical when fine-tuning large LLMs. Summary of the Invention
[0005] To solve the problems existing in the prior art, the present invention provides a method and system for zero-order optimization fine-tuning of large models on random subspaces, aiming to address the challenges of high-dimensional LLM fine-tuning. Specifically, a hierarchical low-rank perturbation method is designed for gradient estimation, specifically for LLM fine-tuning. In each layer, a low-rank perturbation matrix is generated by combining two column-orthogonal matrices and a Gaussian random matrix and used for gradient estimation. This method is different from traditional zero-order optimization methods (such as MeZO) that apply non-low-rank perturbations to the entire model, significantly reducing the variance of gradient estimation and the angular error between the estimated gradient and its expectation. Compared with random subspace zero-order optimization methods (such as S-RGF), NewZero significantly reduces the memory and computational costs by using smaller and layer-specific low-rank perturbation matrices instead of large model-level projection matrices. Additionally, a lazy update strategy is introduced to generate perturbations periodically rather than step by step, further reducing the overhead.
[0006] To achieve the above object, the present invention provides the following solutions:
[0007] A method for zero-order optimization fine-tuning of a large model on a random subspace, the method comprising:
[0008] Obtain task data in the target domain, preprocess the task data in the target domain to obtain a dataset in the format required by the large model;
[0009] Use the obtained dataset to fine-tune a pre-trained large language model;
[0010] Input the data to be processed into the fine-tuned large language model to complete the tasks in the target domain.
[0011] Preferably, using the obtained dataset to fine-tune a pre-trained large language model includes:
[0012] Step 1: Input the obtained dataset into the first layer of the large language model, and use the output result obtained from this layer as the input for the next layer, repeating this process until the final layer is reached to obtain the final output And through the final output Compare with the standard output Y corresponding to the current input data to calculate the loss function value obtained from the current forward propagation;
[0013] Step 2: For the first loop or loops where the iteration number is an integer multiple of the set update number, it is necessary to update the U and V matrices corresponding to each parameter in the large language model;
[0014] Step 3: For cases where there are already U and V matrices or where it is not necessary to update the U and V matrices, calculate the projected gradient;
[0015] Step 4: For each parameter, multiply the parameter matrix by the projection gradient to obtain the gradient matrix, and update the parameter to the current value minus the product of the learning rate and the gradient matrix;
[0016] Step 5: Repeat Steps 1 to 4 for T times to obtain the fine-tuned model.
[0017] Preferably, in Step 2, updating the U and V matrices corresponding to each parameter in the large language model includes:
[0018] Generate a random value, multiply the perturbation parameter given in the hyperparameters by the random value to obtain the full-scale perturbation value of the large language model parameters;
[0019] Add all the parameters to the full-scale perturbation value, calculate the loss function value, and then subtract twice the full-scale perturbation value to calculate the loss function value;
[0020] Subtract the two loss function values, and then divide by twice the perturbation parameter value to obtain the projection gradient;
[0021] Multiply the projection gradient by the parameter matrix to obtain the gradient matrix of the parameter;
[0022] Perform SVD decomposition on the gradient matrix to obtain three matrices P, S, and Q; only take the first R columns of the second dimension of the P matrix as the U matrix of the parameter; take the first R rows of the first dimension of the Q matrix as the V matrix of the parameter, and store the U and V matrices;
[0023] Repeat this operation for each parameter to obtain the U and V matrices of all parameters.
[0024] Preferably, in Step 3, calculating the projection gradient includes:
[0025] For each parameter, generate a random matrix that conforms to the Gaussian distribution, with a size of R x R, denoted as y;
[0026] Calculate the z matrix as the product of U y V, that is, the local perturbation of the parameter;
[0027] Multiply the perturbation parameter by the local perturbation matrix to obtain the perturbation matrix;
[0028] After calculating the perturbation matrices of all parameters, calculate the estimated projection gradient.
[0029] The present invention also provides a system for performing zero-order optimization and fine-tuning of a large model in a random subspace. The system is used to implement any of the above methods. The system includes: an acquisition module, a fine-tuning module, and a task execution module;
[0030] The acquisition module is used to acquire task data in the target domain, preprocess the task data in the target domain, and obtain a data set in the format required by the large model;
[0031] The fine-tuning module is used to fine-tune the pre-trained large language model using the obtained data set;
[0032] The task execution module is used to input the data to be processed into the fine-tuned large language model to complete the tasks in the target domain.
[0033] Preferably, the fine-tuning module includes: a loss calculation unit, a parameter matrix update unit, a projection gradient calculation unit, a gradient update unit, and an iteration unit;
[0034] The loss calculation unit is used to input the obtained data set into the first layer of the large language model, and use the output result obtained from this layer as the input of the next layer, and repeat this process until the final layer is reached to obtain the final output And through the final output Compare with the standard output Y corresponding to the current input data to calculate the loss function value obtained from the current forward propagation;
[0035] The parameter matrix update unit is used for the first loop or the loop whose iteration number is an integer multiple of the set update number, and it is necessary to update the U and V matrices corresponding to each parameter in the large language model;
[0036] The projection gradient calculation unit is used to calculate the projection gradient in the case where there are already U and V matrices or there is no need to update the U and V matrices;
[0037] The gradient update unit is used for each parameter, multiply the parameter matrix by the projection gradient to obtain a gradient matrix, and update the parameter to the current value minus the learning rate multiplied by the gradient matrix;
[0038] The iteration unit is used to repeat the loss calculation unit to the gradient update unit T times to obtain the fine-tuned model.
[0039] Preferably, the parameter matrix update unit includes: a full-scale perturbation value calculation subunit, a loss function calculation subunit, a projection gradient calculation subunit, a gradient matrix calculation subunit, a decomposition subunit, and an iteration subunit;
[0040] The full-scale perturbation value calculation subunit is used to generate a random value, multiply the perturbation parameter by the random value through the perturbation parameter given in the hyperparameters to obtain the full-scale perturbation value of the large language model parameters;
[0041] The loss function calculation subunit is used to add all the parameters with the full-scale perturbation value, calculate the loss function value, and then subtract twice the full-scale perturbation value to calculate the loss function value;
[0042] The projection gradient calculation subunit is used to subtract two loss function values and then divide by twice the perturbation parameter value to obtain the projection gradient;
[0043] The gradient matrix calculation subunit is used to multiply the projection gradient by the parameter matrix to obtain the gradient matrix of the parameter;
[0044] The decomposition subunit is used to perform SVD decomposition on the gradient matrix to obtain three matrices P, S, and Q; only take the first R columns of the second dimension of the P matrix as the U matrix of the parameter; take the first R rows of the first dimension of the Q matrix as the V matrix of the parameter, and store the U and V matrices;
[0045] The iteration subunit is used to repeat this operation for each parameter to obtain the U and V matrices of all parameters.
[0046] Preferably, the projection gradient calculation unit includes: a random matrix calculation subunit, a local perturbation calculation subunit, a perturbation matrix calculation subunit, and an estimated projection gradient calculation subunit;
[0047] The random matrix calculation subunit is used to generate a random matrix that conforms to the Gaussian distribution for each parameter, with a size of R x R, denoted as y;
[0048] The local perturbation calculation subunit is used to calculate the z matrix as the product of U y V, that is, the local perturbation of the parameter;
[0049] The perturbation matrix calculation subunit is used to multiply the perturbation parameter by the local perturbation matrix to obtain the perturbation matrix;
[0050] The estimated projection gradient calculation subunit is used to calculate the estimated projection gradient after calculating the perturbation matrices of all parameters.
[0051] Compared with the prior art, the beneficial effects of the present invention are:
[0052] The present invention proves that NewZero can effectively fine-tune large language models (LLMs), is applicable to a variety of tasks and fine-tuning schemes, and its memory overhead is comparable to the overhead in the model inference stage. Additional experiments show that NewZero can also optimize non-differentiable objectives. Theoretical analysis reveals how NewZero accelerates convergence by reducing the variance of gradient estimation.
[0053] The present invention significantly reduces memory overhead: NewZero only requires memory consumption comparable to the inference stage through a hierarchical low-rank perturbation strategy, greatly alleviating the high memory demand of traditional gradient calculation. Compared with existing random subspace zero-order optimization methods (such as S-RGF), NewZero avoids storing large-scale projection matrices, further reducing memory and computational costs.
[0054] The present invention improves the gradient estimation accuracy: the variance of the proposed gradient estimation method is significantly lower than that of traditional zero-order optimization methods (such as MeZO). The gradient estimation is closer to the true gradient (BP gradient), ensuring the accuracy of the optimization direction and improving the effect of model fine-tuning. BRIEF DESCRIPTION OF THE DRAWINGS
[0055] In order to more clearly illustrate the technical solutions of the present invention, the accompanying drawings required for use in the embodiments will be briefly introduced below. Obviously, the accompanying drawings in the following description are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other accompanying drawings can be obtained based on these drawings.
[0056] Figure 1 It is a schematic flowchart of a method for zero-order optimization fine-tuning of a large model on a random subspace according to an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0057] The technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only some embodiments of the present invention, rather than all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.
[0058] To make the above objects, features, and advantages of the present invention more obvious and understandable, the present invention will be further described in detail below with reference to the accompanying drawings and specific embodiments.
[0059] Embodiment 1
[0060] As can be seen from the background art,
[0061] Fine-tuning large language models (LLMs) has been proven effective in various downstream tasks. However, as the scale of LLMs continues to increase, the memory requirements for backpropagation become increasingly unbearable. Zero-order (ZO) optimization methods provide a memory-saving alternative by estimating gradients using forward propagation, but the variance of their gradient estimation usually grows linearly with the model parameter dimension - which is a major problem for LLMs. Therefore, a random subspace zero-order optimization method is proposed to address the challenges brought by the high dimensionality of LLMs. A low-rank perturbation method suitable for LLMs is designed, which significantly reduces memory consumption while improving training performance. It can be demonstrated that the proposed gradient estimation method can approximate the backpropagation gradient, with a variance lower than that of traditional zero-order methods, and can ensure convergence when combined with stochastic gradient descent (SGD).
[0062] Zero-order (ZO) optimization does not rely on backpropagation (BP), but estimates the gradient through random perturbations. A classic gradient estimation method is the Simultaneous Perturbation Stochastic Approximation (SPSA), which is defined as follows:
[0063]
[0064] where \(L(\mathbf{w}; \mathcal{B})\) is the loss on a mini-batch \(\mathcal{B}\) of size \(B\) uniformly sampled from the training dataset \(\mathcal{D}\), \(\mathbf{z} \in \mathbb{R}^d\) represents a random perturbation sampled from \(\mathcal{N}(0, I_d)\), and \(\epsilon\) is the perturbation magnitude.
[0065] The SPSA in the above equation is an unbiased gradient estimator of the desired gradient . This method only requires two forward passes to estimate the gradient and does not require backpropagation calculations, thus significantly reducing the computational cost and GPU memory consumption. Using this estimated gradient, it is possible to combine with existing first-order optimizers (such as SGD) to develop corresponding ZO optimizers, such as ZO-SGD, which is defined as follows:
[0066]
[0067] where \(\eta_t\) is the learning rate at the \(t\)-th step. In practice, MeZO implements ZO-SGD through in-place operations and uses a single random seed to efficiently regenerate perturbations, thus significantly reducing the memory overhead.
[0068] This scheme can be applied to fields such as healthcare, law, and education to solve the memory bottleneck and the problem of excessive variance in high-dimensional gradient estimation in domain fine-tuning.
[0069] Specifically, as Figure 1 shown, an embodiment of the present invention discloses a method for zero-order optimization fine-tuning of a large model on a random subspace, and the method includes:
[0070] Obtain task data in the target domain, preprocess the task data in the target domain to obtain a dataset in the format required by the large model;
[0071] Use the obtained dataset to fine-tune a pre-trained large language model;
[0072] Input the data to be processed into the fine-tuned large language model to complete the task in the target domain.
[0073] In this embodiment, the object of optimization: a pre-trained large language model, such as Llama2-7b, opt-1.3b, etc.
[0074] The dataset used for fine-tuning: medical data, including several case descriptions and diagnosis results. During the data preprocessing process, the data is divided into training samples, test samples, and validation samples according to a certain ratio.
[0075] Fine-tuning of large models is the process of further training a pre-trained large model using a dataset from a specific domain. It aims to optimize the model's performance on specific tasks, enabling the model to better adapt to and complete tasks in a specific domain.
[0076] In this embodiment, data acquisition and preprocessing:
[0077] 1. Data collection
[0078] Collect relevant datasets in the medical field, such as patient medical records, imaging reports, laboratory test results, etc.
[0079] Data examples:
[0080] Medical record text: Includes the patient's chief complaint, diagnosis, and treatment process.
[0081] Report text: Includes the doctor's diagnosis conclusion and suggestions.
[0082] 2. Data cleaning and formatting
[0083] Remove noise (such as typos, irrelevant characters).
[0084] Unify the data format, and organize the medical record - report pairs into an input - output format for the fine-tuning task.
[0085] 3. Divide the training set and the validation set
[0086] Divide the data into an 80% training set and a 20% validation set according to a certain proportion.
[0087] In this embodiment, using the obtained dataset to fine-tune a pre-trained large language model includes:
[0088] Step 2.1: Select the base model
[0089] Select a pre-trained large language model, such as Llama2-7b, opt-1.3b, etc.
[0090] Step 2.2: Set the fine-tuning parameters
[0091] Set common hyperparameters such as the learning rate, number of training epochs, batch size, etc. as needed. Other hyperparameters such as the update frequency of the parameter matrices U and V also need to be set.
[0092] Step 2.3: Fine-tuning process
[0093] First, load the pre-trained model and its corresponding weights. Then select a suitable loss function and optimizer. For large language models, the cross-entropy loss function is usually used. The optimizer can be sgd, which is a commonly used optimization algorithm for training neural network models. Next, use the selected dataset for fine-tuning training.
[0094] 1. Forward Propagation
[0095] Forward Propagation is the core process for inference or calculation in a neural network. Its purpose is to obtain the output value (such as the predicted value or classification result) by passing and calculating the input data layer by layer through the network. Input data: The training data divided from the dataset, and several pieces determined by the batch size are passed in each time. Take the X obtained in step 1 as the input data and input it into the large language model, that is, input it into the first layer of the neural network, and take the output result of this layer as the input of the next layer, and so on until the final layer is reached to obtain the final output And through the final output And the standard output Y corresponding to the current input data X, calculate the loss function value obtained from this forward propagation.
[0096] 2. For the first loop, or loops where the iteration number is an integer multiple of the set update number, it is necessary to update the U and V matrices corresponding to each parameter in the large model. The update method is to generate a random value, multiply the perturbation parameter given in the hyperparameters by the random value to obtain the full-scale perturbation value for the large model parameters. First, add this perturbation value to all parameters, calculate the loss function value, and then subtract twice this perturbation value on this basis to calculate the loss function value. Subtract the two random function values and divide by twice the perturbation parameter value to obtain the projected gradient. Then multiply this projected gradient by the parameter matrix (that is, each parameter in the large model, and these parameters are stored in the form of a matrix) to obtain the gradient matrix of the parameters. Perform SVD decomposition on the gradient matrix to obtain three matrices P, S, and Q. Only take the first R columns (determined by the hyperparameter gauss rank) in the second dimension of the P matrix as the U matrix of the parameter; take the first R rows in the first dimension of the Q matrix as the V matrix of the parameter, and store the U and V matrices. Repeat this operation for each parameter to obtain the U and V matrices of all parameters.
[0097] 3. For the case where there are already U and V matrices, or when there is no need to update the U and V matrices, perform the following calculations. For each parameter, generate a random matrix that conforms to the Gaussian distribution, with size R x R, denoted as y. Calculate the z matrix as the product of U, y, and V, which is the local perturbation of the parameter. Multiply the perturbed parameter by this local perturbation matrix to obtain the perturbation matrix. After calculating the perturbation matrices for all parameters, calculate the estimated gradient matrix. The specific process is as follows: First, add this perturbation value to all parameters and calculate the loss function value. Then, on this basis, subtract twice this perturbation value and calculate the loss function value. Subtract the two random function values and divide by twice the perturbed parameter value to obtain the projected gradient.
[0098] 4. Gradient update:
[0099] Update the parameters of the model so that the diagnosis generated by the adjusted model is closer to the true diagnosis. For each parameter, multiply the parameter matrix by the projected gradient to obtain the gradient matrix. Update the parameter to the current value minus the value of the learning rate (hyperparameter) multiplied by the gradient matrix.
[0100] 5. Repeat the above steps 1 - 4 for T (the set training times hyperparameter) times to obtain the fine-tuned model. During the process, the validation set can be used to calculate the current accuracy and record the model parameters with the highest accuracy.
[0101] 6. Use the test set to verify the trained model. Check the improvement in the accuracy of the diagnosis generated by the model.
[0102] Specifically, the fine-tuning strategy:
[0103] 1. Low-rank perturbation matrix design: For each layer's parameter matrix of the model Generate a low-rank perturbation matrix
[0104] Generate a random matrix: Randomly generate a small random matrix (r << min(m i , n i ))
[0105] QR decomposition
[0106] Perform svd decomposition on the gradient obtained in the 2 lazy updates to obtain a column-orthogonal matrix and
[0107] Combine matrices: Use the formula to generate a low-rank perturbation matrix.
[0108] 2. Lazy update strategy
[0109] Update U once every F training stepsi and V i , avoiding high overhead caused by frequent updates.
[0110] In each update, the method of mezo is used for gradient estimation. That is, a Z matrix of the original size is generated, two forward propagations are performed, the difference between the two losses is divided by the perturbation size to obtain the estimated gradient value. The gradient value is subjected to svd decomposition to obtain a column orthogonal matrix for each layer of parameters and
[0111] Training process:
[0112] 1. Gradient estimation
[0113] In each training, a small random matrix is randomly generated (r << min(m i , n i ))). Through the stored column orthogonal matrices and Use the formula to generate a low-rank perturbation matrix. Use this matrix for two forward propagations to calculate the matrix estimated value (similar to mezo).
[0114] Update the parameters through the gradient estimated value:
[0115] 2. Training process
[0116] Set the initial learning rate η 0 , and use stochastic gradient descent (SGD) for optimization.
[0117] Periodically evaluate the loss on the validation set and dynamically adjust the learning rate. Finally, a tuned model is obtained through training.
[0118] Recent theoretical studies have explored the use of low-dimensional perturbations in a random subspace to reduce gradient variance and thus improve the convergence rate. The key to the random subspace method lies in generating perturbation vectors located within the subspace P
[0119]
[0120] where P ∈ Rd×q is a random projection matrix whose elements are sampled from N(0, 1), z ∈ Rq is a low-dimensional random perturbation vector sampled from N(0, Iq), and q < d is the dimension of the subspace. Therefore, the gradient estimator in the subspace is defined as:
[0121]
[0122] Large language models (LLMs) are huge, and the parameters for their training and fine-tuning are usually high-dimensional. This results in an overly large size of matrix P, which is q times the dimension d of the model parameters in full-parameter fine-tuning and is also large in other fine-tuning schemes (such as LoRA). Therefore, this method significantly increases the memory requirement and computational complexity. Thus, it is particularly important to develop an efficient subspace construction strategy for LLM fine-tuning to minimize memory consumption.
[0123] Exploring through the update direction in a low-dimensional subspace can effectively reduce the computational and memory overhead. Assume that the parameters of the i-th layer of the model are represented in matrix form as Next, it will be explained how to design its low-rank perturbation matrix
[0124] The present invention proposes a low-rank perturbation strategy applicable to the model parameter matrix of each layer, which is different from the previous random subspace method for the entire model parameters. In each step, the present invention generates a low-dimensional random matrix where r << min(m i , n i ), then performs an SVD decomposition on the gradient matrix obtained by perturbing the original Z to create the projection matrices and Both of these matrices are column-orthogonal matrices. Experiments show that the matrices obtained in this way have better performance compared to Gaussian random projection matrices. Subsequently, the present invention combines these three matrices to generate a low-rank perturbation matrix:
[0125]
[0126] The present invention calculates the loss difference:
[0127]
[0128] Note that multiplying a scalar by a set means multiplying the scalar by each element in the set, and adding two sets means adding the corresponding elements. This is only a mathematical representation. In practice, ρ in the above formula can be calculated through two forward propagations of all layers. Then, the present invention obtains the gradient estimate of the i-th layer:
[0129]
[0130] It can be proven the proximity between the gradient estimate of the present invention and the original gradient calculated through backpropagation (BP) in the first-order (FO) method, and the gradient estimate of the present invention has improvements in variance and angular error compared to the gradient estimate of MeZO.
[0131] Subsequently, the gradient in any first-order optimizer (such as SGD) can be replaced with the gradient estimate in the formula:
[0132]
[0133] In SubZero, the present invention selects SGD as the default optimizer.
[0134] In this embodiment, for the cases to be predicted, the tuned model can be used to generate a diagnosis of the cases. After preprocessing the case description in a similar manner to step 1, data X is obtained. Inputting X into the model to obtain the output diagnosis Y, thus achieving the purpose of diagnosis through the large language model.
[0135] Embodiment 2
[0136] The present invention also provides a system for zero-order optimization and fine-tuning of a large model on a random subspace. The system is used to implement any of the methods described above. The system includes: an acquisition module, a fine-tuning module, and a task execution module;
[0137] The acquisition module is used to acquire task data in the target domain, preprocess the task data in the target domain, and obtain a data set in the format required by the large model;
[0138] The fine-tuning module is used to fine-tune the pre-trained large language model using the obtained data set;
[0139] The task execution module is used to input the data to be processed into the fine-tuned large language model to complete the tasks in the target domain.
[0140] In this embodiment, the fine-tuning module includes: a loss calculation unit, a parameter matrix update unit, a projection gradient calculation unit, a gradient update unit, and an iteration unit;
[0141] The loss calculation unit is used to input the obtained data set into the first layer of the large language model, and use the output result obtained from this layer as the input for the next layer, repeating this process until the final layer is reached to obtain the final output And through the final output Compare with the standard output Y corresponding to the current input data to calculate the loss function value obtained from the current forward propagation;
[0142] The parameter matrix update unit is used to update the U and V matrices corresponding to each parameter in the large language model for the first loop or for loops where the iteration number is an integer multiple of the set update number;
[0143] The projection gradient calculation unit is used to calculate the projection gradient for cases where there are already U and V matrices, or when the U and V matrices do not need to be updated;
[0144] The gradient update unit is used to multiply the parameter matrix by the projected gradient for each parameter to obtain a gradient matrix, and update the parameter to the current value minus the value of the learning rate multiplied by the gradient matrix;
[0145] The iteration unit is used to repeat the loss calculation unit to the gradient update unit T times to obtain a fine-tuned model.
[0146] In this embodiment, the parameter matrix update unit includes: a full perturbation value calculation subunit, a loss function calculation subunit, a projected gradient calculation subunit, a gradient matrix calculation subunit, a decomposition subunit, and an iteration subunit;
[0147] The full perturbation value calculation subunit is used to generate a random value, multiply the perturbation parameter by the random value through the perturbation parameter given in the hyperparameters to obtain the full perturbation value of the large language model parameters;
[0148] The loss function calculation subunit is used to add all the parameters with the full perturbation value, calculate the loss function value, and then subtract twice the full perturbation value to calculate the loss function value;
[0149] The projected gradient calculation subunit is used to subtract the two loss function values and then divide by twice the perturbation parameter value to obtain the projected gradient;
[0150] The gradient matrix calculation subunit is used to multiply the projected gradient by the parameter matrix to obtain the gradient matrix of the parameter;
[0151] The decomposition subunit is used to perform SVD decomposition on the gradient matrix to obtain three matrices P, S, and Q; only take the first R columns of the second dimension of the P matrix as the U matrix of the parameter; take the first R rows of the first dimension of the Q matrix as the V matrix of the parameter, and store the U and V matrices;
[0152] The iteration subunit is used to repeat this operation for each parameter to obtain the U and V matrices of all parameters.
[0153] In this embodiment, the projected gradient calculation unit includes: a random matrix calculation subunit, a local perturbation calculation subunit, a perturbation matrix calculation subunit, and an estimated projected gradient calculation subunit;
[0154] The random matrix calculation subunit is used to generate a random matrix that conforms to the Gaussian distribution for each parameter, with a size of R x R, denoted as y;
[0155] The local perturbation calculation subunit is used to calculate the z matrix as the product of UyV, that is, the local perturbation of the parameter;
[0156] The perturbation matrix calculation subunit is used to multiply the perturbation parameter by the local perturbation matrix to obtain the perturbation matrix;
[0157] After the estimated projection gradient calculation subunit is used to calculate the perturbation matrix of all parameters, the estimated projection gradient is calculated.
[0158] The embodiments described above are only descriptions of the preferred embodiments of the present invention, and do not limit the scope of the present invention. Without departing from the design spirit of the present invention, various deformations and improvements made by those of ordinary skill in the art to the technical solutions of the present invention shall fall within the protection scope determined by the claims of the present invention.
Claims
1. A method for fine-tuning a large model by zero-order optimization on a random subspace, characterized in that: The method comprises: Obtain the task data in the target field, preprocess the task data in the target field, and obtain the data set in the format required by the large model; Use the obtained dataset to fine-tune the pre-trained large language model; The data to be processed is input into the fine-tuned large language model to complete the task in the target field.
2. The method according to claim 1, characterized in that Using the obtained dataset to fine-tune the pre-trained large language model includes: Step 1: Input the obtained data set into the first layer of the large language model, and use the output of this layer as the input of the next layer, repeating the process until the final layer is reached to obtain the final output. And through the final output Compare it with the standard output Y corresponding to the input data this time, and calculate the loss function value obtained in this forward propagation; Step 2: For the first loop or a loop whose number of iterations is an integer multiple of the set update number, it is necessary to update the U and V matrices corresponding to each parameter in the large language model; Step 3: If the U and V matrices are already available or do not need to be updated, calculate the projected gradient; Step 4: For each parameter, multiply the parameter matrix by the projected gradient to obtain the gradient matrix, and update the parameter to the current value minus the learning rate multiplied by the gradient matrix; Step 5: Repeat steps 1 to 4 T times to obtain the fine-tuned model.
3. The method according to claim 2, characterized in that In the step 2, updating the U and V matrices corresponding to each parameter in the large language model includes: Generate a random value, multiply the perturbation parameter given in the hyperparameters by the random value, and obtain the full perturbation value of the large language model parameters; Add the full perturbation value to all parameters, calculate the loss function value, and then subtract twice the full perturbation value to calculate the loss function value; Subtract the two loss function values and divide them by twice the perturbation parameter value to obtain the projected gradient; Multiply the projected gradient by the parameter matrix to obtain the gradient matrix of the parameters; Decompose the gradient matrix by SVD to obtain three matrices P, S, and Q. Take only the first R columns of the second dimension of the P matrix as the U matrix of the parameter. Take the first R rows of the first dimension of the Q matrix as the V matrix of the parameter, and store the U and V matrices. Repeat this operation for each parameter to obtain the U and V matrices of all parameters.
4. The method according to claim 2, characterized in that In the step 3, calculating the projection gradient includes: For each parameter, generate a random matrix that conforms to the Gaussian distribution, with a size of R x R, denoted as y; Calculate the z matrix as the product of U y V, i.e. the local perturbation of the parameters; Multiply the disturbance parameter by the local disturbance matrix to obtain the disturbance matrix; After calculating the perturbation matrices for all parameters, the estimated projected gradients are calculated.
5. A system for performing zero-order optimization fine-tuning on a large model in a random subspace, the system being used to implement the method of any one of claims 1 to 4, characterized in that: The system comprises: an acquisition module, a fine-tuning module and a task execution module; The acquisition module is used to acquire the task data of the target field, pre-process the task data of the target field, and obtain a data set in the format required by the large model; The fine-tuning module is used to fine-tune the pre-trained large language model using the obtained data set; The task execution module is used to input the data to be processed into the fine-tuned large language model to complete the task in the target field.
6. The system according to claim 5, characterized in that The fine-tuning module includes: a loss calculation unit, a parameter matrix update unit, a projection gradient calculation unit, a gradient update unit and an iteration unit; The loss calculation unit is used to input the obtained data set into the first layer of the large language model, and use the output result of this layer as the input of the next layer, and repeat the cycle until the final layer is reached to obtain the final output And through the final output Compare it with the standard output Y corresponding to the input data this time, and calculate the loss function value obtained in this forward propagation; The parameter matrix updating unit is used to update the U and V matrices corresponding to each parameter in the large language model for the first cycle or the cycle whose number of iterations is an integer multiple of the set update number; The projection gradient calculation unit is used to calculate the projection gradient when the U and V matrices are already available or when the U and V matrices do not need to be updated; The gradient updating unit is used for multiplying the parameter matrix by the projected gradient for each parameter to obtain a gradient matrix, and updating the parameter to a value of the current value minus the learning rate multiplied by the gradient matrix; The iteration unit is used to repeat the loss calculation unit to the gradient update unit T times to obtain a fine-tuned model.
7. The system according to claim 6, characterized in that The parameter matrix updating unit includes: a total disturbance value calculation subunit, a loss function calculation subunit, a projection gradient calculation subunit, a gradient matrix calculation subunit, a decomposition subunit and an iteration subunit; The full disturbance value calculation subunit is used to generate a random value, and multiply the disturbance parameter and the random value by the disturbance parameter given in the hyperparameter to obtain the full disturbance value of the large language model parameter; The loss function calculation subunit is used to add the full disturbance value to all parameters, calculate the loss function value, and then subtract twice the full disturbance value to calculate the loss function value; The projected gradient calculation subunit is used to subtract the two loss function values and then divide them by twice the perturbation parameter value to obtain the projected gradient; The gradient matrix calculation subunit is used to multiply the projection gradient with the parameter matrix to obtain the gradient matrix of the parameters; The decomposition subunit is used to perform SVD decomposition on the gradient matrix to obtain three matrices P, S, and Q; only the first R columns of the second dimension of the P matrix are taken as the U matrix of the parameter; the first R rows of the first dimension of the Q matrix are taken as the V matrix of the parameter, and the U and V matrices are stored; The iterative subunit is used to repeat the operation for each parameter to obtain the U and V matrices of all parameters.
8. The system according to claim 6, characterized in that The projected gradient calculation unit includes: a random matrix calculation subunit, a local perturbation calculation subunit, a perturbation matrix calculation subunit and an estimated projected gradient calculation subunit; The random matrix calculation subunit is used to generate a random matrix that conforms to Gaussian distribution for each parameter, with a size of R x R, denoted as y; The local perturbation calculation subunit is used to calculate the z matrix as the product of U y V, that is, the local perturbation of the parameters; The disturbance matrix calculation subunit is used to multiply the disturbance parameter by the local disturbance matrix to obtain the disturbance matrix; The estimated projected gradient calculation subunit is used to calculate the estimated projected gradient after calculating the disturbance matrix of all parameters.
Citation Information
Cited By
Large language model zero-order fine tuning method and system based on low-rank projection matrix learning
CN122045829A