Method and device for inhibiting large model vertical domain fine tuning overfitting and storage medium

By performing random mask processing and multiple mask sampling integration of the low-rank parameter matrix of the LoRA model, the overfitting problem in the fine-tuning of the vertical domain of the large model is solved, and the generalization performance and adaptation efficiency of the model are improved.

CN120258085APending Publication Date: 2025-07-04PEKING UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510155047.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-12
Publication Date
2025-07-04

Smart Images

  • Figure CN120258085A_ABST
    Figure CN120258085A_ABST
Patent Text Reader

Abstract

The invention discloses a method and device for inhibiting large model vertical domain fine tuning overfitting and a storage medium, and belongs to the technical field of large model fine tuning and deep learning optimization. In order to solve the problem of overfitting possibly caused when LoRA is used for efficient fine tuning of parameters, a low-rank matrix decomposition technology introducing random masks is mainly adopted, and model integration is carried out in combination with multiple times of mask sampling. Through the method, the generalization ability of the model can be effectively improved, overfitting is prevented, and the expression ability of the model is kept in a downstream task even under the condition that the data size is small. Compared with a traditional method, the method has the advantages of being easy to implement, efficient and good in generalization performance.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of large model fine-tuning and deep learning optimization, and specifically relates to a method, device and storage medium for suppressing overfitting in the vertical domain fine-tuning of large models. Background Art

[0002] In order to quickly and efficiently adapt large language models to downstream tasks, researchers have proposed a new fine-tuning paradigm represented by the Low-Rank Adaptation (LoRA) method, called Parameter-Efficient Fine-Tuning (PEFT). However, in the process of fine-tuning the model for downstream tasks using the LoRA method, how to select the rank of the matrix in the LoRA method is a difficult problem. Too small a rank will result in a small number of model parameters, limiting the expressive ability of the model, while too large a rank will increase the degree of freedom of the large model and enhance the risk of overfitting of the model on the fine-tuning data.

[0003] The goal of the parameter-efficient fine-tuning paradigm represented by low-rank adaptation methods is to quickly and efficiently adapt large language models to downstream tasks. The LoRA method (see Edward J Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu, Yuanzhi Li, Shean Wang, Lu Wang, and Weizhu Chen. LoRA: Low-rank adaptation of large language models. arXiv preprint arXiv:2106.09685, 2021) decomposes the incremental parameter matrix into the product of two low-rank matrices. The rank of this decomposition is crucial for the LoRA method. Too small a rank may lead to insufficient expressive power, while too large a rank may lead to overfitting. A variant of the LoRA method, AdaLoRA (see Qingru Zhang, Minshuo Chen, Alexander Bukharin, Pengcheng He, Yu Cheng, Weizhu Chen, and Tuo Zhao. Adaptive budget allocation for parameter-efficient fine-tuning. In The Eleventh International Conference on Learning Representations, 2023.) proposes to decompose the incremental weights using a method similar to singular value decomposition (SVD decomposition) and select parameters through importance scoring. However, this selection method also depends on the gradients on the training set, increasing the risk of overfitting.

[0004] The Dropout mechanism (see Geoffrey E Hinton, Nitish Srivastava, Alex Krizhevsky, Ilya Sutskever, and Ruslan R Salakhutdinov. Improving neural networks by preventing co-adaptation of feature detectors. arXiv preprint arXiv:1207.0580, 2012.) is a commonly used technique in deep neural networks to mitigate overfitting. In the standard Dropout method, each neuron in the network has a certain probability of being ignored during training. For specific model architectures, different Dropout techniques have been proposed and introduced, such as Spatial Dropout for convolutional layers (see Jonathan Tompson, Ross Goroshin, Arjun Jain, Yann LeCun, and Christoph Bregler. Efficient object localization using convolutional networks. In Proceedings of the IEEE conference on computer vision and pattern recognition, pages 648–656, 2015.) and Recurrent Dropout for recurrent neural networks (see Stanislau Semeniuta, Aliaksei Severyn, and Erhardt Barth. Recurrent dropout without memory loss. In Proceedings of COLING 2016, the 26th International Conference on Computational Linguistics: Technical Papers, pages 1757–1766, 2016.).

[0005] Currently, in the parameter-efficient fine-tuning paradigm based on LoRA, there is no method to consider reducing the impact of overfitting from both practical and theoretical perspectives. At the same time, there is no Dropout method designed specifically for the LoRA model architecture in the field of Dropout research. The absence of such methods leads to overfitting problems in large models during the adaptation process of downstream tasks, easily due to the relatively small amount of downstream training data, resulting in poor performance and generalization ability on downstream tasks. Summary of the Invention

[0006] The objective of the present invention is to propose a technical solution for suppressing overfitting in the vertical domain fine-tuning of large models. By introducing the Dropout mechanism into the LoRA fine-tuning framework, it aims to reduce the overfitting problem caused by a small amount of training data in downstream tasks, thereby improving the generalization performance and adaptation efficiency of the model in downstream tasks.

[0007] To achieve the above objective, the present invention adopts the following technical solutions:

[0008] A method for suppressing overfitting in the vertical domain fine-tuning of large models, comprising the following steps:

[0009] 1) For each linear layer of the large model, use the LoRA method to construct an incremental parameter matrix through the product of two low-rank parameter matrices A and B;

[0010] 2) Randomly mask the two low-rank parameter matrices A and B according to a preset probability to generate a sparse incremental parameter matrix, obtaining a constructed model;

[0011] 3) Use the training data in the specified domain to fine-tune and train the constructed model;

[0012] 4) In the usage stage, perform multiple sampling masks on the fine-tuned model to obtain multiple groups of different parameters, and integrate the outputs of the fine-tuned model under different parameters.

[0013] Further, in step 1), the large model includes a large language model.

[0014] Further, in step 2), sample mask vectors from the Bernoulli distribution according to a preset probability to mask the two low-rank parameter matrices A and B, and the expression is as follows:

[0015]

[0016] where diag is a function that converts a vector into a diagonal matrix, Bern(1 - p) represents sampling from the Bernoulli distribution according to a preset probability p, m A , m B are mask vectors sampled from the Bernoulli distribution, and is the masked low-rank parameter matrix.

[0017] Further, in step 2), the sparse incremental parameter matrix

[0018] Further, in step 3), the objective function of the fine-tuning training is:

[0019]

[0020] Among them, N is the number of sampling times, is the loss function, x is the input sample, and θ 0 is the pre-trained parameter, and m k is the set of mask vectors for the k-th sampling, and Δθ(m k ) are all trainable parameters after masking. Bern(1 - p) represents sampling from the Bernoulli distribution according to the preset probability p.

[0021] Furthermore, the formula for integrating the output of the fine-tuning model under different parameters in step 4) is:

[0022]

[0023] Among them, o k (x) is the output of the input sample x under the k-th sampling, is the fine-tuning model, are the fine-tuning models under different parameters.

[0024] An apparatus for suppressing overfitting in the vertical domain fine-tuning of a large model includes a memory and a processor. The memory stores a computer program, and when the processor executes the computer program, the steps of the above method are implemented.

[0025] A computer-readable storage medium stores a computer program, and when the computer program is executed, the steps of the above method are implemented.

[0026] The present invention has achieved the following beneficial effects:

[0027] 1. The present invention effectively suppresses the overfitting risk that may occur when using LoRA for parameter-efficient fine-tuning in the vertical domain task, enabling good generalization performance even in the case of scarce training data and avoiding the model from only memorizing the training data.

[0028] 2. The present invention reasonably controls the balance between overfitting and underfitting by introducing random masks in the LoRA low-rank matrix, thereby effectively preventing overfitting in downstream tasks with less data volume and improving the generalization ability of the model.

[0029] 3. The present invention further enhances the model performance through the integration method during testing, and compared with competing products, has the advantages of simple implementation, high performance, and excellent generalization ability. Brief Description of the Drawings

[0030] Figure 1 is a schematic diagram of the present invention for suppressing overfitting in the vertical domain fine-tuning of a large model. Detailed Embodiments

[0031] To make the technical features, advantages or technical effects in the above technical solution of the present invention more obvious and understandable, the following will be described in detail with reference to the accompanying drawings.

[0032] I. Embodiment of the method of the present invention:

[0033] This embodiment specifically proposes a method for suppressing overfitting in the vertical domain fine-tuning of large models. As Figure 1 shown, this is a Dropout scheme based on the low-rank parameter matrix decomposition of LoRA. The rows / columns of the two low-rank matrices in the existing LoRA method are randomly masked with a given probability p, and the original incremental parameter matrix of LoRA is transformed into a sparse parameter matrix to control the overfitting phenomenon that may occur during the fine-tuning process. This method is mainly applied to the scenario of vertical domain fine-tuning of large models, that is, based on training data in a specific domain (such as medical data, financial data, etc.), the large model in the general domain is trained into a vertical domain model corresponding to the domain. The specific processing process of this method is as follows:

[0034] 1) For a given pre-trained large model, for each of its linear layers, using the LoRA method, an incremental parameter matrix ΔW is constructed through the product of two low-rank parameter matrices A and B, which is expressed as follows:

[0035] ΔW = BA

[0036] 2) According to the given probability p, sample the mask vectors m A and m B from the Bernoulli distribution, and mask the two low-rank parameter matrices A and B respectively to obtain the masked low-rank parameter matrices and which are expressed as follows:

[0037]

[0038] Generate the sparse incremental parameter matrix and further obtain the constructed model that applies this sparse incremental parameter matrix.

[0039] 3) Use the training data in a specific domain to fine-tune the transformed model, and optimize the incremental parameter matrix based on the training loss. Denote the set of all masks in the entire transformed model as m, then all the trainable parameters after masking are Δθ(m). Then, for the training objective of the input sample x is:

[0040]

[0041] where N is the number of sampling times, is the loss function, and θ 0is a pre-training parameter, m k is the set of mask vectors for the k-th sampling.

[0042] 4) In the testing (or using) phase, multiple sets of different parameters are obtained by sampling masks multiple times, and the fine-tuned models supported by different parameters are integrated to improve the model's ability to capture information from different aspects. Denote the different fine-tuned models obtained based on the fine-tuned model and the mask m as Then the calculation formula for the output o(x) of the integrated model during testing is:

[0043]

[0044] where, o k (x) is the output of the input sample x under the k-th sampling.

[0045] II. Theoretical analysis of the upper bound of the generalization error of the large model fine-tuning of the method of the present invention:

[0046] This part analyzes the large model fine-tuning process based on the method of the present invention and proposes the corresponding upper bound of the generalization error. This upper bound of the generalization error theoretically proves the effectiveness of the method of the present invention proposed above in suppressing overfitting, which is the theoretical guarantee for the method to play a role in practice. This theory analyzes that the LoRA fine-tuning training method combined with the method of the present invention can be transformed into an optimization problem with sparsity regularization, and further gives the upper bound of the generalization error under this sparsity regularization framework. Based on this upper bound, it is confirmed that the method of the present invention balances the empirical risk minimization training objective and the complexity of the adaptation function by introducing appropriate sparsity into the LoRA tunable parameters during the fine-tuning process, thereby helping to narrow the gap between the empirical risk and the generalization risk and reducing the overfitting of the training data. The theoretical analysis is as follows:

[0047] Given a large model which contains parameters θ 0 , loss function and LoRA parameters Δθ, its optimization objective can be transformed into an optimization problem with sparsity regularization, that is

[0048]

[0049] where θ 0 is the pre-training parameter, Δθ is the incremental parameter to be trained, λ is the regularization strength, p is the Dropout probability, and d is the mask vector sampled from the Bernoulli distribution based on the probability p.

[0050] Define the entire training set as S, and define the remaining data set after removing a certain sample from the data set as S i = S - {x i}, the optimal parameters obtained by optimizing based on the training set S and the loss function are denoted as Based on the above analysis and representation, an upper bound on the generalization error of large model fine-tuning based on the method of the present invention is further proposed.

[0051] Given a large model , Dropout probability p, and regularization strength λ, for a dataset S of size n, let S i denote the dataset obtained by removing the i-th sample from the training set S, and be the optimal parameters that can be obtained based on the optimization objective on S and S i . If the loss function is η-Lipschitz continuous, and is close enough to , its Hessian matrix is positive semi-definite and its eigenvalue decomposition is Udiag(Λ)U -1 , Λ = {Λ1,..., Λ m}, where Λ min is the smallest eigenvalue, then for some constant C and a given δ, the following inequality holds with probability 1 - δ:

[0052]

[0053] where, and are the generalization error and the empirical error respectively. It can be found that a suitable Dropout probability p can effectively compress the gap between the generalization error and the empirical error, reducing the risk of overfitting during model training.

[0054] In addition, the effectiveness of test-time ensembling is further proven theoretically. If the loss function is convex with respect to the output of the last activation layer of the model , then there is:

[0055]

[0056] where x and y are the input sample and the corresponding label respectively, is the distribution formed by the parameters obtained by sampling with different masks. This inequality shows that the test-time ensembling method can further compress the upper bound of the model generalization error, that is, it can further improve the model performance.

[0057] III. Performance Test of the Method of the Present Invention:

[0058] The method of the present invention (LoRADropout) is experimentally tested on multiple different benchmark datasets (GLUE, MMLU, Vicuna), and the test results are shown in Table 1 and Table 2. The test results indicate that the method of the present invention can improve the accuracy of the traditional LoRA method in multiple downstream tasks and alleviate the overfitting phenomenon on the fine-tuning dataset.

[0059] Table 1 Test Results Based on the GLUE Benchmark Dataset and the DeBERTa-V3 Model

[0060]

[0061] Table 2 Test Results Based on the MMLU and Vicuna Benchmark Datasets and the LLaMA2-7B Model

[0062]

[0063] Although the present invention has been disclosed above by way of examples, it is not intended to limit the present invention. Any appropriate modifications or equivalent replacements made by those of ordinary skill in the art to the technical solutions of the present invention shall be covered within the protection scope of the present invention. The protection scope of the present invention shall be defined by the claims.

Claims

1. A method for suppressing overfitting in the vertical domain fine-tuning of large models, characterized in that, Including the following steps: 1) For each linear layer of the large model, use the LoRA method to construct an incremental parameter matrix through the product of two low-rank parameter matrices A and B; 2) Randomly mask the two low-rank parameter matrices A and B according to a preset probability to generate a sparse incremental parameter matrix, obtaining a constructed model; 3) Use the training data in the specified domain to fine-tune and train the constructed model; 4) In the usage stage, perform multiple sampling masks on the fine-tuned model to obtain multiple sets of different parameters, and integrate the outputs of the fine-tuned model under different parameters.

2. The method according to claim 1, wherein The large model in step 1) includes a large language model.

3. The method according to claim 1, characterized in that, In step 2), sample a mask vector from the Bernoulli distribution according to a preset probability to mask the two low-rank parameter matrices A and B, and the expression is as follows: where diag is a function that converts a vector into a diagonal matrix, Bern(1 - p) represents sampling from a Bernoulli distribution according to a preset probability p, and m A , m B is a mask vector sampled from a Bernoulli distribution, and is the masked low-rank parameter matrix.

4. The method according to claim 3, characterized in that The sparse incremental parameter matrix in step 2) 5. The method according to claim 1, wherein The objective function for fine-tuning training in step 3) is as follows: where N is the number of samplings, is the loss function, x is the input sample, and θ 0 are the pre-trained parameters, m k is the set of mask vectors for the k-th sampling, and Δθ(m k ) are all the trainable parameters after masking. Bern(1 - p) represents sampling from the Bernoulli distribution according to the preset probability p.

6. The method according to claim 5, characterized in that The formula for integrating the outputs of the fine-tuned model under different parameters in step 4) is: Among them, o k (x) is the output of the input sample x under the k-th sampling, is the fine-tuning model, is the fine-tuning model under different parameters.

7. An apparatus for suppressing overfitting in fine-tuning of a large model in a vertical domain, characterized in that, Including a memory and a processor, the memory stores a computer program, and when the processor executes the computer program, it implements the steps of the method described in any one of claims 1-6.

8. A computer-readable storage medium, characterized in that, A computer program is stored, and when the computer program is executed, it implements the steps of the method described in any one of claims 1-6.