Vertical domain large model orthogonal fine tuning method based on Givens rotation
By using the Givens rotation vertical large-scale model orthogonal fine-tuning method in large-scale pre-trained models, the problems of inefficient parameter efficiency and difficulty in semantic correlation retention during fine-tuning are solved, efficient parameter fine-tuning and semantic correlation retention are achieved, and the adaptability and general performance of the model are improved.
Patent Information
- Application Number
- CN202510150627.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-11
- Publication Date
- 2025-06-24
AI Technical Summary
Existing large-scale pretrained models face inefficiency in parameter when fine-tuning, and it is difficult to effectively retain common knowledge and concepts in the pretrained model, resulting in overfitting or catastrophic forgetting.
The vertical domain large model orthogonal fine-tuning method (GivensOFT) based on Givens rotation is used to explicitly retain the semantic association relationship in the pre-training stage by multiplying the learned orthogonal matrix on the linear layer of the model, and the parameter complexity is reduced by using Givens rotation.
The semantic association of the pretrained model is effectively retained, parameter efficiency and downstream adaptation flexibility are improved, overfitting and catastrophic forgetting are avoided, and the model's general semantic understanding and logical reasoning capabilities are improved.
Smart Images

Figure CN120197724A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of artificial intelligence, and relates to an orthogonal fine-tuning method for a large model in the vertical domain based on Givens rotation. Background Art
[0002] In recent years, large-scale pre-trained models have shown increasingly powerful performance in tasks such as natural language and vision. Due to the explosive growth of the number of parameters in large models, the current pre-trained large models face the problem of low parameter efficiency during fine-tuning. At the same time, during the fine-tuning process, it is necessary to retain the general knowledge and concepts contained in the pre-trained large model to avoid overfitting or catastrophic forgetting phenomena after fine-tuning, which affect the general semantic understanding and logical reasoning ability of the model.
[0003] With the increasing popularity of large-scale pre-trained models, researchers have begun to pay more and more attention to developing more efficient parameter fine-tuning methods to adapt to downstream tasks. Compared with the entire set of parameters that need to be fine-tuned, Parameter-Efficient Fine-Tuning (PEFT) develops parameter-light adapters for different downstream tasks, thus greatly reducing the training and storage costs of the model. There are three mainstream methods of PEFT: The first is prompt tuning such as P-Tuning V2 (see Liu X, Ji K, Fu Y, et al. P-Tuning: Prompt Tuning Can Be Comparable to Fine-tuning Across Scales and Tasks [C] / / Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 2: Short Papers). 2022: 61-68), Prefix Tuning (see Li X L, Liang P. Prefix-Tuning: Optimizing Continuous Prompts for Generation [C] / / Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2021: 4582-4597.), etc. These methods add additional learnable prefix tokens to the input of the Transformer layer to achieve downstream adaptation; Adapter tuning inserts additional trainable multi-layer perceptron modules after some sub-layers in the original Transformer block for fine-tuning, typical methods such as H-Adapter, P-Adapter, etc. (see Houlsby N, Giurgiu A, Jastrzebski S, et al. Parameter-efficient transfer learning for NLP [C] / / International conference on machine learning.PMLR, 2019: 2790 - 2799; Pfeiffer J, Kamath A, Rücklé A, et al. AdapterFusion: Non-Destructive Task Composition for Transfer Learning[C] / / Proceedings of the 16th Conference of the European Chapter of the Association for Computational Linguistics: Main Volume. 2021: 487 - 503.); and reparameterization tuning such as LoRA, OFT, etc., whose model architecture remains unchanged, only reparameterizes and fine-tunes the update amounts of some model parameters with low parameter overhead, and in the inference stage, the update amounts can be directly incorporated into the original model parameters to achieve zero additional overhead inference (refer to Hu E J, Shen Y, Wallis P, et al. Lora: Low-rank adaptation of large language models[J]. arXiv preprint arXiv:2106.09685, 2021; Qiu Z, Liu W, Feng H, et al. Controlling text-to-image diffusion by orthogonal finetuning[J]. Advances in Neural Information Processing Systems, 2023, 36: 79320 - 79362). He et al. and Mao et al. also tried to integrate these three types of methods into a unified framework and implement an adaptive fine-tuning method selection according to different scenarios (refer to He J, Zhou C, Ma X, et al. Towards a Unified View of Parameter-Efficient Transfer Learning[C] / / International Conference on Learning Representations. 2021; Mao Y, Mathias L, Hou R, et al.UniPELT: A Unified Framework for Parameter-Efficient Language Model Tuning[C] / / Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2022: 6253-6264). There are also some studies that attempt to integrate these three methods into a unified framework. Among these PEFT methods, the reparameterization fine-tuning method has the highest relevance to the present invention.
[0004] Most existing methods lack a mechanism to explicitly control the invariance of the semantic associations in the model during the fine-tuning phase, thus failing to effectively retain the general knowledge and concepts contained in the pre-trained large model, which may cause overfitting and catastrophic forgetting problems of the model on a small corpus. Although the OFT method achieves strict isometry of the hidden layer semantics in terms of angles by using orthogonal fine-tuning to ensure the invariance of semantic associations, its parameter complexity is high and it lacks the flexibility to learn small semantic association offsets in the downstream task corpus, and it still cannot be well adapted in various fine-tuning task scenarios. Summary of the Invention
[0005] Aiming at the problems existing in the prior art, the purpose of the present invention is to provide an orthogonal fine-tuning method for vertical domain large models based on Givens rotation (Givens OFT). The present invention can effectively use orthogonal fine-tuning to ensure the invariance of semantic associations before and after fine-tuning, and at the same time has the advantages of high parameter efficiency and high flexibility in downstream adaptation, and can be well adapted in various vertical domain fine-tuning task scenarios.
[0006] The present invention explicitly retains the semantic association relationship encoded by the model in the pre-training stage by left-multiplying a learnable orthogonal matrix to the linear layer of the model, thereby retaining the general knowledge and capabilities obtained through pre-training. In addition, Givens OFT uses Givens rotation to reduce the parameter complexity of the learnable quasi-orthogonal transformation from the square level of the hidden layer dimension to the linear level, improving the efficiency of the pre-trained large model in adapting to various NLP and vision downstream fields or tasks. This technology is mainly applied to scenarios that require fine-tuning of large-scale pre-trained models, such as vertical domain knowledge injection of large language models, etc.
[0007] The technical solution of the present invention is as follows:
[0008] An orthogonal fine-tuning method for vertical domain large models based on Givens rotation, the steps of which include:
[0009] 1) Select a general base large model M and a corpus dataset D of target vertical domain knowledge;
[0010] 2) Use the corpus dataset D to train the general base large model M, and then use the Givens rotation method to fine-tune all the linear layers in the trained general base large model M to inject the target vertical domain knowledge into the general base large model M, obtaining a target vertical domain large model M';
[0011] Among them, the method of using the Givens rotation method to fine-tune all the linear layers in the general base large model M is as follows: construct a d×d orthogonal matrix for Givens rotation for the current linear layer W to be fine-tuned. When performing the first rotation P1 on the linear layer W, simultaneously rotate all (2k, 2k + 1)-planes in the linear space where the linear layer W is located, where k is all integers between 0 and (d / 2)-1; when performing the second rotation P2 on the linear layer W, simultaneously rotate all the unrotated planes in all (4k, 4k + 2)-planes in the linear space where the linear layer W is located, where k is all integers between 0 and (d / 4)-1; when performing the third rotation P3 on the linear layer W, simultaneously rotate all the unrotated planes in all (8k, 8k + 4)-planes in the linear space where the linear layer W is located, where k is all integers between 0 and (d / 8)-1; and so on. When performing the nth rotation P n on the linear layer W, simultaneously rotate all the unrotated planes in all (2 n k, 2 n k + 2 n -1 )-planes in the linear space where the linear layer W is located.
[0012] Furthermore, change each Givens rotation to a quasi-Givens rotation transformation Among them, (α i , β i ) consists of four learnable parameters; then calculate the inner product of the column vectors of each quasi-Givens rotation transformation respectively, and design a regularization term and add it to the model loss function in the fine-tuning stage; among them, θ i is the rotation angle to be fine-tuned in the Givens rotation, α i =(α 1i , α 2i ), β i =(β 1i , β 2i ), α 1i , α 2i , β 1i, β 2i are learnable parameters in the linear transformation.
[0013] Furthermore, the target vertical domain is the medical field, and the corpus dataset D includes medical literature and medical visit records.
[0014] Furthermore, the target vertical domain is the legal field, and the corpus dataset D includes laws and regulations, and legal case precedents.
[0015] The present invention also provides a consultation service method, and its steps include:
[0016] 1) Select a general base large model M and a corpus dataset D of target vertical domain knowledge;
[0017] 2) Use the corpus dataset D to train the general base large model M, and then use the Givens rotation method to fine-tune all the linear layers in the trained general base large model M to inject the target vertical domain knowledge into the general base large model M to obtain a target vertical domain large model M';
[0018] 3) Input the consultation of the target vertical domain into the target vertical domain large model M' to obtain a corresponding consultation result.
[0019] The advantages of the present invention are as follows:
[0020] This invention greatly improves the retention ability of the pre-trained knowledge of the current pre-trained large model fine-tuning algorithm, so that it will not lose the general performance of the large model due to overfitting or catastrophic forgetting in various downstream tasks, such as logical reasoning ability, semantic understanding ability, etc. In addition, compared with traditional orthogonal fine-tuning, the present invention greatly reduces the parameter overhead of the method, reducing it from O(d 2 ) to O(d), and improves the semantic deviation adaptation ability of the method, making it easier to adapt to the vertical domain fine-tuning scenario of large models in the real world. Description of the Drawings
[0021] Figure 1 is the method flow chart of the present invention. Detailed Embodiments
[0022] The present invention will be further described in detail below with reference to the drawings. The examples given are only used to explain the present invention and are not used to limit the scope of the present invention.
[0023] The traditional orthogonal fine-tuning algorithm multiplies all the linear layers W in the Transformer block of the large model M on the left by a learnable orthogonal matrix R to achieve fine-tuning. For the constraint of orthogonality, the Cayley reparameterization method needs to be used to characterize the learnable orthogonal matrix, and its parameter complexity is O(d2 ), where d is the dimension of the hidden layer. This brings a huge parameter overhead, increasing both the training cost and scalability of the model. Therefore, this method designs to reduce the parameter complexity of characterizing the orthogonal matrix, thereby improving the parameter efficiency of the algorithm.
[0024] First, we introduce the mathematical tool of Givens rotation. As shown in the following matrix, formally, a Givens rotation is a d×d orthogonal matrix, which, based on the identity matrix, at the intersection of the horizontal and vertical coordinates of the i-th and j-th dimensions (i.e., the element G in the i-th row and j-th column of the matrix) ij ) are the cosine and sine values of a certain angle θ respectively, and its geometric meaning is to rotate the plane formed by the i-th and j-th dimensions of the d-dimensional linear space by an angle θ. We denote it as G(i,j;θ).
[0025]
[0026] Theoretically, we have proven that for any orthogonal transformation, we can use d - 1 Givens rotations to rotate any vector in the d-dimensional space to any position on the same sphere as it, that is, we can achieve the effect equivalent to any rotation orthogonal transformation. Specifically, the rotation method is to sequentially rotate the (d - 2, d - 1)-plane, (d - 3, d - 2)-plane,..., (0, 1)-plane of the linear space where the linear layer W to be fine-tuned is located (note: the (a, b)-plane refers to the plane spanned by the a-th and b-th dimensions of the linear space). Mathematically, we characterize the d-dimensional rotation orthogonal matrix R to be learned in the following way:
[0027]
[0028] where each Givens rotation has only one parameter θ that can be learned to fine-tune all linear layers in the large model M r , so that the overall parameter complexity of the method is optimized to O(d).
[0029] However, although this approach significantly reduces the parameter complexity at present, it requires extremely high computational complexity. It needs to calculate the product of d-dimensional sparse matrices d - 1 times in a single fine-tuning, which will greatly increase the computational time overhead. Therefore, we propose a parallel rotation strategy. In each rotation, the uncoupled dimensions are rotated simultaneously. Specifically, when performing the first rotation P1, all (2k, 2k + 1)-planes in the linear space where the linear layer W to be fine-tuned is located are rotated simultaneously, where k is all integers between 0 and (d / 2) - 1; when performing the second rotation P2, all (4k, 4k + 2)-planes that have not been considered are rotated simultaneously, where k is all integers between 0 and (d / 4) - 1; the third time, all (8k, 8k + 4)-planes are rotated simultaneously, where k is all integers between 0 and (d / 8) - 1. And so on. Mathematically, each parallel rotation can be denoted as:
[0030]
[0031] The overall parallel rotation strategy can be denoted as Finally, for each linear layer W to be adjusted in model M, its original forward process is: h = W T x, where x is the d-dimensional input received by the linear layer, h is the output of this linear layer, and r is a subscript used to mark the product matrix. GivensOFT uses the following formula for fine-tuning, and its forward process is changed to:
[0032]
[0033] In this way, the number of matrix multiplications in our forward process is optimized from d - 1 times to log d times, thus significantly reducing the computational overhead of the model. In addition, due to the special properties of the rotation transformation, we can also optimize the matrix multiplication to matrix Hadamard multiplication for calculation, and its computational complexity can be optimized from O(d 3 ) to O(d 2 ), greatly improving the computational efficiency.
[0034] In addition, since strict orthogonal fine-tuning cannot flexibly adapt to the minor semantic shifts between the potential in the downstream corpus and the pre-training stage, we propose a method of soft orthogonal regularization, which no longer requires strict orthogonality, but uses the way of orthogonal regularization to achieve more flexible adaptation of the downstream corpus and tasks while ensuring orthogonality as much as possible. Specifically, based on the above fine-tuning algorithm, each adjustable Givens rotation G i is changed to a quasi-Givens rotation transformation That is:
[0035]
[0036] Among them, (α i , β i ) consists of four learnable parameters. We calculate the inner product of the column vectors of all quasi-Givens transformations respectively, and design a regularization term and add it to the model loss function in the fine-tuning stage. This regularization term can reduce the inner product of the column vectors to encourage, as much as possible, orthogonality. On the premise of ensuring orthogonality as much as possible, the model is allowed to flexibly adapt to small semantic shifts in downstream tasks, thereby improving the flexibility of the model to adapt to downstream tasks. θ i is the fine-tuning parameter required in Givens rotation, and its meaning is the rotation angle of the plane rotation. Now, all non-0 / 1 positions (i.e., cos / sin theta) in Givens rotation are replaced with freely learnable parameters α 1i , α 2i , β 1i , β 2i , and denote α i =(α 1i , α 2i ), β i =(β 1i , β 2i ). In this way, α 1i , α 2i , β 1i , β 2i are learnable parameters in the linear transformation.
[0037] This technology is generally applicable to the research and development of any vertical domain large model. The technical process includes the following steps: First, prepare a general base large model M for fine-tuning, and a corpus dataset D of knowledge in the corresponding vertical domain (such as: medical, legal, etc.), which includes corresponding literature, actual case records, etc. (such as: medical literature, medical records, etc. used in the medical vertical domain; laws, regulations, case precedents, etc. used in the legal vertical domain); Second, use the dataset D to fine-tune all linear layers in the large model M using the GivensOFT method to achieve efficient injection of vertical domain knowledge, while ensuring that the basic capabilities of the original model are maintained as much as possible, so as to obtain the fine-tuned vertical domain large model M'. For the fine-tuned vertical domain large model, it has obtained a more powerful vertical domain knowledge reserve, has the ability to answer professional questions related to the vertical domain compared with the base large model, and at the same time does not lose the powerful text understanding ability and logical reasoning ability of the base large model, and can better serve the corresponding vertical domain tasks, such as intelligent medical consultation in the medical vertical domain, intelligent legal consultation in the legal vertical domain, etc.
[0038] Although specific embodiments of the present invention are disclosed for illustrative purposes, which are intended to assist in understanding the content of the present invention and implementing it accordingly, those skilled in the art can understand that various substitutions, changes, and modifications are possible without departing from the spirit and scope of the present invention and the appended claims. Therefore, the present invention should not be limited to the content disclosed in the preferred embodiments, and the scope of protection claimed by the present invention shall be defined by the scope defined in the claims.
Claims
1. A vertical large model orthogonal fine-tuning method based on Givens rotation, the steps comprising: 1) Select a general base model M and a corpus dataset D of target vertical domain knowledge; 2) Using the corpus dataset D to train the universal base large model M, and then using the Givens rotation method to fine-tune all linear layers in the trained universal base large model M, so as to inject the target vertical domain knowledge into the universal base large model M, and obtain the target vertical domain large model M'; The method for fine-tuning all linear layers in the universal base large model M by using the Givens rotation method is as follows: construct a d×d orthogonal matrix for Givens rotation for the linear layer W to be fine-tuned, and when performing the first rotation P1 on the linear layer W, all (2k, 2k+1)-planes in the linear space where the linear layer W is located are rotated at the same time, where k is all integers between 0 and (d / 2)-1; when performing the second rotation P2 on the linear layer W, all (4k, 4k+2)-planes in the linear space where the linear layer W is located that are not rotated are rotated at the same time, where k is all integers between 0 and (d / 4)-1; when performing the third rotation P3 on the linear layer W, all (8k, 8k+4)-planes in the linear space where the linear layer W is located that are not rotated are rotated at the same time, where k is all integers between 0 and (d / 8)-1; and so on, when performing the nth rotation P4 on the linear layer W, n When all (2 n k, 2 n k+2 n-1 )-The unrotated plane in the plane is rotated.
2. The method according to claim 1, characterized in that: Spin each Givens Change to a quasi-Givens rotation transform Among them, (α i , β i ) consists of four learnable parameters; then for each quasi-Givens rotation transformation Calculate the inner product of its column vectors respectively and design the regularization term Added to the model loss function in the fine-tuning stage; where θ i is the rotation angle required for fine-tuning in Givens rotation, α i =(α 1i , α 2i ), β i =(β 1i , β 2i ), α 1i , α 2i , β 1i , β 2i are learnable parameters in the linear transformation.
3. The method according to claim 1 or 2, characterized in that: The target vertical domain is the medical field, and the corpus dataset D includes medical literature and medical consultation records.
4. The method according to claim 1 or 2, characterized in that: The target vertical domain is the legal field, and the corpus dataset D includes laws, regulations, and legal case precedents.
5. A consulting service method, comprising the steps of: 1) Select a general base model M and a corpus dataset D of target vertical domain knowledge; 2) Using the corpus dataset D to train the universal base large model M, and then using the Givens rotation method to fine-tune all linear layers in the trained universal base large model M, so as to inject the target vertical domain knowledge into the universal base large model M, and obtain the target vertical domain large model M'; 3) Input the consultation of the target vertical domain into the target vertical domain macro model M' to obtain the corresponding consultation result.
6. The method according to claim 5, characterized in that The method for fine-tuning all linear layers in the universal base large model M by using the Givens rotation method is as follows: a d×d orthogonal matrix for Givens rotation is constructed for the linear layer W to be fine-tuned, and when the first rotation P1 is performed on the linear layer W, all (2k, 2k+1)-planes in the linear space where the linear layer W is located are rotated at the same time, where k is all integers between 0 and (d / 2)-1; when the second rotation P2 is performed on the linear layer W, all unrotated planes in the (4k, 4k+2)-planes in the linear space where the linear layer W is located are rotated at the same time, where k is all integers between 0 and (d / 4)-1; when the third rotation P3 is performed on the linear layer W, all unrotated planes in the (8k, 8k+4)-planes in the linear space where the linear layer W is located are rotated at the same time, where k is all integers between 0 and (d / 8)-1; and so on, when the nth rotation P4 is performed on the linear layer W, n When all (2 n k, 2 n k+2 n-1 )-The unrotated plane in the plane is rotated.
7. The method according to claim 6, characterized in that Spin each Givens Change to a quasi-Givens rotation transform Among them, (α i , β i ) consists of four learnable parameters; then for each quasi-Givens rotation transformation Calculate the inner product of its column vectors respectively and design the regularization term Added to the model loss function in the fine-tuning stage; where θ i is the rotation angle required for fine-tuning in Givens rotation, α i =(α 1i , α 2i ), β i =(β 1i , β 2i ), α 1i , α 2i , β 1i , β 2i are learnable parameters in the linear transformation.
8. The method according to claim 5, 6 or 7, characterized in that: The target vertical domain is the medical field, and the corpus dataset D includes medical literature and medical consultation records.
9. The method according to claim 5, 6 or 7, characterized in that: The target vertical domain is the legal field, and the corpus dataset D includes laws, regulations, and legal case precedents.