Knowledge question and answer method and device based on large model fine tuning and storage medium
By adding orthogonal subspace to the large language model and performing two-stage training, the problem of poor effectiveness of conventional efficient parameter fine-tuning methods when injecting new domain knowledge is solved, achieving more efficient target knowledge learning and more accurate knowledge Q&A performance.
Patent Information
- Application Number
- CN202510588607.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-08
- Publication Date
- 2025-06-06
- Estimated Expiration
- 2045-05-08
AI Technical Summary
In the knowledge Q&A scenario, the use of conventional efficient parameter fine-tuning methods to inject new domain knowledge is poor, resulting in limited performance limits of the model in the target domain.
Two-stage training is performed by adding orthogonal subspaces next to the original parameters of the large language model, including the inverse of the orthogonal basis, trainable parameters and orthogonal basis. The first stage freezes the training parameters and fits the orthogonal subspace to the subspace of the target knowledge field data set; the second stage freezes the orthogonal basis and trains the training parameters to complete knowledge injection.
This method allows the model to learn the target knowledge field more efficiently during the training process, improves the accuracy of new domain knowledge questions and answers, and maintains the efficiency of efficient parameter fine-tuning.
Smart Images

Figure CN120106230A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of deep learning and large language models, and in particular to a knowledge question answering method, device and storage medium based on large model fine-tuning. Background Art
[0002] Large language models based on the Transformer structure, represented by Deepseek, ChatGPT, and LLaMa, are being widely used in various industries. Their efficient emergence and generalization capabilities have made them a new research focus in the field of artificial intelligence, replacing traditional algorithms such as LSTM and RNN in the field of natural language processing. In order to obtain better reasoning performance, the current mainstream large language models usually adopt a large parameter design approach. Under this design approach, the number of parameters of large language models is often in the billions, and the leading models in the field usually have hundreds of billions of parameters, which makes large language models have significantly improved capabilities compared to traditional natural language processing methods. However, despite the remarkable achievements, large language models still face many challenges. On the one hand, the huge model size leads to high computing costs and energy consumption, which limits its deployment in resource-limited environments; on the other hand, due to the lack of sufficient diverse training data, the model may produce biased or inaccurate results in some cases. How to overcome these problems and further improve the efficiency and accuracy of large language models has become one of the key directions of current research.
[0003] At present, the mainstream research idea focuses on efficient parameter fine-tuning, which is an idea that aims to study the training of some parameters (usually at the million level) instead of all billions of parameters to achieve the same performance as training all parameters. Due to the inherent shortage of trainable parameters, although the resources required for training are greatly reduced and the training speed is faster, the performance upper limit of efficient parameter fine-tuning is usually not as good as full parameter fine-tuning. For this reason, many researchers at home and abroad are committed to the research of improving the upper limit of efficient parameter fine-tuning. The current mainstream efficient parameter fine-tuning algorithms can be divided into two categories: The first type of algorithm is the low-rank adaptation algorithm (Low-Rank Adaptation, LoRA). (J. H, SHEN Y,WALLIS P, et al. LoRA: Low-Rank Adaptation of Large Language Models.[J].arXiv: Computation and Language,arXiv: Computation and Language, 2021.) is the first work to propose a low-rank adaptation algorithm. This work splits the high-dimensional matrix to be trained into two smaller matrix products, thereby using low-dimensional matrix updates to simulate high-dimensional matrix updates. In addition, (Zhang, Qingru, Minshuo Chen, AlexanderBukharin, Nikos Karampatziakis, Pengcheng He, Yu Cheng, Weizhu Chen and TuoZhao. “AdaLoRA: Adaptive Budget Allocation for Parameter-Efficient Fine-Tuning.” (2023).) The weights are decomposed into two orthogonal matrices and one diagonal matrix using singular value decomposition (SVD). All three weight matrices are learnable. During training, singular values are gradually pruned according to their importance scores (constructed by the moving average of the gradient-weight product size). This adaptive method enables the model to dynamically adjust the rank within each low-rank adaptation module and effectively manage the number of parameters based on the importance of the weight matrix. The efficient parameter fine-tuning algorithm with the low-rank adaptation algorithm as the core maintains the overall high efficiency of training and low consumption of storage resources and training resources, but its performance ceiling is relatively limited and usually lags far behind full parameter fine-tuning.
[0004] The second type of method uses a structure called "adapter". This structure was proposed in the paper (Houlsby, Neil, AndreiGiurgiu, Stanislaw Jastrzebski, Bruna Morrone, Quentin de Laroussilhe, AndreaGesmundo, Mona Attariyan and Sylvain Gelly. "Parameter-Efficient TransferLearning for NLP." ArXiv abs / 1902.00751 (2019): n. pag.). This structure is similar to the feed-forward neural network (FFN) in the Transformer structure of the large language model, with linear and nonlinear layers, which are usually attached to or before and after the original parameters of the model. (Wang, Yaqing, SubhabrataMukherjee, Xiaodong Liu, Jing Gao and Jianfeng Gao. “AdaMix: Mixture-of-Adaptations for Parameter-efficient Model Tuning.” ArXiv abs / 2210.17451(2022): n. pag.) Combining the idea of the mixture of experts (MoE), multiple adapters are connected in parallel to form a group of adapter experts so that they can adapt to a variety of different domain knowledge.
[0005] Both of these main ideas face the problem of insufficient upper limit of fine-tuning capabilities, that is, when the training data set becomes larger, their performance will quickly lag behind full parameter fine-tuning; in addition, when they are combined with hybrid expert models, they need to explicitly specify a router to specify which expert modules are responsible at this time, which affects training efficiency and inference efficiency. Summary of the invention
[0006] In view of the shortcomings of the prior art, the present invention provides a knowledge question answering method, device and storage medium based on large model fine-tuning to solve the problem of poor injection effect when using conventional efficient parameter fine-tuning methods (such as low-rank adaptation methods) to inject new field knowledge (such as legal, medical, financial and other field knowledge) in knowledge question answering scenarios.
[0007] A knowledge question answering method based on large model fine-tuning includes the following steps: (1) adding an orthogonal subspace next to the original parameters of the large language model, wherein the orthogonal subspace includes an orthogonal basis, training parameters, and an inverse of the orthogonal basis; (2) Perform two-stage training. In the first stage, freeze the training parameters and fit the orthogonal subspace to the subspace of the target knowledge domain dataset. In the second stage, freeze the orthogonal basis and train the training parameters in the subspace that has been fitted to the target knowledge domain dataset to complete knowledge injection. (3) Input questions in the target knowledge domain into the trained large language model to generate response answers.
[0008] Step (1) specifically includes: (1-1) The large language model is recorded as , at each layer Without involving its own parameters, a set of orthogonal bases are randomly created As the initial basis of the orthogonal subspace, the orthogonal basis is attached to the original parameters of the model , and also as the entry parameter of the orthogonal subspace; , Indicates the number of layers of the large language model; (1-2) The orthogonal basis The inverse Recorded as , as the exit parameter of the orthogonal subspace, in and Create a set of parameters As the training parameters inside the orthogonal subspace, it is composed of , and The orthogonal subspace composed of three sets of parameters.
[0009] Preferably, in order to reduce the number of sub-orthogonal subspaces, in step (1), the number of added orthogonal subspaces can be compressed by a cross-layer parameter sharing strategy, specifically including: Based on large language model Number of layers The groups are configured: in the shallow layer close to the input, each large language model layer is a group, and an orthogonal subspace is set for each group; in the middle layer, each 2 to 4 large language model layers are a group, and an orthogonal subspace is set for each group; in the deep layer close to the output, each large language model layer is a group, and an orthogonal subspace is set for each group.
[0010] In step (2), in the first stage of training, the training parameters in the orthogonal subspace are Freeze and set to the identity matrix, using a dataset collected from the target knowledge domain Subset of To train the orthogonal subspace of frozen training parameters and adapt it to the space of the target knowledge domain dataset.
[0011] In step (2), in the second stage of training, the training parameters are randomly initialized , the orthogonal basis and Freeze as a basis of the subspace and train the parameters in an orthogonal subspace Conduct training, using the entire target knowledge domain dataset during training .
[0012] Preferably, in the second stage of training, after each set training time step T, a greedy algorithm is used to perform rank contraction to search for a smaller rank within the performance tolerance limit, so that the rank of the orthogonal subspace is gradually reduced.
[0013] A knowledge question and answer device based on large model fine-tuning includes a memory and one or more processors, wherein the memory stores executable code, and when the one or more processors execute the executable code, they are used to implement the above-mentioned knowledge question and answer method based on large model fine-tuning.
[0014] A computer-readable storage medium stores a program, which, when executed by a processor, implements the above-mentioned knowledge question-answering method based on large model fine-tuning.
[0015] Compared with the prior art, the present invention has the following beneficial effects: The present invention maps latent vectors to the target knowledge domain subspace by constructing an explicit orthogonal subspace, and then constructs trainable parameters in the orthogonal subspace, thereby prompting the model to learn more efficiently in the target knowledge domain during training, rather than randomly learning in an uncertain subspace. This method has the advantages of explicit domain determination and strong interpretability, so that the learnable parameters are more concentrated in the target knowledge domain, which can not only maintain the efficiency of efficient parameter fine-tuning, but also further improve the performance upper limit of model training, and ultimately improve the accuracy of new domain knowledge question answering. BRIEF DESCRIPTION OF THE DRAWINGS
[0016] Figure 1 It is a schematic diagram of adding an orthogonal subspace next to the original parameters of the large language model in an embodiment of the present invention.
[0017] Figure 2 This is a training flow chart of an embodiment of the present invention.
[0018] Figure 3 It is a schematic diagram of the structure of the device of the present invention. DETAILED DESCRIPTION
[0019] The present invention is further described in detail below in conjunction with the accompanying drawings and embodiments. It should be pointed out that the embodiments described below are intended to facilitate the understanding of the present invention and do not have any limiting effect on the present invention.
[0020] To address the problems that the existing efficient parameter fine-tuning has limited performance ceiling when fine-tuning large language models and poor interpretability of training process and structure, the present invention adopts a method based on explicit orthogonal subspace to promote the model to learn target domain knowledge, thereby improving the performance ceiling of the efficient parameter fine-tuning method when fine-tuning large language models.
[0021] The core technology of the present invention is the construction method of orthogonal subspace, which enables the trainable parameters in the efficient parameter fine-tuning method to be closely integrated with the target domain and the training process of the large language model.
[0022] The embodiment of the present invention uses LLaMa-2-7B as an example large language model to illustrate the specific implementation steps of the present invention. However, the implementation of the present invention is not limited to this model, which is only used as a reference case.
[0023] A knowledge question answering method based on large model fine-tuning includes the following steps: Step 1, add an orthogonal subspace next to the original parameters of the large language model, the orthogonal subspace includes an orthogonal basis, trainable parameters, and the inverse of the orthogonal basis.
[0024] like Figure 1 As shown, in the structure of LLaMa-2-7B, in the module to be trained, an orthogonal subspace is attached to each parameter matrix to be trained, and an orthogonal subspace is trained for each field that the user wants to train (such as finance, law, etc.). The specific sub-steps include: Step 1-1, the large language model LLaMa-2-7B consists of three parts, namely the embedding layer , the main body of the model of the Transformer architecture And language modeling head , where the Transformer architecture model body It occupies the most important parameter amount of the model, and each layer contains several parameter matrices , these parameter matrices are distributed in the self-attention module (Self-Attention) and the feed-forward neural network module (Feed-Forward Network, FFN). Each parameter matrix will be assigned three layers of orthogonal subspaces belonging to the matrix The latent vector flowing into this parameter matrix and the latent vector flowing out of the parameter matrix The calculation method before allocating orthogonal subspaces is ; The calculation method after allocating orthogonal subspaces is .
[0025] Step 1-2, during initialization, will be orthogonally initialized, and for this we need to derive a set of orthogonal bases, a common approach being singular value decomposition (SVD). Select an arbitrary matrix , whose size is , then its singular value decomposition should be defined as , at this time That is, the required set of orthogonal bases is assigned to . is initialized to the identity matrix I. Will The inverse of is initialized, and Form a set of mutually inverse matrices. In this embodiment, The size is , The size is , The size is , The size is ,Keep and The shape is consistent.
[0026] The present invention herein The selection of is randomly generated, and the randomness of this selection will have a certain impact on the actual final training performance, but after the first stage of warm-up training, the performance calculated by the final mathematical field accuracy is small and can be ignored. Calculate, not The sum of .
[0027] According to the above method, the present invention constructs a three-layer orthogonal subspace, explicitly constructs the optimization space of the target knowledge domain, and prompts the large language model to learn in an interpretable space. After applying the method of the present invention, compared with the low-rank adaptive algorithm, the accuracy of law (calculated based on the test set of the LegalBench dataset) on LLaMa-2-7B is increased from 6.55% (39.87% to 46.42%) to 9.30% (39.87% to 49.17%).
[0028] Step 2: Perform two-stage training. In the first stage, freeze the trainable parameters and fit the orthogonal subspace to the subspace of the target knowledge domain dataset. In the second stage, freeze the orthogonal basis and train the trainable parameters in the subspace that has been fitted to the target knowledge domain dataset to complete knowledge injection.
[0029] like Figure 2As shown in the figure, after two different training steps, the orthogonal subspace is first preheated (stage 1) to make it fit to the target domain subspace, and then the trainable parameters are trained conventionally (stage 2) to inject specific knowledge. The specific sub-steps include: Step 2-1, in the first stage of warm-up training, freeze the trainable parameters , open orthogonal basis and its inverse For warm-up training, using the target knowledge dataset A subset of Train the model so that the orthogonal basis and its inverse Gradually converge to the subspace where the original data set D is located. In this step, it is necessary to apply orthogonal constraints so that the orthogonal basis Maintain orthogonality and apply product constraints so that the orthogonal basis and its inverse Keep the inverse. The two losses are as follows: ; .
[0030] Step 2-2, in the second stage of training, freeze the orthogonal basis and its inverse , open trainable parameters For regular training. Using the complete target knowledge dataset Train the model to make trainable parameters Learn knowledge in the target knowledge domain subspace. In the regular training step, the two losses described in step 2-1 are not applied and .
[0031] Following the above steps, we can obtain a three-layer orthogonal subspace that is well learned in the target domain. After step 2, one example of improved answer quality is as follows: Question: "Company A and Company B signed a contract which included an arbitration clause and selected the China Arbitration Association as the arbitration institution. When a dispute arose, Company A requested arbitration, but Company B stated that the arbitration agreement was invalid, believing that the dispute was beyond the scope of arbitration prescribed by law. In this case, is the arbitration agreement valid?"
[0032] Correct answer (given by the model trained based on the present invention): (Refer to the Law on Litigation and Non-litigation Procedures - Arbitration Law 2017-09-01: Article 17(a)) According to Article 17(a) of the Law on Litigation and Non-litigation Procedures - Arbitration Law 2017-09-01, if the agreed arbitration matters exceed the scope of arbitration prescribed by law, the arbitration agreement will be invalid. Litigation and Non-litigation Procedure Law - Arbitration Law 2017-09-01: "The China Arbitration Association shall formulate arbitration rules in accordance with this Law and the relevant provisions of the Civil Procedure Law.", Litigation and Non-litigation Procedure Law - Arbitration Law 2017-09-01: "Article 16 An arbitration agreement includes an arbitration clause in a contract and an agreement to request arbitration reached in other written forms before or after a dispute occurs. If one party fails to perform the award of an arbitration institution established in accordance with the law, the other party may apply to the people's court with jurisdiction for execution. The people's court to which the application is made shall execute the award. If the respondent provides evidence to prove that the arbitration award should not be executed, the people's court may rule not to execute the award after review and verification by a collegial panel. However, if the parties do not have an arbitration clause in the contract or fail to reach a written arbitration agreement afterwards, or if the matter of the award does not fall within the scope of the arbitration agreement or the arbitration institution has no right to arbitrate, the award shall still be executed. Therefore, in this case, Party A may apply to the people's court with jurisdiction for execution of the award of the arbitration institution, and Party B shall perform the award.
[0033] Comparative answers (given by the model trained by the low-rank adaptive method): Litigation and Non-litigation Procedural Law - Arbitration Law 2017-09-01: "An arbitration agreement shall have the following contents:", Litigation and Non-litigation Procedural Law - Arbitration Law 2017-09-01: "(i) The intention to request arbitration;", Litigation and Non-litigation Procedural Law - Arbitration Law 2017-09-01: "(ii) The matters of arbitration;", Litigation and Non-litigation Procedural Law - Arbitration Law 2017-09-01: "(iii) The selected arbitration commission.", Litigation and Non-litigation Procedural Law - Arbitration Law 2017-09-01: "Article 17 An arbitration agreement shall be invalid if any of the following circumstances exists:", Litigation and Non-litigation Procedural Law - Arbitration Law 2017-09-01: "(i) The agreed matters of arbitration exceed the scope of arbitration prescribed by law;".
[0034] It can be seen that after applying the present invention, the answers given by the trained model are more complete and the reasoning is more sufficient.
[0035] In order to further compress the amount of trainable parameters while keeping the performance loss small enough, the embodiments of the present invention also propose a cross-layer parameter sharing strategy and a greedy rank compression method.
[0036] Cross-layer parameter sharing strategy. Based on the difference in knowledge density of different layers of the large language model, the density of the orthogonal subspace modules that would originally be added to each layer is reduced to one shared by 2 to 4 layers. In this embodiment, in the 32-layer LLaMa model, the Transformer layers numbered 0 to 7 have a higher knowledge density, and each layer maintains an orthogonal subspace module for each parameter matrix; the Transformer layers numbered 8 to 15 have a medium knowledge density, and every two layers of parameter matrices at the same position share an orthogonal subspace module; the Transformer layers numbered 16 to 27 have the lowest knowledge density, and every four layers of parameter matrices at the same position share an orthogonal subspace module; the Transformer layers numbered 28 to 31 complete the formation of the final output format, and the same as the strategy for layers 0 to 7, each parameter matrix maintains an orthogonal subspace module. This strategy as a whole will reduce the total amount of trainable parameters to about 59%, calculated as follows: .
[0037] Greedy rank compression method. Step training, using greedy algorithm to dynamically adjust the rank, if the loss caused by rank shrinkage exceeds the original loss If the value is times, the rank is increased; otherwise, the rank is decreased until the optimal rank is found.
[0038] Based on the same inventive principle, an embodiment of the present invention also provides a knowledge question and answer device based on large model fine-tuning, including a memory and one or more processors, wherein the memory stores executable code, and when the one or more processors execute the executable code, they are used to implement the knowledge question and answer method based on large model fine-tuning mentioned in the above embodiment.
[0039] The embodiment of the knowledge question-answering device based on large model fine-tuning of the present invention can be applied to any device with data processing capabilities, and the device with data processing capabilities can be a device or apparatus such as a computer. The device embodiment can be implemented through software, or through hardware or a combination of software and hardware. Taking software implementation as an example, as a device in a logical sense, it is formed by the processor of any device with data processing capabilities in which it is located reading the corresponding computer program instructions in the non-volatile memory into the internal memory for execution. From the hardware level, if Figure 3 As shown in the figure, it is a hardware structure diagram of any device with data processing capability where the knowledge question answering device based on large model fine-tuning of the present invention is located. Figure 3 In addition to the processor, memory, network interface, and non-volatile memory shown, any device with data processing capabilities in which the apparatus in the embodiments is located may also include other hardware, generally based on the actual functions of the device with data processing capabilities, which will not be described in detail.
[0040] Based on the same inventive principle, an embodiment of the present invention further provides a computer-readable storage medium on which a program is stored. When the program is executed by a processor, the knowledge question-answering method based on large model fine-tuning mentioned in the above embodiment is implemented.
[0041] The embodiments described above provide a detailed description of the technical solutions and beneficial effects of the present invention. It should be understood that the above are only specific embodiments of the present invention and are not intended to limit the present invention. Any modifications, supplements and equivalent substitutions made within the scope of the principles of the present invention should be included in the protection scope of the present invention.
Claims
1. A knowledge question answering method based on large model fine-tuning, characterized in that: The following steps are involved: (1) adding an orthogonal subspace next to the original parameters of the large language model, wherein the orthogonal subspace includes an orthogonal basis, training parameters, and an inverse of the orthogonal basis; (2) Perform two-stage training. In the first stage, freeze the training parameters and fit the orthogonal subspace to the subspace of the target knowledge domain dataset. In the second stage, the orthogonal basis is frozen and the training parameters are trained in the subspace that has been fitted to the target knowledge domain dataset to complete the knowledge injection. (3) Input questions in the target knowledge domain into the trained large language model to generate response answers.
2. The knowledge question answering method based on large model fine-tuning according to claim 1 is characterized in that: Step (1) specifically includes: (1-1) The large language model is recorded as , at each layer Without involving its own parameters, a set of orthogonal bases are randomly created As the initial basis of the orthogonal subspace, the orthogonal basis is attached to the original parameters of the model , and also as the entry parameter of the orthogonal subspace; , Indicates the number of layers of the large language model; (1-2) The orthogonal basis The inverse Recorded as , as the exit parameter of the orthogonal subspace, in and Create a set of parameters As the training parameters inside the orthogonal subspace, it is composed of , and The orthogonal subspace composed of three sets of parameters.
3. The knowledge question answering method based on large model fine-tuning according to claim 1 is characterized in that: In step (1), the number of orthogonal subspaces is compressed by a cross-layer parameter sharing strategy, which specifically includes: Based on large language model Number of layers The groups are configured: in the shallow layer close to the input, each large language model layer is a group, and an orthogonal subspace is set for each group; in the middle layer, each 2 to 4 large language model layers are a group, and an orthogonal subspace is set for each group; in the deep layer close to the output, each large language model layer is a group, and an orthogonal subspace is set for each group.
4. The knowledge question answering method based on large model fine-tuning according to claim 1 is characterized in that: In step (2), in the first stage of training, the training parameters in the orthogonal subspace are Freeze and set to the identity matrix, using a dataset collected from the target knowledge domain Subset of To train the orthogonal subspace of frozen training parameters and adapt it to the space of the target knowledge domain dataset.
5. The knowledge question answering method based on large model fine-tuning according to claim 4 is characterized in that: In step (2), in the second stage of training, the training parameters are randomly initialized , the orthogonal basis and Freeze as a basis of the subspace and train the parameters in an orthogonal subspace Conduct training, using the entire target knowledge domain dataset during training .
6. The knowledge question answering method based on large model fine-tuning according to claim 5 is characterized in that: In the second stage of training, after each set training time step T, a greedy algorithm is used to perform rank contraction and search for a smaller rank within the performance tolerance limit, so that the rank of the orthogonal subspace is gradually reduced.
7. A knowledge question answering device based on large model fine-tuning, characterized in that: It includes a memory and one or more processors, wherein the memory stores executable code, and when the one or more processors execute the executable code, it is used to implement the knowledge question answering method based on large model fine-tuning as described in any one of claims 1-6.
8. A computer-readable storage medium, characterized in that: A program is stored thereon, and when the program is executed by a processor, the knowledge question answering method based on large model fine-tuning as described in any one of claims 1-6 is implemented.
Citation Information
Patent Citations
Financial question and answer text processing method and device, equipment and storage medium
CN117668177A
Small sample target detection and identification method based on self-supervised ion separation space
CN117876868A
Large model parameter fine tuning method, device, equipment, medium and product
CN119398127A
Small sample continuous learning method and system based on continuous knowledge protection decomposition
CN119476410A
Telecommunication service question and answer method, system and server based on large model fine tuning
CN119513266A