Knowledge-intensive task-oriented model reasoning method, device, equipment, medium and product
By deploying knowledge adapters and routers in parallel in large language models and using knowledge graphs for training, the catastrophic forgetting problem of large language models in knowledge-intensive tasks is solved, the accuracy and adaptability of the model are improved, and it is suitable for intelligent question-answering scenarios.
Patent Information
- Application Number
- CN202510699284.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-28
- Publication Date
- 2025-09-30
AI Technical Summary
Large language models suffer from catastrophic forgetting when faced with knowledge-intensive tasks, resulting in inaccurate question-answering results.
Knowledge adapters and knowledge routers are deployed in parallel in the Transformer architecture of the initial large language model. Knowledge graphs are used for training, and the degree of integration of new knowledge is dynamically adjusted. Knowledge integration is driven by knowledge boundary detection and routing strategies, thereby enhancing the model's knowledge coverage and reducing the risk of conflict between new and old knowledge.
It effectively alleviates the problem of catastrophic forgetting, improves the accuracy and adaptability of large language models in knowledge-intensive tasks, and is suitable for intelligent question-answering scenarios that require continuous knowledge updating.
Smart Images

Figure CN120725129A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of artificial intelligence technology, and in particular to a model reasoning method, device, equipment, medium and product for knowledge-intensive tasks. Background Art
[0002] With the rapid development of large-scale model technology, large language models have demonstrated tremendous potential and broad application prospects in natural language processing and other fields, and their excellent cross-domain generation capabilities have attracted considerable attention. However, when faced with knowledge-intensive tasks, such as specialized fields like home-based guest interviews, large language models still fall short. Due to the lack of in-depth knowledge accumulation in fields like home-based guest interviews and the potential for knowledge forgetting, large language models occasionally provide misleading answers and factual deviations when answering specialized questions in this field, creating a "knowledge illusion" that affects the accuracy of the results.
[0003] Existing techniques for integrating domain-specific knowledge into large language models involve embedding a series of lightweight, domain-specific adapters into pre-trained large language models as knowledge repositories. These adapters focus on storing and processing domain-specific knowledge without affecting the generalizability of other parts of the model. While this approach demonstrates good reliability, it still suffers from the problem of catastrophic forgetting, resulting in inaccurate question-answering results derived from large language models. Summary of the Invention
[0004] The present invention provides a model reasoning method, device, equipment, medium and product for knowledge-intensive tasks, which is used to solve the catastrophic forgetting problem of large language models in the existing technology, resulting in inaccurate question and answer results.
[0005] In a first aspect, the present invention provides a model reasoning method for knowledge-intensive tasks, comprising: Obtaining input text; the input text is a question raised based on a specific knowledge-intensive task; Inputting the input text into a target large language model to obtain a question-answering result output by the target large language model; Among them, the target large language model is a knowledge adapter and a knowledge router deployed in parallel in N consecutive target Transformer layers in the Transformer architecture of the initial large language model; any of the knowledge adapters is used to integrate new knowledge into the initial large language model while keeping the parameters of the initial large language model unchanged; any of the knowledge routers is used to dynamically adjust the degree to which the new knowledge is integrated into the initial large language model according to the initial large language model's knowledge mastery of the new knowledge; the initial large language model is trained based on the unknown knowledge of the preset large language model in the knowledge graph.
[0006] In one embodiment, the initial large language model is determined by: Generate multiple-choice questions; the multiple-choice questions are generated based on knowledge triples in the knowledge graph; Inputting the multiple-choice question into a preset large language model to obtain a question-answering result of the preset large language model; If the question-and-answer result does not conform to the standard answer corresponding to the multiple-choice question, determining that the knowledge triple is unknown knowledge of the preset large language model; Model training is performed based on the unknown knowledge of the preset large language model in the knowledge graph to obtain an initial large language model.
[0007] In one embodiment, generating a multiple-choice question includes: Generate a question-answer pair based on the knowledge triples in the knowledge graph; the question-answer pair includes a question related to the knowledge triples and its corresponding standard answer; Determine an entity with the shortest edit distance to the head entity in the knowledge triple, and replace the head entity in the standard answer with the entity to obtain a first interference item; Determining a plurality of second distractors having the shortest edit distances to the standard answer; The standard answer, the first interference item and the plurality of second interference items are integrated into multiple choices, and a multiple choice question is generated based on the question and the multiple choices.
[0008] In one embodiment, the output of any target Transformer layer is determined by: Using the current layer knowledge adapter, the input of the FFN in the current target Transformer layer is fused with the output of the previous layer knowledge adapter to obtain a fusion matrix; Using the current layer knowledge adapter, the fusion matrix is subjected to nonlinear transformation to obtain a nonlinear transformation result; Using the current layer knowledge router, a routing value is generated according to the knowledge mastery of the new knowledge by the input of the FFN; the routing value is used to regulate the strength of knowledge fusion; multiplying the routing value, the nonlinear transformation result and the routing value to obtain a weighted result; The weighted result is fused with the output of the FFN to obtain the output of the current target Transformer layer.
[0009] In one embodiment, the target large language model is fine-tuned in the following manner: By optimizing the first objective function of the routing optimization phase, the model parameters of the target large language model are adjusted to obtain a first fine-tuning model; By optimizing the second objective function in the question-answer pair training phase, the model parameters of the first fine-tuning model are adjusted to obtain a second fine-tuning model; By optimizing the third objective function of the relationship classification training phase, the model parameters of the second fine-tuning model are adjusted to obtain a fine-tuned target large language model.
[0010] In one embodiment, the expression of the first objective function is as follows: ; in, represents the first objective function; Represents the input of FFN; Represents the computational function of a multilayer perceptron; Represents a knowledge sample; represents the routing label of the knowledge sample; BCE() represents the binary cross entropy loss function; E() represents the expected calculation function.
[0011] In one embodiment, the expression of the second objective function is as follows: ; in, represents the second objective function; Represents a sample of questions, represents the model’s predicted answer, is the standard answer sample; CE() represents the cross entropy loss function; E() represents the expected calculation function.
[0012] In one embodiment, the expression of the third objective function is as follows: ; in, represents the third objective function; represents the loss caused by relation classification, Indicates its corresponding weight; Token loss representing model inference; ; in, represents the temperature hyperparameter; Represents a set of relations; Represents the relationship in knowledge triple sample i; Indicates that the relationship set contains The rest of the relationship; Represents the relationship vector representation obtained by concatenating the head entity vector representation and the tail entity vector representation corresponding to the knowledge triple sample i; and Respectively and Aligned into a unified dimensional space; E() represents the expected calculation function; k represents the knowledge statement fragment; ; Among them, k represents the knowledge statement fragment sample; Represents the knowledge statement fragment predicted by the model; CE() represents the cross entropy loss function; E() represents the expectation calculation function.
[0013] In a second aspect, the present invention further provides a model reasoning device for knowledge-intensive tasks, comprising: An acquisition module is used to acquire input text; the input text is a question raised based on a specific knowledge-intensive task; A model inference module, configured to input the input text into a target large language model and obtain a question-answering result output by the target large language model; Among them, the target large language model is a knowledge adapter and a knowledge router deployed in parallel in N consecutive target Transformer layers in the Transformer architecture of the initial large language model; any of the knowledge adapters is used to integrate new knowledge into the initial large language model while keeping the parameters of the initial large language model unchanged; any of the knowledge routers is used to dynamically adjust the degree to which the new knowledge is integrated into the initial large language model according to the initial large language model's knowledge mastery of the new knowledge; the initial large language model is trained based on the unknown knowledge of the preset large language model in the knowledge graph.
[0014] In a third aspect, the present invention provides an electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein when the processor executes the program, the steps of the model reasoning method for knowledge-intensive tasks as described above are implemented.
[0015] In a fourth aspect, the present invention also provides a non-transitory computer-readable storage medium, on which a computer program is stored, and when the computer program is executed by a processor, the steps of the model reasoning method for knowledge-intensive tasks as described above are implemented.
[0016] In a fifth aspect, the present invention also provides a computer program product, which includes a computer program, which can be stored on a non-transitory computer-readable storage medium, and when the computer program is executed by the processor, implements the steps of the model reasoning method for knowledge-intensive tasks as described in any one of the above.
[0017] The model reasoning method, device, equipment, medium and product for knowledge-intensive tasks provided by the present invention first perform training based on the unknown knowledge of the large language model preset in the knowledge graph, thereby enhancing the coverage of basic knowledge. Then, the knowledge adapter and knowledge router are deployed in parallel in the Transformer architecture of the trained initial large language model. The adapter injects new knowledge without changing the parameters of the original model, and the router dynamically adjusts the fusion strength according to the model's mastery of the new knowledge. This not only effectively expands the knowledge boundary of the large language model, but also reduces the risk of conflict between new and old knowledge through a weighted mechanism of knowledge mastery perception. This structural improvement can effectively alleviate the problem of catastrophic forgetting while maintaining the original capabilities of the model, thereby significantly improving the accuracy and adaptability of the target large language model in knowledge-intensive tasks, and is suitable for intelligent question-answering scenarios that require continuous knowledge updating. BRIEF DESCRIPTION OF THE DRAWINGS
[0018] In order to more clearly illustrate the technical solutions in the present invention or the prior art, a brief introduction will be given below to the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.
[0019] Figure 1 It is a flowchart of the model reasoning method for knowledge-intensive tasks provided by the present invention.
[0020] Figure 2 This is a schematic diagram of knowledge boundary detection provided by the present invention.
[0021] Figure 3 It is an overall schematic diagram of the routing strategy driven knowledge integration provided by the present invention.
[0022] Figure 4 It is a detailed schematic diagram of the routing strategy-guided knowledge adapter provided by the present invention.
[0023] Figure 5 It is a structural diagram of the model reasoning device for knowledge-intensive tasks provided by the present invention.
[0024] Figure 6 It is a structural schematic diagram of the electronic device provided by the present invention. DETAILED DESCRIPTION
[0025] To make the objectives, technical solutions, and advantages of the present invention more clear, the technical solutions of the present invention will be clearly and completely described below in conjunction with the accompanying drawings. Obviously, the embodiments described are only some of the embodiments of the present invention, not all of them. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts shall fall within the scope of protection of the present invention.
[0026] The terms "first," "second," and the like in the present invention are used to distinguish similar objects and are not used to describe a particular order or precedence. It should be understood that the terms used in this manner are interchangeable where appropriate, so that the embodiments of the present invention can be implemented in orders other than those illustrated or described herein.
[0027] The following combination Figures 1-6 The present invention describes the model reasoning method, device, equipment, medium and product for knowledge-intensive tasks.
[0028] It should be noted that the model reasoning method for knowledge-intensive tasks provided by the embodiment of the present invention is implemented based on a model reasoning device for knowledge-intensive tasks.
[0029] The model reasoning method for knowledge-intensive tasks provided by the embodiment of the present invention is suitable for designing knowledge boundary detection and routing strategy-driven knowledge integration solutions by utilizing the rich structure of the knowledge graph, quantifiable and updatable knowledge units, without affecting the knowledge already mastered by the large model. This enables the large model to efficiently integrate and incorporate knowledge that has not yet been mastered in knowledge-intensive fields, alleviate the catastrophic forgetting problem of the large model when processing knowledge-intensive tasks of home customers, and have excellent generalization performance in other vertical fields.
[0030] The embodiment of the present invention takes a model reasoning device for knowledge-intensive tasks as an execution subject and describes a model reasoning method for knowledge-intensive tasks.
[0031] Combine Figure 1 , Figure 1 It is a flowchart of the model reasoning method for knowledge-intensive tasks provided by the present invention.
[0032] like Figure 1 As shown, the model reasoning method for knowledge-intensive tasks includes the following steps: Step 101: Get input text; Step 102: Input the input text into a target large language model to obtain a question-answering result output by the target large language model.
[0033] Specifically, users ask a question based on a specific knowledge-intensive task and can also enter appropriate prompts. These prompts guide the model to output the desired content. Based on the question and corresponding prompts, users enter text into the interactive input box of the target large language model. Knowledge-intensive tasks span a wide range of fields, such as home customer service support management, academic research, and engineering technology. For operators, these primarily involve business scenarios related to intelligent network platform operations, which are being applied to home customer service support management.
[0034] The target large language model is called to analyze and infer the user's input text, and the question-and-answer results for questions raised in knowledge-intensive tasks are output through an interactive page.
[0035] The target large language model is trained by introducing a knowledge boundary detection mechanism and routing strategy-driven knowledge integration into the existing large language model. This process focuses on implicitly encoding some representative knowledge from the knowledge graph into the parameters of the large language model, without requiring the knowledge graph to participate in the inference phase. The goal is to use the routing strategy to integrate the key knowledge from the knowledge graph into the large language model, thereby enhancing the large language model's ability to handle knowledge-intensive tasks such as home visitor integration.
[0036] Specifically, for a given large language model parameter With a set of knowledge triples The core goal is to adjust the parameters of large language models Make fine adjustments to obtain fine-tuned optimized model parameters , aiming to not affect the existing knowledge of the large language model Incorporating previously unknown knowledge To improve the efficiency and accuracy of this process, only new domain knowledge that is not currently mastered by the large language model is injected, such as: .
[0037] Knowledge boundary detection and routing strategy-driven knowledge integration are the two core steps to build an efficient and flexible knowledge integration framework.
[0038] First, knowledge triples of the knowledge graph are utilized to generate knowledge statements and multiple-choice questions and answers using predefined relationship templates to accurately detect unknown areas outside the knowledge boundary of the large language model. Once the unknown knowledge area is accurately located, the preset large language model is trained using the unknown knowledge to obtain an initial large language model that efficiently integrates the unknown knowledge area.
[0039] Subsequently, an innovative tool, the knowledge adapter, is introduced into the initial large language model. This knowledge adapter runs in parallel with the original Transformer layer of the large language model and is trained to store new knowledge. Furthermore, the core of this knowledge integration framework is the knowledge routing strategy, whose purpose is to strategically determine whether new knowledge from the knowledge adapter should be used. Throughout the entire process, only the knowledge adapter and knowledge routing mechanism are fine-tuned, while the original Transformer parameters of the large language model remain unchanged. This effectively improves the efficiency and effectiveness of knowledge integration while ensuring model stability. At the same time, it efficiently integrates new domain knowledge without compromising other unrelated knowledge samples.
[0040] The model reasoning method for knowledge-intensive tasks provided by the present invention first performs training based on the unknown knowledge of the large language model preset in the knowledge graph, thereby enhancing the coverage of basic knowledge. Then, the knowledge adapter and the knowledge router are deployed in parallel in the Transformer architecture of the trained initial large language model. The adapter injects new knowledge without changing the parameters of the original model, and the router dynamically adjusts the fusion strength according to the model's mastery of the new knowledge. This not only effectively expands the knowledge boundary of the large language model, but also reduces the risk of conflict between new and old knowledge through the weighted mechanism of knowledge mastery perception. This structural improvement can effectively alleviate the problem of catastrophic forgetting while maintaining the original capabilities of the model, thereby significantly improving the accuracy and adaptability of the target large language model in knowledge-intensive tasks, and is suitable for intelligent question-answering scenarios that require continuous knowledge updating.
[0041] In some embodiments, the initial large language model is determined by: Generate multiple-choice questions; the multiple-choice questions are generated based on knowledge triples in the knowledge graph; Inputting the multiple-choice question into a preset large language model to obtain a question-answering result of the preset large language model; If the question-and-answer result does not conform to the standard answer corresponding to the multiple-choice question, determining that the knowledge triple is unknown knowledge of the preset large language model; Model training is performed based on the unknown knowledge of the preset large language model in the knowledge graph to obtain an initial large language model.
[0042] Specifically, given the inefficiency of fine-tuning a large language model on the overall knowledge graph, the goal of this example is to identify and integrate only the unknown knowledge of the large language model. To overcome the difficulty of evaluating open-ended questions, this example converts triples into multiple-choice questions, which allows for accurate evaluation of the unknown knowledge of the large language model and then integrates it into the model. This strategy achieves efficient knowledge integration, using multiple-choice training data to improve performance in specific areas, as shown in the schematic diagram. Figure 2 As shown, Figure 2 This is a schematic diagram of knowledge boundary detection provided by the present invention.
[0043] The step of generating multiple-choice questions specifically includes the following steps: Generate a question-answer pair based on the knowledge triples in the knowledge graph; the question-answer pair includes a question related to the knowledge triples and its corresponding standard answer; Determine an entity with the shortest edit distance to the head entity in the knowledge triple, and replace the head entity in the standard answer with the entity to obtain a first interference item; Determining a plurality of second distractors having the shortest edit distances to the standard answer; The standard answer, the first interference item and the plurality of second interference items are integrated into multiple choices, and a multiple choice question is generated based on the question and the multiple choices.
[0044] Specifically, for each knowledge triple in the knowledge graph<h,r,t> , h represents the head entity, r represents the relationship, and t represents the tail entity. Using advanced large models and preset relationship templates, each knowledge triple is converted into multiple-choice questions and knowledge statement fragments.
[0045] For example, the knowledge triple <set-top box fault code 20402, representative meaning, set-top box account is in arrears> is restated as a question with a golden answer: "What does set-top box error 20402 mean? Answer: set-top box account is in arrears", and a knowledge statement fragment: "set-top box fault code 20402 represents the meaning that the set-top box account is in arrears."
[0046] When exploring the application of large language models in open-ended question-answering scenarios, a significant challenge lies in accurately evaluating the accuracy of long text answers generated by them. To address this challenge, this example transforms the traditional open-ended question-answering model into a multiple-choice question-answering format as an effective way to improve evaluation efficiency and accuracy.
[0047] First, based on each knowledge triple in the knowledge graph, a question-answer pair is generated. Each question-answer pair includes the question related to the knowledge triple and its corresponding standard answer. In addition to constructing a single correct answer, multiple distractors are also designed to construct a complete test set with multiple options.
[0048] Then, calculate the edit distance between each entity in the knowledge graph and the head entity in the knowledge triple. The edit distance refers to the minimum number of single-character editing operations required to convert one string to another. Single-character editing operations generally include insertion, deletion, and replacement. Replace the head entity in the standard answer with the entity with the shortest edit distance to obtain the first interference item. If there are multiple entities with the shortest edit distance, you can randomly select an entity with the shortest edit distance to replace the head entity in the standard answer to form the first interference item, or you can replace the head entity in the standard answer with these multiple entities with the shortest edit distance to form multiple first interference items.
[0049] At the same time, an advanced large-scale model is used to generate multiple candidate items that are easily confused with the standard answer. The edit distance between each candidate and the standard answer is then calculated, and a preset number of candidates with the shortest edit distance are selected. Based on the weight of the edit distance, the preset number of candidates are randomly sampled to obtain multiple target candidates, which are used as the second interference items.
[0050] It is understandable that the first interference item is designed to test the large language model's ability to identify semantic differences; while the second interference item is designed to increase the complexity and challenge of the problem and further test the large language model's deep understanding and reasoning ability.
[0051] The standard answer, the first distractor, and the second distractor are combined into a multiple-choice question. Furthermore, to eliminate the potential impact of option order on the large language model's judgment, these distractors are randomly and unorderedly arranged and appended to the question in a standard (A), (B), (C), (D), ... format. This ensures fairness and objectivity in the evaluation process, aiming to improve the accuracy and efficiency of large language model performance evaluation in the context of open-ended question answering.
[0052] After constructing a series of high-quality multiple-choice questions, each is accurately fed into a pre-defined large language model, which then outputs the corresponding question-answering results. The pre-defined large language models were selected for their specific capabilities, particularly those that demonstrate excellent instruction-following abilities and are effective at handling multiple-choice challenges.
[0053] Given the tendency of large language models to output lengthy question-and-answer results, regular expression technology is used to accurately extract the pre-set large language model's explicit response to each option within the lengthy text. If the extraction is unsuccessful or the extracted content does not match the standard answer, the model's response is considered incorrect. The knowledge triples corresponding to the multiple-choice question are then considered unknown knowledge to the pre-set large language model and serve as a benchmark for evaluating model performance.
[0054] By deeply analyzing the answers of the preset large language model, we can clearly define the knowledge boundaries of the current model in knowledge-intensive fields and clearly distinguish the knowledge sets that have been mastered by the preset large language model. and the unknown set of knowledge yet to be discovered .
[0055] The preset large language model is trained through the unknown knowledge set that has yet to be discovered to obtain an initial large language model that integrates knowledge in unknown fields.
[0056] The embodiments of the present invention effectively explore the knowledge shortcomings of the preset large language model. The initial large language model formed by targeted training can better cover the knowledge graph knowledge and enhance the accuracy and reliability of the model in related knowledge fields.
[0057] Afterwards, in order to continuously expand the knowledge boundaries of the initial large language model, a routing strategy is designed to drive new knowledge integration modules to obtain the target large language model. The purpose is to cleverly integrate knowledge that has not yet been mastered into its knowledge system, while ensuring that this process does not weaken the model's performance in the existing knowledge field.
[0058] In some embodiments, the output of any of the target Transformer layers is determined by: Using the current layer knowledge adapter, the input of the FFN in the current target Transformer layer is fused with the output of the previous layer knowledge adapter to obtain a fusion matrix; Using the current layer knowledge adapter, the fusion matrix is subjected to nonlinear transformation to obtain a nonlinear transformation result; Using the current layer knowledge router, a routing value is generated according to the knowledge mastery of the new knowledge by the input of the FFN; the routing value is used to regulate the strength of knowledge fusion; multiplying the routing value, the nonlinear transformation result and the routing value to obtain a weighted result; The weighted result is fused with the output of the FFN to obtain the output of the current target Transformer layer.
[0059] Specifically, the routing strategy driven knowledge integration focuses on solving the efficiency bottleneck and knowledge forgetting problems in the training process of large language models. It innovatively designs a new architecture that integrates intelligent routing strategies and knowledge adapters, so that the newly added knowledge can be efficiently stored in the parallel knowledge adapter module while keeping the original parameters of the model stable. In addition, a relationship classification task based on entity knowledge fragments is designed to further improve the generalization ability of the language representation of large language models. The overall diagram is as follows Figure 3 As shown, Figure 3 It is an overall schematic diagram of the routing strategy driven knowledge integration provided by the present invention.
[0060] To optimize parameter utilization, a parallel knowledge adapter module was introduced as an expansion unit for knowledge learning. This design enables effective acquisition and absorption of new knowledge while preserving the original model parameters. Given the exceptional knowledge storage capabilities of the feedforward neural network (FFN) in language models based on the Transformer architecture, the knowledge adapter was strategically deployed in parallel after the last N Transformer layers in the T-layer structure, and therefore after the last N FFN layers.
[0061] Specifically, for each selected knowledge adapter layer ,in , the input of the FFN layer Output of the previous layer knowledge adapter Intelligent combination is performed to promote the smooth transfer and deep integration of knowledge in the model, as shown in the following formula: ; Where m is the length of the input sequence and d is the hidden dimension. If the current layer knowledge adapter is the first layer, the output of the previous layer knowledge adapter is Set to a vector of all zeros.
[0062] Knowledge Adapter Layer Utilization The downward projection of Transformed into bottleneck dimension The specified lower dimensional space facilitates the learning of new patterns with minimal additional space. The second is the nonlinear activation function , and apply upward projection for: .
[0063] Then, the present invention directly and effectively integrates the output of the current layer knowledge adapter with the original output of the FFN. The process is briefly described as follows: .
[0064] Then, is passed to the next Transformer attention layer as the input of the next Transformer layer. You can also randomly select some Directly enter the final linear layer and softmax function layer for processing.
[0065] This example integrates knowledge starting from the bottom layer of the Transformer. This strategy can effectively improve overall knowledge integration performance. This is because the top layer of the Transformer excels at calibrating and refining knowledge, while the bottom layer naturally has a stronger ability to absorb and integrate novel information. Specifically, the bottom layer focuses more on capturing fine-grained knowledge, while the top layer tends to process more abstract, coarse-grained knowledge representations. Therefore, integrating new knowledge from the bottom layer provides a more solid foundation for the model, allowing new knowledge to be integrated more smoothly and effectively.
[0066] In order to ensure that the added bypass knowledge adapter module does not confuse the large language model with its existing knowledge, a routing-oriented strategy is introduced to more effectively inject knowledge-intensive task domain knowledge from the knowledge adapter into the large language model. Intuitively, for a given question, the routing strategy first evaluates whether the large language model already has the knowledge reserves required to answer the question. If it is determined to be unknown, the routing strategy will intelligently enhance the The degree of fusion of knowledge in the language model provides the necessary supplementary information for the large language model to enhance its answering ability; on the contrary, if the large language model has mastered the relevant knowledge, the corresponding reduction The design of routing strategy has a solid theoretical foundation, considering that checking the internal state of a large language model can determine whether it is aware of the current problem. Figure 4 As shown, Figure 4 It is a detailed schematic diagram of the routing strategy-guided knowledge adapter provided by the present invention.
[0067] In specific implementation, a routing value is extracted from the input of the FFN layer. This value serves as a key indicator for regulating the strength of knowledge fusion. Its calculation method is as follows: ; in, Indicates that the routing module is implemented as a The activation function is a multi-layer perceptron, and the Mean function averages the vector along the length of the sequence. This allows the routing value Mapping to Range , indicating that the large language model is based on its FFN layers Therefore, the routing strategy helps the large language model learn new knowledge without forgetting what it already knows.
[0068] However, if the routing policy only encounters new knowledge during fine-tuning, it will have difficulty recognizing existing knowledge. To address this issue, the fine-tuning dataset also includes a small number of examples representing knowledge already mastered by the large language model. Before fine-tuning, the routing policy is first pre-trained on the binary injection task using a balanced mixture of known and unknown examples. The loss function for the routing policy is a binary cross-entropy loss function, as shown below: ; in, is a knowledge sample, routing label It is 1 for new knowledge and 0 for previously acquired knowledge.
[0069] Finally, we can obtain an additively filtered adapter vector, which is combined with the original FFN output, which can selectively incorporate knowledge from the adapter into the fixed base model as follows: .
[0070] The embodiments of the present invention achieve precise control and effective integration of new knowledge, improve the model's ability to absorb and process new knowledge, enable the model to use new knowledge more flexibly and accurately in complex task processing, and enhance the performance and adaptability of the model.
[0071] In some embodiments, the target large language model is fine-tuned by: By optimizing the first objective function of the routing optimization phase, the model parameters of the target large language model are adjusted to obtain a first fine-tuning model; By optimizing the second objective function in the question-answer pair training phase, the model parameters of the first fine-tuning model are adjusted to obtain a second fine-tuning model; By optimizing the third objective function of the relationship classification training phase, the model parameters of the second fine-tuning model are adjusted to obtain a fine-tuned target large language model.
[0072] Specifically, we use the unknown knowledge identified during the knowledge boundary detection phase to fine-tune key modules such as the knowledge adapter and routing module. This highly efficient fine-tuning approach to training the knowledge adapter is time-efficient and resource-efficient. This fine-tuning process is meticulously divided into three phases: routing optimization, question-answer pair training, and relation classification (RC) training. These phases are closely linked and collectively drive the achievement of the following objective function: ; in, It is the first objective function, used to optimize model parameters in the routing optimization phase; is the second objective function, used to optimize the model parameters in the QA training phase; It is the third objective function, which is used to optimize model parameters in the relation classification RC training stage.
[0073] In the first stage, a balanced dataset containing known and unknown knowledge samples is used to generate a series of question-based instructions (including questions and prompt words) q and standard answers y. These samples are then applied to the fine-tuning process of the target large language model. By optimizing the first objective function of the routing tuning stage, the model parameters of the target large language model are adjusted to obtain the first fine-tuned model.
[0074] The expression of the first objective function is as follows: ; in, represents the first objective function; Represents the input of FFN; Represents the computational function of a multilayer perceptron; Represents a knowledge sample; represents the routing label of the knowledge sample; BCE() represents the binary cross entropy loss function; E() represents the expected calculation function.
[0075] In the first phase, a balanced dataset containing known and unknown knowledge samples is used to tune and optimize the routing module.
[0076] In the second stage, a series of question-based instructions q and standard answers y are also used to apply these samples to the fine-tuning process of the target large language model. By optimizing the first objective function of the routing tuning stage, after the routing tuning has been completed, the model parameters of the first fine-tuning model are adjusted to obtain the second fine-tuning model.
[0077] During the QA training phase, the parallel knowledge adapter is fine-tuned by leveraging the unknown knowledge identified during the knowledge detection phase. This process incorporates specific instructions based on the question, ensuring that the training process is closely centered around the actual question. At the same time, the standard answer is set as a gold standard to evaluate and optimize the model's response quality. The QA loss is similar to the traditional training loss used in Transformer-based language models, specifically optimized for instructions within the specific domain of home customer comprehensive query.
[0078] The expression of the second objective function is as follows: ; in, represents the second objective function; Represents a sample of questions, represents the model’s predicted answer, is the standard answer sample; CE() represents the cross entropy loss function; E() represents the expected calculation function.
[0079] Of course, during the QA training phase, the training set can also include some yes / no QA samples to enhance the model's versatility across various question types. For example, the question "Does the set-top box fault code 20402 mean that the set-top box account is in arrears?" has a standard answer of "yes" / "no."
[0080] In the second stage, we focus on using the QA loss function to deeply refine the model, aiming to seamlessly integrate the detected unknown knowledge into the model and enhance its ability to handle complex question-answering tasks.
[0081] In the third stage, in order to improve the versatility of routing strategy driven knowledge integration, a relation classification task is used, which can promote the adapter to understand the relational facts within the text and improve the generalization of the representation learned by the knowledge adapter. This task applies knowledge statement fragment samples containing entities Randomly select knowledge triples in the knowledge graph , and its corresponding knowledge statement fragment sample , the token embedding output of the last layer knowledge adapter The above mentioned entities are average pooled, that is, the head entity and the tail entity are average pooled to obtain their vector representations. and Then, splice and As a relation vector representation .
[0082] To improve the understanding of relational facts, As positive samples and knowledge graphs, in addition to The other relations are taken as negative samples, and InfoNCE loss is subsequently adopted to make the positive samples closer and the negative samples farther away, so as to better distinguish the positive and negative relations. The formula is as follows: ; in, Acts as a temperature hyperparameter; Represents a set of relations; Represents the relationship in knowledge triple sample i; Indicates that the relationship set contains The rest of the relationship; Represents the relationship vector representation obtained by concatenating the head entity vector representation and the tail entity vector representation corresponding to the knowledge triple sample i; and Respectively and Aligned to a uniform dimensional space; E() represents the expected calculation function.
[0083] In addition, the traditional training loss used in the Transformer model is also adopted, and a certain token loss will be caused when splicing between each two Transformer layers: ; Among them, k represents the knowledge statement fragment sample; Represents the knowledge statement fragment predicted by the model; CE() represents the cross entropy loss function; E() represents the expectation calculation function.
[0084] Therefore, the expression of the third objective function is as follows: ; in, Represents the weight corresponding to the relationship classification loss function.
[0085] In the third stage, knowledge statement fragments and knowledge triples were introduced as training sets to further expand and improve the model, enhancing the cross-domain and cross-scenario adaptability of the large language model, so that it can still maintain excellent performance and stability when facing more extensive and complex application scenarios.
[0086] The embodiment of the present invention gradually adjusts the model parameters by sequentially optimizing the objective functions of the three stages: routing optimization, question-answer pair training, and relationship classification training. This phased fine-tuning strategy can specifically improve the performance of the model in different task links, continuously optimizing the model in routing decisions, question-answer accuracy, and relationship classification capabilities, and ultimately obtaining a fine-tuned target large language model that performs better and is more adaptable in a variety of knowledge-intensive tasks.
[0087] In summary of the above schemes, the present invention mainly focuses on implicitly encoding some representative knowledge of the knowledge graph into the parameters of the large language model, and the knowledge graph does not need to be involved in the reasoning stage. Specifically, a knowledge boundary detection mechanism is first designed to identify the cognitive limitations of the backbone model of the large language model on a specific sampled knowledge graph. For each triple in the knowledge graph, a predefined relationship template is used to generate knowledge statement fragments and multiple-choice questions, and regular expressions are used to extract the answers of the large language model in the multiple-choice questions. The knowledge that the large language model has not mastered is constructed into the unknown knowledge question and answer library. Then, in order to solve the efficiency bottleneck and knowledge forgetting problems in the training process, a domain knowledge adapter architecture is introduced. The architecture has a built-in intelligent routing strategy to ensure that the newly added knowledge can be efficiently stored in the parallel adapter module, while keeping the original parameters of the large language model stable and unchanged, thereby significantly improving the training efficiency and effectively alleviating the knowledge forgetting phenomenon. Finally, to further improve the generalization ability of the model's language representation, a relationship classification task based on entity knowledge fragments is designed. The relationship type is predicted based on the corresponding token embedding output in the adapter's head entity and tail entity knowledge fragments, which helps the adapter understand the relationship facts within the text, thereby giving the model stronger semantic understanding and reasoning capabilities.
[0088] The model reasoning device for knowledge-intensive tasks provided by the present invention is described below. The model reasoning device for knowledge-intensive tasks described below and the model reasoning method for knowledge-intensive tasks described above can refer to each other.
[0089] Reference Figure 5 , Figure 5 It is a structural diagram of the model reasoning device for knowledge-intensive tasks provided by the present invention.
[0090] The model reasoning device for knowledge-intensive tasks includes: The acquisition module 510 is used to acquire input text; the input text is a question raised based on a specific knowledge-intensive task.
[0091] A model inference module 520 is configured to input the input text into a target large language model and obtain a question-answering result output by the target large language model; Among them, the target large language model is a knowledge adapter and a knowledge router deployed in parallel in N consecutive target Transformer layers in the Transformer architecture of the initial large language model; any of the knowledge adapters is used to integrate new knowledge into the initial large language model while keeping the parameters of the initial large language model unchanged; any of the knowledge routers is used to dynamically adjust the degree to which the new knowledge is integrated into the initial large language model according to the initial large language model's knowledge mastery of the new knowledge; the initial large language model is trained based on the unknown knowledge of the preset large language model in the knowledge graph.
[0092] The model inference device for knowledge-intensive tasks provided by the present invention first performs training based on the unknown knowledge of the large language model preset in the knowledge graph, thereby enhancing the coverage of basic knowledge. Then, the knowledge adapter and the knowledge router are deployed in parallel in the Transformer architecture of the trained initial large language model. The adapter injects new knowledge without changing the parameters of the original model, and the router dynamically adjusts the fusion strength according to the model's mastery of the new knowledge. This not only effectively expands the knowledge boundary of the large language model, but also reduces the risk of conflict between new and old knowledge through the weighted mechanism of knowledge mastery perception. This structural improvement can effectively alleviate the problem of catastrophic forgetting while maintaining the original capabilities of the model, thereby significantly improving the accuracy and adaptability of the target large language model in knowledge-intensive tasks, and is suitable for intelligent question-answering scenarios that require continuous knowledge updating.
[0093] Furthermore, the model reasoning device for knowledge-intensive tasks is also used to: Generate multiple-choice questions; the multiple-choice questions are generated based on knowledge triples in the knowledge graph; Inputting the multiple-choice question into a preset large language model to obtain a question-answering result of the preset large language model; If the question-and-answer result does not conform to the standard answer corresponding to the multiple-choice question, determining that the knowledge triple is unknown knowledge of the preset large language model; Model training is performed based on the unknown knowledge of the preset large language model in the knowledge graph to obtain an initial large language model.
[0094] Furthermore, the model reasoning device for knowledge-intensive tasks is also used to: Generate a question-answer pair based on the knowledge triples in the knowledge graph; the question-answer pair includes a question related to the knowledge triples and its corresponding standard answer; Determine an entity with the shortest edit distance to the head entity in the knowledge triple, and replace the head entity in the standard answer with the entity to obtain a first interference item; Determining a plurality of second distractors having the shortest edit distances to the standard answer; The standard answer, the first interference item and the plurality of second interference items are integrated into multiple choices, and a multiple choice question is generated based on the question and the multiple choices.
[0095] Furthermore, the model reasoning device for knowledge-intensive tasks is also used to: Using the current layer knowledge adapter, the input of the FFN in the current target Transformer layer is fused with the output of the previous layer knowledge adapter to obtain a fusion matrix; Using the current layer knowledge adapter, the fusion matrix is subjected to nonlinear transformation to obtain a nonlinear transformation result; Using the current layer knowledge router, a routing value is generated according to the knowledge mastery of the new knowledge by the input of the FFN; the routing value is used to regulate the strength of knowledge fusion; multiplying the routing value, the nonlinear transformation result and the routing value to obtain a weighted result; The weighted result is fused with the output of the FFN to obtain the output of the current target Transformer layer.
[0096] Furthermore, the model reasoning device for knowledge-intensive tasks is also used to: By optimizing the first objective function of the routing optimization phase, the model parameters of the target large language model are adjusted to obtain a first fine-tuning model; By optimizing the second objective function in the question-answer pair training phase, the model parameters of the first fine-tuning model are adjusted to obtain a second fine-tuning model; By optimizing the third objective function of the relationship classification training phase, the model parameters of the second fine-tuning model are adjusted to obtain a fine-tuned target large language model.
[0097] It should be noted that the model reasoning device for knowledge-intensive tasks provided by the present invention can execute the model reasoning method for knowledge-intensive tasks described in any of the above embodiments during specific operation, which will not be elaborated in this embodiment.
[0098] Figure 6 Schematic diagram of the structure of the electronic device provided by the present invention, such as Figure 6As shown, the electronic device may include: a processor 610 , a communications interface 620 , a memory 630 and a communication bus 640 , wherein the processor 610 , the communications interface 620 and the memory 630 communicate with each other via the communication bus 640 . The processor 610 can call the logic instructions in the memory 630 to execute a model reasoning method for knowledge-intensive tasks, the method including: obtaining input text; the input text is a question raised based on a specific knowledge-intensive task; inputting the input text into the target large language model to obtain a question-answering result output by the target large language model; wherein the target large language model is a Transformer architecture of an initial large language model in which N consecutive target Transformer layers respectively deploy a knowledge adapter and a knowledge router in parallel; any of the knowledge adapters is used to incorporate new knowledge into the initial large language model while keeping the parameters of the initial large language model unchanged; any of the knowledge routers is used to dynamically adjust the degree to which the new knowledge is incorporated into the initial large language model according to the initial large language model's knowledge mastery of the new knowledge; the initial large language model is trained based on the unknown knowledge of the preset large language model in the knowledge graph.
[0099] Furthermore, the logic instructions in the aforementioned memory 630 can be implemented as software functional units and, when sold or used as independent products, can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, or the portion that contributes to the prior art, or a portion of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing a computer device (which can be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in various embodiments of the present invention. The aforementioned storage medium includes various media capable of storing program code, such as a USB flash drive, a mobile hard drive, a read-only memory (ROM), a random access memory (RAM), a magnetic disk, or an optical disk.
[0100] On the other hand, the present invention also provides a computer program product, which includes a computer program stored on a non-transitory computer-readable storage medium, and the computer program includes program instructions. When the program instructions are executed by a computer, the computer can execute the model reasoning method for knowledge-intensive tasks provided by the above-mentioned embodiments, the method including: obtaining input text; the input text is a question raised based on a specific knowledge-intensive task; inputting the input text into a target large language model to obtain a question-and-answer result output by the target large language model; wherein the target large language model is a Transformer architecture of an initial large language model in which N consecutive target Transformer layers are respectively deployed in parallel with a knowledge adapter and a knowledge router; any of the knowledge adapters is used to incorporate new knowledge into the initial large language model while keeping the parameters of the initial large language model unchanged; any of the knowledge routers is used to dynamically adjust the degree to which the new knowledge is incorporated into the initial large language model according to the knowledge mastery of the new knowledge by the initial large language model; the initial large language model is trained based on the unknown knowledge of the preset large language model in the knowledge graph.
[0101] On the other hand, the present invention also provides a non-transitory computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, is implemented to execute the model reasoning method for knowledge-intensive tasks provided in the above-mentioned embodiments, the method comprising: obtaining input text; the input text is a question raised based on a specific knowledge-intensive task; inputting the input text into a target large language model to obtain a question-and-answer result output by the target large language model; wherein the target large language model is a Transformer architecture of an initial large language model in which N consecutive target Transformer layers respectively deploy a knowledge adapter and a knowledge router in parallel; any of the knowledge adapters is used to incorporate new knowledge into the initial large language model while keeping the parameters of the initial large language model unchanged; any of the knowledge routers is used to dynamically adjust the degree to which the new knowledge is incorporated into the initial large language model according to the initial large language model's knowledge mastery of the new knowledge; the initial large language model is trained based on the unknown knowledge of the preset large language model in the knowledge graph.
[0102] The device embodiments described above are merely illustrative, wherein the units described as separate components may or may not be physically separate, and the components shown as units may or may not be physical units. Some or all of the modules may be selected according to actual needs to achieve the purpose of the present embodiment. Those skilled in the art can understand and implement the present invention without inventive effort.
[0103] Through the above description of the embodiments, those skilled in the art will clearly understand that each embodiment can be implemented using software plus a necessary general-purpose hardware platform, or of course, hardware. Based on this understanding, the essence of the above technical solution, or the portion that contributes to the prior art, can be embodied in the form of a software product. This computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, a magnetic disk, or an optical disk, and includes a number of instructions for causing a computer device (such as a personal computer, server, or network device) to execute the methods described in each embodiment or certain portions of the embodiments.
[0104] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or make equivalent replacements for some of the technical features therein. However, these modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the various embodiments of the present invention.
Claims
1. A model reasoning method for knowledge-intensive tasks, characterized by: The model reasoning method for knowledge-intensive tasks includes: Obtaining input text; the input text is a question raised based on a specific knowledge-intensive task; Inputting the input text into a target large language model to obtain a question-answering result output by the target large language model; Among them, the target large language model is a knowledge adapter and a knowledge router deployed in parallel in N consecutive target Transformer layers in the Transformer architecture of the initial large language model; any of the knowledge adapters is used to integrate new knowledge into the initial large language model while keeping the parameters of the initial large language model unchanged; any of the knowledge routers is used to dynamically adjust the degree to which the new knowledge is integrated into the initial large language model according to the initial large language model's knowledge mastery of the new knowledge; the initial large language model is trained based on the unknown knowledge of the preset large language model in the knowledge graph.
2. The model reasoning method for knowledge-intensive tasks according to claim 1, characterized in that: The initial large language model is determined in the following manner: Generate multiple-choice questions; the multiple-choice questions are generated based on knowledge triples in the knowledge graph; Inputting the multiple-choice question into a preset large language model to obtain a question-answering result of the preset large language model; If the question-and-answer result does not conform to the standard answer corresponding to the multiple-choice question, determining that the knowledge triple is unknown knowledge of the preset large language model; Model training is performed based on the unknown knowledge of the preset large language model in the knowledge graph to obtain an initial large language model.
3. The model reasoning method for knowledge-intensive tasks according to claim 2, characterized in that: The generating of multiple choice questions includes: Generate a question-answer pair based on the knowledge triples in the knowledge graph; the question-answer pair includes a question related to the knowledge triples and its corresponding standard answer; Determine an entity with the shortest edit distance to the head entity in the knowledge triple, and replace the head entity in the standard answer with the entity to obtain a first interference item; Determining a plurality of second distractors having the shortest edit distances to the standard answer; The standard answer, the first interference item and the plurality of second interference items are integrated into multiple choices, and a multiple choice question is generated based on the question and the multiple choices.
4. The model reasoning method for knowledge-intensive tasks according to claim 3, characterized in that: The output of any target Transformer layer is determined as follows: Using the current layer knowledge adapter, the input of the FFN in the current target Transformer layer is fused with the output of the previous layer knowledge adapter to obtain a fusion matrix; Using the current layer knowledge adapter, the fusion matrix is subjected to nonlinear transformation to obtain a nonlinear transformation result; Using the current layer knowledge router, a routing value is generated according to the knowledge mastery of the new knowledge by the input of the FFN; The routing value is used to regulate the strength of knowledge fusion; multiplying the routing value, the nonlinear transformation result and the routing value to obtain a weighted result; The weighted result is fused with the output of the FFN to obtain the output of the current target Transformer layer.
5. The model reasoning method for knowledge-intensive tasks according to claim 3, characterized in that: The target large language model is fine-tuned in the following way: By optimizing the first objective function of the routing optimization phase, the model parameters of the target large language model are adjusted to obtain a first fine-tuning model; By optimizing the second objective function in the question-answer pair training phase, the model parameters of the first fine-tuning model are adjusted to obtain a second fine-tuning model; By optimizing the third objective function of the relationship classification training phase, the model parameters of the second fine-tuning model are adjusted to obtain a fine-tuned target large language model.
6. The model reasoning method for knowledge-intensive tasks according to claim 5, characterized in that: The expression of the first objective function is as follows: ; in, represents the first objective function; Represents the input of FFN; Represents the computational function of a multilayer perceptron; Represents a knowledge sample; represents the routing label of the knowledge sample; BCE() represents the binary cross entropy loss function; E() represents the expected calculation function.
7. The model reasoning method for knowledge-intensive tasks according to claim 5, characterized in that: The expression of the second objective function is as follows: ; in, represents the second objective function; Represents a sample of questions, represents the model’s predicted answer, is the standard answer sample; CE() represents the cross entropy loss function; E() represents the expected calculation function.
8. The model reasoning method for knowledge-intensive tasks according to claim 5, characterized in that: The expression of the third objective function is as follows: ; in, represents the third objective function; represents the loss caused by relation classification, Indicates its corresponding weight; Token loss representing model inference; ; in, represents the temperature hyperparameter; Represents a set of relations; Represents the relationship in knowledge triple sample i; Indicates that the relationship set contains The rest of the relationship; Represents the relationship vector representation obtained by concatenating the head entity vector representation and the tail entity vector representation corresponding to the knowledge triple sample i; and Respectively and Aligned into a unified dimensional space; E() represents the expected calculation function; k represents the knowledge statement fragment; ; Among them, k represents the knowledge statement fragment sample; Represents the knowledge statement fragment predicted by the model; CE() represents the cross entropy loss function; E() represents the expectation calculation function.
9. A model reasoning device for knowledge-intensive tasks, characterized in that: include: Acquisition module, used to obtain input text; The input text is a question raised based on a specific knowledge-intensive task; A model inference module, configured to input the input text into a target large language model and obtain a question-answering result output by the target large language model; The target large language model is a Transformer architecture of an initial large language model in which N consecutive target Transformer layers are respectively deployed in parallel with a knowledge adapter and a knowledge router; any of the knowledge adapters is used to incorporate new knowledge into the initial large language model while keeping the parameters of the initial large language model unchanged; Any of the knowledge routers is used to dynamically adjust the degree to which the new knowledge is integrated into the initial large language model based on the initial large language model's knowledge of the new knowledge; the initial large language model is trained based on the unknown knowledge of the preset large language model in the knowledge graph.
10. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein: When the processor executes the computer program, the steps of the model reasoning method for knowledge-intensive tasks as described in any one of claims 1 to 8 are implemented.
11. A non-transitory computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps of the model reasoning method for knowledge-intensive tasks as claimed in any one of claims 1 to 8 are implemented.
12. A computer program product, comprising a computer program, characterized in that: When the computer program is executed by a processor, the steps of the model reasoning method for knowledge-intensive tasks as claimed in any one of claims 1 to 8 are implemented.