Large language model knowledge editing method and interactive method based on knowledge representation decoupling
By decoupling and coupling target knowledge, the problem of the impact of large language models on fine-grained irrelevant knowledge during the update process is solved, and knowledge updates with high accuracy and high retention rates are achieved, which improves the reliability of the model and the credibility of the output.
Patent Information
- Application Number
- CN202510607173.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-13
- Publication Date
- 2025-09-26
- Estimated Expiration
- 2045-05-13
AI Technical Summary
When updating target knowledge, existing large language models often fail to effectively protect fine-grained irrelevant knowledge that is semantically related to the target knowledge but logically independent, resulting in inaccurate or inconsistent text content generated by the model. Traditional editing methods only focus on the invariance of coarse-grained irrelevant knowledge during the editing process, affecting other knowledge.
By designing knowledge decouplers and couplers, the target knowledge is decoupled from the highly coupled knowledge vector and updated to new knowledge, while constraining the target knowledge-irrelevant representation to remain unchanged. Methods such as contrastive learning loss, knowledge constraint loss and representation reconstruction loss are used to ensure the effectiveness and accuracy of the decoupling process.
It achieves the protection of fine-grained irrelevant knowledge during the knowledge updating process, improves the knowledge updating accuracy and fine-grained knowledge retention rate of large language models, and enhances the reliability of the model and the credibility of the output content.
Smart Images

Figure CN120181205B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of artificial intelligence technology, and in particular relates to a large language model knowledge editing method and an interactive method based on knowledge representation decoupling. Background Art
[0002] The statements in this section merely provide background information related to the present invention and do not necessarily constitute prior art.
[0003] Large language models have been a significant advancement in natural language processing in recent years. Pre-trained on massive amounts of text data, these models are capable of generating high-quality text content and are widely used in fields such as text generation, machine translation, and question-answering systems. However, due to the dynamic nature of world knowledge and the inevitable noise in model training data, the knowledge embedded in large language models can be inaccurate or outdated. To address these issues, knowledge editing technology has emerged. It aims to efficiently and accurately update the knowledge embedded in large language models, ensuring that the models can generate more accurate text content that better reflects real-world information, thereby improving the accuracy and reliability of the models.
[0004] Traditional knowledge editing methods can be mainly divided into three categories: memory-based methods, optimization-based methods, and locate-then-edit methods. Memory-based methods use external memory modules or specific architectures to store and manage new knowledge. These methods do not modify the internal parameters of the model, but instead update knowledge by adding additional neurons or using a retrieval mechanism based on dynamic query and knowledge injection. Optimization-based methods integrate new knowledge by modifying model parameters without relying on external storage. These methods often introduce constraints to prevent overfitting problems during the optimization process, or design super-network architectures to achieve efficient knowledge implantation during parameter tuning. The locate-then-edit method analyzes the parameter distribution characteristics of the large language model, accurately identifies the key parameter set corresponding to the target knowledge, and then implements targeted parameter correction to achieve knowledge updating.
[0005] However, existing methods often fail to effectively protect fine-grained irrelevant knowledge that is semantically related to the target knowledge but logically independent when updating the target knowledge. This problem may lead to inaccurate or inconsistent text content generated by the model in practical applications, and existing methods may mistakenly change other knowledge related to the entity "Twitter". Since this incorrectly changed knowledge shares the same entity with the target knowledge, it has a high semantic relevance to the target knowledge. Traditional editing methods only focus on the invariance of coarse-grained irrelevant knowledge during the editing process, and this knowledge has a low semantic relevance to the edited knowledge. The root cause of this problem is that existing methods usually represent knowledge as highly coupled internal model vectors during the editing process, without clearly distinguishing different knowledge. This coupled representation makes it inevitable to affect other knowledge when updating the target knowledge. Summary of the Invention
[0006] In order to solve the technical problems existing in the above-mentioned background technology, the present invention provides a large language model knowledge editing method and an interactive method based on knowledge representation decoupling. By decoupling the target knowledge from the highly coupled knowledge vector and updating the target knowledge to new knowledge, while avoiding affecting other fine-grained irrelevant knowledge, the impact of the knowledge editing process on other knowledge and capabilities of the large language model is reduced, and the reliability of the large language model is enhanced.
[0007] In order to achieve the above object, the present invention adopts the following technical solutions:
[0008] A first aspect of the present invention provides a large language model knowledge editing method based on knowledge representation decoupling, which includes:
[0009] Acquire knowledge representation and new knowledge to be injected;
[0010] For knowledge representation, the knowledge decoupler is used to obtain target knowledge related representation and target knowledge irrelevant representation;
[0011] The new knowledge to be injected is encoded into the target knowledge-related representation, and then the target knowledge-related representation and the target knowledge-irrelevant representation are coupled into a new knowledge representation through the knowledge coupler;
[0012] Based on the new knowledge representation, the parameters of the large language model are updated, and the new knowledge representation is injected into the large language model. In the process of updating the parameters of the large language model, the target knowledge-irrelevant representation is constrained to remain unchanged.
[0013] Furthermore, the new knowledge representation is expressed as:
[0014] ;
[0015] ;
[0016] in, 、 and Respectively represent the subject, relation and object in the target knowledge triple, Indicates the use of and Constructed text prompt words, Represents the target knowledge related representation, represents the target knowledge-independent representation, In order to obtain the optimizable variables of the new knowledge representation, the knowledge coupler Rec represents the target knowledge Representation independent of target knowledge Constructing new knowledge representation , Represents the knowledge representation of the knowledge decoupler input during the forward propagation of the large language model Replace with new knowledge representation , Represents the output when the knowledge representation of the knowledge decoupler input is replaced by the new knowledge representation probability.
[0017] Furthermore, the knowledge decoupler and the knowledge coupler adopt a contrastive learning loss function during training:
[0018] ;
[0019] in, represents the information noise contrast estimation loss function, Represents the target knowledge related representation, represents the target knowledge-independent representation, represents the knowledge representation of the knowledge decoupler input, and Respectively and The negative sample set.
[0020] Furthermore, the knowledge decoupler and the knowledge coupler adopt knowledge constraint loss during training:
[0021] ;
[0022] Among them, the knowledge constraint loss converts the target knowledge The information is encoded into the target knowledge related representation , fine-grained irrelevant knowledge Information encoded into irrelevant representations middle, 、 and Respectively represent the subject, relation and object in the target knowledge triple, and They represent the replacement of the knowledge representation of the knowledge decoupler input with the target knowledge related representation during the forward propagation of the large language model. Representation independent of target knowledge , Indicates that the input of the large language model is and Output when constructing text prompt words The probability of Indicates that the input of the large language model is and Output when constructing text prompt words probability.
[0023] Furthermore, the knowledge decoupler and the knowledge coupler adopt representation reconstruction loss during training:
[0024] ;
[0025] in, represents the knowledge representation of the knowledge decoupler input, represents the reconstructed representation obtained by the knowledge decoupler, Calculates the Frobenius norm.
[0026] Furthermore, the knowledge decoupler utilizes the relational representation to decouple the knowledge representation into a target knowledge-related representation and a target knowledge-irrelevant representation.
[0027] Furthermore, parameter update constraints are adopted during the update of the large language model parameters:
[0028] ;
[0029] in, is the Frobenius norm calculation, is the large language model parameter matrix to be updated, The new knowledge is saved in the form of key-value pairs in the parameter matrix, and the new knowledge is injected into the large language model in the form of key-value pairs.
[0030] Furthermore, fine-grained irrelevant knowledge is used to maintain constraints during the update of the large language model parameters:
[0031] ;
[0032] in, is the Frobenius norm calculation, is the large language model parameter matrix to be updated, represents the knowledge representation of the knowledge decoupler input, To preserve the key-value pairs of new knowledge in the parameter matrix, we use linear transformations that are irrelevant to the representation. Constrain the parameter update process to ensure the consistency of fine-grained irrelevant knowledge before and after parameter update.
[0033] Furthermore, during the update of the large language model parameters, coarse-grained irrelevant knowledge is used to maintain constraints:
[0034] ;
[0035] in, is the Frobenius norm calculation, is the large language model parameter matrix to be updated, ( ) is the key-value pair representation of coarse-grained irrelevant knowledge, which constrains the consistency of coarse-grained irrelevant knowledge before and after parameter update.
[0036] A second aspect of the present invention provides an interaction method, comprising:
[0037] Get user questions;
[0038] Based on the user question, the answer is obtained through the large language model obtained by the large language model knowledge editing method based on knowledge representation decoupling as described in the first aspect.
[0039] Compared with the prior art, the present invention has the following beneficial effects:
[0040] The present invention decouples the target knowledge from the highly coupled knowledge vector and updates the target knowledge to new knowledge while avoiding affecting other fine-grained irrelevant knowledge, thereby reducing the impact of the knowledge editing process on other knowledge and capabilities of the large model and enhancing the reliability of the large model.
[0041] Compared with existing knowledge editing methods, the present invention achieves fine-grained separation and multi-level protection of knowledge representation, and thus the method of this embodiment achieves significant improvements in core indicators such as knowledge update accuracy and fine-grained knowledge retention rate.
[0042] This invention can be widely used in dynamic knowledge-intensive scenarios such as financial information updates and medical knowledge base maintenance. It can effectively maintain the knowledge consistency of large language models, enhance the credibility of output content, provide technical guarantees for model reliability in high-risk areas, and thus enhance the practical application value of artificial intelligence systems. BRIEF DESCRIPTION OF THE DRAWINGS
[0043] The accompanying drawings, which constitute a part of the present invention, are used to provide a further understanding of the present invention. The exemplary embodiments of the present invention and their descriptions are used to explain the present invention and do not constitute improper limitations on the present invention.
[0044] Figure 1This is a workflow diagram of a knowledge editing module based on decoupling according to the first embodiment of the present invention;
[0045] Figure 2 This is a workflow diagram of the knowledge decoupling module in the first embodiment of the present invention. DETAILED DESCRIPTION
[0046] To make the objectives, technical solutions and advantages of the embodiments of the present invention more clear, the technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention.
[0047] It should be noted that the following detailed descriptions are illustrative and intended to provide further explanation of the present invention. Unless otherwise specified, all technical and scientific terms used herein have the same meaning as commonly understood by those skilled in the art to which the present invention belongs.
[0048] Example 1
[0049] This embodiment provides a large language model knowledge editing method based on knowledge representation decoupling.
[0050] The large language model knowledge editing method based on knowledge representation decoupling provided in this embodiment aims to solve the problem of fine-grained editing of embedded knowledge in large language models, and designs a knowledge decoupling module and a knowledge editing module based on decoupling.
[0051] like Figure 2 As shown in the figure, the knowledge decoupling module consists of a knowledge decoupler and a knowledge coupler. The knowledge decoupler decouples the target knowledge-related representation and the target knowledge-irrelevant representation from the entity to be edited, and the knowledge coupler recouples the two representations and subsequently constrains the effectiveness of the decoupling through a loss function.
[0052] like Figure 1 As shown in Figure 3, the knowledge editing module based on decoupled representation obtains new knowledge representation by updating the relevant representation of target knowledge and recoupling it with irrelevant representation, and keeps the fine-grained irrelevant knowledge unchanged when the parameters are updated.
[0053] (1) Knowledge decoupler and knowledge coupler workflow.
[0054] The knowledge decoupling module consists of a knowledge decoupler (referred to as decoupler) and a knowledge coupler (referred to as coupler). The decoupler is used to decouple the target knowledge related representation and target knowledge irrelevant representation from the entity to be edited. For example, the two pieces of knowledge "Twitter's CEO is Linda Yaccarino" and "Twitter's headquarters is in San Francisco" both involve the entity "Twitter", but due to the semantic similarity, previous editing methods cannot maintain the integrity of the knowledge. To solve this problem, the knowledge decoupler uses knowledge representation to and relational representation As input, the relational representation is used to decouple the edited knowledge into a representation related to the target knowledge. Representations that are irrelevant to the target knowledge (irrelevant knowledge) :
[0055] ;
[0056] ;
[0057] in, represents the relevant knowledge decoupler, represents the irrelevant knowledge decoupler, represents the linear transformation during the decoupling process, is a nonlinear activation function. Both the knowledge representation and the relationship representation come from the internal state vector representation of the large language model when processing input text (e.g., "The CEO of Twitter is"). The knowledge representation is the vector representation of the edited entity (e.g., "Twitter"), and the relationship representation is the vector representation of the edited relationship (e.g., "The CEO is").
[0058] Then, the knowledge coupler uses the decoupled representation to complete the representation reconstruction. It is used to ensure that no excessive information is lost during the decoupling process when constructing the loss function later:
[0059] .
[0060] in, represents the linear transformation in the coupling process, Represents a knowledge coupler.
[0061] (2) Knowledge decoupler and knowledge coupler training process.
[0062] (201) Decoupling loss based on contrastive learning.
[0063] In order to enhance the semantic information of the decoupled representation, ensure that the decoupled representation and Encoded knowledge representation The effective information in the data is obtained by introducing the mutual information maximization criterion and maximizing the decoupling representation through the information noise contrast estimation (InfoNCE) loss function. and and knowledge representation In addition, in order to improve the semantic discrimination of the decoupled representation, and They serve as negative samples in the contrastive learning loss function to minimize their mutual information.
[0064] To this end, this embodiment constructs the following contrastive learning loss function:
[0065] ;
[0066] ;
[0067] in, Indicates N with the subject A collection of representations constructed from different entities, and Respectively and negative sample set; in InfoNCE, the loss function maximizes the anchor sample S and the positive sample Mutual information of the sample, minimizing With negative samples The mutual information of Representing a collection A sample of is the vector similarity measurement function, is the temperature parameter.
[0068] (202) Knowledge Constraint Loss.
[0069] On this basis, in order to meet the fine-grained control requirements of knowledge update, in order to ensure that the decoupled representation effectively encodes the target knowledge and fine-grained knowledge The semantic information of , imposes the embedded knowledge representation constraints and target knowledge representation constraints on it. Respectively represent the subject, relation and object in the target knowledge triple, Represents a knowledge triple with the same subject as the target knowledge. Specifically, before training, Wikipedia is used to search and construct the text of fine-grained unrelated knowledge for each target knowledge, and input it into the large language model together with the target knowledge. The decoupled target knowledge representation and embedded knowledge representation are used to replace the original knowledge representation of a specific layer, and the object corresponding to each knowledge is output, that is, and .
[0070] Based on this, the following knowledge constraint loss is constructed:
[0071] .
[0072] in, It represents the process of replacing the original knowledge representation in the computation process with the specified representation. Specifically, and It represents the probability of outputting the corresponding object after replacing the original knowledge representation with the target knowledge representation and the embedded knowledge representation respectively.
[0073] (203) Characterizing the reconstruction loss.
[0074] Finally, in order to ensure the information integrity of the decoupling and coupling process, the mean square error loss function is used to establish and minimize the subject original knowledge representation and reconstructing knowledge representation The error measure between them is constructed as follows:
[0075] ;
[0076] Finally, as an end-to-end training process, the knowledge decoupler and knowledge coupler are trained using the following loss function:
[0077] ;
[0078] in, and is the weight coefficient of the loss function.
[0079] (3) Knowledge editing module based on decoupling.
[0080] In order to achieve fine-grained editing of large language model knowledge, the knowledge update algorithm based on decoupling decouples the knowledge representation into target knowledge representation and irrelevant knowledge representation, ensuring that the update operation effectively injects new knowledge into the model while avoiding the impact on fine-grained irrelevant knowledge. Specifically, by utilizing the target knowledge representation and irrelevant knowledge representation ,This embodiment constructs the following three parameter update constraints to ensure that the knowledge to be updated can be effectively injected into the large language model, while ensuring that fine-grained irrelevant knowledge is not damaged.
[0081] (301) Injection of new knowledge.
[0082] To achieve the goal of knowledge In order to accurately inject knowledge representation, this embodiment first optimizes the target knowledge representation after decoupling the knowledge representation, and encodes the new knowledge to be injected into the target knowledge representation. In this process, the representation of fine-grained embedded knowledge (i.e., fine-grained irrelevant knowledge) remains unchanged, and then the two are coupled into a new knowledge representation through the knowledge coupler. :
[0083] ;
[0084] ;
[0085] in, Represents the target knowledge related representation, represents the target knowledge-independent representation. In this process, To ensure that new knowledge is obtained by optimizing the change amount of target knowledge related representation After the target knowledge related representation and irrelevant representation are reconstructed into new knowledge representation ; 、 and Respectively represent the subject, relation and object in the target knowledge triple, Indicates the use of and Constructed text prompt words, Represents the target knowledge related representation, Represents target knowledge-independent representation, Rec represents target knowledge-related representation Representation independent of target knowledge Constructing new knowledge representation , Represents the knowledge representation of the knowledge decoupler input during the forward propagation of the large language model Replace with new knowledge representation , Indicates that the large language model has an input of and Output when constructing text prompt words The probability of Represents the output when the knowledge representation of the knowledge decoupler input is replaced by the new knowledge representation The probability of Indicates optimizing the large language model to output The loss function is In order to construct a new knowledge representation, we first add learnable variables to the target knowledge representation. , and then use Rec to construct the target knowledge related representation and irrelevant representation into the initial new knowledge representation, and then use the original knowledge representation to replace the new knowledge representation during the forward propagation of the large language model, and optimize the variable through the target loss function , using the final optimized Compute new knowledge representations.
[0086] On this basis, the following parameter update constraints are constructed to inject information from the new knowledge representation into the parameters of the large language model:
[0087] ;
[0088] in, is the large language model parameter matrix to be updated, The new knowledge is stored in the parameter matrix in the form of key-value pairs. This formula injects new knowledge into the model in the form of key-value pairs.
[0089] (302) Fine-grained irrelevant knowledge retention.
[0090] In order to keep fine-grained irrelevant knowledge unaffected, constrain target irrelevant knowledge Keeping it unchanged, construct the following optimization objectives:
[0091] .
[0092] The target-independent knowledge obtained by this constraint decoupling has the smallest change before and after editing. In order to reduce the amount of calculation, the nonlinear activation function is ignored and the following constraint function is obtained:
[0093] ;
[0094] in, is the large language model parameter matrix to be updated, represents the knowledge representation of the knowledge decoupler input. This formula uses the linear transformation of irrelevant representation Constrain the parameter update process to ensure the consistency of fine-grained irrelevant knowledge before and after parameter update.
[0095] (303) Coarse-grained irrelevant knowledge retention.
[0096] Current research shows that knowledge representations obtained by sampling text from a corpus can approximate coarse-grained embedded knowledge, which can, to a certain extent, ensure that the general capabilities or knowledge of large language models are not affected. Therefore, this embodiment adopts the following optimization objectives to constrain the impact of update operations on coarse-grained embedded knowledge:
[0097] ;
[0098] in, is the large language model parameter matrix to be updated, ( ) is the key-value pair representation of coarse-grained irrelevant knowledge. This formula constrains the consistency of coarse-grained irrelevant knowledge before and after parameter update.
[0099] In summary, based on the three knowledge preservation functions mentioned above, we optimize to construct the final :
[0100] .
[0101] In summary, the large language model knowledge editing method based on knowledge representation decoupling provided in this embodiment includes the following steps:
[0102] Step 1: Obtain knowledge representation and new knowledge to be injected;
[0103] Step 2: For knowledge representation, obtain target knowledge related representation and target knowledge irrelevant representation through knowledge decoupler;
[0104] Step 3: Encode the new knowledge to be injected into the target knowledge-related representation, and then couple the target knowledge-related representation and the target knowledge-irrelevant representation into a new knowledge representation through the knowledge coupler;
[0105] Step 4: Based on the new knowledge representation, the parameters of the large language model are updated through the three-point knowledge preservation function, and the new knowledge representation is injected into the large language model. During the process of updating the parameters of the large language model, the target knowledge-independent representation is constrained to remain unchanged.
[0106] The large language model knowledge editing method based on knowledge representation decoupling provided in this embodiment aims to apply the knowledge representation decoupling theory to the field of large language model knowledge updating in a breakthrough manner, and innovatively proposes a knowledge decoupling module and a knowledge editing algorithm based on decoupling.
[0107] The large language model knowledge editing method based on knowledge representation decoupling provided in this embodiment addresses the fine-grained knowledge interference problem caused by knowledge coupling in traditional methods, and achieves fine-grained updates of embedded knowledge in large language models through the technical path of decoupling-reconstruction-knowledge injection. Compared with existing knowledge editing methods, the method of this embodiment has achieved significant improvements in core indicators such as knowledge update accuracy and fine-grained knowledge retention rate due to the realization of fine-grained separation and multi-level protection of knowledge representation. At the same time, the method of this embodiment can be widely used in dynamic knowledge-intensive scenarios such as financial information updates and medical knowledge base maintenance. It can effectively maintain the knowledge consistency of large language models, enhance the credibility of output content, provide technical guarantees for model reliability in high-risk areas, and thus enhance the practical application value of artificial intelligence systems.
[0108] Example 2
[0109] This embodiment provides an interaction method, which specifically includes:
[0110] Get user questions;
[0111] Based on the user question, an answer is obtained through the large language model obtained by the large language model knowledge editing method based on knowledge representation decoupling as described in Example 1.
[0112] The foregoing description is merely a preferred embodiment of the present invention and is not intended to limit the present invention. Those skilled in the art will readily appreciate that various modifications and variations of the present invention are possible. Any modifications, equivalent substitutions, or improvements made within the spirit and principles of the present invention are intended to be within the scope of protection of the present invention.
Claims
1. A large language model knowledge editing method based on knowledge representation decoupling, characterized by: include: Acquire knowledge representation and new knowledge to be injected; For knowledge representation, the knowledge decoupler is used to obtain target knowledge related representation and target knowledge irrelevant representation; The new knowledge to be injected is encoded into the target knowledge-related representation, and then the target knowledge-related representation and the target knowledge-irrelevant representation are coupled into a new knowledge representation through the knowledge coupler; Based on the new knowledge representation, the parameters of the large language model are updated and the new knowledge representation is injected into the large language model. During the update of the parameters of the large language model, the target knowledge-independent representation is constrained to remain unchanged. The new knowledge representation is expressed as: ; ; in, 、 and Respectively represent the subject, relation and object in the target knowledge triple, Indicates the use of and Constructed text prompt words, Represents the target knowledge related representation, represents the target knowledge-independent representation, In order to obtain the optimizable variables of the new knowledge representation, the knowledge coupler Rec represents the target knowledge Representation independent of target knowledge Constructing new knowledge representation , Represents the knowledge representation of the knowledge decoupler input during the forward propagation of the large language model Replace with new knowledge representation , Represents the output when the knowledge representation of the knowledge decoupler input is replaced by the new knowledge representation probability.
2. The large language model knowledge editing method based on knowledge representation decoupling according to claim 1 is characterized in that: The knowledge decoupler and knowledge coupler adopt contrastive learning loss function during training: ; in, represents the information noise contrast estimation loss function, Represents the target knowledge related representation, represents the target knowledge-independent representation, represents the knowledge representation of the knowledge decoupler input, and Respectively and The negative sample set.
3. The large language model knowledge editing method based on knowledge representation decoupling according to claim 1 is characterized in that: The knowledge decoupler and knowledge coupler adopt knowledge constraint loss during training: ; Among them, the knowledge constraint loss converts the target knowledge The information is encoded into the target knowledge related representation , fine-grained irrelevant knowledge Information encoded into irrelevant representations middle, 、 and Respectively represent the subject, relation and object in the target knowledge triple, and They represent the replacement of the knowledge representation of the knowledge decoupler input with the target knowledge related representation during the forward propagation of the large language model. Representation independent of target knowledge , Indicates that the input of the large language model is and Output when constructing text prompt words The probability of Indicates that the input of the large language model is and Output when constructing text prompt words probability.
4. The large language model knowledge editing method based on knowledge representation decoupling according to claim 1 is characterized in that: The knowledge decoupler and knowledge coupler adopt representation reconstruction loss during training: ; in, represents the knowledge representation of the knowledge decoupler input, represents the reconstructed representation obtained by the knowledge decoupler, Calculates the Frobenius norm.
5. The large language model knowledge editing method based on knowledge representation decoupling according to claim 1 is characterized in that: The knowledge decoupler uses relational representation to decouple knowledge representation into target knowledge related representation and target knowledge unrelated representation.
6. The large language model knowledge editing method based on knowledge representation decoupling according to claim 1 is characterized in that: The parameter update constraints are used in the process of updating the parameters of the large language model: ; in, is the Frobenius norm calculation, is the large language model parameter matrix to be updated, The new knowledge is saved in the form of key-value pairs in the parameter matrix, and the new knowledge is injected into the large language model in the form of key-value pairs.
7. The large language model knowledge editing method based on knowledge representation decoupling according to claim 1 is characterized in that: The large language model parameters are updated using fine-grained irrelevant knowledge to maintain constraints: ; in, is the Frobenius norm calculation, is the large language model parameter matrix to be updated, represents the knowledge representation of the knowledge decoupler input, To preserve the key-value pairs of new knowledge in the parameter matrix, we use linear transformations that are irrelevant to the representation. Constrain the parameter update process to ensure the consistency of fine-grained irrelevant knowledge before and after parameter update.
8. The large language model knowledge editing method based on knowledge representation decoupling according to claim 1 is characterized in that: The coarse-grained irrelevant knowledge is used to maintain constraints during the update of the large language model parameters: ; in, is the Frobenius norm calculation, is the large language model parameter matrix to be updated, ( ) is the key-value pair representation of coarse-grained irrelevant knowledge, which constrains the consistency of coarse-grained irrelevant knowledge before and after parameter update.
9. An interactive method, characterized in that: include: Get user questions; Based on the user question, an answer is obtained through the large language model obtained by the large language model knowledge editing method based on knowledge representation decoupling according to any one of claims 1 to 8.
Citation Information
Patent Citations
Knowledge injection and training method and system for knowledge enhancement pre-training language model
CN116450839A
Large language model knowledge editing method and device based on pattern matching
CN119204091A