A Method and System for Knowledge Editing of Large Language Models
By grading neuron importance and updating the two-stage gradient of the large language model, the accuracy and efficiency of knowledge editing in the existing technology are solved, and efficient and robust knowledge editing effect is achieved.
Patent Information
- Application Number
- CN202510386461.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-31
- Publication Date
- 2025-07-29
- Estimated Expiration
- 2045-03-31
AI Technical Summary
The knowledge editing methods of existing large language models are insufficient in terms of accuracy, efficiency, and adaptability, especially when large-scale knowledge updates are computationally expensive and may undermine existing knowledge.
By rating the importance of model neurons, selecting key neuron sets, clustering knowledge, and adopting a two-stage gradient update method, including shared updates and personalized adjustments, accurately locate target knowledge and reduce computational overhead.
It realizes efficient and accurate knowledge editing, significantly reduces computing overhead, ensures the stability and robustness of the model, and is suitable for large-scale knowledge update scenarios.
Smart Images

Figure CN119886073B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of artificial intelligence technology, and specifically to a method and system for editing knowledge of large language models. Background Art
[0002] In recent years, large language models (LLMs) have made remarkable progress in the field of natural language processing. Through the training on large-scale Internet data, these models can not only generate high-quality text but also store and retrieve a large amount of factual knowledge. However, the dynamicity and accuracy issues of these knowledge bases pose great challenges to the practical applications of large language models. The knowledge in LLMs usually comes from static training corpora, which may contain outdated, inaccurate or even contradictory information, resulting in the model providing incorrect or irrelevant answers during reasoning. For example, a model may wrongly answer who the legal representative of a certain company is currently, or provide expired medical guidelines.
[0003] The traditional solution is to retrain the entire large model, which requires re-collecting and cleaning a large amount of data and consuming huge computing resources. For large language models with billions or even hundreds of billions of parameters (such as GPT-3, Llama, etc.), the computing cost of retraining is often unaffordable. In addition, retraining may also destroy the useful knowledge already stored in the model, resulting in a decline in the overall performance of the model.
[0004] To avoid the high cost of retraining, various knowledge editing methods have been proposed in recent years. These methods aim to modify specific knowledge by locally updating model parameters while maintaining the integrity of other knowledge. These methods can be roughly classified into the following categories: (1) Fine-tuning-based methods: Fine-tuning methods directly train specific layers or all parameters of the model to update the target knowledge. This type of method is simple and intuitive, but its limitation is that the update range is too large, which may introduce the problem of knowledge forgetting, and the computational cost is still high. (2) Memory module-based methods: Some studies have introduced external memory modules (such as SERAC) for explicitly storing and retrieving new knowledge. Although this method avoids modifying the internal parameters of the model, it relies on the implementation of external memory, resulting in slow retrieval speed and difficulty in dealing with the deep correlation of knowledge during the model inference process. (3) Locally parameter update-based methods: In recent years, some "Locate-and-Edit" methods, such as ROME, MEMIT, etc., have become popular. These methods analyze the parameter distribution of large language models, locate the key layers or neurons storing specific knowledge, and update them accordingly. But they have two main limitations. One is the limitation of the fixed layer assumption: Existing methods determine a fixed set of key layers after model training, but the storage locations of different knowledge in large models may be different. The fixed layer assumption cannot adapt to the diversity of storage patterns, resulting in insufficient editing accuracy. The second point is the high computational overhead: For large-scale knowledge updates, existing methods often need to update the parameters of knowledge item by item, and the gradient calculation process is inefficient. Especially in scenarios where tens of thousands of knowledge items need to be updated, the time cost is extremely high.
[0005] In summary, there is still much room for improvement in the accuracy, efficiency, and adaptability of current knowledge editing methods for large language models. Summary of the Invention
[0006] In this embodiment, a knowledge editing method, system, electronic device, and storage medium for large language models are provided to solve the problems of poor accuracy, efficiency, and adaptability of current knowledge editing methods for large language models.
[0007] In a first aspect, an embodiment of the present invention provides a knowledge editing method for large language models, and the knowledge editing method for large language models includes:
[0008] Score the importance of model neurons to obtain a score value for each neuron;
[0009] Calculate the contribution value based on the score, and select a set of key neurons participating in knowledge editing according to a preset contribution threshold;
[0010] Select key neurons and cluster the knowledge to be edited;
[0011] The sample closest to the cluster center is used as the anchor sample, and all instances within the cluster share an update vector based on the anchor sample for the first-stage update. An additional personalized adjustment is made to each instance for the second-stage update.
[0012] In an alternative embodiment, the importance of the model neurons is scored to obtain a score value for each neuron, including:
[0013] The score value of each neuron is calculated by any one of the activation value score, the weight importance score, and the residual sensitivity score.
[0014] In an alternative embodiment, the contribution value is calculated based on the score, and a set of key neurons participating in knowledge editing is selected according to a preset contribution threshold, including:
[0015] The neuron sets are sorted in the way of cumulative contribution value to determine the neuron set that meets a specific contribution threshold.
[0016] In an alternative embodiment, the key neurons are selected and the knowledge to be edited is clustered, including:
[0017] Each knowledge instance and its associated neuron set are represented as a key-value pair;
[0018] Using the Jaccard similarity as a metric, calculate the similarity between the neuron sets of two knowledge instances;
[0019] According to the Jaccard similarity, the similar knowledge instances are aggregated into the same cluster.
[0020] In an alternative embodiment, the sample closest to the cluster center is used as the anchor sample, and all instances within the cluster share an update vector based on the anchor sample for the first-stage update, including:
[0021] Determine the cluster center;
[0022] Within each cluster, calculate the Jaccard distance between each sample and the cluster center, and select the sample closest to the cluster center as the anchor sample;
[0023] All instances within the cluster share the same update vector calculated based on the anchor sample to minimize the average loss between the anchor sample and its target output;
[0024] Apply the calculated shared update vector to all instances within the cluster.
[0025] In an alternative embodiment, an additional personalized adjustment is made to each instance for the second-stage update, including:
[0026] For each instance within a cluster, perform additional personalized adjustments on the basis of its shared update to minimize the instance-specific adjustment term for the residual error after the shared update;
[0027] Add the shared update vector in the first stage to the personalized adjustment vector in the second stage to form a complete gradient descent update vector.
[0028] In an optional embodiment, the optimization objective for minimizing the instance-specific adjustment term for the residual error after the shared update is:
[0029] ;
[0030] wherein, represents the optimal personalized adjustment amount for sample , represents the hidden representation of sample calculated at layer L, represents the input of a single instance within the cluster, represents the instance 's target output, represents the personalized gradient adjustment variable for instance .
[0031] Compared with the prior art, the beneficial effects of the large language model knowledge editing method of the present invention are as follows:
[0032] This method realizes the accurate positioning of target knowledge by introducing explanation-based key neuron recognition, thereby providing higher accuracy and efficiency in knowledge editing in large language models. Secondly, through the combination of knowledge clustering and two-stage gradient update, the computational overhead during the editing process is effectively reduced, while ensuring the stability and robustness of the model. In addition, this method theoretically reveals the deep connection between interpretability analysis and the knowledge storage distribution of neural networks, providing a new technical perspective for understanding model editing tasks. Empirical evaluation shows that this method is significantly superior to existing methods in knowledge editing tasks on multiple public datasets, fully demonstrating its strong practical application potential and wide applicability. Generally speaking, this method provides an efficient, accurate and robust solution for the knowledge editing task of large language models, which helps to promote the application and development of large language models in dynamic knowledge update and personalized knowledge adjustment.
[0033] In a second aspect, an embodiment of the present invention provides a large language model knowledge editing system, including:
[0034] A scoring module for scoring the importance of model neurons to obtain a scoring value for each neuron;
[0035] A dynamic neuron selection module, configured to calculate contribution values based on scores and select a set of key neurons participating in knowledge editing according to a preset contribution threshold;
[0036] A knowledge clustering module, configured to select key neurons and cluster the knowledge to be edited;
[0037] A two-stage gradient update module, configured to use the sample closest to the cluster center as an anchor sample, and all instances within the cluster share an update vector based on the anchor sample for the first-stage update, and perform additional personalized adjustments on each instance for the second-stage update.
[0038] In a third aspect, an embodiment of the present invention provides an electronic device, including a processor, a communication interface, a memory, and a bus. Among them, the processor, the communication interface, and the memory complete communication with each other through the bus, and the processor can call the logical instructions in the memory to execute the steps of the method provided in the first aspect.
[0039] In a fourth aspect, an embodiment of the present invention provides a non-transitory computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, it implements the steps of the large language model knowledge editing method described in the first aspect.
[0040] Compared with the prior art, the beneficial effects of the large language model knowledge editing system, electronic device, and storage medium of the present invention are the same as those of the large language model knowledge editing method described in the first aspect, so they will not be elaborated here. Description of the Drawings
[0041] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, the following drawings are some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.
[0042] Figure 1 It is a flowchart of the large language model knowledge editing method in the embodiment of the present invention;
[0043] Figure 2 It is a framework diagram of the large language model knowledge editing method in the embodiment of the present invention;
[0044] Figure 3 It is a schematic diagram of the process framework of the large language model knowledge editing method in the embodiment of the present invention;
[0045] Figure 4 It is a structural block diagram of the large language model knowledge editing system in the embodiment of the present invention;
[0046] Figure 5 This is a structural block diagram of an electronic device in an embodiment of the present invention. Detailed implementation manners
[0047] To more clearly understand the purpose, technical solution, and advantages of the present application, the present application will be described and explained below with reference to the accompanying drawings and embodiments.
[0048] Unless otherwise defined, the technical terms or scientific terms involved in the present application shall have the general meaning understood by those with ordinary skills in the technical field to which the present application belongs. In the present application, words such as "a", "one", "a kind of", "the", "these", etc. do not indicate a limitation in quantity, and they can be singular or plural. The terms "include", "comprise", "have" and any variants thereof involved in the present application are intended to cover non-exclusive inclusion; for example, a process, method, system, product, or device including a series of steps or modules (units) is not limited to the listed steps or modules (units), but may include unlisted steps or modules (units), or may include other steps or modules (units) inherent in these processes, methods, products, or devices. The terms "connect", "be connected", "couple" and the like involved in the present application are not limited to physical or mechanical connections, but may include electrical connections, whether directly or indirectly connected. The term "plurality" involved in the present application means two or more. "And / or" describes the association relationship of associated objects, indicating that three relationships may exist. For example, "A and / or B" may mean: A exists alone, A and B exist simultaneously, and B exists alone. Usually, the character " / " indicates that the objects before and after are in an "or" relationship. The terms "first", "second", "third", etc. involved in the present application are only used to distinguish similar objects and do not represent a specific order for the objects.
[0049] In an embodiment of the present invention, a large language model knowledge editing method (TripleE Edit) is provided. Figure 1 This is a flowchart of the large language model knowledge editing method of the present invention, as Figure 1 and Figure 2 shown, and this process includes the following steps:
[0050] S100. Score the importance of the model neurons to obtain a score value for each neuron;
[0051] Specifically, scoring the importance of the model neurons to obtain a score value for each neuron includes:
[0052] Calculating the score value of each neuron by any one of activation value scoring, weight importance scoring, and residual sensitivity scoring.
[0053] Activation value scoring: Evaluate the importance of a neuron in knowledge storage by calculating the strength of the neuron's activation of the input knowledge. The formula is: Among them, represents the neuron on the input the activation value.
[0054] Weight importance scoring: Identify the neurons that play a dominant role in information transmission by analyzing the strength of the weight connections between neurons. Its scoring formula is: Among them, represents the neuron and the connection weight between.
[0055] Residual sensitivity scoring: Evaluate the impact of each neuron in knowledge storage by analyzing the contribution of the residual flow to the model output. The formula is as follows: Among them, represents the layer in the neuron the activation value, is the neuron the output weight.
[0056] S200. Calculate the contribution value based on the scoring, and select the key neuron set participating in knowledge editing according to the preset contribution threshold;
[0057] Specifically, calculate the contribution value based on the scoring, and select the key neuron set participating in knowledge editing according to the preset contribution threshold, including:
[0058] Sort the neuron sets in the way of cumulative contribution value to determine the neuron set that meets the specific contribution threshold.
[0059] It should be noted that this method adopts a dynamic selection mechanism to adaptively select the most relevant neuron set in different knowledge editing tasks. The specific steps are as follows: Select any one of the above three methods, calculate the scoring value of each neuron, and sort according to the scoring, so as to determine the neuron set used for knowledge editing tasks. Sort with the cumulative contribution value to determine the neurons that meet the specific contribution threshold I , the formula is:
[0060] ;
[0061] Among them, is the scoring value of the th neuron, is the contribution threshold. For the batch editing scenario, this method aggregates the neuron scores of all tasks to uniformly identify the neuron set that makes the greatest contribution to multiple tasks, further optimizing the editing efficiency.
[0062] Through the above dynamic selection mechanism, this method can accurately identify the target neurons while minimizing the interference to other knowledge. Compared with the previous mainstream method of selecting several layers in a coarse-grained manner, this method directly selects from the neuron accuracy. After selecting one of the methods AC, WE, and RE for the task, according to different performances, the different adaptabilities of each method to the environmental settings and input attributes can be deduced. This shows that this method implants interpretability into the steps of selecting specific neurons, and can trace back to find out which neurons are crucial for specific attributes, so as to select effective neurons for the next editing task.
[0063] S300. Select the key neurons and cluster the knowledge to be edited;
[0064] After identifying the target neurons, this method efficiently edits the model through an optimized parameter update mechanism. The core goal is to reduce the gradient calculation overhead through batch optimization techniques while ensuring the efficiency and stability of editing. Different types of knowledge may correspond to different neuron distributions. This method (TripleE Edit) uses a clustering algorithm to group the knowledge to be edited, integrating the update operations of similar knowledge together to achieve batch update. The main goal of knowledge clustering is to group knowledge with similar attributes into the same cluster according to the characteristics of different knowledge instances, so as to reduce the conflicts that may occur during the knowledge editing process.
[0065] Specifically, selecting the key neurons and clustering the knowledge to be edited includes:
[0066] Represent each knowledge instance and its associated neuron set as a key-value pair;
[0067] Use Jaccard similarity as a metric to calculate the similarity between the neuron sets of two knowledge instances;
[0068] According to the Jaccard similarity, aggregate similar knowledge instances into the same cluster.
[0069] Exemplarily, first for each knowledge instance its corresponding neuron set is represented as a key-value pair. This representation can reflect the association between the knowledge instance and its storage location in the model. Then use Jaccard similarity as the similarity measurement method between knowledge instances to calculate the ratio difference between the intersection and union of the neuron sets of two knowledge instances. The specific formula is:
[0070] ;
[0071] Among them, and respectively represent the neuron sets corresponding to the knowledge instances and Then, similar knowledge instances are aggregated into the same clusters.
[0072] The optimization objective of clustering is to minimize the Jaccard distance between data points and their corresponding cluster centers. The formula is as follows:
[0073] ;
[0074] Among them, represents the th cluster, is the center point of this cluster. For each cluster, the sample closest to the cluster center is selected as the anchor sample.
[0075] This anchor sample is regarded as the knowledge instance that best represents this cluster. Finally, the knowledge instances are divided into multiple mini - batches for the clustering result, and each batch is refined according to its specific attributes, so as to achieve higher pertinence and efficiency in knowledge editing. Through the above steps, knowledge clustering ensures that the knowledge instances within each cluster have high internal similarity and provides a basis for grouped optimization in subsequent editing tasks using large models.
[0076] S400: Use the sample closest to the cluster center as the anchor sample. The instances within all clusters share an update vector based on the anchor sample for the first - stage update, and additional personalized adjustments are made to each instance for the second - stage update.
[0077] This step is a two - stage gradient update, aiming to improve the computational efficiency in the sequential batch editing process. By dividing the gradient optimization process of each cluster into two stages - a shared stage and a personalized fine - tuning stage, the utilization of shared information is maximized while retaining the unique adjustments of each instance.
[0078] Specifically, use the sample closest to the cluster center as the anchor sample. The instances within all clusters share an update vector based on the anchor sample for the first - stage update, including:
[0079] Determine the cluster center;
[0080] Within each cluster, calculate the Jaccard distance between each sample and the cluster center, and select the sample closest to the cluster center as the anchor sample;
[0081] All instances within a cluster share the same update vector calculated based on the anchor sample to minimize the average loss between the anchor sample and its target output;
[0082] The calculated shared update vector is applied to all instances within the cluster.
[0083] For example, first cluster the anchor samples Based on the gradient of , 4 rounds of unified updates are performed (a total of 5 rounds of updates in the whole process). All instances in the cluster share this update vector , because the anchor sample is regarded as the representative of the cluster and can capture the core features of the cluster. With its target output The average loss of , calculate the shared update vector , the formula is:
[0084] ;
[0085] Where, represents the calculated shared update vector, represents the adjustment amount or gradient update variable used for optimization, P G Represents the probability output of a probability distribution or model G under specific circumstances, represents the hidden representation of the anchor sample calculated on layer L.
[0086] The updates at this stage capture the common patterns within the clusters, lay the foundation for subsequent personalized updates, and reduce the burden of repeated calculations within the clusters.
[0087] Furthermore, additional personalized adjustments are made to each instance for the second phase of updates, including:
[0088] For each instance in the cluster, additional personalized adjustments are made based on its shared update to minimize the instance-specific adjustment of the residual error after the shared update;
[0089] The shared update vector from the first stage is added to the personalized adjustment vector from the second stage to form the complete gradient descent update vector.
[0090] For example, the second stage performs additional personalized adjustments for instances within each cluster to account for subtle differences between instances. Perform 1 round of fine-tuning based on the shared update. The optimization goal is to minimize the instance-specific adjustment of the residual error after the shared update , the formula is:
[0091] ;
[0092] In the formula, represents the optimal personalized adjustment amount for the sample , represents the hidden representation of the sample calculated in layer L, represents a single instance input within the cluster, represents the instance 's target output, represents the instance 's personalized gradient adjustment variable.
[0093] This finely tunes the details of each instance, ensuring that the model can adapt to the unique characteristics of each instance. Finally, adding the two-stage vectors together gives the complete gradient descent update vector as:
[0094] ;
[0095] In the first stage, shared updates reduce most of the redundant calculations, avoiding the high computational cost of traditional instance-by-instance update methods. The second-stage personalized fine-tuning ensures that the uniqueness of each instance is preserved and personalized features are not lost due to shared updates. This method focuses on the identified important neurons, avoiding global modifications to all model parameters and reducing potential disruptions to model stability and integrity. Through the two-stage gradient update method, this method achieves efficient and accurate model editing, is particularly suitable for large-scale editing scenarios, significantly reduces computational overhead while ensuring high-quality knowledge editing and model stability.
[0096] This method achieves efficient and accurate knowledge editing of large language models. Experiments show that this method improves the editing efficiency by several times or more, while significantly reducing the risk of model performance degradation, providing an efficient and practical solution for the dynamic knowledge update of large models.
[0097] Specifically, as Figure 3 shown, Figure 3 shows a schematic diagram of the two-stage specific process optimized by the mechanism of grouping similar knowledge using a clustering algorithm and performing batch optimization and the two-step gradient descent method.
[0098] Clustering Phase: In this phase, all knowledge instances to be edited are grouped by calculating the similarity between their feature representations. A clustering algorithm (such as K-means) is used to cluster similar knowledge instances together. Batch Optimization Phase: Once the knowledge is grouped, instances within the same cluster will share an update vector. In this way, redundant calculations between different instances are reduced, significantly improving the computational efficiency. Shared Update Phase: Knowledge instances within all clusters share a unified update vector. By performing gradient updates on the central samples within the cluster (usually samples representing the features of the class), the common features of the knowledge in this cluster can be captured. This process usually performs multiple rounds of gradient descent. Personalized Fine-tuning Phase: After the shared update, each knowledge instance within the cluster is fine-tuned individually to ensure that their specific features are further optimized. The update of each instance is based on the shared update vector and usually performs fewer rounds of gradient descent. The fine-tuning in this phase ensures that the personalized differences of each instance are retained while avoiding over-adjustment. Through this two-phase optimization process, TripleE Edit can not only efficiently perform large-scale knowledge updates but also ensure that the personalized features of each instance are carefully adjusted, maintaining the consistency of the model in multiple tasks.
[0099] TripleE Edit proposed by the present invention has significant advantages and positive effects. First, by introducing the identification of key neurons based on interpretability, TripleE Edit achieves accurate positioning of the target knowledge, thus providing higher accuracy and efficiency in knowledge editing in large language models. Second, through the combination of knowledge clustering and two-phase gradient updates, TripleE Edit effectively reduces the computational overhead during the editing process while ensuring the stability and robustness of the model. In addition, TripleE Edit theoretically reveals the deep connection between interpretability analysis and the knowledge storage distribution in neural networks, providing a new technical perspective for understanding model editing tasks. Empirical evaluations show that TripleE Edit significantly outperforms existing methods in knowledge editing tasks on multiple public datasets, fully demonstrating its strong practical application potential and wide applicability. Generally speaking, TripleE Edit provides an efficient, accurate, and robust solution for the knowledge editing tasks of large language models, contributing to the application and development of large language models in dynamic knowledge updates and personalized knowledge adjustments.
[0100] To verify the effectiveness of the present invention, experimental evaluations of TripleE Edit were conducted on the public datasets Counterfact and ZsRE, and experiments were carried out on different architectures such as GPT-2 XL, GPT-J, and Llama3. The editing quality of the model was evaluated by comparing with existing editing methods in terms of editing effectiveness (Eff), editing generalization (Gen), editing locality (Spe), language fluency (Flu), and knowledge consistency (Cons). The main results are shown in Table 1:
[0101] Table 1 Results of Continuous Editing Experiment
[0102]
[0103] TripleE Edit performs excellently in terms of editing effectiveness. The experimental results on the Counterfact dataset show that the editing success rates of the three methods of selecting neurons in TripleE Edit all reach over 95%. In addition, the experiments on the ZsRE dataset indicate that TripleE Edit can quickly adapt to diverse editing tasks and outperforms the comparison methods in all metrics. Therefore, in the scenario of continuous knowledge editing, our method has more advantages in various evaluation metrics, with better generalization and stability. The editing speed is significantly improved compared with the current knowledge editing methods;
[0104] Table 2 Comparative Results of Continuous Editing Experiment Time
[0105]
[0106] TripleE Edit significantly reduces the computational overhead through knowledge clustering and two-stage gradient update. The experimental data shows (Table 2) that in the editing task using GPT-J, the average editing time of TripleE Edit is only 178 seconds, significantly lower than 764 seconds of ROME and 334 seconds of MEMIT, achieving several times of efficiency improvement. Compared with traditional methods, TripleE Edit can not only complete the same-scale editing tasks in a shorter time but also ensure the high quality and stability of the editing results. Therefore, TripleE Edit demonstrates great practical value and efficiency advantages in continuous knowledge editing, especially suitable for scenarios that require rapid update and processing of a large amount of knowledge.
[0107] An embodiment of the present invention also provides a large language model knowledge editing system, which is used to implement the above method embodiments. Those that have been described will not be repeated here. The following terms "module", "unit", "sub-unit", etc. can be a combination of software and / or hardware that can achieve a predetermined function. Although the system described in the following embodiments is preferably implemented in software, implementation in hardware or a combination of software and hardware is also possible and contemplated.
[0108] As Figure 4 shown, Figure 4 is a structural block diagram of the large language model knowledge editing system in the present invention. The system includes:
[0109] A scoring module 101, configured to score the importance of model neurons to obtain a scoring value for each neuron;
[0110] A dynamic neuron selection module 102, configured to calculate a contribution value based on the score and select a set of key neurons participating in knowledge editing according to a preset contribution threshold;
[0111] A knowledge clustering module 103, configured to select key neurons and cluster the knowledge to be edited;
[0112] A two-stage gradient update module 104, configured to use the sample closest to the clustering center as an anchor sample, and all instances within the cluster share an update vector based on the anchor sample for the first-stage update, and perform additional personalized adjustments on each instance for the second-stage update.
[0113] Figure 5 is a structural block diagram of the electronic device provided by the embodiment of the present invention. As Figure 5 shown, the electronic device may include: a processor 610, a communication interface 620, a memory 630, and a communication bus 640. Among them, the processor 610, the communication interface 620, and the memory 630 communicate with each other through the communication bus 640. The processor 610 can call the logical instructions in the memory 630 to execute the following methods:
[0114] Score the importance of model neurons to obtain a scoring value for each neuron;
[0115] Calculate a contribution value based on the score and select a set of key neurons participating in knowledge editing according to a preset contribution threshold;
[0116] Select key neurons and cluster the knowledge to be edited;
[0117] The sample closest to the cluster center is used as the anchor sample, and all instances within the cluster share an update vector based on the anchor sample for the first-stage update. An additional personalized adjustment is made to each instance for the second-stage update.
[0118] In addition, when the logical instructions in the above-mentioned memory 630 are implemented in the form of software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or a part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in various embodiments of the present invention. The aforementioned storage medium includes: various media such as USB flash drives, mobile hard disks, read-only memories (ROM, Read-Only Memory), random access memories (RAM, Random Access Memory), magnetic disks, or optical discs that can store program codes.
[0119] The embodiments of the present invention also provide a non-transitory computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, it is configured to execute the methods provided in the above-mentioned embodiments, for example, including:
[0120] Score the importance of the model neurons to obtain the score value of each neuron;
[0121] Calculate the contribution value based on the score, and select a key neuron set participating in knowledge editing according to a preset contribution threshold;
[0122] Select the key neurons and cluster the knowledge to be edited;
[0123] The sample closest to the cluster center is used as the anchor sample, and all instances within the cluster share an update vector based on the anchor sample for the first-stage update. An additional personalized adjustment is made to each instance for the second-stage update.
[0124] Through the description of the above embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by means of software plus a necessary general hardware platform, and of course, it can also be implemented by hardware. Based on this understanding, the essence of the above technical solution, or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, magnetic disk, optical disk, etc., and includes several instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute the methods of each embodiment or some parts of the embodiments.
[0125] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions recorded in the foregoing embodiments, or perform equivalent replacements for some of the technical features; and these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.
Claims
1. A method for editing knowledge of a large language model, characterized in that, The large language model knowledge editing method includes: Scoring the importance of model neurons to obtain a score value for each neuron; Calculating contribution values based on the scores and selecting a set of key neurons participating in knowledge editing according to a preset contribution threshold; Selecting key neurons and clustering the knowledge to be edited; Using the sample closest to the cluster center as the anchor sample, and all instances within the cluster sharing an update vector based on the anchor sample for the first-stage update, and performing additional personalized adjustments on each instance for the second-stage update.
2. The large language model knowledge editing method according to claim 1, wherein Scoring the importance of model neurons to obtain a score value for each neuron, including: Calculating the score value for each neuron through any one of activation value scoring, weight importance scoring, and residual sensitivity scoring.
3. The large language model knowledge editing method according to claim 1, characterized in that Calculating contribution values based on the scores and selecting a set of key neurons participating in knowledge editing according to a preset contribution threshold, including: Sorting the neuron sets in the way of cumulative contribution values to determine the neuron set that meets a specific contribution threshold.
4. The large language model knowledge editing method according to claim 1, wherein Selecting key neurons and clustering the knowledge to be edited, including: Representing each knowledge instance and its related neuron set as a key-value pair; Using Jaccard similarity as a metric to calculate the similarity between the neuron sets of two knowledge instances; Aggregating similar knowledge instances into the same cluster according to Jaccard similarity.
5. The large language model knowledge editing method according to claim 1, wherein Using the sample closest to the cluster center as the anchor sample, and all instances within the cluster sharing an update vector based on the anchor sample for the first-stage update, including: Determining the cluster center; Within each cluster, calculating the Jaccard distance between each sample and the cluster center, and selecting the sample closest to the cluster center as the anchor sample; All instances within the cluster share the same update vector calculated based on the anchor sample to minimize the average loss between the anchor sample and its target output; Sharing and applying the calculated update vector to all instances within the cluster.
6. The large language model knowledge editing method according to claim 1, wherein Performing additional personalized adjustments on each instance for the second-stage update, including: For each instance within the cluster, performing additional personalized adjustments on the basis of its shared update to minimize the instance-specific adjustment term of the residual error after the shared update; Adding the shared update vector in the first stage and the personalized adjustment vector in the second stage to form a complete gradient descent update vector.
7. The method for editing knowledge of a large language model according to claim 6, wherein The optimization objective of minimizing the instance-specific adjustment term of the residual error after the shared update is: ; Wherein, represents the input sample being edited, represents the corresponding edited target output, represents the optimal personalized adjustment amount for , represents the hidden representation of calculated at layer L, represents the distribution of the model output , represents the negative log-likelihood loss calculated based on the model, represents the shared update vector, represents the personalized gradient adjustment variable of 8. A large language model knowledge editing system, characterized in that Including: A scoring module for scoring the importance of model neurons to obtain a score value for each neuron; A dynamic neuron selection module for calculating contribution values based on the scores and selecting a set of key neurons participating in knowledge editing according to a preset contribution threshold; A knowledge clustering module for selecting key neurons and clustering the knowledge to be edited; A two-stage gradient update module for using the sample closest to the cluster center as the anchor sample, and all instances within the cluster sharing an update vector based on the anchor sample for the first-stage update, and performing additional personalized adjustments on each instance for the second-stage update.
9. An electronic device, comprising a memory, a processor, and a computer program stored on the memory and executable on the processor, characterized in that, When the processor executes the program, it implements the large language model knowledge editing method according to any one of claims 1 to 7.
10. A non-transitory computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by the processor, it implements the steps of the large language model knowledge editing method according to any one of claims 1 to 7.
Citation Information
Patent Citations
Method for training large language model, storage medium and equipment
CN118297136A
Multi-modal large language model neuron attribution method and related equipment
CN118761476A