Large model distillation knowledge editing method fusing context constraints

Through the multi-teacher distillation model and self-induced distribution alignment mechanism, the problem of the existing knowledge editing methods degradation during large-scale continuous editing is solved, and an efficient, robust and scalable knowledge update effect is achieved.

CN120196724APending Publication Date: 2025-06-24NORTHEASTERN UNIV CHINA
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510447218.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-10
Publication Date
2025-06-24

AI Technical Summary

Technical Problem

The performance of existing knowledge editing methods has significantly decreased during large-scale continuous editing, and the lack of dynamic constraint mechanisms has led to knowledge conflicts and old knowledge coverage, and the calculation overhead is large, making it difficult to support large-scale editing.

Method used

The multi-teacher distillation model is adopted to guide student model training through dynamic context information, trainable low-rank matrix and fine-tuned loss function, and a self-induced distribution alignment mechanism is introduced to optimize student model to achieve efficient, robust and scalable knowledge updates.

Benefits of technology

It realizes efficient and robust knowledge updates in large-scale continuous editing, the inference speed is close to that of the teacher model, and the parameter design supports thousands-level knowledge editing to adapt to the needs of continuous updates.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120196724A_ABST
    Figure CN120196724A_ABST
Patent Text Reader

Abstract

The invention provides a large model distillation knowledge editing method fusing context constraints, and relates to the technical field of natural language processing. According to the method, by generating the dynamic context information of the question input by the user, the context correlation of the question input by the user is improved; constructing a multi-teacher distillation model, and reserving the universality of a plurality of teacher models by setting a knowledge editing range and a corresponding problem processing method; a trainable low-rank matrix and a fine-tuning loss function are introduced to optimize a loss function of the multi-teacher distillation model, fine-tuning training is performed on a student model, a self-induced distribution alignment mechanism is introduced to optimize the trained student model, the optimized student model does not need real-time retrieval or long prompt, the reasoning speed is close to that of a teacher model, and the reasoning efficiency is improved. And the method has expandability, efficient parameter design supports thousand-scale knowledge editing, and continuous updating requirements are met.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of natural language processing, and particularly relates to a large model distillation knowledge editing method integrating context constraints. Background Art

[0002] Large language models (LLMs) have shown excellent performance in various natural language processing tasks. Knowledge editing technology is an important research direction in the fields of natural language processing (NLP) and artificial intelligence in recent years, aiming to modify the knowledge in large language models in an efficient and accurate manner to address issues such as outdated, incorrect, or biased knowledge.

[0003] However, traditional knowledge editing methods (such as MEMIT, ROME, FT-L) have poor continuous editing stability, and their performance significantly deteriorates during continuous editing. For example, the accuracy of MEMIT drops by more than 30% after 10 edits, and the language fluency (PPL) deteriorates to more than 25. The reason is that these methods lack a dynamic constraint mechanism, and each edit independently updates the parameters, resulting in knowledge conflicts and overwriting of old knowledge. For example, FT-L directly fine-tunes specific layers to maximize the target probability, but does not limit the parameter change range, ultimately causing an overall distribution shift in the model. There is also a trade-off contradiction between generalization ability (processing related queries) and locality (not affecting unrelated queries), and it is difficult to achieve both generalization and locality. For example, IKE improves generalization through in-context learning (the generalization rate on the zsRE dataset reaches 89%), but the mis-editing rate for out-of-scope queries is as high as 12.5%. Although MEND reduces the mis-editing rate to 3%, the generalization index is only 78%. This is because traditional methods rely on fixed retrieval strategies or static parameter updates and cannot dynamically distinguish the query scope.

[0004] In addition, knowledge editing methods based on external memory (such as SERAC) need to maintain an independent knowledge editing library, with high computational overhead, making it difficult to support large-scale editing. Adding 1,000 pieces of knowledge requires an additional storage of 1.2GB, and the inference latency increases by 40%. Context learning methods such as IKE need to splice long examples (average length 500 tokens) in the input, resulting in a total editing time of more than 200 hours for 10,000 edits. Full-scale fine-tuning (such as FT-M) updates all parameters, with a single edit taking up to 30 minutes and being costly.

[0005] Knowledge editing methods based on parameter modification (such as ROME) directly optimize the cross-entropy loss of a single sample, causing the model to overfit the target distribution and resulting in a decline in language quality. Experiments show that the PPL value of the generated text increases from the initial 15.3 to 28.7, and the repetition rate increases by 22%. The reason is that its loss function does not consider context distribution alignment, resulting in the output deviating from the natural language pattern.

[0006] Based on the parameter injection method (such as T-Patcher), the parameter modification ratio is high, which affects the model robustness. It is necessary to expand the model parameter quantity by 5%-10%, resulting in a 35% decrease in the inference speed. MEMIT realizes the update by editing the MLP layer weights, but each single edit needs to modify 0.8% of the parameters. After 10,000 edits, the model stability is significantly reduced (the variance increases by 63%).

[0007] Siyuan Qi et al. proposed the Consistent Context Editing (ICE) method in "Learning Knowledge from Self-Induced Distributions" to enhance the robustness of knowledge editing by optimizing the context distribution rather than a single target distribution. Specifically, ICE introduces a context loss, requiring the output distributions of the model with and without context cues to be consistent, thus internalizing new knowledge. This method is superior to traditional fine-tuning in terms of accuracy, locality, and language quality, but mainly targets single or few-edit scenarios and does not optimize the efficiency of large-scale continuous editing. In addition, the ICE method relies on an external model to generate context. If the generated content has noise or hallucinations, it may affect the editing effect.

[0008] Although the ICE method improves the robustness of single edits, it is not optimized for large-scale editing and is prone to performance degradation (such as a more than 30% decrease in accuracy) during continuous updates. DistillMIKE reduces the computational overhead through distillation, but it requires a pre-trained retrieval module, and the accuracy of the range classifier directly affects the editing effect.

[0009] DistillMIKE proposed by Shanbao Qiao et al. extends context-based editing (IKE) to large-scale tasks and introduces knowledge distillation to reduce the computational overhead. Its core includes: Selective Retrieval Enhancement: Apply retrieval-enhanced IKE only to "in-range" queries, matching relevant facts from the editing memory as context cues, while "out-of-range" queries directly use the responses of the unedited base model to balance generalization and locality; Low-Rank Adapter (LoRA) Distillation: Implicitly inject context cues into the model parameters through multi-teacher distillation (IKE and the base model), eliminating the need for long prompts during inference and only requiring a small number of parameters to be updated. Experiments show that DistillMIKE has performance close to MIKE in 10,000-edit tasks and is significantly superior to baselines such as MEMIT and PMET in terms of generalization. However, it still relies on a pre-trained retrieval enhancement module and a range classifier, with a relatively high actual deployment complexity. In addition, the reduction in the number of examples leads to a decline in the generalization ability, indicating its sensitivity to the construction of context cues. Summary of the Invention

[0010] The technical problem to be solved by the present invention is to provide a large model distillation knowledge editing method integrating context constraints in view of the above-mentioned deficiencies of the prior art, construct a multi-teacher distillation model, guide the training process of the student model through dynamic context information, introducing a trainable low-rank matrix and a fine-tuning loss function, and introduce a self-induced distribution alignment mechanism to optimize the trained student model, so as to achieve efficient, robust and scalable knowledge update.

[0011] To solve the above technical problems, the technical solutions adopted by the present invention are as follows:

[0012] The present invention provides a large model distillation knowledge editing method integrating context constraints, including the following steps:

[0013] Step 1: Obtain the user input question and generate dynamic context information of the user input question;

[0014] Obtain the user input question q, and dynamically retrieve from the external knowledge base the set of editing facts E q ={e1, e2,..., e n}, where e1, e2,..., e n are the 1st, 2nd,..., nth editing facts, generate the explanatory context c q of the user input question q; calculate the priorities of the editing facts in the set of editing facts E q , and filter out the highest-priority editing fact e q , and generate the dynamic context information [c q , e q , q] of the user input question q;

[0015] Step 2: Set the knowledge editing scope and the corresponding question processing method, determine the knowledge editing scope to which the user input question belongs, and use the question processing method corresponding to the knowledge editing scope to process the user input question;

[0016] Set the knowledge editing scope based on the external knowledge base, including within the knowledge editing scope and outside the knowledge editing scope; the knowledge within the knowledge editing scope is the knowledge belonging to the external knowledge base, and the knowledge outside the knowledge editing scope is the knowledge not belonging to the external knowledge base;

[0017] The question processing methods corresponding to the knowledge editing scope include:

[0018] For the user input question within the knowledge editing scope, use the dynamic context information [c q , e q , q] as the input question;

[0019] For the user input question outside the knowledge editing scope, use the original user input question q as the input question;

[0020] Use the scope classifier to determine the knowledge editing scope to which the user input question q belongs, and process the user input question using the question processing method corresponding to the knowledge editing scope;

[0021] Step 3: Construct a multi-teacher distillation model;

[0022] Step 3.1: Construct a multi-teacher distillation model, including multiple teacher models and one student model;

[0023] The first teacher model is used to guide the student model to learn the dynamic context information of the input question within the knowledge editing scope and fuse the dynamic context information into the parameters of the student model; the first teacher model obtains the user input question q belonging to the knowledge editing scope, and uses the ICE method to fuse the dynamic context information [c q ,e q ,q], and outputs the first teacher model distribution P teacher1 (y|q), where y is the answer to the user input question q;

[0024] The second teacher model is used to constrain the weight parameters of the student model outside the knowledge editing scope, so that the student model retains the knowledge outside the knowledge editing scope; the second teacher model obtains the user input question q belonging to outside the knowledge editing scope and outputs the second teacher model distribution where θ0 is the parameter of the second teacher model;

[0025] The student model receives the first teacher model distribution P teacher1 (y|q) and the second teacher model distribution and outputs the student model distribution where θ s is the parameter of the student model;

[0026] Step 3.2: Establish a distillation loss function for the multi-teacher distillation model;

[0027] Based on the weighted KL divergence of the student model distribution, the first teacher model distribution, and the second teacher model distribution, establish a distillation loss function L distill for the multi-teacher distillation model, as shown in the following formula:

[0028]

[0029] where α is the judgment probability output by the scope classifier, D KL (·) is the KL divergence operation, λ c is the conflict penalty weight, and x old is the knowledge outside the knowledge editing scope;

[0030] Step 4: Introduce a trainable low-rank matrix and a fine-tuning loss function into the multi-teacher distillation model to optimize the loss function of the multi-teacher distillation model, and perform fine-tuning training on the student model to obtain a trained student model;

[0031] Step 4.1: Freeze the parameters θ0 of the second teacher model, and introduce a trainable low-rank matrix into the second teacher network, as shown in the following formula:

[0032] ΔW = A·B T

[0033] where, r is the rank of the low-rank matrix ΔW, d is the eigen-dimension of the parameter matrix, and r << d;

[0034] Step 4.2: Establish a fine-tuning loss function for the multi-teacher distillation model to optimize the loss function of the multi-teacher distillation model, and establish the total loss function of the multi-teacher distillation model;

[0035] Establish a fine-tuning loss function L FT , as shown in the following formula:

[0036]

[0037] where, y * is the target knowledge;

[0038] Based on the weighted fine-tuning loss function L FT of the multi-teacher distillation model and the distillation loss function L distill of the multi-teacher distillation model, obtain the total loss function L total of the multi-teacher distillation model, as shown in the following formula:

[0039] L total = L distill + β·L FT

[0040] where, β is the balance coefficient;

[0041] Step 5: Introduce a self-induced distribution alignment mechanism to optimize the trained student model, and dynamically adjust the target weights for optimizing the student model according to the historical editing times to obtain an optimized student model;

[0042] Step 5.1: During the continuous knowledge editing process of the student model, introduce a self-induced distribution alignment mechanism to perform optimization training on the trained student model;

[0043] Introduce a self-induced distribution alignment mechanism, and establish an optimization loss function L align based on the JS divergence, as shown in the following formula:

[0044]

[0045] Among them, D JS (·) is the JS divergence operation;

[0046] Step 5.2: Dynamically adjust and optimize the target weight of the student model according to the historical editing times to obtain an optimized student model;

[0047] Dynamically adjust and optimize the target weight γ(t) of the student model according to the historical editing times t, as shown in the following formula:

[0048] γ(t) = γ0·e -ηt γ(t) = γ0·e -ηt

[0049] Among them, γ0 is the initial weight, and η is the decay rate;

[0050] Gradually reduce the dependence of the student model on old knowledge based on dynamic weight adjustment to obtain an optimized student model;

[0051] Step 6: Use the optimized student model to process the user input question and output the answer to the user input question.

[0052] The beneficial effects produced by adopting the above technical solutions are as follows: A large model distillation knowledge editing method integrating context constraints provided by the present invention improves the context relevance of user input questions by generating dynamic context information of user input questions; constructs a multi-teacher distillation model, and retains the generality of multiple teacher models by setting the knowledge editing scope and corresponding question processing methods; introduces a trainable low-rank matrix and a fine-tuning loss function to optimize the loss function of the multi-teacher distillation model, fine-tune and train the student model, and introduces a self-induced distribution alignment mechanism to optimize the trained student model. The optimized student model does not require real-time retrieval or long prompts, the inference speed is close to that of the teacher model, and it has scalability. The parameter-efficient design supports knowledge editing on a scale of thousands and adapts to the continuous update requirements. BRIEF DESCRIPTION OF THE DRAWINGS

[0053] Figure 1 It is a flowchart of a large model distillation knowledge editing method integrating context constraints provided by an embodiment of the present invention. DETAILED DESCRIPTION OF THE INVENTION

[0054] The following combines the drawings and embodiments to further describe in detail the specific implementation manners of the present invention. The following embodiments are used to illustrate the present invention, but are not used to limit the scope of the present invention.

[0055] A knowledge editing method for large model distillation that integrates context constraints in this embodiment combines context consistency constraints and knowledge distillation, proposes an optimization objective for dynamically adjusting context distribution alignment, and designs a lightweight distillation architecture to improve the efficiency of large-scale editing and the stability of continuous updates and reduce computational overhead, as Figure 1 shown, including the following steps:

[0056] Step 1: Obtain the user input question and generate dynamic context information for the user input question;

[0057] Obtain the user input question q, and dynamically retrieve from the external knowledge base the set of editing facts E q ={e1, e2,..., e n}, where e1, e2,..., e n are the 1st, 2nd,..., nth editing facts, generate the explanatory context c q for the user input question q; calculate the priorities of the editing facts in the set of editing facts E q , and filter out the highest-priority editing fact e q , and generate the dynamic context information [c q , e q , q] for the user input question q;

[0058] Step 2: Set the knowledge editing scope and the corresponding problem processing method, determine the knowledge editing scope to which the user input question belongs, and use the problem processing method corresponding to the knowledge editing scope to process the user input question;

[0059] Set the knowledge editing scope based on the external knowledge base, including within the knowledge editing scope and outside the knowledge editing scope; the knowledge within the knowledge editing scope is the knowledge belonging to the external knowledge base, and the knowledge outside the knowledge editing scope is the knowledge not belonging to the external knowledge base;

[0060] The problem processing methods corresponding to the knowledge editing scope include:

[0061] For the user input question within the knowledge editing scope, use the dynamic context information [c q , e q , q] as the input question;

[0062] For the user input question outside the knowledge editing scope, use the original user input question q as the input question;

[0063] Use the scope classifier to determine the knowledge editing scope to which the user input question q belongs, and use the problem processing method corresponding to the knowledge editing scope to process the user input question;

[0064] Step 3: Construct a multi-teacher distillation model;

[0065] Step 3.1: Construct a multi-teacher distillation model, including multiple teacher models and a student model;

[0066] The first teacher model is used to guide the student model to learn the dynamic context information of the input question within the knowledge editing range and fuse the dynamic context information into the parameters of the student model; the first teacher model obtains the user input question q within the knowledge editing range, and uses the ICE method to fuse the dynamic context information [c q ,e q ,q], and outputs the first teacher model distribution P teacher1 (y|q), where y is the answer to the user input question q;

[0067] The second teacher model is used to constrain the weight parameters of the student model outside the knowledge editing range, so that the student model retains the knowledge outside the knowledge editing range; the second teacher model obtains the user input question q outside the knowledge editing range and outputs the second teacher model distribution where θ0 is the parameter of the second teacher model;

[0068] The student model receives the first teacher model distribution P teacher1 (y|q) and the second teacher model distribution and outputs the student model distribution where θ s is the parameter of the student model;

[0069] Step 3.2: Establish a distillation loss function for the multi-teacher distillation model;

[0070] Based on the weighted KL divergence of the student model distribution, the first teacher model distribution, and the second teacher model distribution, establish a distillation loss function L distill of the multi-teacher distillation model, as shown in the following formula:

[0071]

[0072] where α is the judgment probability output by the range classifier, D KL (·) is the KL divergence operation, λ c is the conflict penalty weight, and x old is the knowledge outside the knowledge editing range;

[0073] By minimizing the KL divergence between the output of the student model and the first teacher model and the second teacher model, the student model retains the knowledge of the external knowledge base;

[0074] Step 4: Introduce a trainable low-rank matrix and a fine-tuning loss function into the multi-teacher distillation model to optimize the loss function of the multi-teacher distillation model, and perform fine-tuning training on the student model to obtain a trained student model;

[0075] Step 4.1: Freeze the parameters θ0 of the second teacher model, and introduce a trainable low-rank matrix into the second teacher network, as shown in the following formula:

[0076] ΔW = A·B T

[0077] where, r is the rank of the low-rank matrix ΔW, d is the eigen-dimension of the parameter matrix, and r << d;

[0078] Introducing the trainable low-rank matrix ΔW into the second teacher network reduces the computational overhead;

[0079] Step 4.2: Establish a fine-tuning loss function for the multi-teacher distillation model to optimize the loss function of the multi-teacher distillation model, and establish the total loss function of the multi-teacher distillation model;

[0080] Establish a fine-tuning loss function L FT , as shown in the following formula:

[0081]

[0082] where, y * is the target knowledge;

[0083] Based on the weighted fine-tuning loss function L FT of the multi-teacher distillation model and the distillation loss function L distill of the multi-teacher distillation model, obtain the total loss function L total of the multi-teacher distillation model, as shown in the following formula:

[0084] L total = L distill + β·L FT

[0085] where, β is the balance coefficient; in this embodiment, the balance coefficient β is set to 0.3.

[0086] Step 5: Introduce a self-induced distribution alignment mechanism to optimize the trained student model, and dynamically adjust the target weights for optimizing the student model according to the historical editing times to obtain an optimized student model;

[0087] Step 5.1: During the continuous knowledge editing process of the student model, introduce a self-induced distribution alignment mechanism to perform optimization training on the trained student model;

[0088] Introduce a self-induced distribution alignment mechanism and establish an optimized loss function \(L\) based on the JS divergence align , as shown in the following formula:

[0089]

[0090] where \(D\) JS (·) is the JS divergence operation;

[0091] The optimized loss function \(L\) based on the JS divergence align forces the output distributions of the student model with and without context information to be consistent, avoiding catastrophic forgetting of the student model.

[0092] Step 5.2: Dynamically adjust the target weight for optimizing the student model according to the historical edit count to obtain an optimized student model;

[0093] Dynamically adjust the target weight \(\gamma(t)\) for optimizing the student model according to the historical edit count \(t\), as shown in the following formula:

[0094] \(\gamma(t)=\gamma_0\cdot e\) -ηt \(\gamma(t)=\gamma_0\cdot e\) -ηt

[0095] where \(\gamma_0\) is the initial weight and \(\eta\) is the decay rate;

[0096] Gradually reduce the dependence of the student model on old knowledge based on the dynamic weight adjustment to obtain an optimized student model;

[0097] Step 6: Use the optimized student model to process the user input question and output the answer to the user input question.

[0098] In this embodiment, the optimized student model is used to test the knowledge editing accuracy, generalization, and locality on the zsRE, CounterFact, and MQuAKE datasets, and is compared with the parameter efficiency, editing speed, and multi-hop reasoning ability of methods such as ICE, DistillMIKE, and MEMIT. Tests on the zsRE and CounterFact datasets show that in terms of effectiveness, the effectiveness of the optimized student model in this embodiment reaches 98.5%, an improvement of 12.7% compared to the MEMIT method; in terms of generalization, the generalization index of the optimized student model in this embodiment reaches 94.2%, an improvement of 8.3% compared to the IKE method; in terms of locality, the locality of the optimized student model in this embodiment maintains a high level of 93.8%, avoiding excessive modification of irrelevant knowledge. Especially in the multi-hop question-answering scenario, the generalization correction accuracy of related knowledge is improved by 35% compared to traditional methods, demonstrating the deep knowledge association ability of the optimized student model.

[0099] In this embodiment, in 10,000 knowledge editing tasks, the optimized student model has a reasoning speed increased by more than 40% compared with the traditional IKE method, and the memory occupancy is reduced by 60%. Through a trainable low-rank matrix, context cues are implicitly injected into the model parameters, reducing the parameter update amount to 0.1% of full-scale fine-tuning while maintaining an editing accuracy of 97.3%. Through the self-induced distribution alignment mechanism, in 10 consecutive rounds of iterative editing, the knowledge forgetting rate is lower than 2.1%, which is significantly lower than the 47.6% knowledge forgetting rate of the traditional fine-tuning method. The language fluency PPL value is stable at 15.3 ± 0.5, which is more than 30% better than the baseline method. The self-induced distribution alignment mechanism enables the student model to maintain the integrity of the original knowledge system when absorbing new knowledge and avoid catastrophic forgetting.

[0100] In this embodiment, by fine-tuning the student model, the GPU video memory requirement is reduced from 48GB to 24GB (based on the A6000 graphics card), the single-inference latency of the student model is reduced from 850ms to 210ms, and only 15 minutes of fine-tuning is required for adding 1,000 new pieces of knowledge, with the efficiency improved by 200 times compared with full-scale retraining. Using a dual-teacher network to construct a multi-teacher distillation model improves the knowledge correction accuracy by 19% and reduces the output distribution variance by 63% based on KL divergence constraints, significantly enhancing the coherence of the generated text.

[0101] In this embodiment, through the combination of theoretical innovation and engineering optimization, breakthroughs are achieved in the three dimensions of knowledge editing accuracy, efficiency, and scalability, providing an efficient and reliable solution for the knowledge update of large-scale language models. Experiments prove that it still maintains the sublinear time growth characteristic when processing ten-thousand-level editing tasks, laying a foundation for industrial applications.

[0102] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit them. Although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements for some or all of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the scope defined by the claims of the present invention.

Claims

1. A large model distillation knowledge editing method integrating context constraints, characterized by: The following steps are involved: Step 1: Obtain the user input question and generate dynamic context information of the user input question; Step 2: Set the knowledge editing scope and the corresponding problem processing method, determine the knowledge editing scope to which the user input question belongs, and use the problem processing method corresponding to the knowledge editing scope to process the user input question; Step 3: Build a multi-teacher distillation model; Step 4: Introduce a trainable low-rank matrix and a fine-tuning loss function into the multi-teacher distillation model to optimize the loss function of the multi-teacher distillation model, perform fine-tuning training on the student model, and obtain a trained student model; Step 5: Introduce the self-induced distribution alignment mechanism to optimize the trained student model, dynamically adjust the target weight of the optimized student model according to the number of historical edits, and obtain the optimized student model; Step 6: Use the optimized student model to process the user input question and output the answer to the user input question.

2. According to claim 1, a large model distillation knowledge editing method integrating context constraints is characterized by: The specific method of step 1 is: Get the user input question q, and dynamically retrieve the edit fact set E related to the user input question q from the external knowledge base q ={e1,e2,...,e n }, where e1, e2, …, e n Generate explanatory context c of user input question q for the 1st, 2nd, ..., nth edit facts q ; Calculate the edit fact set E q The priority of each edit fact in the filter, the highest priority edit fact e q , generating dynamic context information of user input question q [c q ,e q ,q].

3. According to claim 2, a large model distillation knowledge editing method integrating context constraints is characterized by: The specific method of step 2 is: Set the knowledge editing scope based on the external knowledge base, including knowledge within the knowledge editing scope and knowledge outside the knowledge editing scope; the knowledge within the knowledge editing scope is the knowledge belonging to the external knowledge base, and the knowledge outside the knowledge editing scope is the knowledge not belonging to the external knowledge base; The problem handling methods corresponding to the knowledge editing scope include: For user input questions within the knowledge editing scope, dynamic context information [c q ,e q ,q] as the input question; For user input questions outside the scope of knowledge editing, the original user input question q is used as the input question; Use the range classifier to determine the knowledge editing range to which the user input question q belongs, and use the problem processing method corresponding to the knowledge editing range to process the user input question.

4. According to claim 3, a large model distillation knowledge editing method integrating context constraints is characterized by: The step 3 comprises: Step 3.1: Build a multi-teacher distillation model, including multiple teacher models and one student model; Step 3.2: Establish the distillation loss function of the multi-teacher distillation model.

5. According to claim 4, a large model distillation knowledge editing method integrating context constraints is characterized by: The specific method of step 3.1 is: The first teacher model is used to guide the student model to learn the dynamic context information of the input question within the knowledge editing scope, and integrate the dynamic context information into the parameters of the student model; The first teacher model obtains the user input question q that belongs to the knowledge editing scope and uses the ICE method to fuse dynamic context information [c q ,e q ,q], output the first teacher model distribution P teacher1 (y|q), where y is the answer to question q entered by the user; The second teacher model is used to constrain the weight parameters of the student model outside the knowledge editing scope, so that the student model retains the knowledge outside the knowledge editing scope; the second teacher model obtains the user input question q that belongs to the knowledge editing scope, and outputs the second teacher model distribution Among them, θ0 is the parameter of the second teacher model; The student model receives the first teacher model distribution P teacher1 (y|q) and the second teacher model distribution Output student model distribution Among them, θ s are the parameters of the student model.

6. The large model distillation knowledge editing method integrating context constraints according to claim 5 is characterized by: The specific method of step 3.2 is: Based on the weighted KL divergence of the student model distribution, the first teacher model distribution, and the second teacher model distribution, the distillation loss function L of the multi-teacher distillation model is established. distill , as shown in the following formula: Among them, α is the judgment probability output by the range classifier, D KL (·) is the KL divergence operation, λ c is the conflict penalty weight, x old Knowledge outside the scope of knowledge editing.

7. The large model distillation knowledge editing method integrating context constraints according to claim 6 is characterized by: The step 4 comprises: Step 4.1: Freeze the parameters θ0 of the second teacher model and introduce a trainable low-rank matrix into the second teacher network, as shown in the following formula: ΔW=A·B T in, r is the rank of the low-rank matrix ΔW, d is the characteristic dimension of the parameter matrix, r<<d; Step 4.2: Establish a multi-teacher distillation model fine-tuning loss function to optimize the loss function of the multi-teacher distillation model and establish the total loss function of the multi-teacher distillation model; Establish a multi-teacher distillation model to fine-tune the loss function L FT , as shown in the following formula: Among them, y * For target knowledge; Weighted multi-teacher distillation model fine-tuning loss function L FT And the distillation loss function L of the multi-teacher distillation model distill Get the total loss function L of the multi-teacher distillation model total , as shown in the following formula: L total =L distill +β·L FT Among them, β is the balance coefficient.

8. The large model distillation knowledge editing method integrating context constraints according to claim 7 is characterized by: The step 5 comprises: Step 5.1: In the process of continuous knowledge editing of the student model, a self-induced distribution alignment mechanism is introduced to optimize the trained student model; Step 5.2: Dynamically adjust the target weight of the optimized student model according to the number of historical edits to obtain the optimized student model.

9. The large model distillation knowledge editing method integrating context constraints according to claim 8, characterized in that: The specific method of step 5.1 is: Introduce the self-induced distribution alignment mechanism and establish the optimization loss function L based on JS divergence align , as shown in the following formula: Among them, D JS (·) is the JS divergence operation.

10. The large model distillation knowledge editing method integrating context constraints according to claim 9, characterized in that: The specific method of step 5.2 is: The target weight γ(t) of the student model is dynamically adjusted and optimized according to the number of historical edits t, as shown in the following formula: γ(t)=γ0·e -ηt γ(t)=γ0·e -ηt Among them, γ0 is the initial weight, η is the decay rate; Based on dynamic weight adjustment, the student model's dependence on old knowledge is gradually reduced to obtain an optimized student model.