Big language model knowledge editing system based on mixed experts

By introducing a hybrid expert architecture and keyword attention router in the large language model, combined with semantic-based data batch processing method, the problem of difficulty in balancing local modification and generalization capabilities of large language models when knowledge updates is solved, and high accuracy and high balanced knowledge editing effects are achieved.

CN120197694APending Publication Date: 2025-06-24NANJING UNIV OF AERONAUTICS & ASTRONAUTICS
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510160370.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-13
Publication Date
2025-06-24

AI Technical Summary

Technical Problem

When the existing large language models update knowledge, it is difficult to balance the model's local modification of specific inputs and the generalization ability of broad inputs, which leads to the model's understanding and application of other relevant knowledge when updating specific knowledge, and it is difficult for the existing technology to achieve a balance of high generalization and locality.

Method used

A large language model knowledge editing system based on hybrid experts is adopted, which includes a single-layer hybrid expert model editing adapter, keyword attention router, and semantic-based data batch processing module. By introducing additional expert models, using attention mechanisms and semantic similarity grouping techniques, model knowledge is dynamically edited to ensure accurate updates to specific knowledge without affecting the generalization ability of the model.

Benefits of technology

It realizes high accuracy and high balance in large language model knowledge editing, ensuring that when updating model knowledge, there is no adverse impact on other inputs of the model, significantly improving the accuracy and efficiency of editing, and reducing the demand for computing resources.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120197694A_ABST
    Figure CN120197694A_ABST
Patent Text Reader

Abstract

The invention discloses a big language model knowledge editing system based on mixed experts, and belongs to the field of natural language processing. According to the method, the hybrid expert architecture and the keyword attention router are combined, so that the model knowledge is dynamically updated under the condition that the original parameters of the large language model are kept unchanged; according to the single-layer bypass hybrid expert adapter provided by the invention, only single-layer additional experts are introduced into a model, and input with similar knowledge requirements is routed to the same expert through a keyword attention router, so that the experts can efficiently distinguish and process different types of knowledge information; the invention further provides a data batch processing method based on semantics, similar instances are grouped in the training stage, specialization of an expert model is promoted, and knowledge learning preference of a large language model is better met. According to the method, excellent performance is shown on models of various types and scales and in various editing tasks, and balance between generalization ability and local optimization is achieved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of natural language processing, and particularly to a large language model knowledge editing system based on mixture of experts. Background Art

[0002] With the rapid development of artificial intelligence technology, large language models have become one of the core technologies in the field of natural language processing. These models learn a vast amount of world knowledge through pre-training and can access and utilize this knowledge through natural language prompts. However, the dynamism of the real world requires these models to be updated regularly to correct outdated information or incorporate new knowledge. Frequently retraining or fine-tuning large language models to incorporate these updates is usually impractical as it requires substantial resources and time. To address this issue, the concept of knowledge editing (also known as model editing) has been introduced. The goal of knowledge editing is to effectively modify the output of a large language model for target knowledge while maintaining overall performance. In recent years, a variety of knowledge editing techniques have been proposed, but these methods either perform poorly overall or struggle to strike a balance between generalization and locality. Specifically, the evaluation of knowledge editing involves three dimensions: reliability, generalization, and locality. Existing methods generally perform poorly along these dimensions and cannot achieve a balance between high generalization and locality.

[0003] Previous research has revealed a core problem in large language models during model editing: when fine-tuning or updating the model, it is often difficult to balance local modifications to the model for specific inputs and its generalization ability for a wide range of inputs. This dilemma stems from the fact that when large language models process knowledge updates, they often need to find a delicate balance between maintaining the stability of the original knowledge structure and absorbing new knowledge. In traditional model editing methods, the editing process often leads to a decline in the performance of the model for non-target inputs or affects the model's understanding and application of other related knowledge when updating specific knowledge. For example, when the model is asked to update information about a country's president, ideally, the model should only display the updated information for queries related to that president, while for other unrelated queries, such as the capital or language of the country, its answers should remain unchanged. However, existing technologies often struggle to achieve this and frequently exhibit the "ripple effect", where a small update can cause extensive and unexpected changes in the model's knowledge structure.

[0004] In the process of knowledge update and editing of large language models, how to efficiently and precisely edit specific knowledge while maintaining the model's generalization ability is an urgent problem to be solved. Summary of the Invention

[0005] The present invention provides a large language model knowledge editing system based on a mixture of experts, which can effectively edit the large language model while ensuring that when updating the model knowledge, it has no adverse impact on other inputs of the model, achieving high accuracy and high balance in model editing.

[0006] An embodiment of the present invention provides a large language model knowledge editing system based on a mixture of experts, including:

[0007] A single-layer mixture of experts model editing adapter, which is used to introduce additional experts in a single layer of the initial model during the model editing process while keeping the original parameters unchanged;

[0008] A keyword attention router, which is set in the single-layer mixture of experts model editing adapter and is used to utilize the attention mechanism to direct the inputs with similar knowledge requirements to the same expert, while enabling the expert to distinguish similar inputs;

[0009] A semantic-based data batching module, which is used to promote the specialization of the expert model by grouping semantically similar instances during the training phase to conform to the knowledge learning preferences of the large language model.

[0010] Optionally, in an embodiment of the present invention, the single-layer mixture of experts model editing adapter is further used to introduce multiple parallel expert models through a bypass, embed the multiple expert models into the feed-forward network of the large language model, and only edit a single model layer in the large language model; through a gating decision mechanism, dynamically allocate the input data to different experts. The gating decision mechanism determines the participation degree of each expert according to different input information and generates the final output through a weighted aggregation method; during the forward propagation process, combine the frozen original parameters with the output of the newly added experts and generate the final result through a weighted fusion method.

[0011] Optionally, in an embodiment of the present invention, the keyword attention router is further used to identify the keywords in the input text using named entity recognition technology and assign corresponding values to each token, where the keyword tokens are assigned a value of 1 and the remaining tokens are assigned a value of 0; convert the input tokens into vector representations through the word embedding layer of the large language model, and calculate the correlation between each token and all experts based on the attention mechanism to obtain the attention value of each token; according to the calculated attention value, assign the input tokens to the relevant experts for processing to ensure that the input tokens with similar knowledge requirements are directed to the same expert.

[0012] Optionally, in an embodiment of the present invention, the semantic-based data batch processing module is further configured to calculate the semantic similarity between editing data instances. By using the word embedding layer of the model to be edited, each editing data instance is converted into a high-dimensional vector representation, and the cosine similarity is used to measure the similarity between vectors, thereby determining the semantic proximity of each data instance. A clustering algorithm is used to group the editing data.

[0013] Optionally, in an embodiment of the present invention, the word embedding layer is used to map the input discrete tokens to continuous text vectors in the semantic space.

[0014] Optionally, in an embodiment of the present invention, the clustering algorithm is used to divide the editing data instances into multiple clusters according to the calculated semantic similarity, and each cluster contains several instances with similar semantics.

[0015] Optionally, in an embodiment of the present invention, the expert model is based on a feed-forward network structure.

[0016] The large language model knowledge editing system based on hybrid experts in the embodiments of the present invention has the following beneficial effects:

[0017] First, the present invention realizes the precise update of specific knowledge in the large language model by introducing a hybrid expert architecture and a keyword attention router. This targeted editing method not only improves the model update efficiency but also minimizes the impact of the editing process on the overall performance of the model, thereby significantly improving the editing accuracy without sacrificing the model's generalization ability. Second, the present invention allows for flexible updating of the model's knowledge without retraining the entire model, which enables the model to quickly adapt to new knowledge or information without incurring high computational costs. In addition, through the semantic-based data batch processing method, the present invention further promotes the specialization of experts, enabling the model to be more precise and efficient in processing specific types of tasks, further enhancing the performance of model editing. In terms of computational resource optimization, the present invention significantly reduces the demand for computational resources while maintaining high performance, enabling efficient model editing even in resource-constrained environments, thereby reducing the cost of model updates. Finally, the experimental results show that the present invention exhibits excellent performance in various editing tasks, and can achieve high-accuracy and balanced editing in both different types and different scales of models, significantly superior to the prior art.

[0018] The additional aspects and advantages of the present invention will be partially given in the following description, partially become apparent from the following description, or be understood through the practice of the present invention. Description of the Drawings

[0019] The above and / or additional aspects and advantages of the present invention will become apparent and readily understood from the following description of embodiments in conjunction with the accompanying drawings, wherein:

[0020] Figure 1 Schematic diagram of a large language model knowledge editing system based on a mixture of experts according to an embodiment of the present invention. Detailed implementation manners

[0021] Embodiments of the present invention will be described in detail below. Examples of the embodiments are shown in the accompanying drawings, where the same or similar reference numerals denote the same or similar elements or elements having the same or similar functions throughout. The embodiments described below by referring to the accompanying drawings are exemplary and are intended to explain the present invention and should not be construed as limiting the present invention.

[0022] Figure 1 Schematic diagram of a large language model knowledge editing system based on a mixture of experts according to an embodiment of the present invention.

[0023] As Figure 1 shown, the large language model knowledge editing system based on a mixture of experts includes:

[0024] A single-layer mixture of experts model editing adapter for introducing additional experts into a single layer of the initial model during the model editing process while keeping the original parameters unchanged.

[0025] A keyword attention router disposed in the single-layer mixture of experts model editing adapter for using the attention mechanism to direct inputs requiring similar knowledge to the same expert while enabling the expert to distinguish similar inputs.

[0026] A semantic-based data batching module for promoting the specialization of the expert model by grouping semantically similar instances during the training phase to conform to the knowledge learning preferences of the large language model.

[0027] During the model editing process, the present invention enhances the accuracy of the model for specific knowledge updates by introducing additional experts into a single layer of the model while keeping the original parameters unchanged, while minimizing the impact on the general capabilities of the model.

[0028] In one embodiment of the present invention, the single-layer hybrid expert model editing adapter is further configured to introduce multiple parallel expert models through a bypass, embed the multiple experts into the feed-forward network of the large language model, and only edit a single model layer in the large language model. Through a gating decision mechanism, the input data is dynamically allocated to different experts. The gating decision mechanism determines the participation degree of each expert according to different input information, and generates the final output through weighted aggregation. During the forward propagation process, the adapter combines the frozen original parameters with the newly added expert outputs and generates the final result through weighted fusion.

[0029] In an embodiment of the present invention, the gating decision mechanism can determine the participation degree of each expert according to different input information, and generate the final output through weighted aggregation.

[0030] This mechanism improves the generalization performance of model editing by identifying keywords in the input and using these keywords as the basis for routing decisions.

[0031] In one embodiment of the present invention, the keyword attention router is further configured to use named entity recognition technology to identify keywords in the input text and assign corresponding values to each token, where the keyword tokens are assigned a value of 1 and the remaining tokens are assigned a value of 0; convert the input tokens into vector representations through the word embedding layer of the large language model, and calculate the correlation between each token and all experts based on the attention mechanism to obtain the attention value of each token; according to the calculated attention values, allocate the input tokens to the relevant experts for processing, ensuring that input tokens with similar knowledge requirements are directed to the same expert.

[0032] In an embodiment of the present invention, named entity recognition technology can identify keywords from the input text, such as locations, people, dates, organizations, etc.

[0033] In an embodiment of the present invention, the word embedding layer can map the input discrete tokens into continuous text vectors in the semantic space.

[0034] In one embodiment of the present invention, the semantic-based data batching module is further configured to calculate the semantic similarity between editing data instances. By using the word embedding layer of the model to be edited, each editing data instance is converted into a high-dimensional vector representation, and similarity metrics such as cosine similarity are used to measure the similarity between these vectors, thereby determining the semantic proximity of each data instance; use a clustering algorithm to group the editing data.

[0035] In an embodiment of the present invention, the clustering algorithm can divide the editing data instances into multiple clusters according to the calculated semantic similarity, and each cluster contains several instances with similar semantics.

[0036] In an embodiment of the present invention, the large language model is based on a Transformer structure, and the expert model is based on a feedforward network structure.

[0037] The following is a specific embodiment of the large language model knowledge editing system based on mixed experts of the present invention.

[0038] like Figure 1 As shown in the figure, the large language model knowledge editing system based on hybrid experts is divided into three parts: 1. Single-layer hybrid expert model editing adapter; 2. Keyword attention router; 3. Semantic-based data batch processing.

[0039] Single-layer hybrid expert model editing adapter:

[0040] The present invention integrates multiple parallel experts into the feedforward network (FFN) of a large language model through a bypass mechanism, thereby retaining the original parameters of the model and enhancing the locality of model editing. In addition, this adaptive modification is only applied to one layer in the model. Specifically, E experts are integrated, denoted as f1, f2, ..., f E , the input word is x. The present invention firstly uses a gating network to Assign different input tokens to different experts:

[0041]

[0042] Where i represents the i-th value in the vector, and R(·) is the routing strategy of the gating network. pk (·) will set all values ​​except the first k values ​​to zero. Then, by weighted aggregation of the calculation results of each expert on the input x, the final output h is obtained:

[0043]

[0044] Gated Network Determines the output h of the e-th expert i It should be noted that for experts, whose calculations can be omitted to save computing resources.

[0045] Overall, the forward process of this single-layer hybrid expert editing adapter combined with the frozen original parameters W0 can be expressed as:

[0046]

[0047] Among them, λ is a non-negative weighting coefficient used to balance the integration of old knowledge and new knowledge.

[0048] Keyword Attention Router:

[0049] The router first uses named entity recognition (NER) technology to identify the keyword in the input. All tokens identified by NER are regarded as keywords, including locations, people, dates, organizations, etc. Then, through the attention mechanism, the router assigns the word tokens corresponding to the keywords to the corresponding experts, ensuring that inputs with similar knowledge requirements are directed to the same expert. This method effectively captures and maintains the semantic associations of knowledge in the input data, thus enhancing the reliability and generalization ability of model editing.

[0050] Specifically, given an input sentence Use NER technology to identify the keyword x key . Assign the value 1 to the keyword and 0 to the rest of the tokens. These values will be used in the subsequent steps to help the model pay more attention to the keyword:

[0051]

[0052] Through the word embedding layer of the large language model, obtain the vector representation e i of the word token x i ∈ R d . Use the embedding layer of the model to be edited as the embedding layer of the router. Given E experts, input e i into the attention mechanism as follows:

[0053] Q = W q e i , K = W k e i , V = W v e i

[0054] A(e i ) = softmax(QK T )V

[0055] where Wq, Wk, Wv ∈ R E×d are learnable parameters. To further help the gating network focus on the keyword, the routing function can be expressed as:

[0056]

[0057] Semantic-based data batching:

[0058] This method effectively promotes the specialization of experts in the model editing process by intelligently grouping data during the training phase, better conforming to the knowledge learning preferences of the large language model. The core idea of this method is to group semantically similar editing instances together to more efficiently utilize the learning ability of the model and optimize the effect of model editing.

[0059] Specifically, first, it is necessary to calculate the semantic similarity between edited data instances. The word embedding layer of the model is used to convert the text into a vector representation in a high-dimensional space. Then, cosine similarity is used to calculate the similarity between these vectors, thereby measuring the semantic proximity of the data instances. Finally, a clustering algorithm (such as K-means) is used to group the edited data. The clustering algorithm divides the data instances into several clusters according to the similarity, and each cluster contains several instances that are semantically similar. The advantage of this grouping strategy is that when the model processes this data, it can form a dedicated processing strategy for the instances in each cluster, enhancing the specialization ability of the expert model. The expert model can concentrate on processing semantically similar data and improve the processing effect for specific knowledge.

[0060] The large language model knowledge editing system based on hybrid experts proposed according to the embodiments of the present invention combines a hybrid expert architecture and a keyword attention router, realizing the dynamic update of model knowledge while keeping the original parameters of the large language model unchanged. The present invention proposes a single-layer bypass hybrid expert adapter, which only introduces a single layer of additional experts into the model and routes the inputs with similar knowledge requirements to the same expert through the keyword attention router, enabling the experts to efficiently distinguish and process different types of knowledge information. In addition, the present invention also proposes a semantic-based data batch processing method, which promotes the specialization of the expert model by grouping similar instances during the training phase, better conforming to the knowledge learning preferences of the large language model. Experimental results show that the present invention exhibits excellent performance on various types and scales of models, as well as various editing tasks, achieving a balance between generalization ability and local optimization.

[0061] In the description of this specification, the description referring to terms such as "one embodiment", "some embodiments", "example", "specific example", or "some examples" means that the specific features, structures, materials, or characteristics described in connection with the embodiment or example are included in at least one embodiment or example of the present invention. In this specification, the schematic representations of the above terms do not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials, or characteristics described can be combined in a suitable manner in any one or N embodiments or examples. In addition, without contradiction, those skilled in the art can combine and combine the different embodiments or examples described in this specification and the features of different embodiments or examples.

[0062] Furthermore, the terms "first" and "second" are used for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the quantity of the indicated technical features. Thus, features defined with "first" and "second" may explicitly or implicitly include at least one of such features. In the description of the present invention, the meaning of "N" is at least two, such as two, three, etc., unless otherwise specifically defined.

Claims

1. A large language model knowledge editing system based on hybrid experts, characterized in that: include: A single-layer mixed-expert model editing adapter is used in the model editing process to introduce additional experts in a single layer of the initial model while keeping the original parameters unchanged; A keyword attention router, disposed in the single-layer hybrid expert model editing adapter, for utilizing an attention mechanism to direct inputs requiring similar knowledge to the same expert while enabling the expert to distinguish between similar inputs; The semantic-based data batch processing module is used to promote the specialization of the expert model by grouping semantically similar instances during the training phase to meet the knowledge learning preference of the large language model.

2. The system according to claim 1, characterized in that The single-layer hybrid expert model editing adapter is further used to introduce multiple parallel expert models through bypass, embed multiple expert models into the feedforward network of the large language model, and only edit a single model layer in the large language model; dynamically distribute input data to different experts through a gated decision mechanism, and the gated decision mechanism determines the degree of participation of each expert according to different input information, and generates the final output through weighted aggregation; in the forward propagation process, the frozen original parameters are combined with the newly added expert output, and the final result is generated by weighted fusion.

3. The system according to claim 1, characterized in that The keyword attention router is further used to identify keywords in the input text using named entity recognition technology and assign corresponding values ​​to each token, where the keyword token is assigned a value of 1 and the remaining tokens are assigned a value of 0; the input tokens are converted into vector representations through the word embedding layer of the large language model, and based on the attention mechanism, the correlation between each token and all experts is calculated to obtain the attention value of each token; According to the calculated attention value, the input tokens are assigned to the experts related to them for processing, ensuring that input tokens with similar knowledge requirements are directed to the same experts.

4. The system according to claim 1, characterized in that The semantic-based data batch processing module is further used to calculate the semantic similarity between edited data instances, convert each edited data instance into a high-dimensional vector representation by using the word embedding layer of the model to be edited, and use cosine similarity to measure the similarity between vectors, thereby determining the semantic proximity of each data instance, and using a clustering algorithm to group the edited data.

5. The system according to claim 4, characterized in that The word embedding layer is used to map the input discrete word units into continuous text vectors in the semantic space.

6. The system according to claim 4, characterized in that The clustering algorithm is used to divide the edit data instances into multiple clusters according to the calculated semantic similarity, and each cluster contains a number of instances with similar semantics.

7. The system according to any one of claims 1 to 6, characterized in that: The expert model is based on a feedforward network structure.

Citation Information

Cited By

  • Voice and music collaborative generation method and system based on dynamic mixed attention and expert architecture, terminal equipment and medium

    CN121545490A