Processing method and device for updating model knowledge by optimizing MLP weight

By optimizing the MLP weight in the large language model, the problems of high computational complexity and cost during knowledge update in the existing technology are solved, and a more efficient and flexible knowledge update process is achieved.

CN120218209AActive Publication Date: 2025-06-27BEIJING DP TECH CO LTD

Patent Information

Application Number
CN202510337895.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-21
Publication Date
2025-06-27
Estimated Expiration
2045-03-21

AI Technical Summary

Technical Problem

Existing large language models have a high computational complexity and high computational volume when knowledge is updated, resulting in long training cycles and high cost, especially when small batches or individual knowledge is updated.

Method used

By optimizing multi-layer feedforward neural network (MLP) weights, only the weight parameters of the MLP layer are optimized, rather than the overall model parameters. Specific steps include: creating forward and counterfactual questions-answer text pairs for the batch knowledge entries specified by the user; extracting irrelevant knowledge from the pre-trained knowledge corpus; weight optimization of the MLP layer based on these data records, and performing counterfactual and irrelevant knowledge evaluation.

Benefits of technology

It reduces the computational complexity and calculation amount, shortens the training cycle, improves update efficiency and flexibility, and reduces update costs.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120218209A_ABST
    Figure CN120218209A_ABST
Patent Text Reader

Abstract

The embodiment of the invention relates to a processing method and device for updating model knowledge by optimizing MLP weight. The method comprises the following steps: creating corresponding first and second data records for knowledge entries of an update knowledge set specified by a user; extracting pre-training knowledge corpora irrelevant to the updated knowledge set from a pre-training knowledge corpus of a target model specified by a user to form a third data record set; weight parameter optimization is carried out on all MLP layers of the target model based on all the first data records; performing anti-fact problem evaluation on the updated target model based on all the second data records; performing irrelevant knowledge problem evaluation on the updated target model based on the third data record set; and identifying whether the first evaluation result and the second evaluation result are both passed, if not, continuing to optimize until the latest first evaluation result and the latest second evaluation result are both passed, and if yes, feeding back updating completion to the current user. The updating efficiency can be improved, and the updating cost can be reduced.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of data processing, and particularly relates to a processing method and device for updating model knowledge by optimizing MLP weights. Background Art

[0002] Large Language Models (LLMs) can memorize a vast amount of knowledge information and perform various Natural Language Processing (NLP) tasks, such as text generation, machine translation, intelligent question answering, text classification, etc. To update the knowledge of a pre-trained large language model, it is generally achieved through model fine-tuning. Briefly, the conventional fine-tuning scheme is as follows: collect the knowledge information to be updated to construct a fine-tuning corpus, and optimize the overall model parameters of the large language model based on the fine-tuning corpus until convergence. This conventional scheme is relatively conservative, and the optimization object is the overall model parameters. Although it can achieve the maximized optimization effect only in terms of mathematical principles, there are obvious defects in the actual processing process: the parameter scale of any current large language model is extremely large. If fine-tuning is performed according to the full model parameter optimization method of the conventional scheme, the resulting computational complexity and computational amount will be very large, thus making it difficult to shorten the training cycle and reduce the training cost. Such an update cost is difficult to accept when performing small-batch or individual knowledge updates.

[0003] Most of the current common large language models (such as GPT, BERT, LLaMA, T5, etc.) are mostly implemented based on the Transformer model, and their internal components include multiple MLP networks (also called MLP layers), and each MLP layer corresponds to a set of weight parameters. Many studies have shown that large language models will learn a vast amount of knowledge corpora during the pre-training stage and mainly achieve the processing effect of knowledge memorization by optimizing the weight parameters of each MLP layer. Since this is the case, it means that only optimizing the MLP weights can also achieve the purpose of updating knowledge, and can greatly reduce the computational complexity and computational amount, thereby achieving the purpose of shortening the training cycle, improving the update efficiency and flexibility, and reducing the update cost. And how to achieve the processing effect of updating knowledge by optimizing the MLP weights is exactly the technical problem to be solved by the present invention. Summary of the Invention

[0004] The objective of the present invention is to provide a processing method, apparatus, electronic device, and computer-readable storage medium for updating model knowledge by optimizing MLP weights in view of the deficiencies of the prior art. The present invention uses a batch of knowledge entries specified by the user as the update knowledge set; counts the total number of knowledge entries in the update knowledge set to obtain a total number N1; uses the large language model currently specified by the user (a large language model implemented based on the Transformer model structure and having completed pre-training and NLP task fine-tuning) as the target model; creates a first data record consisting of positive question-answer text pairs and a second data record consisting of counterfactual question-answer text pairs for each knowledge entry in the update knowledge set, and extracts a third data record set of pre-training knowledge corpus irrelevant to all knowledge entries from the pre-training knowledge corpus of the target model; optimizes the weight parameters of all MLP layers of the target model based on all the first data records; performs counterfactual question evaluation on the updated target model based on all the second data records to obtain a first evaluation result; performs irrelevant knowledge question evaluation on the updated target model based on the third data record set to obtain a second evaluation result; and when both the first and second evaluation results are passed, feedback to the user that the model knowledge update is complete. The present invention only optimizes the MLP weights. Through the present invention, the computational complexity and amount of calculation can be reduced, the training cycle can be shortened, the update efficiency and flexibility can be improved, and the update cost can be reduced.

[0005] To achieve the above objective, a first aspect of an embodiment of the present invention provides a processing method for updating model knowledge by optimizing MLP weights, the method comprising:

[0006] Using a batch of knowledge entries specified by the user as the update knowledge set; counting the total number of knowledge entries in the update knowledge set to obtain a total number N1; using the large language model currently specified by the user as the target model; the large language model specified by the user is a large language model implemented based on the Transformer model structure and having completed pre-training and NLP task fine-tuning, the NLP tasks include at least text generation tasks, machine translation tasks, intelligent question-answering tasks, text classification tasks; the total number N1 is a positive integer, at least 1; the update knowledge set consists of N1 knowledge entries (s i ,r i ,o i ), where 1 ≤ knowledge index i ≤ N1, s i 、r i 、o i are respectively the knowledge theme, theme-object relationship, and knowledge object of the knowledge triple;

[0007] For each of the knowledge entries (s i ,r i ,oi ) Create a first data record corresponding to a question-answer text pair; and for each of the knowledge entries (s i ,r i ,o i ) Create a second data record corresponding to a counterfactual question-answer text pair; and extract from the pre-trained knowledge corpus corresponding to the target model the pre-trained knowledge corpus unrelated to all the knowledge entries (s i ,r i ,o i ) to form a corresponding third data record set;

[0008] Optimize the weight parameters of all MLP layers of the target model based on N1 of the first data records;

[0009] Evaluate the counterfactual questions of the updated target model based on N1 of the second data records to obtain a corresponding first evaluation result; and evaluate the unrelated knowledge questions of the updated target model based on the third data record set to obtain a corresponding second evaluation result; both the first and second evaluation results include pass and fail;

[0010] Identify whether both the first and second evaluation results are pass; if not, continue to optimize until both the latest first and second evaluation results are pass evaluations; if so, feedback to the current user that the model knowledge update is complete.

[0011] Preferably, all internal modules of the target model are divided into two major sections: a preprocessing section and a forward inference section; specifically: within the target model, the internal modules for tokenizing the model input text and the internal modules for embedding and encoding the token sequence are incorporated into the preprocessing section; within the target model, all internal modules involved in the forward inference of NLP tasks based on the initial vector output by the preprocessing section are incorporated into the forward inference section;

[0012] The preprocessing section is used to tokenize the model input text to obtain a corresponding token sequence; and perform embedding and encoding processing on the token sequence according to the embedding encoding rules of the target model and denote the obtained embedding encoding vector as the initial vector H0 and send it to the forward inference section; the token sequence consists of multiple tokens; the initial vector H0 consists of multiple token initial vectors h0; the token initial vectors h0 correspond one-to-one with the tokens;

[0013] The forward inference section is used to perform forward inference based on the input initial vector H0 to obtain a corresponding generated text and output it; the total number of MLP layers in the forward inference section is denoted as the total number N2, and each MLP layer is labeled with a corresponding M jlayer; the total number N2 is a positive integer; 1 ≤ layer index j ≤ N2; each of the M j The input and output vectors of each layer during the inference process are denoted as the corresponding process vector H in,j Process vector H out,j ; The process vector H in,j Consists of multiple tokenization process vectors h in,j ; The process vector H out,j Consists of multiple tokenization process vectors h out,j ; The tokenization process vector h in,j 、h out,j Corresponds one-to-one with the tokenization or the initial tokenization vector h0;

[0014] Each of the M j The inference process of the layer is:

[0015] H out,j = W out,j σ(W in,j γ(H in,j ))

[0016] where σ is a preset activation function, γ is a preset normalization function, and W in,j 、W out,j are the input and output layer weights of the M j layer respectively;

[0017] The first data record corresponds one-to-one with the knowledge entry (s i , r i , o i ); The first data record includes a first question text and a first answer label; The first question text is a natural language question text containing the corresponding knowledge topic s i and the topic-object relationship r i ; The first answer label matches the corresponding knowledge object o i ;

[0018] The second data record corresponds one-to-one with the knowledge entry (s i , r i , o i ); The second data record includes a second question text and a second answer label; The second question text is a classification question text that has a counterfactual relationship with the corresponding knowledge entry (s i , r i , o i ); The second answer label is the answer text corresponding to the second question text, and its answer content includes two classification results: yes and no;

[0019] The pre-trained knowledge corpus includes a plurality of the pre-trained knowledge corpora (s tr ,r tr ,o tr ); s tr 、r tr 、o tr are respectively the knowledge topic, topic-object relationship, and knowledge object of the knowledge triple;

[0020] The third data record set includes a plurality of third data records; the third data record includes a third question text and a third answer label; each of the third data records corresponds to a pre-trained knowledge corpus (s i ,r i ,o i ) that has nothing to do with all the knowledge entries (s tr ,r tr ,o tr ); the third question text is a natural language question text that includes the corresponding knowledge topic (s tr and the topic-object relationship r tr ; the third answer label matches the corresponding knowledge object o tr .

[0021] Preferably, for each of the knowledge entries (s i ,r i ,o i ) of the updated knowledge set, a question-answer text pair is created to form a corresponding first data record, specifically including:

[0022] Configure a question text generation instruction template for the target model, denoted as the first instruction template; the configurable parameters of the first instruction template include a topic configuration parameter, a topic-object relationship configuration parameter, and an object configuration parameter; the first instruction template is a formatted instruction text template; the first instruction template is used to require the target model to generate a question text with the object configuration parameter as the expected answer using the topic configuration parameter and the topic-object relationship configuration parameter as the question text elements;

[0023] Take each of the knowledge entries (s i ,r i ,o i ) of the updated knowledge set as the corresponding current knowledge entry one by one; and use the knowledge topic s i , the topic-object relationship r i and the knowledge object o ias the corresponding current knowledge topic, current topic-object relationship, and current knowledge object; and set the topic configuration parameter, the topic-object relationship configuration parameter, and the object configuration parameter of the first instruction template to the corresponding current knowledge topic, the current topic-object relationship, and the current knowledge object; and use the first instruction template with parameter settings completed as the corresponding first instruction text to input into the target model for question text generation processing, and use the question text obtained from this processing as the corresponding first question text; and use the current knowledge object as the corresponding first answer label; and form a corresponding first data record from the first question text and the first answer label corresponding to the current knowledge entry.

[0024] Preferably, for each of the knowledge entries (s i , r i , o i ) create a counterfactual question-answer text pair to form a corresponding second data record, specifically including:

[0025] Configure another question text generation instruction template for the target model, denoted as the second instruction template; the configurable parameters of the second instruction template include a topic configuration parameter, a topic-object relationship configuration parameter, and an object configuration parameter; the second instruction template is a formatted instruction text template; the second instruction template is used to require the target model to form a current fact triple from the topic configuration parameter, the topic-object relationship configuration parameter, and the object configuration parameter, and generate a classification question text that has a counterfactual relationship with the current fact triple, and generate a corresponding binary classification answer text for this classification question text, and require that this binary classification answer text can only be yes or no;

[0026] Take each of the knowledge entries (s i , r i , o i ) of the updated knowledge set one by one as the corresponding current knowledge entry; and use the knowledge topic s i , the topic-object relationship r i , and the knowledge object o ias the corresponding current knowledge topic, current topic-object relationship, and current knowledge object; and set the topic configuration parameter, topic-object relationship configuration parameter, and object configuration parameter of the second instruction template as the corresponding current knowledge topic, current topic-object relationship, and current knowledge object; and use the second instruction template with parameter settings completed as the corresponding second instruction text to input into the target model for question and answer text generation processing, and use the classified question text and binary classification answer text obtained from this processing as the corresponding second question text and second answer label; and form a corresponding second data record from the second question text and second answer label corresponding to the current knowledge entry.

[0027] Preferably, extract the pre-trained knowledge corpus irrelevant to all the knowledge entries (s i ,r i ,o i ) from the pre-trained knowledge corpus corresponding to the target model to form a corresponding third data record set, specifically including:

[0028] For each knowledge topic s in the pre-trained knowledge corpus tr that is irrelevant to all the knowledge entries (s i ,r i ,o i ), mark the pre-trained knowledge corpus (s i ,r tr ,o tr ,o tr ) as irrelevant corpus;

[0029] And use each irrelevant corpus as the corresponding current knowledge entry one by one; and use the knowledge topic s tr , the topic-object relationship r tr , and the knowledge object o tr of the current knowledge entry as the corresponding current knowledge topic, current topic-object relationship, and current knowledge object; and set the topic configuration parameter, topic-object relationship configuration parameter, and object configuration parameter of the first instruction template as the corresponding current knowledge topic, current topic-object relationship, and current knowledge object; and use the first instruction template with parameter settings completed as the corresponding first instruction text to input into the target model for question text generation processing, and use the question text obtained from this processing as the corresponding third question text; and use the current knowledge object as the corresponding third answer label; and form a corresponding third data record from the third question text and third answer label corresponding to the current knowledge entry;

[0030] And all the obtained third data records form the corresponding third data record set.

[0031] Preferably, optimizing the weight parameters of all MLP layers of the target model based on N1 of the first data records specifically includes:

[0032] Step 601, record the current overall model parameters of the target model as parameter θ;

[0033] Step 602, use the first problem text of each of the first data records as the current model input text to process the target model, and record the token sequence and the initial vector H0 output by the preprocessing section during this processing as the corresponding token sequence U i and the initial vector H 0,i , and record the process vectors H j input and output by each of the M in,j layers during this processing as the corresponding process vectors H out,j , H in,j,i , and cache all the process vectors H out,j,i during this processing; in,j,i , H out,j,i

[0034] Among them, the token sequence U i corresponds one-to-one with the knowledge entry (s i , r i , o i ) and is composed of multiple tokens u i , and one of them corresponds to the knowledge topic s i ; the initial vector H 0,i corresponds one-to-one with the knowledge entry (s i , r i , o i ) and is composed of multiple token initial vectors h 0,i , and the token initial vector h 0,i corresponds one-to-one with the token u i ; the process vectors H in,j,i , H out,j,i correspond one-to-one with the knowledge entry (s i , r i , o i ); the process vector H in,j,i is composed of multiple token process vectors h in,j,i ; the H out,j,i is composed of multiple token process vectors h out,j,i ;

[0035] Step 603, for each of the knowledge entries (si , r i , o i The process vector H of the corresponding last layer out,j=N2,i The last of the tokenization process vectors is extracted as the corresponding end-layer tail word vector And for each of the end-layer tail word vectors a global increment δ is set whose vector length and vector feature dimension are both the same as the current end-layer tail word vector and are kept consistent i ; And all the global increments δ i are initialized to all-zero vectors;

[0036] Step 604, and based on the negative log-likelihood loss function and N1 groups of the end-layer tail word vectors the global increment δ i and the knowledge object o i a corresponding optimization objective function is set; and in the direction of minimizing the optimization objective function, the optimal solution of each global increment δ i is solved and the solution result is used as the corresponding optimal global increment And based on each end-layer tail word vector and its corresponding optimal global increment the corresponding target tail word vector e is calculated tag,i ;

[0037] Among them, the optimization objective function is:

[0038]

[0039] L NLL is the negative log-likelihood loss function, is the conditional probability that when the overall model parameters of the target model are the parameter θ, after adjusting the corresponding end-layer tail word vectors through each global increment δ i the generated text of the target model is the corresponding knowledge object o ; i The calculation method of the target tail word vector e

[0040] is: tag,i

[0041]

[0042] Step 605, a counter C initialized to 1 is set;

[0043] Step 606, for each of the knowledge entries (s i , r i,o i ) The corresponding first problem text is input into the target model as the current model input text, and for each of the M j layer input and output process vectors H in,j,i 、H out,j,i are cached again;

[0044] Step 607, the M j layer where the layer index j matches the counter C is used as the current MLP layer;

[0045] Step 608, among the N2 most recently cached process vectors H i ,r i ,o i ) corresponding to each knowledge entry (s in,j,i , the process vector H i corresponding to the knowledge topic s in,j=C,i and the word segmentation process vector h subject,C,i corresponding to the current MLP layer are denoted as the corresponding topic vector h subject,C,i ; and based on each of the topic vectors h in,j=C and the corresponding input layer weight W j=C,i the latest single-layer key vector k j=C,i is estimated;

[0046] Among them, the estimation method of the single-layer key vector k j=C,i k in,j=C =σ(W subject,C,i γ(h i ));

[0048] σ is a preset activation function, and γ is a preset normalization function;

[0049] Step 609, among the N2 most recently cached process vectors H i ,r i ,o out,j,i ) corresponding to each knowledge entry (s out,j=C,i , the process vector H out,j=C,i corresponding to the current MLP layer, the last word segmentation process vector h C,i is denoted as the single-layer end word vector e tag,i ; and based on the target end word vector e C,i and the single-layer end word vector e j,i the corresponding single-layer local increment △δ j,i is estimated;

[0050] Among them, the estimation method of the single-layer local increment △δ j=C,i is:

[0051]

[0052] Step 610: Compose a corresponding key vector matrix K from N1 of the single-layer key vectors k j=C,i ; and compose a corresponding local increment matrix R from N1 of the single-layer local increments △δ j=C ; and estimate a corresponding single-layer increment weight △W based on the key vector matrix K j=C,i ; and the local increment matrix R j=C ; j=C and estimate a corresponding single-layer increment weight △W j=C ; out,j=C ;

[0053] wherein, the estimation method of the single-layer increment weight △W out,j=C is as follows:

[0054]

[0055] where X j=C is the covariance matrix parameter of the preset j = C layer;

[0056] Step 611: Update the output layer weight W of the current MLP layer based on the single-layer increment weight △W out,j=C ; and update the parameter θ of the target model based on the new output layer weight W out,j=C ; out,j=C New W

[0057] = Old W out,j=C + ΔW out,j=C ; out,j=C ;

[0058] Step 612: Increment the counter C by 1; and identify whether the incremented counter C exceeds the total number N2; if not, return to step 606; if so, confirm that this optimization is complete.

[0059] Preferably, the counterfactual problem evaluation of the updated target model based on N1 of the second data records to obtain a corresponding first evaluation result specifically includes:

[0060] Successively use each of the second data records as the corresponding current record; use the second question text of the current record as the current model input text to be processed by the target model, and use the generated text obtained from the current model processing as the corresponding current answer; identify whether the current answer matches the second answer label of the current record. If it does not match, set the corresponding first Q&A result as non-matching. If it matches, set the corresponding first Q&A result as matching; identify whether all of the obtained N1 first Q&A results are matching. If so, set the corresponding first evaluation result as passed. Otherwise, set the corresponding first evaluation result as failed.

[0061] Preferably, the obtaining of the corresponding second evaluation result by performing an irrelevant knowledge question evaluation on the updated target model based on the third data record set specifically includes:

[0062] Step 81, perform a round of traversal on all the third data records in the third data record set; during this round of traversal, use the currently traversed third data record as the corresponding current record; use the third question text of the current record as the current model input text to be processed by the target model, and use the generated text obtained from the current model processing as the corresponding first predicted answer; form a corresponding first prediction-label pair from the first predicted answer and the third answer label of the current record; at the end of this round of traversal, input all the obtained first prediction-label pairs into a preset first model loss function for calculation to obtain the corresponding first loss value;

[0063] Wherein, the first model loss function is implemented based on a cross-entropy loss function or a negative log-likelihood loss function;

[0064] Step 82, identify whether the first loss value meets a preset first loss value range; if it meets, set the corresponding second evaluation result as passed; if it does not meet, set the corresponding second evaluation result as failed.

[0065] A second aspect of the embodiments of the present invention provides an apparatus for implementing the processing method for updating model knowledge by optimizing MLP weights described in the first aspect above. The apparatus includes: a data receiving module, a preprocessing module, an MLP optimization module, an optimization evaluation module, and an update feedback module;

[0066] The data receiving module is used to take the batch of knowledge entries specified by the user as the updated knowledge set; count the total number of knowledge entries in the updated knowledge set to obtain the total number N1; and take the large language model specified by the current user as the target model; the large language model specified by the user is a large language model implemented based on the Transformer model structure and has completed pre-training and NLP task fine-tuning, and the NLP tasks include at least text generation tasks, machine translation tasks, intelligent question answering tasks, and text classification tasks; the total number N1 is a positive integer and at least 1; the updated knowledge set consists of N1 knowledge entries (s i ,r i ,o i ), where 1 ≤ knowledge index i ≤ N1, s i 、r i 、o i are the knowledge subject, subject-object relationship, and knowledge object of the knowledge triple respectively.

[0067] The preprocessing module is used to create a question-answer text pair for each of the knowledge entries (s i ,r i ,o i ) in the updated knowledge set to form a corresponding first data record; create a counterfactual question-answer text pair for each of the knowledge entries (s i ,r i ,o i ) to form a corresponding second data record; and extract the pre-trained knowledge corpus unrelated to all the knowledge entries (s i ,r i ,o i ) from the pre-trained knowledge corpus corresponding to the target model to form a corresponding third data record set.

[0068] The MLP optimization module is used to optimize the weight parameters of all MLP layers of the target model based on the N1 first data records.

[0069] The optimization evaluation module is used to perform a counterfactual question evaluation on the updated target model based on the N1 second data records to obtain a corresponding first evaluation result; and perform an unrelated knowledge question evaluation on the updated target model based on the third data record set to obtain a corresponding second evaluation result; both the first and second evaluation results include pass and fail.

[0070] The update feedback module is used to identify whether both the first and second evaluation results are pass; if not, continue to optimize until both the latest first and second evaluation results are pass; if so, feedback to the current user that the model knowledge update is complete.

[0071] In a third aspect of the embodiments of the present invention, an electronic device is provided, including: a memory, a processor, and a transceiver;

[0072] The processor is used to be coupled with the memory, read and execute instructions in the memory to implement the method steps described in the first aspect above;

[0073] The transceiver is coupled with the processor, and the processor controls the transceiver to perform message sending and receiving.

[0074] In a fourth aspect of the embodiments of the present invention, a computer-readable storage medium is provided. The computer-readable storage medium stores computer instructions. When the computer instructions are executed by a computer, the computer is caused to execute the instructions of the method described in the first aspect above.

[0075] The embodiments of the present invention provide a processing method, apparatus, electronic device, and computer-readable storage medium for updating model knowledge by optimizing MLP weights. As can be seen from the above content, the embodiments of the present invention use the batch knowledge entries specified by the user as the updated knowledge set; count the total number of knowledge entries in the updated knowledge set to obtain the total number N1; use the currently user-specified large language model (a large language model implemented based on the Transformer model structure and having completed pre-training and NLP task fine-tuning) as the target model; create a first data record consisting of positive question-answer text pairs and a second data record consisting of counterfactual question-answer text pairs for each knowledge entry in the updated knowledge set, and extract the pre-training knowledge corpus irrelevant to all knowledge entries from the pre-training knowledge corpus of the target model to form a third data record set; optimize the weight parameters of all MLP layers of the target model based on all the first data records; perform counterfactual question evaluation on the updated target model based on all the second data records to obtain a first evaluation result; perform irrelevant knowledge question evaluation on the updated target model based on the third data record set to obtain a second evaluation result; and when both the first and second evaluation results are passed, feedback to the user that the model knowledge update is complete. The embodiments of the present invention only optimize the MLP weights. Through the embodiments of the present invention, not only the computational complexity and amount of calculation are reduced, the training cycle is shortened, and the update cost is reduced, but also the update efficiency and update flexibility are improved. Description of the Drawings

[0076] Figure 1 It is a schematic diagram of a processing method for updating model knowledge by optimizing MLP weights provided in Embodiment 1 of the present invention;

[0077] Figure 2 It is a module structure diagram of a processing apparatus for updating model knowledge by optimizing MLP weights provided in Embodiment 2 of the present invention;

[0078] Figure 3 This is a schematic structural diagram of an electronic device provided in Embodiment 3 of the present invention. Detailed implementation manners

[0079] In order to make the objectives, technical solutions, and advantages of the present invention clearer, the present invention will be further described in detail below with reference to the accompanying drawings. Apparently, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.

[0080] Embodiment 1 of the present invention provides a processing method for updating model knowledge by optimizing MLP weights, as Figure 1 shown in the schematic diagram of a processing method for updating model knowledge by optimizing MLP weights provided in Embodiment 1 of the present invention. The method mainly includes the following steps:

[0081] Step 1: Use the batch of knowledge entries specified by the user as the updated knowledge set; count the total number of knowledge entries in the updated knowledge set to obtain the total number N1; and use the large language model currently specified by the user as the target model.

[0082] Here, the large language model in the embodiment of the present invention is a large language model implemented based on the Transformer model structure and has completed pre-training and NLP task fine-tuning. The NLP tasks of the large language model at least include text generation tasks, machine translation tasks, intelligent question and answer tasks, and text classification tasks. The total number N1 is a positive integer, at least 1; the updated knowledge set consists of N1 knowledge entries (s i , r i , o i ), where 1 ≤ knowledge index i ≤ N1, and s i , r i , o i are the knowledge topic, topic-object relationship, and knowledge object of the knowledge triple, respectively.

[0083] It should be noted that all internal modules of the target model in the embodiment of the present invention are divided into two major sections: a preprocessing section and a forward inference section; specifically: 1) In the target model, the internal module for tokenizing the model input text and the internal module for embedding and encoding the token sequence are included in the preprocessing section; 2) In the target model, all internal modules involved in the forward inference of NLP tasks based on the initial vector output by the preprocessing section are included in the forward inference section.

[0084] The preprocessing section in the embodiments of the present invention is used to perform word segmentation on the input text of the model to obtain a corresponding word segmentation sequence; and perform embedding encoding on the word segmentation sequence according to the embedding encoding rules of the target model, and record the obtained embedding encoding vector as the initial vector H0 and send it to the forward inference section; where the word segmentation sequence is composed of multiple word segments; the initial vector H0 is composed of multiple word segment initial vectors h0; the word segment initial vector h0 corresponds one-to-one with the word segment.

[0085] The forward inference section in the embodiments of the present invention is used to perform forward inference based on the input initial vector H0 to obtain a corresponding generated text and output it. Among them, the total number of MLP layers in the forward inference section is denoted as the total number N2, and each MLP layer is marked as the corresponding M j layer; the total number N2 is a positive integer; 1 ≤ layer index j ≤ N2.

[0086] In each M j layer of the embodiments of the present invention, the input and output vectors during the inference process are denoted as the corresponding process vector H in,j process vector H out,j . Among them, the process vector H in,j is composed of multiple word segment process vectors h in,j ; the process vector H out,j is composed of multiple word segment process vectors h out,j ; the word segment process vectors h in,j , h out,j correspond one-to-one with the word segment or the word segment initial vector h0.

[0087] In each M j layer of the embodiments of the present invention, the inference process is as follows:

[0088] H out,j = W out,j σ(W in,j γ(H in,j ));

[0089] Among them, σ is a preset activation function, γ is a preset normalization function, and W in,j , W out,j are the input and output layer weights of the M j layer respectively.

[0090] Step 2, create a question-answer text pair for each knowledge entry (s i , r i , o i ) in the knowledge set to form a corresponding first data record; and for each knowledge entry (s i , r i , o i)Create a counterfactual question-answer text pair to form a corresponding second data record; and extract the pre-trained knowledge corpus corresponding to the target model that is unrelated to all knowledge entries (s i ,r i ,o i )to form a corresponding third data record set;

[0091] Specifically, it includes: Step 21, for each knowledge entry (s i ,r i ,o i )in the updated knowledge set, create a question-answer text pair to form a corresponding first data record;

[0092] Among them, the first data record corresponds one-to-one with the knowledge entry (s i ,r i ,o i ); the first data record includes a first question text and a first answer label; the first question text is a natural language question text containing the corresponding knowledge theme s i and the theme-object relationship r i ; the first answer label matches the corresponding knowledge object o i ;

[0093] Specifically, it includes: Step 211, configure a question text generation instruction template for the target model, denoted as the first instruction template;

[0094] Here, the configurable parameters of the first instruction template in the embodiment of the present invention include a theme configuration parameter, a theme-object relationship configuration parameter, and an object configuration parameter; the first instruction template is a formatted instruction text template; the first instruction template is used to require the target model to generate a question text with the object configuration parameter as the expected answer using the theme configuration parameter and the theme-object relationship configuration parameter as the question text elements;

[0095] Step 212, take each knowledge entry (s i ,r i ,o i )in the updated knowledge set as the corresponding current knowledge entry one by one; and use the knowledge theme s i , the theme-object relationship r i and the knowledge object o ias the corresponding current knowledge topic, current topic-object relationship, and current knowledge object; and set the topic configuration parameter, topic-object relationship configuration parameter, and object configuration parameter of the first instruction template to the corresponding current knowledge topic, current topic-object relationship, and current knowledge object; and use the first instruction template with the parameter settings as the corresponding first instruction text to input into the target model for question text generation processing, and use the question text obtained from this processing as the corresponding first question text; and use the current knowledge object as the corresponding first answer label; and form a corresponding first data record from the first question text and the first answer label corresponding to the current knowledge entry;

[0096] Step 22, and for each knowledge entry (s i ,r i ,o i ) create a counterfactual question-answer text pair to form the corresponding second data record;

[0097] wherein, the second data record corresponds one-to-one with the knowledge entry (s i ,r i ,o i ); the second data record includes a second question text and a second answer label; the second question text is a classification question text that has a counterfactual relationship with the corresponding knowledge entry (s i ,r i ,o i ); the second answer label is the answer text corresponding to the second question text, and its answer content includes two classification results: yes and no;

[0098] Specifically, it includes: Step 221, configure another question text generation instruction template for the target model, denoted as the second instruction template;

[0099] Here, the configurable parameters of the second instruction template in the embodiment of the present invention include a topic configuration parameter, a topic-object relationship configuration parameter, and an object configuration parameter; the second instruction template is a formatted instruction text template; the second instruction template is used to require the target model to form a current fact triple from the topic configuration parameter, the topic-object relationship configuration parameter, and the object configuration parameter, and generate a classification question text that has a counterfactual relationship with the current fact triple, and generate a corresponding binary classification answer text for the classification question text, and require that the binary classification answer text can only be yes or no;

[0100] Step 222, take each knowledge entry (s i ,r i ,o i ) in the updated knowledge set one by one as the corresponding current knowledge entry; and use the knowledge topic s i of the current knowledge entry, the topic-object relationship r iand knowledge object o i as the corresponding current knowledge topic, current topic-object relationship, and current knowledge object; and set the topic configuration parameter, topic-object relationship configuration parameter, and object configuration parameter of the second instruction template to the corresponding current knowledge topic, current topic-object relationship, and current knowledge object; and use the second instruction template with the parameter settings as the corresponding second instruction text to input into the target model for question and answer text generation processing, and use the classified question text and binary classification answer text obtained from this processing as the corresponding second question text and second answer label; and form a corresponding second data record from the second question text and second answer label corresponding to the current knowledge entry;

[0101] For example, the second instruction text is "Given that the topic configuration parameter = "cat", the topic-object relationship configuration parameter = "likes to eat", and the object configuration parameter = "canned food", please form a current fact triple from the topic configuration parameter, topic-object relationship configuration parameter, and object configuration parameter, and generate a classified question text that has a counterfactual relationship with the current fact triple, and generate a corresponding binary classification answer text for this classified question text, and require that the binary classification answer text can only be yes or no";

[0102] After inputting the second instruction text into the target model, the classified question text given by the target model that has a counterfactual relationship with the current fact triple is "Doesn't the cat like to eat canned food?", and the corresponding binary classification answer text is "No";

[0103] Step 23, and extract from the pre-trained knowledge corpus corresponding to the target model the pre-trained knowledge corpus that has nothing to do with all knowledge entries (s i ,r i ,o i ) to form a corresponding third data record set;

[0104] Among them, the pre-trained knowledge corpus includes multiple pre-trained knowledge corpora (s tr ,r tr ,o tr );s tr 、r tr 、o tr are respectively the knowledge topic, topic-object relationship, and knowledge object of the knowledge triple;

[0105] The third data record set includes multiple third data records; the third data record includes a third question text and a third answer label; each third data record corresponds to a pre-trained knowledge corpus (s i ,r i ,o i ) that has nothing to do with all knowledge entries (s tr ,r tr ,otr ); The third question text is a natural language question text containing the corresponding knowledge topic (s tr and the topic-object relationship r tr ; The third answer label matches the corresponding knowledge object o tr ;

[0106] Specifically, it includes: Step 231, mark the pre-trained knowledge corpus that has nothing to do with each knowledge topic s tr in all knowledge entries (s i , r i , o i ) as irrelevant corpus; i The pre-trained knowledge corpus (s tr , r tr , o tr ) that has nothing to do with all knowledge topics s is marked as irrelevant corpus;

[0107] Step 232, take each irrelevant corpus as the corresponding current knowledge entry one by one; and take the knowledge topic s tr , the topic-object relationship r tr and the knowledge object o tr of the current knowledge entry as the corresponding current knowledge topic, current topic-object relationship and current knowledge object; and set the topic configuration parameter, topic-object relationship configuration parameter and object configuration parameter of the first instruction template as the corresponding current knowledge topic, current topic-object relationship and current knowledge object; and take the first instruction template with parameter settings as the corresponding first instruction text and input it into the target model for question text generation processing, and take the question text obtained from this processing as the corresponding third question text; and take the current knowledge object as the corresponding third answer label; and form a corresponding third data record from the third question text and the third answer label corresponding to the current knowledge entry;

[0108] Step 233, and form a corresponding third data record set from all the obtained third data records.

[0109] Step 3, optimize the weight parameters of all MLP layers of the target model based on N1 first data records;

[0110] Specifically, it includes: Step 3-1, record the current overall model parameters of the target model as parameter θ;

[0111] Step 3-2, take the first question text of each first data record as the current model input text and input it into the target model for processing, and record the token sequence and the initial vector H0 output by the preprocessing section during this processing as the corresponding token sequence U i and the initial vector H 0,i , and during this processing, each M jThe process vectors H for layer input and output in,j , H out,j are denoted as the corresponding process vectors H in,j,i , H out,j,i , and all the process vectors H in,j,i , H out,j,i in the current processing are cached;

[0112] Here, the word segmentation sequence U of the embodiment of the present invention i corresponds one-to-one with the knowledge entry (s i , r i , o i ), and is composed of multiple word segments u i , where one corresponds to the knowledge topic s i ; the initial vector H 0,i corresponds one-to-one with the knowledge entry (s i , r i , o i ), and is composed of multiple initial vectors h of word segments 0,i , and the initial vector h of the word segment 0,i corresponds one-to-one with the word segment u i ; the process vectors H in,j,i , H out,j,i correspond one-to-one with the knowledge entry (s i , r i , o i ); the process vector H in,j,i is composed of multiple process vectors h of word segments in,j,i ; H out,j,i is composed of multiple process vectors h of word segments out,j,i ;

[0113] Step 3-3: Extract the last word segment process vector i , r i , o i ) of the last layer process vector corresponding to each knowledge entry as the corresponding end layer tail word vector and set a global increment δ with a vector length and vector feature dimension both consistent with the current end layer tail word vector for each end layer tail word vector ; and initialize all the global increments δ i to all-zero vectors; i ;

[0114] Here, the reason for extracting the last word segment process vector of the last layer process vector as the corresponding end layer tail word vector e i, because the last MLP layer of the large language model implemented based on the Transformer model structure is generally the second or third last processing network before the model output, and the next layer network of the last MLP layer generally performs specific text generation or classification operations based on the last token vector of the output vector of the last MLP layer; that is to say, if the model output is expected to change, then the change amount can be synchronously reflected in the change amount of this last layer end token vector e i of;

[0115] Step 3-4, and based on the negative log-likelihood loss function and N1 groups of last layer end token vectors Global increment δ i and knowledge object o i Set a corresponding optimization objective function; and towards the direction of minimizing the optimization objective function, solve for the optimal solution of each global increment δ i and use the solution result as the corresponding optimal global increment And based on each last layer end token vector and its corresponding optimal global increment Calculate the corresponding target end token vector e tag,i ;

[0116] Here, the optimization objective function of the embodiment of the present invention is:

[0117]

[0118] Wherein, L NLL is the negative log-likelihood loss function, is when the overall model parameters of the target model are the parameters θ, and through each global increment δ i adjust the corresponding last layer end token vector afterwards, the generated text of the target model is the corresponding knowledge object o i conditional probability;

[0119] The calculation method of the target end token vector e tag,i in the embodiment of the present invention is:

[0120]

[0121] It should be noted that the current step 3-4 is to obtain each knowledge item (s i , r i , o i ) by solving the optimization objective function, and the overall change amount generated on the corresponding last layer end token vector is the optimal global increment and from each last layer end token vector and its corresponding optimal global increment Calculate the target output quantity that the knowledge update corresponding to each knowledge item (s i , r i , o i ) needs to reach on the model output, that is, the target tail word vector e tag,i ;

[0122] It should also be noted that the optimization method for the N2 M j layers in the embodiments of the present invention is a cyclic iteration method of optimizing one layer at a time. Specifically: first, when the overall model parameters are the original parameters θ, optimize the output layer weight W j=1 of the M out,j=1 layer and update the overall model parameters of the target model based on the optimization result; then, on the basis of the target model after the previous update, optimize the output layer weight W j=2 of the M out,j=2 layer and update the overall model parameters of the target model again based on the optimization result; and so on. After such cyclic iteration N2 times, the full optimization of the N2 output layer weights W j of the N2 M out,j layers is completed; Steps 3-5 to 3-12 below are the specific implementation processes of this cyclic iteration method;

[0123] It should also be noted that the single-layer optimization process for each M j layer is simply: first, process the first question text corresponding to each knowledge item (s i , r i , o i ) based on the current model parameters of the target model, and obtain a batch of cached parameters during this processing: process vector H in,j,i 、process vector H out,j,i ; then, calculate N1 single-layer key vectors k j based on the N1 segmentation process vectors h in,j,i corresponding to the N1 knowledge topics s i in the N1 process vectors H in,j,i of the current M C,i layer, where C is the layer index of the current M j layer; then, extract the last segmentation process vector h j of the N1 process vectors H out,j,i of the current M out,j,i layer as the corresponding tail word vector to obtain N1 single-layer tail word vectors e C,i , and calculate N1 single-layer local increments △δ tag,i based on the target tail word vector e C,i and the N1 single-layer tail word vectors e j,i ; and use the current Mj N1 single - layer key vectors k of the layer C,i and N1 single - layer local increments △δ j,i Based on these, calculate the current M j layer's output - layer weight increment, that is, the single - layer increment weight △W out,j ; and based on the current M j layer's single - layer increment weight △W out,j perform a parameter update on the current M j layer and the target model; Steps 3 - 5 to 3 - 12 below will elaborate on the single - layer optimization process of the embodiments of the present invention;

[0124] Step 3 - 5, set a counter C initialized to 1;

[0125] Step 3 - 6, take the first - question text corresponding to each knowledge entry (s i , r i , o i ) as the input text of the current model and input it into the target model for processing, and cache the process vectors H j input and output during this processing for each M in,j,i layer again; out,j,i cache again;

[0126] Step 3 - 7, take the M j layer where the layer index j matches the counter C as the current MLP layer;

[0127] Step 3 - 8, among the N2 latest - cached process vectors H i , r i , o i ) corresponding to each knowledge entry, take the tokenization process vector h in,j,i corresponding to the knowledge topic s i and the current MLP layer as the corresponding topic vector h in,j=C,i ; and calculate the latest single - layer key vector k subject,C,i based on each topic vector h subject,C,i and the corresponding input - layer weight W in,j=C for estimation; j=C,i perform estimation;

[0128] Among them, the estimation method of the single - layer key vector k j=C,i is:

[0129] k j=C,i = σ(W in,j=c γ(h subject,C,i ));

[0130] σ is a preset activation function, and γ is a preset normalization function;

[0131] Step 3-9: For each knowledge entry (s i , r i , o i ), among the N2 most recently cached process vectors H out,j,i corresponding to it, the process vector H out,j=C,i corresponding to the current MLP layer, the last tokenization process vector h out,j=C,i is denoted as the single-layer end token vector e C,i ; and based on the target end token vector e tag,i and the single-layer end token vector e C,i , estimate the corresponding single-layer local increment △δ j,i ;

[0132] Among them, the estimation method of the single-layer local increment △δ j,i is as follows:

[0133]

[0134] Here, in the embodiments of the present invention, when calculating the single-layer local increment △δ j,i , first calculate a difference component through (e tag,i - e C,i ), and then set a weight according to a weight setting method of "lighter far and heavier near" based on the layer index j of the current MLP layer Then, obtain the single-layer local increment △δ j,i of the current MLP layer by weighting the current difference component with this weight; The weight setting method of "lighter far and heavier near" mentioned here actually means that the closer to the last MLP layer, the greater the weight, and vice versa, the farther from the last MLP layer, the smaller the weight;

[0135] Step 3-10: Consist of N1 single-layer key vectors k j=C,i to form a corresponding key vector matrix K j=C ; and consist of N1 single-layer local increments △δ j=C,i to form a corresponding local increment matrix R j=C ; and estimate the corresponding single-layer increment weight △W j=C based on the key vector matrix K j=C and the local increment matrix R out,j=C ;

[0136] Here, the estimation method of the single-layer increment weight △W out,j=C implemented in the present invention is as follows:

[0137]

[0138] Among them, X j=C is the covariance matrix parameter of the preset j = C layer;

[0139] To fully understand the estimation method in the current step 3-10, the following briefly explains the derivation process of this estimation method:

[0140] (The first step) First, for the optimization of each layer, we expect the optimized output layer weights to be able to take into account the N1 knowledge entries (s i , r i , o i ) updated this time and other pre-trained knowledge corpora in the pre-trained knowledge corpus that are irrelevant to the knowledge set updated this time (abbreviated as irrelevant corpora). Therefore, the embodiment of the present invention sets a single-layer optimization objective function as:

[0141]

[0142] Among them, is the output layer weight after the M j layer is optimized; x is the corpus index of the irrelevant corpus, and N x is the total number of irrelevant corpora, and k j,x is the single-layer key vector of each M j layer saved at the end of pre-training for each irrelevant corpus, and e j,x is the single-layer tail word vector of each M j layer saved at the end of pre-training for each irrelevant corpus; and is the tail word vector obtained by adding the single-layer local increment △δ j,i to the single-layer tail word vector e j,i cached most recently this time. The single-layer tail word vector can be regarded as the target vector for the optimization of each layer;

[0143] Here, the embodiment of the present invention believes that the non-linear association relationship between the knowledge theme and the knowledge object is actually realized by a mapping relationship similar to a key-value pair given by the layer output weight W j of each M out,j layer; solving the single-layer optimization objective function composed of N x historical key-value pairs (k j,x , e j,x ) and N1 latest key-value pairs (k j,I , e j,i ), the obtained can not only satisfy the mapping relationship of the N1 knowledge themes s i -knowledge object o i this time, but also ensure that the mapping relationship of the knowledge theme-knowledge object of the N x irrelevant corpora does not change;

[0144] (The second step) From the N x single-layer key vectors kj,x Form the corresponding key vector Consisting of N1 single-layer key vectors k j,i Form the corresponding key vector K j , consisting of N x Historically stored single-layer end-word vectors e j,x Form the corresponding value vector Similarly, consisting of N1 latest single-layer end-word vectors Form the corresponding value vector V j ; and use the Normal Equation to solve the above single-layer optimization objective function, and we can get:

[0145]

[0146] Perform the calculation on the above formula to get:

[0147]

[0148] (Step 3) Denote the output layer weight before optimization of the M j layer as weight W out-old,j , and denote the output layer weight increment before and after optimization as the single-layer increment weight △W out,j , and we can get:

[0149]

[0150] Then, the expression output in the second step can be further expressed as:

[0151]

[0152] Expand the above formula to get:

[0153]

[0154] (Step 4) Because So the expanded formula output in the third step can be further expressed as:

[0155]

[0156] Subtract from both sides of the above formula, and we can get:

[0157]

[0158] That is:

[0159] That is also:

[0160] (Step 5) From M jThe N1 latest cached single-layer tail word vectors e of the layer j,i constitute the corresponding value vector V' j , which is composed of M j The N1 single-layer local increments Δδ of the layer j,i constitute the corresponding change amount R j , then the value vector Value vector V' j and the change amount R j The relational expression is: V j = V j '+ R j . Substituting this relational expression into the expression output in the fourth step, we can get:

[0161]

[0162] (Step 6) Because V j ' = W out-old,j K j , so the expression output in the fifth step can be further changed to:

[0163]

[0164] That is:

[0165] (Step 7) Each key vector is a historically known quantity. In the embodiment of the present invention, based on N2 known key vectors N2 corresponding covariance matrix parameters are set for N2 M j layers In this way, the estimated expression of the single-layer incremental weight ΔW out,j given in the current step 3-10 can be obtained:

[0166] ΔW out,j = R j K j T (X j + K j K j T ) -1 ;

[0167] Step 3-11, update the output layer weight W out,j=C of the current MLP layer based on the single-layer incremental weight ΔW out,j=C ; and update the parameters θ of the target model based on the new output layer weight W out,j=C ;

[0168] The new W out,j=C = the old W out,j=C + ΔWout,j=C ;

[0169] Step 3-12: Increment the counter C by 1; and identify whether the incremented counter C exceeds the total number N2. If it does not exceed, return to Step 3-6; if it has exceeded, confirm the end of this optimization.

[0170] Step 4: Perform a counterfactual question evaluation on the updated target model based on N1 second data records to obtain a corresponding first evaluation result; and perform an irrelevant knowledge question evaluation on the updated target model based on the third data record set to obtain a corresponding second evaluation result;

[0171] Specifically, it includes: Step 41: Perform a counterfactual question evaluation on the updated target model based on N1 second data records to obtain a corresponding first evaluation result;

[0172] Among them, the first evaluation result includes passed and failed;

[0173] Specifically, it includes: successively taking each second data record as the corresponding current record; and taking the second question text of the current record as the current model input text to input into the target model for processing, and taking the generated text obtained from this model processing as the corresponding current answer; and identifying whether the current answer matches the second answer label of the current record. If it does not match, set the corresponding first Q&A result as unmatched, if it matches, set the corresponding first Q&A result as matched; and identify whether all N1 first Q&A results obtained are matched. If so, set the corresponding first evaluation result as passed, if not, set the corresponding first evaluation result as failed;

[0174] Step 42: Perform an irrelevant knowledge question evaluation on the updated target model based on the third data record set to obtain a corresponding second evaluation result;

[0175] Among them, the second evaluation result includes passed and failed;

[0176] Specifically, it includes: Step 421: Perform a round of traversal on all third data records in the third data record set; and during this round of traversal, take the currently traversed third data record as the corresponding current record; and take the third question text of the current record as the current model input text to input into the target model for processing, and take the generated text obtained from this model processing as the corresponding first predicted answer; and form a corresponding first prediction-label pair from the first predicted answer and the third answer label of the current record; and at the end of this round of traversal, input all the obtained first prediction-label pairs into a preset first model loss function for calculation to obtain a corresponding first loss value;

[0177] Here, the first model loss function of the embodiments of the present invention is implemented based on the cross-entropy loss function or the negative log-likelihood loss function;

[0178] Step 422, and identify whether the first loss value satisfies a preset first loss value range; if it satisfies, set the corresponding second evaluation result as passed; if it does not satisfy, set the corresponding second evaluation result as not passed.

[0179] Here, the first loss value range of the embodiments of the present invention is a preset loss value range.

[0180] Step 5, identify whether both the first and second evaluation results are passed; if not, continue to optimize until both the latest first and second evaluation results are passed; if so, feedback to the current user that the model knowledge update is complete.

[0181] Here, if the first and second evaluation results are not all passed, it is necessary to continue to optimize until both the latest first and second evaluation results are passed. The so-called continuous optimization here actually means returning to Step 3 for re-optimization, then re-evaluating through Step 4 again, and finally re-identifying through Step 5 whether both the latest first and second evaluation results are passed.

[0182] It should also be noted that the embodiments of the present invention default to achieving the purpose of updating knowledge by optimizing all MLP layers of the target model; in addition, the embodiments of the present invention can also pre-identify the key MLP layers that affect knowledge editing in the target model, and then achieve the purpose of updating knowledge by only optimizing the key MLP layers. In this way, the computational complexity and computational volume will be lower, the training cycle will be shorter, the update cost will be lower, and the update efficiency and update flexibility will be higher.

[0183] Figure 2 It is a module structure diagram of a processing device for updating model knowledge by optimizing MLP weights provided in the second embodiment of the present invention. This device is a terminal device or a server that implements the foregoing method embodiments, or can also be a device that enables the foregoing terminal device or server to implement the foregoing method embodiments. For example, this device can be a device or a chip system of the foregoing terminal device or server. As Figure 2 shown, this device includes: a data receiving module 201, a preprocessing module 202, an MLP optimization module 203, an optimization evaluation module 204, and an update feedback module 205.

[0184] The data receiving module 201 is used to take the specified batch of knowledge entries by the user as the updated knowledge set; and count the total number of knowledge entries in the updated knowledge set to obtain the total number N1; and take the large language model specified by the current user as the target model; the large language model specified by the user is a large language model implemented based on the Transformer model structure and has completed pre-training and NLP task fine-tuning, and the NLP tasks include at least text generation tasks, machine translation tasks, intelligent question-answering tasks, and text classification tasks; the total number N1 is a positive integer and at least 1; the updated knowledge set consists of N1 knowledge entries (s i ,r i ,o i ), where 1 ≤ knowledge index i ≤ N1, s i 、r i 、o i are the knowledge topic, topic-object relationship, and knowledge object of the knowledge triple respectively.

[0185] The preprocessing module 202 is used to create a question-answer text pair for each knowledge entry (s i ,r i ,o i ) of the updated knowledge set to form a corresponding first data record; and create a counterfactual question-answer text pair for each knowledge entry (s i ,r i ,o i ) to form a corresponding second data record; and extract the pre-training knowledge corpus unrelated to all knowledge entries (s i ,r i ,o i ) from the pre-training knowledge corpus corresponding to the target model to form a corresponding third data record set.

[0186] The MLP optimization module 203 is used to optimize the weight parameters of all MLP layers of the target model based on N1 first data records.

[0187] The optimization evaluation module 204 is used to perform counterfactual question evaluation on the updated target model based on N1 second data records to obtain the corresponding first evaluation result; and perform irrelevant knowledge question evaluation on the updated target model based on the third data record set to obtain the corresponding second evaluation result; both the first and second evaluation results include pass and fail.

[0188] The update feedback module 205 is used to identify whether both the first and second evaluation results are pass; if not, continue to optimize until both the latest first and second evaluation results are pass evaluations; if so, feedback to the current user that the model knowledge update is complete.

[0189] A processing device for optimizing model knowledge by updating the weights of an MLP provided in an embodiment of the present invention may execute the method steps in the above method embodiment, and its implementation principle and technical effects are similar, which will not be elaborated here.

[0190] It should be noted that it should be understood that the division of each module of the above device is only a logical function division. In actual implementation, it can be fully or partially integrated into a physical entity, or physically separated. And these modules can all be implemented in the form of software called by processing elements; they can also all be implemented in the form of hardware; or some modules can be implemented in the form of software called by processing elements, and some modules can be implemented in the form of hardware. For example, the data receiving module can be a separately established processing element, or can be integrated in a certain chip of the above device. In addition, it can also be stored in the memory of the above device in the form of program code, and called and executed by a certain processing element of the above device to perform the functions of the above determined module. The implementation of other modules is similar. In addition, all or part of these modules can be integrated together or can be independently implemented. The processing element described here can be an integrated circuit with signal processing capabilities. In the implementation process, each step of the above method or each of the above modules can be completed by the integrated logic circuit in the processor element or the instruction in the form of software.

[0191] For example, the above modules can be one or more integrated circuits configured to implement the above method, such as: one or more Application Specific Integrated Circuits (ASICs), or, one or more Digital Signal Processors (DSPs), or, one or more Field Programmable Gate Arrays (FPGAs), etc. Again, when a certain module above is implemented in the form of a processing element scheduling program code, the processing element can be a general-purpose processor, such as a Central Processing Unit (CPU) or other processors that can call program code. Again, these modules can be integrated together and implemented in the form of a System-on-a-chip (SOC).

[0192] In the above embodiments, it can be implemented in whole or in part by software, hardware, firmware, or any combination thereof. When implemented using software, it can be implemented in whole or in part in the form of a computer program product. The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, the processes or functions described in the foregoing method embodiments are generated in whole or in part. The above computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable devices. The above computer instructions can be stored in a computer-readable storage medium, or transmitted from one computer-readable storage medium to another computer-readable storage medium. For example, the above computer instructions can be transmitted from one website, computer, server, or data center to another website, computer, server, or data center by wire (such as coaxial cable, optical fiber, Digital Subscriber Line (DSL)) or wireless (such as infrared, wireless, Bluetooth, microwave, etc.). The above computer-readable storage medium can be any available medium that a computer can access or a data storage device such as a server or data center that includes one or more integrated available media. The above available medium can be a magnetic medium (such as a floppy disk, hard disk, magnetic tape), an optical medium (such as a DVD), or a semiconductor medium (such as a solid state disk (SSD)), etc.

[0193] Figure 3 FIG. 4 is a schematic structural diagram of an electronic device provided in Embodiment 3 of the present invention. The electronic device can be a terminal device or a server for implementing the method of the foregoing embodiments, or a terminal device or a server for implementing the method of the foregoing embodiments connected to the foregoing terminal device or server. As Figure 3 shown, the electronic device may include: a processor 301 (such as a CPU), a memory 302, and a transceiver 303; the transceiver 303 is coupled to the processor 301, and the processor 301 controls the transceiver operations of the transceiver 303. Various instructions can be stored in the memory 302 for completing various processing functions and implementing the processing steps described in the foregoing method embodiments. Preferably, the electronic device according to the embodiments of the present invention further includes: a power supply 304, a system bus 305, and a communication port 306. The system bus 305 is used to implement communication connections between components. The above communication port 306 is used for the electronic device to connect and communicate with other peripherals.

[0194] In Figure 3The system bus 305 mentioned above can be a Peripheral Component Interconnect (PCI) bus, an Extended Industry Standard Architecture (EISA) bus, or the like. The system bus can be divided into an address bus, a data bus, a control bus, etc. For the sake of simplicity, only a thick line is used to represent it in the figure, but it does not mean that there is only one bus or one type of bus. The communication interface is used to implement communication between the database access device and other devices (such as clients, read-write libraries, and read-only libraries). The memory may include Random Access Memory (RAM), and may also include non-volatile memory, such as at least one disk memory.

[0195] The above-mentioned processor can be a general-purpose processor, including a Central Processing Unit (CPU), a Network Processor (NP), a Graphics Processing Unit (GPU), etc.; it can also be a Digital Signal Processor (DSP), an Application Specific Integrated Circuit (ASIC), a Field Programmable Gate Array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components.

[0196] It should be noted that the embodiments of the present invention also provide a computer-readable storage medium, in which instructions are stored. When it runs on a computer, it causes the computer to execute the methods and processing procedures provided in the above embodiments.

[0197] An embodiment of the present invention provides a processing method, apparatus, electronic device, and computer-readable storage medium for updating model knowledge by optimizing MLP weights. As can be seen from the above, the embodiment of the present invention uses the batch knowledge entries specified by the user as the updated knowledge set; counts the total number of knowledge entries in the updated knowledge set to obtain the total number N1; uses the currently user-specified large language model (a large language model implemented based on the Transformer model structure and having completed pre-training and NLP task fine-tuning) as the target model; creates a first data record composed of positive question-answer text pairs and a second data record composed of counterfactual question-answer text pairs for each knowledge entry in the updated knowledge set, and extracts the pre-training knowledge corpus unrelated to all knowledge entries from the pre-training knowledge corpus of the target model to form a third data record set; optimizes the weight parameters of all MLP layers of the target model based on all the first data records; performs counterfactual question evaluation on the updated target model based on all the second data records to obtain a first evaluation result; performs unrelated knowledge question evaluation on the updated target model based on the third data record set to obtain a second evaluation result; and feedbacks to the user that the model knowledge update is completed when both the first and second evaluation results pass. The embodiment of the present invention only optimizes the MLP weights. Through the embodiment of the present invention, not only the computational complexity and amount of calculation are reduced, the training cycle is shortened, and the update cost is reduced, but also the update efficiency and update flexibility are improved.

[0198] The steps of the methods or algorithms described in connection with the embodiments disclosed herein may be implemented in hardware, software modules executed by a processor, or a combination of both. The software modules may be placed in a random access memory (RAM), memory, read-only memory (ROM), electrically programmable ROM, electrically erasable programmable ROM, registers, hard disk, removable disk, CD-ROM, or any other form of storage medium known in the art.

[0199] The specific embodiments described above further elaborate on the purpose, technical solutions, and beneficial effects of the present invention. It should be understood that the above are only specific embodiments of the present invention and are not used to limit the protection scope of the present invention. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principles of the present invention shall be included in the protection scope of the present invention.

Claims

1. A processing method for updating model knowledge by optimizing MLP weights, characterized in that: The method comprises: The batch of knowledge items specified by the user is used as the updated knowledge set; the total number of knowledge items in the updated knowledge set is counted to obtain a total number N1; and the large language model specified by the current user is used as the target model; the large language model specified by the user is a large language model based on the Transformer model structure and has completed pre-training and NLP task fine-tuning. The NLP task includes at least text generation task, machine translation task, intelligent question-answering task, and text classification task; the total number N1 is a positive integer and the minimum is 1; the updated knowledge set consists of N1 knowledge items (s i ,r i ,o i ), 1≤knowledge index i≤N1, s i 、r i , o i They are the knowledge subject, subject-object relationship, and knowledge object of the knowledge triple; Each of the knowledge items (s i ,r i ,o i ) creates a question-answer text pair to form a corresponding first data record; and for each of the knowledge items (s i ,r i ,o i ) creates a second data record corresponding to a counterfactual question-answer text pair; and extracts from the pre-trained knowledge corpus corresponding to the target model all the knowledge items (s i ,r i ,o i ) irrelevant pre-training knowledge corpus to form a corresponding third data record set; Optimizing weight parameters of all MLP layers of the target model based on the N1 first data records; Based on the N1 second data records, the updated target model is evaluated for counterfactual questions to obtain a corresponding first evaluation result; and based on the third data record set, the updated target model is evaluated for irrelevant knowledge questions to obtain a corresponding second evaluation result; the first and second evaluation results both include pass and fail; It is identified whether the first and second evaluation results are both passed; if not, the optimization is continued until the latest first and second evaluation results are both passed; if so, feedback is given to the current user that the model knowledge update is complete.

2. The method for updating model knowledge by optimizing MLP weights according to claim 1, characterized in that: All internal modules of the target model are divided into two major sections: a preprocessing section and a forward reasoning section; specifically: in the target model, the internal modules for word segmentation of the model input text and the internal modules for embedding encoding of the word segmentation sequence are included in the preprocessing section; in the target model, all internal modules involved in the process of forward reasoning of the NLP task according to the initial vector output by the preprocessing section are included in the forward reasoning section; The preprocessing module is used to perform word segmentation processing on the model input text to obtain a corresponding word segmentation sequence; and perform embedding coding processing on the word segmentation sequence according to the embedding coding rule of the target model and record the obtained embedded coding vector as the initial vector H0 and send it to the forward reasoning module; the word segmentation sequence is composed of multiple word segmentations; the initial vector H0 is composed of multiple word segmentation initial vectors h0; the word segmentation initial vector h0 corresponds to the word segmentation one by one; The forward reasoning module is used to perform forward reasoning based on the input initial vector H0 to obtain the corresponding generated text and output it; the total number of MLP layers in the forward reasoning module is recorded as the total number N2, and each MLP layer is marked as the corresponding M j layer; the total number N2 is a positive integer; 1≤layer index j≤N2; each of the M j The input and output vectors of the layer during the inference process are recorded as the corresponding process vector H in,j Process vector H out,j ; The process vector H in,j By multiple word segmentation process vectors h in,j Composition; the process vector H out,j By multiple word segmentation process vectors h out,j Composition: The word segmentation process vector h in,j 、h out,j One-to-one correspondence with the word segmentation or the word segmentation initial vector h0; Each of the M j The reasoning process of the layer is: H out,j =W out,j σ(W in,j c(H in,j )), Among them, σ is the preset activation function, γ is the preset normalization function, and W in,j , W out,j are respectively j The input and output layer weights of the layer; The first data record and the knowledge item (s i ,r i ,o i ) one-to-one correspondence; the first data record includes a first question text and a first answer label; the first question text is a containing the corresponding knowledge subject s i and the subject-object relationship r i The natural language question text; the first answer tag and the corresponding knowledge object o i match; The second data record and the knowledge item (s i ,r i ,o i ) one-to-one correspondence; the second data record includes a second question text and a second answer label; the second question text is a corresponding knowledge entry (s i ,r i ,o i ) a classification question text having a counterfactual relationship; the second answer tag is an answer text corresponding to the second question text, and the answer content includes two classification results of yes and no; The pre-training knowledge corpus includes a plurality of the pre-training knowledge corpora (s tr ,r tr ,o tr );s tr 、r tr , o tr They are the knowledge subject, subject-object relationship, and knowledge object of the knowledge triple; The third data record set includes a plurality of third data records; the third data record includes a third question text and a third answer label; each of the third data records corresponds to one corresponding to all the knowledge items (s i ,r i ,o i ) are irrelevant to the pre-training knowledge corpus (s tr ,r tr ,o tr ); the third question text is a text containing the corresponding knowledge topic (s tr and the subject-object relationship r tr The natural language question text; the third answer tag and the corresponding knowledge object o tr match.

3. The method for updating model knowledge by optimizing MLP weights according to claim 2, characterized in that: The knowledge items (s i ,r i ,o i ) creates a first data record corresponding to a question-answer text pair, specifically including: A question text generation instruction template is configured for the target model and is recorded as a first instruction template; the configurable parameters of the first instruction template include subject configuration parameters, subject-object relationship configuration parameters and object configuration parameters; the first instruction template is a formatted instruction text template; the first instruction template is used to require the target model to generate a question text with the object configuration parameters as the expected answer using the subject configuration parameters and the subject-object relationship configuration parameters as question text elements; Each of the knowledge items (s i ,r i ,o i ) one by one as the corresponding current knowledge item; and the knowledge subject s of the current knowledge item i , the subject-object relationship i and the knowledge object o i as the corresponding current knowledge subject, current subject-object relationship and current knowledge object; and set the subject configuration parameters, the subject-object relationship configuration parameters and the object configuration parameters of the first instruction template to the corresponding current knowledge subject, the current subject-object relationship and the current knowledge object; and input the first instruction template with completed parameter settings as the corresponding first instruction text into the target model for question text generation processing and use the question text obtained in this processing as the corresponding first question text; and use the current knowledge object as the corresponding first answer label; and the first question text and the first answer label corresponding to the current knowledge entry form a corresponding first data record.

4. The method for updating model knowledge by optimizing MLP weights according to claim 2, characterized in that: The above are the knowledge items (s i ,r i ,o i ) creates a second data record corresponding to a counterfactual question-answer text pair, specifically including: Another question text generation instruction template is configured for the target model and is recorded as a second instruction template; the configurable parameters of the second instruction template include subject configuration parameters, subject-object relationship configuration parameters and object configuration parameters; the second instruction template is a formatted instruction text template; the second instruction template is used to require the target model to form a current fact triplet by the subject configuration parameters, the subject-object relationship configuration parameters and the object configuration parameters, and to generate a classification question text that has a counterfactual relationship with the current fact triplet, and to generate a corresponding binary classification answer text for the classification question text, and to require that the binary classification answer text can only be yes or no; Each of the knowledge items (s i ,r i ,o i ) one by one as the corresponding current knowledge item; and the knowledge subject s of the current knowledge item i , the subject-object relationship i and the knowledge object o i as the corresponding current knowledge subject, current subject-object relationship and current knowledge object; and the subject configuration parameters, the subject-object relationship configuration parameters and the object configuration parameters of the second instruction template are set to the corresponding current knowledge subject, the current subject-object relationship and the current knowledge object; and the second instruction template with completed parameter setting is input into the target model as the corresponding second instruction text for question and answer text generation processing and the classified question text and the binary answer text obtained in this processing are used as the corresponding second question text and the second answer label; and the second question text and the second answer label corresponding to the current knowledge entry constitute a corresponding second data record.

5. The method for updating model knowledge by optimizing MLP weights according to claim 2, characterized in that: The method extracts all the knowledge items (s i ,r i ,o i ) constitutes a corresponding third data record set, specifically including: Each of the knowledge topics s in the pre-training knowledge corpus tr With all the said knowledge items (s i ,r i ,o i ) of the subject matter i The pre-training knowledge corpus (s tr ,r tr ,o tr ) is recorded as irrelevant corpus; and taking each of the irrelevant corpora as the corresponding current knowledge items one by one; and taking the knowledge subject s of the current knowledge items tr , the subject-object relationship tr and the knowledge object o tr as the corresponding current knowledge subject, current subject-object relationship and current knowledge object; and set the subject configuration parameters, subject-object relationship configuration parameters and object configuration parameters of the first instruction template to the corresponding current knowledge subject, current subject-object relationship and current knowledge object; and input the first instruction template with completed parameter settings as the corresponding first instruction text into the target model for question text generation processing and use the question text obtained in this processing as the corresponding third question text; and use the current knowledge object as the corresponding third answer label; and the third question text and the third answer label corresponding to the current knowledge entry form a corresponding third data record; And all the obtained third data records form a corresponding third data record set.

6. The method for updating model knowledge by optimizing MLP weights according to claim 2, characterized in that: The optimizing weight parameters of all MLP layers of the target model based on the N1 first data records specifically includes: Step 601, record the current overall model parameter of the target model as parameter θ; Step 602: Input the first question text of each first data record as the current model input text into the target model for processing, and record the word segmentation sequence and the initial vector H0 output by the preprocessing module in this processing as the corresponding word segmentation sequence U i and the initial vector H 0,i , and each of the M j The process vector H of layer input and output in,j , H out,j The corresponding process vector H is in,j,i , H out,j,i , and all the process vectors H in this processing process in,j,i , H out,j,i Cache. Among them, the word segmentation sequence U i With the knowledge item (s i ,r i ,o i ) one by one, consisting of multiple participles u i Composition, one of which is related to the subject of knowledge i Corresponding to; the initial vector H 0,i With the knowledge item (s i ,r i ,o i ) one by one, consisting of multiple word segmentation initial vectors h 0,i The word segmentation initial vector h 0,i With the participle u i One-to-one correspondence; the process vector H in,j,i , H out,j,i With the knowledge item (s i ,r i ,o i ) one-to-one correspondence; the process vector H in,j,i By multiple word segmentation process vectors h in,j,i Composition: H out,j,i By multiple word segmentation process vectors h out,j,i composition; Step 603: each of the knowledge items (s i ,r i ,o i ) corresponds to the last layer of the process vector The last word segmentation process vector Extract it as the corresponding last-layer tail word vector And for each of the last layer tail word vectors Set a vector length and vector feature dimension that are the same as the current last layer tail word vector Keep a consistent global increment δ i ; and all the global increments δ i All are initialized to all zero vectors; Step 604, based on the negative log-likelihood loss function and the N1 group of the last layer tail word vector The global increment δ i and the knowledge object o i Set a corresponding optimization objective function; and adjust each of the global increments δ in the direction of making the optimization objective function reach the minimum value; i The optimal solution is solved and the solution result is used as the corresponding optimal global increment And based on each of the last word vectors and its corresponding optimal global increment Calculate the corresponding target tail word vector e tag,i ; Wherein, the optimization objective function is: L NLL is the negative log-likelihood loss function, When the overall model parameter of the target model is the parameter θ, the global increments δ i Adjust the corresponding last layer tail word vector Afterwards, the generated text of the target model is the corresponding knowledge object o i The conditional probability of The target tail word vector e tag,i The calculation method is: Step 605, setting a counter C initialized to 1; Step 606: each of the knowledge items (s i ,r i ,o i ) corresponding to the first question text is used as the current model input text and input into the target model for processing, and each of the M j The process vector H of layer input and output in,j,i , H out,j,i Cache again; Step 607: Match the layer index j with the counter C. j Layer as the current MLP layer; Step 608: each of the knowledge items (s i ,r i ,o i ) corresponding to the N2 latest cached process vectors H in,j,i In the knowledge topics i And the word segmentation process vector h corresponding to the current MLP layer in,j=C,i Denoted as the corresponding topic vector h subject,C,i ; and according to each of the subject vectors h subject,C,i and the corresponding input layer weight W in,j=C Calculate the latest single-level key vector k j=C,i Make estimates; Among them, the single-layer key vector k j=C,i The estimation method is: k j=C,i =σ(W in,j=C c(h subject,C,i )); σ is the preset activation function, γ is the preset normalization function; Step 609: each of the knowledge items (s i ,r i ,o i ) corresponding to the N2 latest cached process vectors H out,j,i In the process vector H corresponding to the current MLP layer, out,j=C,i The last word segmentation process vector h out,j=C,i Recorded as a single-layer tail word vector e C,i ; and according to the target tail word vector e tag,i and the single-layer tail word vector e C,i Estimate the corresponding single-layer local increment △δ j,i ; Among them, the single layer local increment △δ j,i The estimation method is: Step 610: N1 single-layer key vectors k j=C,i Form a corresponding key vector matrix K j=C ; and by N1 said single layer local increment △δ j=C,i Form a corresponding local increment matrix R j=C ; and based on the key vector matrix K j=C and the local increment matrix R j=C Estimate the corresponding single-layer incremental weight △W out,j=C ; Among them, the single-layer incremental weight △W out,j=C The estimation method is: Among them, X j=C is the preset j=Cth layer covariance matrix parameter; Step 611: based on the single-layer incremental weight ΔW out,j=C The output layer weight W of the current MLP layer out,j=C Update; and based on the new output layer weight W out,j=C Updating the parameter θ of the target model; New W out,j=C = old W out,j=C +ΔW out,j=C ; Step 612, add 1 to the counter C; and identify whether the counter C after adding 1 exceeds the total number N2; if not, return to step 606; if it has exceeded, confirm that the optimization is finished.

7. The method for updating model knowledge by optimizing MLP weights according to claim 2, characterized in that: The performing counterfactual question evaluation on the updated target model based on the N1 second data records to obtain a corresponding first evaluation result specifically includes: Treat each of the second data records as the corresponding current record in turn; and input the second question text of the current record as the current model input text into the target model for processing, and use the generated text obtained by this model processing as the corresponding current answer; and identify whether the current answer matches the second answer label of the current record. If not, set the corresponding first question and answer result as not matching; if matching, set the corresponding first question and answer result as matching; and identify whether the N1 first question and answer results obtained are all matching. If so, set the corresponding first evaluation result as passed; otherwise, set the corresponding first evaluation result as failed.

8. The method for updating model knowledge by optimizing MLP weights according to claim 2, characterized in that: The step of performing irrelevant knowledge question evaluation on the updated target model based on the third data record set to obtain a corresponding second evaluation result specifically includes: Step 81, perform a round of traversal on all the third data records of the third data record set; and in this round of traversal, use the third data record currently traversed as the corresponding current record; and input the third question text of the current record as the current model input text into the target model for processing, and use the generated text obtained by this model processing as the corresponding first predicted answer; and form a corresponding first prediction-label pair by the first predicted answer and the third answer label of the current record; and at the end of this round of traversal, bring all the first prediction-label pairs obtained into the preset first model loss function to calculate and obtain the corresponding first loss value; Wherein, the first model loss function is implemented based on a cross entropy loss function or a negative log-likelihood loss function; Step 82, and identify whether the first loss value meets the preset first loss value range; if so, set the corresponding second evaluation result to pass; if not, set the corresponding second evaluation result to fail.

9. A device for executing the processing method for updating model knowledge by optimizing MLP weights according to any one of claims 1 to 8, characterized in that: The device comprises: a data receiving module, a preprocessing module, an MLP optimization module, an optimization evaluation module and an update feedback module; The data receiving module is used to take the batch of knowledge items specified by the user as the updated knowledge set; and to obtain a total number N1 by counting the total number of knowledge items in the updated knowledge set; and to take the large language model specified by the current user as the target model; the large language model specified by the user is a large language model implemented based on the Transformer model structure and has completed pre-training and NLP task fine-tuning, and the NLP task at least includes text generation tasks, machine translation tasks, intelligent question-answering tasks, and text classification tasks; the total number N1 is a positive integer and the minimum is 1; the updated knowledge set consists of N1 knowledge items (s i ,r i ,o i ), 1≤knowledge index i≤N1, s i 、r i , o i They are the knowledge subject, subject-object relationship, and knowledge object of the knowledge triple; The preprocessing module is used to update each of the knowledge items (s i ,r i ,o i ) creates a question-answer text pair to form a corresponding first data record; and for each of the knowledge items (s i ,r i ,o i ) creates a second data record corresponding to a counterfactual question-answer text pair; and extracts from the pre-trained knowledge corpus corresponding to the target model all the knowledge items (s i ,r i ,o i ) irrelevant pre-training knowledge corpus to form a corresponding third data record set; The MLP optimization module is used to optimize weight parameters of all MLP layers of the target model based on N1 first data records; The optimization evaluation module is used to perform a counterfactual question evaluation on the updated target model based on the N1 second data records to obtain a corresponding first evaluation result; and to perform an irrelevant knowledge question evaluation on the updated target model based on the third data record set to obtain a corresponding second evaluation result; the first and second evaluation results both include pass and fail; The update feedback module is used to identify whether the first and second evaluation results are both passed; if not, continue to optimize until the latest first and second evaluation results are both passed; if so, feedback to the current user that the model knowledge update is complete.

10. An electronic device, characterized in that: include: memory, processors, and transceivers; The processor is used to couple with the memory, read and execute instructions in the memory, so as to implement the method according to any one of claims 1 to 8; The transceiver is coupled to the processor, and the processor controls the transceiver to send and receive messages.

11. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores computer instructions, and when the computer instructions are executed by a computer, the computer is enabled to execute the method according to any one of claims 1 to 8.

Citation Information

Patent Citations

  • Method and device for processing question and answer tasks in combination with knowledge graph

    CN119357405A

  • Information processing apparatus, information processing method, and computer-readable recording medium

    US20210064825A1

Cited By

  • Label quality control method and device with active learning ability

    CN120782332A