Knowledge editing method and device for autoregression large language model

By identifying key layers with strong correlation with knowledge information in the autoregressive large language model, and only low-rank matrix parameters are optimized for these layers during the knowledge editing process, the problems of low-efficiency and high cost of knowledge editing in the existing technology are solved, and a more efficient knowledge editing process is achieved.

CN120218023AActive Publication Date: 2025-06-27BEIJING DP TECH CO LTD

Patent Information

Application Number
CN202510337893.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-21
Publication Date
2025-06-27
Estimated Expiration
2045-03-21

AI Technical Summary

Technical Problem

The existing autoregressive large language model has long training cycle, low editing efficiency and high editing cost during the knowledge editing process.

Method used

By finding the attention/MLP layer with strong correlation with knowledge information in the autoregressive large language model as the key layer, and only these key layers are optimized for parameterization every time the knowledge is edited, and the low-rank matrix is ​​injected for processing.

Benefits of technology

It effectively reduces the computational complexity, reduces the total amount of optimization parameters, shortens the training cycle, improves editing efficiency, and reduces editing costs.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120218023A_ABST
    Figure CN120218023A_ABST
Patent Text Reader

Abstract

The embodiment of the invention relates to a knowledge editing method and device for an autoregressive large language model. The method comprises the steps that the autoregressive large language model serves as a target model; performing question-answer text pair conversion on each knowledge entry in a pre-training knowledge base of the target model; processing each text pair according to three types of model reasoning modes (single normal reasoning, single scrambling reasoning and multiple repair reasoning on the premise of scrambling) to obtain a first prediction text, a second prediction text and a third prediction information set; performing primary key layer estimation based on the first prediction text, the second prediction text and the third prediction information set of each text pair; key layer final judgment is carried out according to all the estimated key layers, and parameter resetting is carried out on the target model in a mode of implanting low-rank matrix parameters into all the key layers; and forming an implantation parameter set by all implanted low-rank matrix parameters, and only updating the implantation parameter set in each knowledge editing process. The editing efficiency can be improved, and the editing cost can be reduced.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of data processing, and particularly relates to a knowledge editing method and device for an autoregressive large language model. Background Art

[0002] Autoregressive Large Language Models can perform various Natural Language Processing (NLP) tasks, such as text generation, machine translation, intelligent question answering, text classification, etc. Currently, most common autoregressive large language models (such as GPT, BERT, LLaMA, T5, etc.) are mostly implemented based on the Transformer model. That is to say, if these common autoregressive large language models are disassembled, it will be found that their internal structure mainly consists of a series of Attention modules and Multilayer Perceptron (MLP) modules. Representing these Attention / MLP modules hierarchically, it is a series of Attention / MLP layers, and each Attention / MLP layer corresponds to a set of independent layer weight parameters.

[0003] During the pre-training stage, the autoregressive large language model learns a large amount of corpus knowledge and establishes a mapping relationship between all learned knowledge features and the layer weight parameters of all Attention / MLP layers to achieve the storage of knowledge information by the model. If knowledge information stored in the autoregressive large language model needs to be edited (updated, modified, or optimized), it can be achieved through model fine-tuning. The conventional model fine-tuning method is as follows: collect the knowledge information to be edited to construct a fine-tuning corpus set, and train and optimize the overall model parameters of the autoregressive large language model based on the fine-tuning corpus set until the model converges. That is to say, the optimization object of this conventional knowledge editing method is the overall model parameters of the autoregressive large language model, and the parameter scale of any current common autoregressive large language model is very large, which means that each knowledge editing will face the following problems: a long training period, low editing efficiency, and high editing cost.

[0004] To improve this situation, we attempt to find out the Attention / MLP layers with strong relevance to the knowledge information in the model as key layers through the present invention, and only optimize the parameters of the located key layers during each knowledge editing, and use the method of injecting a low-rank matrix to process when optimizing the key layers. In this way, the computational complexity can be effectively reduced, the total amount of optimized parameters can be reduced, so as to achieve the purpose of shortening the training period, improving the editing efficiency, and reducing the editing cost. Summary of the Invention

[0005] The objective of the present invention is to provide a knowledge editing method, apparatus, electronic device, and computer-readable storage medium for an autoregressive large language model in view of the deficiencies of the prior art. The present invention regards any autoregressive large language model implemented based on the Transformer model structure and having completed pre-training and NLP task fine-tuning as the target model; and performs question-answer text pair conversion on each knowledge entry in the pre-training knowledge base of the target model to obtain a corresponding text pair; and processes each text pair according to three types of model inference methods (single normal inference, single scrambled inference, multiple repair inferences under the premise of scrambling) to obtain three types of prediction information (the first predicted text, the second predicted text, the third prediction information set), and performs a key layer estimation once based on the three types of prediction information (the first predicted text, the second predicted text, the third prediction information set) of each text pair to obtain a corresponding first key layer set; and performs a key layer final judgment based on all the obtained first key layer sets to obtain a second key layer set; and resets the model parameters of the target model by implanting low-rank matrix parameters into the layer weight parameters of all the second key layers, and forms a corresponding implanted parameter set from all the implanted low-rank matrix parameters; and in each subsequent knowledge editing process, achieves the knowledge editing effect by only fine-tuning the implanted parameter set. Through the present invention, the computational complexity can be effectively reduced, the total amount of optimization parameters can be reduced, so as to achieve the purpose of shortening the training cycle, improving the editing efficiency, and reducing the editing cost.

[0006] To achieve the above objective, a first aspect of an embodiment of the present invention provides a knowledge editing method for an autoregressive large language model, the method including:

[0007] Regarding an autoregressive large language model implemented based on the Transformer model structure and having completed pre-training and NLP task fine-tuning as the target model; and denoting each attention layer or MLP layer in the inference section of the target model as the corresponding inference layer A i , and counting the total number of the inference layer A i to obtain the total number N A ; 1 ≤ layer index i ≤ N A ; The internal structure of the target model is divided into two major sections: a preprocessing section and the inference section;

[0008] Configuring a question-and-answer instruction template for the target model; and performing question-answer text pair conversion on each knowledge entry in the pre-training knowledge base of the target model to obtain a corresponding first text pair;

[0009] Take each of the first text pairs as the corresponding current text pair one by one; input the question text of the current text pair into the Q&A instruction template to generate the current instruction text; input the current instruction text as the model input text into the preprocessing section for preprocessing to obtain the corresponding initial vector H0; input the initial vector H0 into the inference section for a forward inference once and take the generated text output during that time as the first predicted text, and for all the inference layers A during that inference process i cache the process vectors output by i , and form a vector matrix M from all the cached process vectors; scramble the initial vector H0 to obtain a scrambled vector; input the scrambled vector into the inference section for a forward inference once and take the generated text output during that time as the second predicted text; then input the scrambled vector into the inference section for a round of N A forward inferences, and during each inference process in this round, perform a process vector correction based on the vector matrix M to obtain a set of third prediction information composed of N A third prediction information; and perform a key layer estimation based on the first and second predicted texts and the set of third prediction information corresponding to the current text pair to obtain the corresponding first key layer set;

[0010] Perform a final key layer judgment based on all the obtained first key layer sets to obtain a second key layer set; and reset the model parameters of the target model by implanting low-rank matrix parameters into the layer weight parameters of all the second key layers;

[0011] Form a corresponding implanted parameter set from all the implanted low-rank matrix parameters; and only update the implanted parameter set during each knowledge editing process.

[0012] Preferably, the NLP tasks at least include text generation and intelligent Q&A;

[0013] The preprocessing section is used to preprocess the model input text. Specifically, it performs word segmentation on the model input text to obtain a word segmentation sequence, and performs embedding encoding processing on each word in the word segmentation sequence according to the set embedding encoding rules to obtain the corresponding initial vector H0 and send it to the inference section; the inference section is used to perform forward inference based on the input initial vector H0 to obtain the corresponding generated text and output it; the inference section includes N A inference layers A i ; the layer weight parameters of each inference layer A i are denoted as weight W i ;

[0014] The initial vector H0 is composed of multiple sub-vectors h 0,j , where 1 ≤ sub-vector index j ≤ NJ , N J is the total number of word segments of the word segment sequence corresponding to the initial vector H0;

[0015] The output vectors of each of the inference layers A corresponding to the initial vector H0 i are denoted as process vectors H i ; the process vector H i consists of multiple sub-vectors h i,j , and the sub-vectors h i,j correspond one-to-one with the sub-vectors h 0,j ;

[0016] The shape of the vector matrix M corresponding to the initial vector H0 is N A ×N J , and it is composed of N A matrix rows and N J matrix columns; the vector matrix M includes N A ×N J matrix units m i,j ; the matrix rows correspond one-to-one with the process vector H i , and also correspond one-to-one with the inference layer A i , and the matrix columns correspond one-to-one with the word segments of the word segment sequence; the matrix unit m i,j corresponds one-to-one with the sub-vector h i,j ;

[0017] The Q&A instruction template is a formatted instruction text template; the configurable parameters of the Q&A instruction template include question configuration parameters, and the question configuration parameters are a text parameter; the Q&A instruction template is used to prompt the target model to perform text generation processing on the corresponding answer to the question configuration parameters;

[0018] The pre-trained knowledge base includes multiple knowledge entries; the text format of the knowledge entries is the knowledge triple text format [subject text, subject-object relationship text, object text];

[0019] The first text pair corresponds one-to-one with the knowledge entry; the first text pair includes the question text and the answer text; the question text is converted from the subject text and the subject-object relationship text of the corresponding knowledge entry; the answer text is converted from the object text of the corresponding knowledge entry;

[0020] The third prediction information set consists of N A third prediction information; the third prediction information includes the third prediction text and the first correction layer index; the first correction layer index corresponds one-to-one with the inference layer A i ;

[0021] The first set of key layers consists of multiple first key layers; each of the first key layers corresponds to one of the inference layers A i ;

[0022] The second set of key layers consists of multiple second key layers; each of the second key layers corresponds to one of the inference layers A i 。

[0023] Preferably, the step of substituting the question text of the current text pair into the Q&A instruction template to generate the current instruction text specifically includes:

[0024] Setting the question configuration parameter of the Q&A instruction template to the question text of the current text pair to obtain a corresponding template setting text, and using the template setting text as the corresponding current instruction text.

[0025] Preferably, the step of scrambling the initial vector H0 to obtain a scrambled vector specifically includes:

[0026] Taking the word segmentation sequence corresponding to the initial vector H0 as the current word segmentation sequence; identifying the subject word segmentation in the current word segmentation sequence; and taking the sub-vector h in the initial vector H0 corresponding to the subject word segmentation 0,j as the subject vector h subject ; generating a random Gaussian noise vector ε with the same vector length and vector feature dimension as the subject vector h subject ; and scrambling the subject vector h based on the random Gaussian noise vector ε subject to obtain the corresponding scrambled vector, scrambled vector = h subject + ε.

[0027] Preferably, the step of then inputting the scrambled vector into the inference section for a round of N A forward inferences and performing a process vector correction based on the vector matrix M during each inference process in this round to obtain a third prediction information set consisting of N A third prediction information specifically includes:

[0028] Step 51, setting a first counter initialized to 1;

[0029] Step 52, inputting the scrambled vector into the inference section for a forward inference; and pausing once every time passing through one of the inference layers A i during this forward inference process; and at each pause, taking the currently passed inference layer A i as the corresponding current inference layer; and taking the process vector H output by the current inference layer ias the corresponding current process vector; and identify whether the layer index i of the current inference layer matches the first counter; if not, resume and continue to transfer the current process vector to the next internal component in the inference section; if it matches, use the matrix row in the vector matrix M corresponding to the current inference layer as the current matrix row, and based on the total number of word segments N corresponding to the vector matrix M J generate a random integer R whose value is between N J / 2 and N J and randomly select R matrix units m from the current matrix row i,j to replace the R corresponding sub-vectors h in the current process vector i,j to obtain a new current process vector, then resume and continue to transfer the new current process vector to the next internal component in the inference section; and at the end of the current forward inference, use the generated text output by the current inference as a corresponding third prediction text, configure a corresponding first correction layer index for the current third prediction text, set the current first correction layer index as the first counter, and form a corresponding third prediction information from the current third prediction text and the corresponding first correction layer index;

[0030] Step 53, increment the first counter by 1; and identify whether the incremented first counter exceeds the total number N A if so, go to Step 54; if not, return to Step 52;

[0031] Step 54, form a corresponding third prediction information set from the obtained N A third prediction information.

[0032] Preferably, the first key layer set is obtained by performing a key layer estimation on the corresponding first and second prediction texts and the third prediction information set based on the current text, specifically including:

[0033] Step 61, use the answer text of the current text pair as the first label text;

[0034] Step 62, use the first label text and the first prediction text as the corresponding current label text and current prediction text to input into a preset first model loss function for calculation and use the loss value output by the function at that time as the corresponding loss value a1;

[0035] Among them, the first model loss function is used to calculate the loss based on the input current label text and the current prediction text and output the corresponding loss value; the first model loss function is implemented based on the cross-entropy loss function or the negative log-likelihood loss function; the value range of the loss value a1 is between 0 and 1;

[0036] Step 63, use the first label text and the second prediction text as the corresponding current label text and current prediction text to input into the first model loss function for calculation, and use the loss value output by the function at this time as the corresponding loss value a2;

[0037] Among them, the value range of the loss value a2 is between 0 and 1;

[0038] Step 64, use the first label text and each third prediction text in the third prediction information set as the corresponding current label text and current prediction text to input into the first model loss function for calculation, and use the loss value output by the function at this time as the corresponding loss value a3;

[0039] Among them, the value range of each loss value a3 is between 0 and 1;

[0040] Step 65, when the loss value a1 meets the preset model convergence loss value range and the loss value a1 is less than the loss value a2, calculate the corresponding loss repair rate r based on the loss value a2 and each loss value a3; and record each loss repair rate r not lower than the preset repair rate threshold as the corresponding preferred repair rate; and use the inference layer A corresponding to the first correction layer index corresponding to each preferred repair rate i As a corresponding first key layer; and form the first key layer set from all the obtained first key layers;

[0041] Among them, the calculation method of the loss repair rate r is:

[0042] The repair rate threshold is a ratio value greater than 0.

[0043] Preferably, the obtaining the second key layer set through final judgment of the key layers based on all the obtained first key layer sets specifically includes:

[0044] Merge all the first key layer sets to obtain a corresponding first set; group the same first key layers in the first set into a group and denote it as a corresponding first group; count the total number of the first key layers in each first group and use the statistical result as the corresponding key layer score; use the first key layers corresponding to the key layer scores greater than a preset score threshold as a corresponding second key layer; and form a corresponding second key layer set from all the obtained second key layers.

[0045] Preferably, resetting the model parameters of the target model by implanting low-rank matrix parameters into the layer weight parameters of all second key layers specifically includes:

[0046] Denote the weight W corresponding to each second key layer i as the corresponding old weight W old ; set two corresponding low-rank matrices A and B for the old weight W old ; initialize the low-rank matrix A based on a random Gaussian distribution, and initialize the low-rank matrix B as a zero matrix; form a corresponding incremental weight △W = AB from the initialized low-rank matrices A and B; and form a corresponding new weight W old from the old weight W new and the incremental weight △W, where W old = W new + △W; reset the weight W i corresponding to the current second key layer based on the current new weight W

[0047] Preferably, only update the implanted parameter set during each knowledge editing process, specifically including:

[0048] Step 91, when knowledge editing is required each time, first receive the updated knowledge entry set for that time;

[0049] wherein, the updated knowledge entry set includes multiple updated knowledge entries; the text format of the updated knowledge entry is the knowledge triple text format;

[0050] Step 92, then generate a corresponding first training question based on the subject text and the subject-object relationship text of each updated knowledge entry; generate a corresponding first label answer based on the object text of each updated knowledge entry; and form a corresponding first data record from each first training question and the corresponding first label answer.

[0051] Step 93, perform another round of traversal on all the first data records; during this round of traversal, take the currently traversed first data record as the corresponding current data record; bring the first training question of the current data record into the Q&A instruction template to generate the corresponding first instruction text; take the first instruction text as the model input text and input it into the target model for processing, and take the generated text output by this processing as the corresponding first predicted answer; and form a corresponding first prediction-label pair from the first predicted answer and the first labeled answer of the current data record.

[0052] Step 94, then bring all the obtained first prediction-label pairs into a preset second model loss function; and based on a preset first model optimizer, modulate the parameters of the implanted parameter set of the target model in the direction of minimizing the second model loss function until the model loss converges.

[0053] Among them, the second model loss function is implemented based on the cross-entropy loss function or the negative log-likelihood loss function; the first model optimizer includes at least the Adam optimizer and the SGD optimizer.

[0054] A second aspect of the embodiments of the present invention provides an apparatus for implementing the knowledge editing method of the autoregressive large language model described in the first aspect above. The apparatus includes: a first preprocessing module, a second preprocessing module, a key layer estimation module, a key layer localization and low-rank matrix parameter implantation module, and a knowledge editing module.

[0055] The first preprocessing module is used to take an autoregressive large language model implemented based on the Transformer model structure and having completed pre-training and NLP task fine-tuning as the target model; and denote each attention layer or MLP layer in the inference section of the target model as the corresponding inference layer A i , and count the total number of the inference layer A i to obtain the total number N A ; 1 ≤ layer index i ≤ N A ; The internal structure of the target model is divided into two major sections: a preprocessing section and the inference section.

[0056] The second preprocessing module is used to configure a Q&A instruction template for the target model; and perform question-answer text pair conversion on each knowledge entry in the pre-trained knowledge base of the target model to obtain the corresponding first text pair.

[0057] The key layer prediction module is used to take each of the first text pairs as the corresponding current text pair one by one; bring the question text of the current text pair into the Q&A instruction template to generate the current instruction text; take the current instruction text as the model input text and input it into the preprocessing section for preprocessing to obtain the corresponding initial vector H0; input the initial vector H0 into the inference section for a forward inference once, and take the generated text output this time as the first prediction text, and for all the inference layers A during this inference process i cache the process vectors output during the process, and form a vector matrix M from all the cached process vectors; scramble the initial vector H0 to obtain a scrambled vector; input the scrambled vector into the inference section for a forward inference once, and take the generated text output this time as the second prediction text; and then input the scrambled vector into the inference section for a round of N A forward inferences, and perform a process vector correction based on the vector matrix M during each inference process in this round, so as to obtain a set of third prediction information composed of N A third prediction texts; and perform a key layer prediction once based on the first and second prediction texts corresponding to the current text pair and the set of third prediction information to obtain the corresponding first key layer set;

[0058] The key layer positioning and low-rank matrix parameter implantation module is used to perform a final judgment on the key layers according to all the obtained first key layer sets to obtain a second key layer set; and reset the model parameters of the target model by implanting low-rank matrix parameters into the layer weight parameters of all the second key layers;

[0059] The knowledge editing module is used to form a corresponding implanted parameter set from all the implanted low-rank matrix parameters; and only update the implanted parameter set during each knowledge editing process.

[0060] A third aspect of the embodiments of the present invention provides an electronic device, including: a memory, a processor, and a transceiver;

[0061] The processor is used to be coupled with the memory, read and execute the instructions in the memory to implement the method steps described in the first aspect above;

[0062] The transceiver is coupled with the processor, and the processor controls the transceiver to perform message sending and receiving.

[0063] A fourth aspect of the embodiments of the present invention provides a computer-readable storage medium, and the computer-readable storage medium stores computer instructions. When the computer instructions are executed by a computer, the computer is caused to execute the instructions of the method described in the first aspect above.

[0064] An embodiment of the present invention provides a knowledge editing method, device, electronic device, and computer-readable storage medium for an autoregressive large language model. As can be seen from the above, the embodiment of the present invention regards any autoregressive large language model implemented based on the Transformer model structure and completed pre-training and NLP task fine-tuning as the target model; and converts each knowledge entry in the pre-training knowledge base of the target model into a corresponding question-answer text pair; and processes each text pair according to three types of model inference methods (single normal inference, single scrambled inference, multiple repair inferences under the premise of scrambling) to obtain three types of prediction information (the first predicted text, the second predicted text, the third prediction information set), and performs a key layer estimation once based on the three types of prediction information (the first predicted text, the second predicted text, the third prediction information set) of each text pair to obtain a corresponding first key layer set; and performs a key layer final judgment based on all the obtained first key layer sets to obtain a second key layer set; and resets the model parameters of the target model by implanting low-rank matrix parameters into the layer weight parameters of all the second key layers, and forms a corresponding implanted parameter set from all the implanted low-rank matrix parameters; and in each subsequent knowledge editing process, only fine-tune the implanted parameter set to achieve the knowledge editing effect. The embodiment of the present invention effectively reduces the computational complexity, reduces the total amount of optimized parameters, thereby shortening the training cycle, improving the editing efficiency, and reducing the editing cost. BRIEF DESCRIPTION OF THE DRAWINGS

[0065] Figure 1 FIG. is a schematic diagram of a knowledge editing method for an autoregressive large language model provided in Embodiment 1 of the present invention;

[0066] Figure 2 FIG. is a module structure diagram of a knowledge editing device for an autoregressive large language model provided in Embodiment 2 of the present invention;

[0067] Figure 3 FIG. is a schematic diagram of the structure of an electronic device provided in Embodiment 3 of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0068] In order to make the objectives, technical solutions, and advantages of the present invention clearer, the present invention will be further described in detail below with reference to the accompanying drawings. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all of the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.

[0069] Embodiment 1 of the present invention provides a knowledge editing method for an autoregressive large language model, as Figure 1As shown in the schematic diagram of a knowledge editing method for an autoregressive large language model provided in the first embodiment of the present invention, the method mainly includes the following steps:

[0070] Step 1, use an autoregressive large language model implemented based on the Transformer model structure and completed with pre-training and NLP task fine-tuning as the target model; and denote each attention layer or MLP layer in the inference section of the target model as the corresponding inference layer A i and count the total number of inference layer A i to obtain the total number N A .

[0071] where 1 ≤ layer index i ≤ N A ; the NLP tasks include at least text generation tasks and intelligent question answering tasks.

[0072] Here, the internal structure of the target model in the embodiment of the present invention is divided into two major sections: a preprocessing section and an inference section.

[0073] The preprocessing section is used to preprocess the model input text. Specifically, the model input text is tokenized to obtain a token sequence, and each token in the token sequence is embedded and encoded according to the set embedding encoding rules to obtain the corresponding initial vector H0 and sent to the inference section. The initial vector H0 is composed of multiple sub-vectors h 0,j where 1 ≤ sub-vector index j ≤ N J , N J is the total number of tokens in the token sequence corresponding to the initial vector H0.

[0074] The inference section is used to perform forward inference according to the input initial vector H0 to obtain the corresponding generated text and output it; the inference section contains N A inference layer A i ; the layer weight parameters of each inference layer A i are denoted as weight W i . After inputting an initial vector H0 into the inference section, each inference layer A i inside will generate a corresponding output vector, and this output vector is denoted as the process vector H i ; the process vector H i is specifically composed of multiple sub-vectors h i,j and the sub-vector h i,j corresponds one-to-one with the sub-vector h 0,j .

[0075] Step 2, configure a question-and-answer instruction template for the target model; and perform question-answer text pair conversion on each knowledge entry in the pre-trained knowledge base of the target model to obtain the corresponding first text pair;

[0076] Specifically, it includes: Step 21, configuring a question-and-answer instruction template for the target model;

[0077] Here, the question-and-answer instruction template of the embodiment of the present invention is a formatted instruction text template; the configurable parameters of the question-and-answer instruction template include question configuration parameters, and this question configuration parameter is a text parameter; the question-and-answer instruction template is used to prompt the target model to perform text generation processing on the corresponding answer to the question configuration parameter;

[0078] Step 22, and performing question-answer text pair conversion on each knowledge entry in the pre-trained knowledge base of the target model to obtain corresponding first text pairs;

[0079] Here, the pre-trained knowledge base of the embodiment of the present invention includes multiple knowledge entries; the text format of the knowledge entry is the knowledge triple text format [subject text, subject-object relationship text, object text];

[0080] The first text pairs of the embodiment of the present invention correspond one by one to the knowledge entries; the first text pair includes a question text and an answer text; the question text is converted from the subject text and the subject-object relationship text of the corresponding knowledge entry; the answer text is converted from the object text of the corresponding knowledge entry; for example, given a knowledge entry [subject text = "snow", subject-object relationship text = "color", object text = "white"], then the question text converted from the subject text = "snow" and the subject-object relationship text = "color" is "What is the color of snow?", and the answer text converted from the object text = "white" is "white".

[0081] Step 3, taking each first text pair as the corresponding current text pair one by one; and bringing the question text of the current text pair into the question-and-answer instruction template to generate the current instruction text; and taking the current instruction text as the model input text and inputting it into the preprocessing section for preprocessing to obtain the corresponding initial vector H0; and inputting the initial vector H0 into the inference section for a forward inference once and taking the generated text output during that time as the first prediction text, and caching the process vectors output by all inference layers A i during the inference process of that time and forming a vector matrix M from all cached process vectors; and scrambling the initial vector H0 to obtain a scrambled vector; and inputting the scrambled vector into the inference section for a forward inference once and taking the generated text output during that time as the second prediction text; and then inputting the scrambled vector into the inference section for a round of N A forward inferences and performing a process vector correction based on the vector matrix M during each inference process of this round to obtain a third prediction information set composed of N A third prediction information; and performing a key layer estimation based on the first, second prediction texts and the third prediction information set corresponding to the current text pair to obtain the corresponding first key layer set;

[0082] Specifically, it includes: Step 31, taking each first text pair as the corresponding current text pair one by one;

[0083] Step 32, substituting the question text of the current text pair into the Q&A instruction template to generate the current instruction text;

[0084] Specifically, it includes: setting the question configuration parameter of the Q&A instruction template as the question text of the current text pair to obtain a corresponding template setting text, and using the template setting text as the corresponding current instruction text;

[0085] Step 33, taking the current instruction text as the model input text and inputting it into the preprocessing section for preprocessing to obtain the corresponding initial vector H0;

[0086] Step 34, inputting the initial vector H0 into the inference section for a forward inference once, and taking the generated text output in that instance as the first prediction text, and caching the process vectors output by all inference layers A i during that inference process, and forming a vector matrix M from all the cached process vectors;

[0087] Here, after inputting an initial vector H0 into the inference section, each inference layer A i inside it will output a corresponding process vector H i ; each process vector H i is composed of multiple sub-vectors h i,j ; a vector matrix M can be formed from all the cached process vectors H i ; the shape of the vector matrix M is N A ×N J , consisting of N A matrix rows and N J matrix columns. The matrix rows correspond one-to-one with the process vectors H i and also with the inference layers A i one-to-one. The matrix columns correspond one-to-one with the word segments in the word segmentation sequence corresponding to the initial vector H0; the vector matrix M includes N A ×N J matrix units m i,j , and the matrix units m i,j correspond one-to-one with the sub-vectors h i of all the cached process vectors H i,j ;

[0088] It should be noted that the specific data content of each matrix unit m i,j of this vector matrix M is set based on the data content of all the process vectors H i cached during the forward inference process of the current Step 34;

[0089] Step 35, and scramble the initial vector H0 to obtain a scrambled vector;

[0090] Specifically, it includes: taking the word segmentation sequence corresponding to the initial vector H0 as the current word segmentation sequence; identifying the subject word segmentation in the current word segmentation sequence; and taking the sub-vector h 0,j in the initial vector H0 corresponding to the subject word segmentation as the subject vector h subject ; generating a random Gaussian noise vector ε with the same vector length and vector feature dimension as the subject vector h subject ; and scrambling the subject vector h subject based on the random Gaussian noise vector ε to obtain the corresponding scrambled vector, where the scrambled vector = h subject + ε;

[0091] It should be noted that the scrambling method in the current step 35 achieves the purpose of scrambling by adding noise to the subject word segmentation sub-vector in the initial vector H0; in addition, the embodiments of the present invention can also adopt other scrambling methods for processing. For example, all word segmentations with the part-of-speech of noun and / or pronoun and / or adjective in the current word segmentation sequence corresponding to the initial vector H0 are recorded as scrambled word segmentations, and the sub-vectors h 0,j corresponding to each scrambled word segmentation are recorded as the preliminary screening sub-vectors, and a preset extraction mode (including random mode and full mode) is identified. If the extraction mode is the random mode, some sub-vectors are randomly selected from all the preliminary screening sub-vectors as the sub-vectors to be scrambled. If the extraction mode is the full mode, all the preliminary screening sub-vectors are used as the sub-vectors to be scrambled, and a random Gaussian noise vector u with the same vector length and vector feature dimension as the current sub-vector to be scrambled is generated for each sub-vector to be scrambled, and a corresponding scrambled vector = the sub-vector to be scrambled + u is generated based on each sub-vector to be scrambled and its corresponding random Gaussian noise vector u;

[0092] Step 36, input the scrambled vector into the inference section for a forward inference once and take the generated text output at that time as the second predicted text;

[0093] Step 37, and input the scrambled vector into the inference section for N A times of forward inference, and perform a process vector correction based on the vector matrix M during each inference process in this round to obtain a third predicted information set composed of N A third predicted information;

[0094] Specifically, it includes: Step 371, setting a first counter initialized to 1;

[0095] Step 372, input the scrambled vector into the inference section for a forward inference; and during this forward inference process, every time passing through an inference layer A ipause once; and at each pause, take the currently passed inference layer A i as the corresponding current inference layer; and take the process vector H output by the current inference layer i as the corresponding current process vector; and identify whether the layer index i of the current inference layer matches the first counter; if not, resume and continue to transfer the current process vector to the next internal component in the inference block; if it matches, take the matrix row in the vector matrix M corresponding to the current inference layer as the current matrix row, and based on the total number of word segments N J corresponding to the vector matrix M J generate a random integer R with a value between N J / 2 and N i,j and randomly select R matrix elements m i,j from the current matrix row to replace the R corresponding sub-vectors h i,j in the current process vector to obtain a new current process vector, then resume and continue to transfer the new current process vector to the next internal component in the inference block; and at the end of the current forward inference, take the generated text output by the current inference as a corresponding third prediction text, configure a corresponding first correction layer index for the current third prediction text, set the current first correction layer index as the first counter, and form a corresponding third prediction information from the current third prediction text and the corresponding first correction layer index;

[0096] Step 373, increment the first counter by 1; and identify whether the incremented first counter exceeds the total number N A ; if so, go to Step 374; if not, return to Step 372;

[0097] Step 374, form a corresponding third prediction information set from the obtained N A third prediction information;

[0098] Here, the third prediction information set of the embodiment of the present invention consists of N A third prediction information; the third prediction information includes a third prediction text and a first correction layer index; the first correction layer index corresponds one-to-one with the inference layer A i ;

[0099] Step 38, and perform a key layer estimation on the corresponding first, second prediction texts and the third prediction information set based on the current text to obtain a corresponding first key layer set;

[0100] wherein, the first key layer set consists of multiple first key layers; each first key layer corresponds to an inference layer A i ;

[0101] Specifically including: Step 381, take the answer text of the current text pair as the first label text;

[0102] Step 382: Input the first label text and the first predicted text as the corresponding current label text and current predicted text into the preset first model loss function for calculation, and take the loss value output by the function at this time as the corresponding loss value a1;

[0103] Here, the first model loss function in the embodiment of the present invention is used to calculate the loss based on the input current label text and current predicted text and output the corresponding loss value; this first model loss function is implemented based on the cross-entropy loss function or the negative log-likelihood loss function; the value range of the loss value a1 is between 0 and 1;

[0104] Step 383: Input the first label text and the second predicted text as the corresponding current label text and current predicted text into the first model loss function for calculation, and take the loss value output by the function at this time as the corresponding loss value a2;

[0105] Here, the value range of the loss value a2 in the embodiment of the present invention is between 0 and 1;

[0106] Step 384: Input the first label text and each third predicted text in the third predicted information set as the corresponding current label text and current predicted text into the first model loss function for calculation, and take the loss value output by the function at this time as the corresponding loss value a3;

[0107] Here, the value range of each loss value a3 in the embodiment of the present invention is between 0 and 1;

[0108] Step 385: When the loss value a1 meets the preset model convergence loss value range and the loss value a1 is less than the loss value a2, calculate the corresponding loss repair rate r based on the loss value a2 and each loss value a3; and record each loss repair rate r that is not lower than the preset repair rate threshold as the corresponding preferred repair rate; and record the inference layer A corresponding to the first correction layer index corresponding to each preferred repair rate i as a corresponding first key layer; and form a first key layer set from all the obtained first key layers;

[0109] Here, the calculation method of the loss repair rate r in the embodiment of the present invention is:

[0110]

[0111] The model convergence loss value range in the embodiment of the present invention is a preset numerical range, and the repair rate threshold is a preset ratio value and this ratio value is greater than 0.

[0112] Step 4: Based on all the obtained first key layer sets, perform a final judgment on the key layers to obtain a second key layer set; and reset the model parameters of the target model by implanting low-rank matrix parameters into the layer weight parameters of all the second key layers.

[0113] Specifically, it includes: Step 41: Based on all the obtained first key layer sets, perform a final judgment on the key layers to obtain a second key layer set.

[0114] Among them, the second key layer set consists of multiple second key layers; each second key layer corresponds to an inference layer A i ;

[0115] Specifically, it includes: Merge all the first key layer sets to obtain a corresponding first set; group the same first key layers in the first set into a group and denote it as the corresponding first grouping; count the total number of first key layers in each first grouping and use the statistical result as the corresponding key layer score; use the first key layers corresponding to the key layer scores greater than the preset score threshold as a corresponding second key layer; and form a corresponding second key layer set from all the obtained second key layers.

[0116] Here, the score threshold in the embodiment of the present invention is a preset positive integer score value.

[0117] Step 42: Reset the model parameters of the target model by implanting low-rank matrix parameters into the layer weight parameters of all the second key layers.

[0118] Specifically, it includes: Denote the weight W corresponding to each second key layer i as the corresponding old weight W old ; set two corresponding low-rank matrices A and B for the old weight W old ; initialize the low-rank matrix A based on a random Gaussian distribution, and initialize the low-rank matrix B as a zero matrix; form a corresponding incremental weight △W = AB from the initialized low-rank matrix A and low-rank matrix B; form a corresponding new weight W old from the old weight W new = W old + △W; reset the weight W new corresponding to the current second key layer based on the current new weight W i ; and form a corresponding low-rank matrix parameter from the matrix parameters of the low-rank matrix A and low-rank matrix B corresponding to the current second key layer.

[0119] Step 5: Form a corresponding implanted parameter set from all the implanted low-rank matrix parameters; and only update the implanted parameter set during each knowledge editing process.

[0120] Specifically, it includes: Step 51, forming a corresponding implanted parameter set from all the implanted low-rank matrix parameters;

[0121] Step 52, and only updating the implanted parameter set during each knowledge editing process;

[0122] Specifically, it includes: Step 521, when knowledge editing is required each time, first receiving the updated knowledge entry set for that time;

[0123] Here, the updated knowledge entry set in the embodiments of the present invention includes multiple updated knowledge entries; the text format of the updated knowledge entries is the knowledge triple text format;

[0124] Step 522, then generating a corresponding first training question based on the subject text and the subject-object relationship text of each updated knowledge entry; and generating a corresponding first labeled answer based on the object text of each updated knowledge entry; and forming a corresponding first data record from each first training question and the corresponding first labeled answer;

[0125] Step 523, then traversing all the first data records once; and during this round of traversal, taking the currently traversed first data record as the corresponding current data record; and bringing the first training question of the current data record into the Q&A instruction template to generate the corresponding first instruction text; and taking the first instruction text as the model input text and inputting it into the target model for processing, and taking the generated text output by this processing as the corresponding first predicted answer; and forming a corresponding first prediction-label pair from the first predicted answer and the first labeled answer of the current data record;

[0126] Step 524, then bringing all the obtained first prediction-label pairs into the preset second model loss function; and based on the preset first model optimizer, modulating the implanted parameter set of the target model in the direction of minimizing the second model loss function until the model loss converges;

[0127] Here, the second model loss function in the embodiments of the present invention is implemented based on the cross-entropy loss function or the negative log-likelihood loss function; the first model optimizer in the embodiments of the present invention at least includes the Adam optimizer and the SGD optimizer. It should be noted that when modulating the implanted parameter set of the target model, no modulation is required for the original model parameters of the target model (i.e., other parameters in the current model parameters except the implanted parameter set).

[0128] Figure 2The module structure diagram of a knowledge editing device for an autoregressive large language model provided in the second embodiment of the present invention. This device is a terminal device or a server for implementing the foregoing method embodiment, or can be a device that enables the foregoing terminal device or server to implement the foregoing method embodiment. For example, this device can be a device or a chip system of the foregoing terminal device or server. As Figure 2 shown, the device includes: a first preprocessing module 201, a second preprocessing module 202, a key layer estimation module 203, a key layer positioning and low-rank matrix parameter implantation module 204, and a knowledge editing module 205.

[0129] The first preprocessing module 201 is used to take an autoregressive large language model implemented based on the Transformer model structure and completed pre-training and NLP task fine-tuning as the target model; and denote each attention layer or MLP layer in the inference section of the target model as the corresponding inference layer A i , and count the total number of inference layer A i to obtain the total number N A ; 1 ≤ layer index i ≤ N A ; The internal structure of the target model is divided into two major sections: a preprocessing section and an inference section.

[0130] The second preprocessing module 202 is used to configure a question-answer instruction template for the target model; and perform question-answer text pair conversion on each knowledge entry in the pre-trained knowledge base of the target model to obtain the corresponding first text pair.

[0131] The key layer estimation module 203 is used to take each first text pair as the corresponding current text pair one by one; and input the question text of the current text pair into the question-answer instruction template to generate the current instruction text; and input the current instruction text into the preprocessing section as the model input text for preprocessing to obtain the corresponding initial vector H0; and input the initial vector H0 into the inference section for a forward inference once and take the generated text output at that time as the first prediction text, and cache the process vectors output by all inference layer A i during that inference process and form a vector matrix M from all cached process vectors; and scramble the initial vector H0 to obtain a scrambled vector; and input the scrambled vector into the inference section for a forward inference once and take the generated text output at that time as the second prediction text; and then input the scrambled vector into the inference section for a round of N A times of forward inference and perform a process vector correction based on the vector matrix M during each inference process in this round to obtain a third prediction information set composed of N A third prediction information; and perform a key layer estimation based on the first, second prediction texts and the third prediction information set corresponding to the current text pair to obtain the corresponding first key layer set.

[0132] The key layer positioning and low-rank matrix parameter implantation module 204 is used to perform final judgment on the key layers based on all the obtained first key layer sets to obtain a second key layer set; and reset the model parameters of the target model by implanting low-rank matrix parameters into the layer weight parameters of all the second key layers.

[0133] The knowledge editing module 205 is used to form a corresponding implanted parameter set from all the implanted low-rank matrix parameters; and only update the implanted parameter set during each knowledge editing process.

[0134] A knowledge editing device for an autoregressive large language model provided by an embodiment of the present invention can execute the method steps in the above method embodiment, and its implementation principle and technical effect are similar, which will not be elaborated here.

[0135] It should be noted that it should be understood that the division of each module of the above device is only a logical function division. In actual implementation, it can be fully or partially integrated into a physical entity, or physically separated. And these modules can all be implemented in the form of software called by a processing element; they can also all be implemented in the form of hardware; or some modules can be implemented in the form of software called by a processing element, and some modules can be implemented in the form of hardware. For example, the first preprocessing module can be a separately established processing element, or can be integrated in a certain chip of the above device. In addition, it can also be stored in the memory of the above device in the form of program code, and called and executed by a certain processing element of the above device to perform the functions of the above determined modules. The implementation of other modules is similar. In addition, these modules can be fully or partially integrated together, or can be independently implemented. The processing element described here can be an integrated circuit with signal processing capabilities. During the implementation process, each step of the above method or each of the above modules can be completed by the integrated logic circuit in the processor element or the instruction in the form of software.

[0136] For example, the above-mentioned modules may be one or more integrated circuits configured to implement the above methods, such as: one or more Application Specific Integrated Circuits (ASICs), or, one or more Digital Signal Processors (DSPs), or, one or more Field Programmable Gate Arrays (FPGAs), etc. For another example, when a certain above-mentioned module is implemented in the form of a processing element scheduler code, the processing element may be a general-purpose processor, such as a Central Processing Unit (CPU) or other processors that can call program code. For another example, these modules may be integrated together and implemented in the form of a System-on-a-chip (SOC).

[0137] In the above embodiments, it can be implemented in whole or in part by software, hardware, firmware, or any combination thereof. When implemented using software, it can be implemented in whole or in part in the form of a computer program product. The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, the processes or functions described in the foregoing method embodiments are generated in whole or in part. The above computer may be a general-purpose computer, a special-purpose computer, a computer network, or other programmable devices. The above computer instructions may be stored in a computer-readable storage medium, or transmitted from one computer-readable storage medium to another computer-readable storage medium. For example, the above computer instructions may be transmitted from a website, computer, server, or data center to another website, computer, server, or data center in a wired manner (such as coaxial cable, optical fiber, Digital Subscriber Line (DSL)) or wirelessly (such as infrared, wireless, Bluetooth, microwave, etc.). The above computer-readable storage medium may be any available medium that a computer can access or a data storage device such as a server or data center that includes one or more integrated available media. The above available medium may be a magnetic medium (such as a floppy disk, hard disk, magnetic tape), an optical medium (such as a DVD), or a semiconductor medium (such as a solid state disk (SSD)), etc.

[0138] Figure 3 This is a schematic structural diagram of an electronic device provided in Embodiment 3 of the present invention. The electronic device may be a terminal device or a server that implements the method of the foregoing embodiments, or may be a terminal device or a server that is connected to the foregoing terminal device or server and implements the method of the foregoing embodiments. As Figure 3As shown, the electronic device may include: a processor 301 (such as a CPU), a memory 302, and a transceiver 303; the transceiver 303 is coupled to the processor 301, and the processor 301 controls the transceiver operations of the transceiver 303. Various instructions may be stored in the memory 302 for completing various processing functions and implementing the processing steps described in the method of the foregoing embodiments. Preferably, the electronic device according to the embodiment of the present invention further includes: a power supply 304, a system bus 305, and a communication port 306. The system bus 305 is used to implement communication connections between components. The above-mentioned communication port 306 is used for the electronic device to connect and communicate with other peripherals.

[0139] As mentioned in Figure 3 the system bus 305 may be a Peripheral Component Interconnect (PCI) bus, an Extended Industry Standard Architecture (EISA) bus, or the like. The system bus may be divided into an address bus, a data bus, a control bus, etc. For the sake of simplicity of representation, only a thick line is used in the figure, but it does not mean that there is only one bus or one type of bus. The communication interface is used to implement communication between the database access device and other devices (such as clients, read-write libraries, and read-only libraries). The memory may include a Random Access Memory (RAM), and may also include a non-volatile memory, such as at least one disk memory.

[0140] The above-mentioned processor may be a general-purpose processor, including a Central Processing Unit (CPU), a Network Processor (NP), a Graphics Processing Unit (GPU), etc.; it may also be a Digital Signal Processor (DSP), an Application Specific Integrated Circuit (ASIC), a Field Programmable Gate Array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components.

[0141] It should be noted that the embodiment of the present invention further provides a computer-readable storage medium, in which instructions are stored, and when they run on a computer, the computer is caused to execute the methods and processing procedures provided in the above embodiments.

[0142] An embodiment of the present invention provides a method, apparatus, electronic device, and computer-readable storage medium for knowledge editing of an autoregressive large language model. As can be seen from the above, the embodiment of the present invention regards any autoregressive large language model implemented based on the Transformer model structure and having completed pre-training and NLP task fine-tuning as the target model; and converts each knowledge entry in the pre-training knowledge base of the target model into a corresponding question-answer text pair; and processes each text pair according to three types of model inference methods (single normal inference, single scrambled inference, multiple repair inferences under the premise of scrambling) to obtain three types of prediction information (the first predicted text, the second predicted text, the third prediction information set), and performs a key layer estimation once based on the three types of prediction information (the first predicted text, the second predicted text, the third prediction information set) of each text pair to obtain a corresponding first key layer set; and performs a key layer final judgment based on all the obtained first key layer sets to obtain a second key layer set; and resets the model parameters of the target model by implanting low-rank matrix parameters into the layer weight parameters of all the second key layers, and forms a corresponding implanted parameter set from all the implanted low-rank matrix parameters; and achieves the knowledge editing effect by only fine-tuning the implanted parameter set in each subsequent knowledge editing process. The embodiment of the present invention reduces the computational complexity, reduces the total amount of optimized parameters, thereby shortening the training cycle, improving the editing efficiency, and reducing the editing cost.

[0143] The steps of the methods or algorithms described in connection with the embodiments disclosed herein may be implemented in hardware, software modules executed by a processor, or a combination thereof. The software modules may be placed in random access memory (RAM), memory, read-only memory (ROM), electrically programmable ROM, electrically erasable programmable ROM, registers, hard disk, removable disk, CD-ROM, or any other form of storage medium known in the art.

[0144] The specific embodiments described above further elaborate on the purpose, technical solutions, and beneficial effects of the present invention. It should be understood that the above description is only the specific embodiments of the present invention and is not used to limit the protection scope of the present invention. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principles of the present invention shall be included in the protection scope of the present invention.

Claims

1. A knowledge editing method for an autoregressive large language model, characterized in that: The method comprises: A large autoregressive language model based on the Transformer model structure and pre-trained and fine-tuned on NLP tasks is used as the target model; and each attention layer or MLP layer in the inference section of the target model is recorded as the corresponding inference layer A i , and the reasoning layer A i The total number of N is calculated A ; 1≤layer index i≤N A ; The internal structure of the target model is divided into two major sections: a preprocessing section and the reasoning section; Configuring a question-answer instruction template for the target model; and converting each knowledge item in the pre-trained knowledge base of the target model into a question-answer text pair to obtain a corresponding first text pair; Each of the first text pairs is used as the corresponding current text pair one by one; the question text of the current text pair is brought into the question-answer instruction template to generate the current instruction text; the current instruction text is used as the model input text and input into the preprocessing module for preprocessing to obtain the corresponding initial vector H0; the initial vector H0 is input into the reasoning module for a forward reasoning and the generated text output at that time is used as the first predicted text, and all the reasoning layers A in the reasoning process at that time are predicted. i The output process vector is cached and a vector matrix M is formed by all cached process vectors; the initial vector H0 is scrambled to obtain a scrambled vector; the scrambled vector is input into the reasoning block for a forward reasoning and the generated text output at that time is used as the second predicted text; and the scrambled vector is input into the reasoning block for a round of N A The forward reasoning is repeated and a process vector correction is performed based on the vector matrix M in each reasoning process of this round to obtain a vector matrix consisting of N A A third prediction information set composed of a third prediction information; and based on the current text, a key layer estimation is performed on the corresponding first and second prediction texts and the third prediction information set to obtain a corresponding first key layer set; Performing key layer final judgment according to all the first key layer sets obtained to obtain a second key layer set; and resetting the model parameters of the target model by implanting low-rank matrix parameters into the layer weight parameters of all the second key layers; A corresponding implantation parameter set is formed by all the implanted low-rank matrix parameters; and only the implantation parameter set is updated during each knowledge editing process.

2. The knowledge editing method of the autoregressive large language model according to claim 1, characterized in that: The NLP tasks include at least text generation and intelligent question answering; The preprocessing module is used to preprocess the model input text, specifically: perform word segmentation processing on the model input text to obtain a word segmentation sequence, and perform embedding coding processing on each word of the word segmentation sequence according to the set embedding coding rule to obtain the corresponding initial vector H0 and send it to the reasoning module; the reasoning module is used to perform forward reasoning based on the input initial vector H0 to obtain the corresponding generated text and output it; the reasoning module includes N A The reasoning layer A i ; Each of the reasoning layers A i The layer weight parameter is recorded as weight W i ; The initial vector H0 is composed of multiple sub-vectors h 0,j Composition, 1≤subvector index j≤N J , N J is the total number of segmented words in the segmented word sequence corresponding to the initial vector H0; Each of the inference layers A corresponding to the initial vector H0 i The output vector is recorded as process vector H i ; The process vector H i By multiple sub-vectors h i,j The sub-vector h i,j With the subvector h 0,j One to one correspondence; The shape of the vector matrix M corresponding to the initial vector H0 is N A ×N J , by N A matrix rows, N J The vector matrix M includes N A ×N J Matrix cells m i,j ; Matrix rows and the process vector H i One-to-one correspondence, also with the reasoning layer A i One-to-one correspondence, the matrix columns correspond to the segmented words in the segmented word sequence one-to-one; the matrix unit m i,j With the subvector h i,j One to one correspondence; The question-and-answer instruction template is a formatted instruction text template; the configurable parameters of the question-and-answer instruction template include a question configuration parameter, and the question configuration parameter is a text parameter; the question-and-answer instruction template is used to prompt the target model to perform text generation processing on the corresponding answer to the question configuration parameter; The pre-trained knowledge base includes a plurality of the knowledge items; the text format of the knowledge items is a knowledge triple text format [subject text, subject-object relationship text, object text]; The first text pair corresponds to the knowledge item one by one; the first text pair includes the question text and the answer text; the question text is converted from the subject text and the subject-object relationship text of the corresponding knowledge item; the answer text is converted from the object text of the corresponding knowledge item; The third prediction information set is composed of N A The third prediction information includes a third prediction text and a first correction layer index; the first correction layer index and the reasoning layer A i One to one correspondence; The first key layer set is composed of multiple first key layers; each of the first key layers corresponds to one reasoning layer A i ; The second key layer set is composed of a plurality of the second key layers; each of the second key layers corresponds to one reasoning layer A i .

3. The knowledge editing method of the autoregressive large language model according to claim 2, characterized in that: The step of bringing the question text of the current text pair into the question-answer instruction template to generate the current instruction text specifically includes: The question configuration parameter of the question-and-answer instruction template is set to the question text of the current text pair to obtain a corresponding template setting text, and the template setting text is used as the corresponding current instruction text.

4. The knowledge editing method of the autoregressive large language model according to claim 2, characterized in that: The scrambling the initial vector H0 to obtain a scrambled vector specifically includes: The word segmentation sequence corresponding to the initial vector H0 is used as the current word segmentation sequence; and the subject participle in the current word segmentation sequence is identified; and the subvector h corresponding to the subject participle in the initial vector H0 is 0,j As the subject vector h subject ; and generate a vector length and vector feature dimension that are the same as the subject vector h subject Keeping the same random Gaussian noise vector ε; and based on the random Gaussian noise vector ε, subject Scramble to obtain the corresponding scrambling vector, scrambling vector = h subject +ε.

5. The knowledge editing method of the autoregressive large language model according to claim 2, characterized in that: Then the scrambled vector is input into the reasoning module for a round of N A The forward reasoning is repeated and a process vector correction is performed based on the vector matrix M in each reasoning process of this round to obtain a vector matrix consisting of N A The third prediction information set is composed of third prediction information, specifically including: Step 51, setting a first counter initialized to 1; Step 52: Input the scrambled vector into the reasoning module to perform a forward reasoning; and in this forward reasoning process, each time the reasoning layer A is passed, i Pause once; and at each pause, the current reasoning layer A i As the corresponding current reasoning layer; and the process vector H output by the current reasoning layer i as the corresponding current process vector; and identify whether the layer index i of the current reasoning layer matches the first counter; if not, release the pause and continue to transmit the current process vector to the next internal component in the reasoning block; if matched, take the matrix row corresponding to the current reasoning layer in the vector matrix M as the current matrix row, and based on the total number of segmentations N corresponding to the vector matrix M J Generate a value in N J / 2 to N J A random integer R between the current and the current matrix row is randomly selected R matrix cells m i,j For the R corresponding sub-vectors h in the current process vector i,j Replace the current process vector to obtain a new one, then release the pause and continue to transmit the new one to the next internal component in the reasoning module; and at the end of this forward reasoning, use the generated text output by this reasoning as a corresponding third predicted text, configure a corresponding first correction layer index for the current third predicted text, set the current first correction layer index to the first counter, and form a corresponding third prediction information by the current third predicted text and the corresponding first correction layer index; Step 53: add 1 to the first counter; and check whether the first counter after adding 1 exceeds the total number N. A Identify; if yes, go to step 54; if no, return to step 52; Step 54, based on the obtained N A The third prediction information constitutes a corresponding third prediction information set.

6. The knowledge editing method of the autoregressive large language model according to claim 2, characterized in that: The step of performing a key layer estimation on the corresponding first and second predicted texts and the third prediction information set based on the current text to obtain a corresponding first key layer set specifically includes: Step 61, using the answer text of the current text pair as the first label text; Step 62, inputting the first label text and the first predicted text as the corresponding current label text and current predicted text into a preset first model loss function for calculation and using the loss value output by the function as the corresponding loss value a1; The first model loss function is used to calculate the loss according to the input current label text and the current predicted text and output the corresponding loss value; the first model loss function is implemented based on the cross entropy loss function or the negative log-likelihood loss function; the loss value a1 ranges from 0 to 1; Step 63, inputting the first label text and the second predicted text as the corresponding current label text and the current predicted text into the first model loss function for calculation and using the loss value output by the function as the corresponding loss value a2; Wherein, the value range of the loss value a2 is between 0 and 1; Step 64, inputting the first label text and each of the third predicted texts in the third prediction information set as the corresponding current label text and the current predicted text into the first model loss function for calculation and using the loss value output by the function as the corresponding loss value a3; Wherein, the value range of each loss value a3 is between 0 and 1; Step 65, when the loss value a1 satisfies the preset model convergence loss value range and the loss value a1 is less than the loss value a2, the corresponding loss repair rate r is calculated based on the loss value a2 and each of the loss values ​​a3; and each of the loss repair rates r that is not less than the preset repair rate threshold is recorded as the corresponding preferred repair rate; and the first correction layer index corresponding to each of the preferred repair rates is recorded as the corresponding inference layer A i as a corresponding first key layer; and all the obtained first key layers constitute the first key layer set; The calculation method of the loss repair rate r is as follows: The repair rate threshold is a ratio value greater than 0.

7. The knowledge editing method of the autoregressive large language model according to claim 2, characterized in that: The step of performing a final key layer determination according to all the first key layer sets obtained to obtain a second key layer set specifically includes: All the first key layer sets are merged to obtain a corresponding first collection; the same first key layers in the first collection are grouped together as a corresponding first group; the total number of the first key layers in each of the first groups is counted and the statistical result is used as the corresponding key layer score; the first key layer corresponding to each of the key layer scores greater than a preset score threshold is used as a corresponding second key layer; and all the obtained second key layers form a corresponding second key layer set.

8. The knowledge editing method of the autoregressive large language model according to claim 2, characterized in that: The method of resetting the model parameters of the target model by implanting low-rank matrix parameters into the layer weight parameters of all second key layers specifically includes: The weight W corresponding to each of the second key layers i Denote the corresponding old weight W old ; and the old weight W old Set two corresponding low-rank matrices A and B; initialize the low-rank matrix A based on random Gaussian distribution, and initialize the low-rank matrix B to an all-zero matrix; and form the corresponding incremental weight △W=AB from the initialized low-rank matrices A and B; and old , the incremental weight △W constitutes the corresponding new weight W new =W old +△W; and based on the current new weight W new The weight W corresponding to the current second key layer i Reset; and the corresponding low-rank matrix parameters are composed of the matrix parameters of the low-rank matrices A and B corresponding to the current second key layer.

9. The knowledge editing method of the autoregressive large language model according to claim 2, characterized in that: In each knowledge editing process, only the implantation parameter set is updated, which specifically includes: Step 91, each time knowledge editing is required, first receiving the updated knowledge item set of that time; Wherein, the updated knowledge item set includes a plurality of updated knowledge items; the text format of the updated knowledge items is the knowledge triple text format; Step 92, generating a corresponding first training question based on the subject text and the subject-object relationship text of each updated knowledge item; generating a corresponding first label answer based on the object text of each updated knowledge item; and forming a corresponding first data record by each first training question and the corresponding first label answer; Step 93, perform another round of traversal on all the first data records; and in this round of traversal, use the first data record currently traversed as the corresponding current data record; and bring the first training question of the current data record into the question-answer instruction template to generate the corresponding first instruction text; and input the first instruction text as the model input text into the target model for processing and use the generated text output by this processing as the corresponding first predicted answer; and the first predicted answer and the first label answer of the current data record form a corresponding first prediction-label pair; Step 94, bringing all the obtained first prediction-label pairs into a preset second model loss function; and based on a preset first model optimizer, modulating the implanted parameter set of the target model in a direction that minimizes the second model loss function until the model loss converges; Among them, the second model loss function is implemented based on the cross entropy loss function or the negative log-likelihood loss function; the first model optimizer includes at least an Adam optimizer and an SGD optimizer.

10. A device for executing the knowledge editing method of the autoregressive large language model according to any one of claims 1 to 9, characterized in that: The device comprises: a first preprocessing module, a second preprocessing module, a key layer estimation module, a key layer positioning and low-rank matrix parameter implantation module, and a knowledge editing module; The first preprocessing module is used to use a large autoregressive language model that is implemented based on the Transformer model structure and has completed pre-training and NLP task fine-tuning as the target model; and each attention layer or MLP layer in the inference section of the target model is recorded as the corresponding inference layer A i , and the reasoning layer A i The total number of N is calculated A ; 1≤layer index i≤N A ; The internal structure of the target model is divided into two major sections: a preprocessing section and the reasoning section; The second preprocessing module is used to configure a question-answer instruction template for the target model; and convert each knowledge item in the pre-trained knowledge base of the target model into a question-answer text pair to obtain a corresponding first text pair; The key layer prediction module is used to take each of the first text pairs as the corresponding current text pair one by one; and bring the question text of the current text pair into the question-answer instruction template to generate the current instruction text; and input the current instruction text as the model input text into the preprocessing module for preprocessing to obtain the corresponding initial vector H0; and input the initial vector H0 into the reasoning module for a forward reasoning and use the generated text output at that time as the first predicted text, and perform forward reasoning on all the reasoning layers A in the reasoning process at that time. i The output process vector is cached and a vector matrix M is formed by all cached process vectors; the initial vector H0 is scrambled to obtain a scrambled vector; the scrambled vector is input into the reasoning block for a forward reasoning and the generated text output at that time is used as the second predicted text; and the scrambled vector is input into the reasoning block for a round of N A The forward reasoning is repeated and a process vector correction is performed based on the vector matrix M in each reasoning process of this round to obtain a vector matrix consisting of N A A third prediction information set composed of a third prediction information; and based on the current text, a key layer estimation is performed on the corresponding first and second prediction texts and the third prediction information set to obtain a corresponding first key layer set; The key layer positioning and low-rank matrix parameter implantation module is used to perform key layer final judgment according to all the first key layer sets obtained to obtain a second key layer set; and reset the model parameters of the target model by implanting low-rank matrix parameters into the layer weight parameters of all the second key layers; The knowledge editing module is used to form a corresponding implantation parameter set consisting of all the implanted low-rank matrix parameters; and only update the implantation parameter set during each knowledge editing process.

11. An electronic device, characterized in that: include: memory, processors, and transceivers; The processor is used to couple with the memory, read and execute instructions in the memory, so as to implement the method according to any one of claims 1 to 9; The transceiver is coupled to the processor, and the processor controls the transceiver to send and receive messages.

12. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores computer instructions, and when the computer instructions are executed by a computer, the computer is enabled to execute the method according to any one of claims 1 to 9.

Citation Information

Patent Citations

  • Editing method and device for pre-training language model

    CN117851613A

  • Knowledge distillation fine tuning method, device and equipment of large language model and storage medium

    CN118839749A

  • Method and device for processing question and answer tasks in combination with knowledge graph

    CN119357405A

Cited By

  • Knowledge graph-based knowledge question-answering method, medium and equipment

    CN121524297A

  • A knowledge editing method based on super network parameter generation

    CN122528834A

  • A knowledge editing method based on super network parameter generation

    CN122528834B