Processing method and device for parameter editing of pre-trained large model
By configuring a two-layer nonlinear network editor gi for pre-trained large models, using reverse gradient calculation and small sample training, the problems of low update efficiency and high cost of pre-trained large models when the data set is small, and efficient and low-cost parameter updates are achieved.
Patent Information
- Application Number
- CN202510332991.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-20
- Publication Date
- 2025-07-04
- Estimated Expiration
- 2045-03-20
AI Technical Summary
In the parameter update of pre-trained large models, especially when the data set is small, conventional fine-tuning methods are prone to cause overfitting problems, resulting in low update efficiency and high cost.
Using the editor gi based on two-layer nonlinear networks, the linear layers of the pre-trained large model are edited parameters, and the reverse gradient calculation and training of a small number of correlation and uncorrelated samples are achieved efficiently updating the model weight parameters.
Improves the flexibility and efficiency of parameter updates, reduces the update cost, and is suitable for parameter update tasks with a small number of input-output transformations.
Smart Images

Figure CN120255930A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of data processing, and particularly relates to a processing method and device for parameter editing of a pre-trained large model. Background Art
[0002] A pre-trained large language model (Pre-trained Large Language Models), also known as a pre-trained large model, refers to a large language model that has completed pre-training on a vast amount of corpus. Some model parameters of the pre-trained large model are similar to knowledge and have a time effect, and need to be updated in a timely manner. For example: the correct answer to a certain question x is answer y at time period t, but is changed to answer y at time period t + 1. * Theoretically, the model parameters should be updated in a timely manner at the start of time period t + 1 so that it can generate the updated y based on the input x. * .
[0003] Currently, the parameter update of the pre-trained large model mainly relies on the fine-tuning (Finetune) technology. However, when the scale of the dataset is small, the conventional fine-tuning method is prone to overfitting problems. In this case, even if only the output y of one input x needs to be adjusted to y. * , multiple input-output pairs also need to be additionally collected and combined with the current (x, y * ) to form a large dataset for fine-tuning. That is to say, this conventional update method has relatively low update efficiency and relatively high update cost when dealing with the parameter update task caused by a small number of input-output transformations (x, y → y * ). Summary of the Invention
[0004] The purpose of the present invention is to provide a processing method, device, electronic device, and computer-readable storage medium for parameter editing of a pre-trained large model in view of the defects of the prior art. The present invention denotes the pre-trained large model that needs to perform parameter editing as the target model, and configures a weight parameter editing model denoted as editor g for each linear layer of the target model. i ; and after receiving the input-output transformation dataset input by the user , select a set of transformation data from the current dataset D set as the base sample, and set a small number of related samples (x , y k ), and an unrelated sample (x k , y un , y un ) for the base sample, and form a training dataset D tr with all the samples; and based on the training dataset D tr and the target model for all editors gi Train; and after the training is completed, based on the dataset D set and all editors g i Edit all the weight parameters w of the target model i The dataset D set may include at least one set of transformation data. The editor g given in the present invention i is a simple editing model implemented based on a two-layer non-linear network (each layer of the non-linear network consists of a fully connected layer, an activation layer, and a residual connection module). It has a small number of model parameters, high training efficiency, and low training cost. The present invention uses the trained editor g i The technical solution for editing the parameters of the target model is a parameter editing scheme based on reverse gradient calculation. This scheme has a small amount of calculation, high editing flexibility, high editing efficiency, and low editing cost. When dealing with the parameter update task caused by a small number of input-output transformations (x, y → y * ), using the solution of the present invention can improve the update flexibility, improve the update efficiency, and reduce the update cost.
[0005] To achieve the above object, a first aspect of an embodiment of the present invention provides a processing method for editing parameters of a pre-trained large model, the method including:
[0006] Denote the pre-trained large model that needs to be parameter-edited as the target model; the target model includes N w layers of linear layers, N w is a positive integer greater than 1; the overall model parameters of the target model are denoted as θ, and the weight parameters and bias parameters of each linear layer are denoted as w i , b i , 1 ≤ layer index i ≤ N w , w i , b i ∈θ; during the forward prediction process of the target model, the input vector of each linear layer is denoted as u i , the gradient of the model loss L with respect to the output vector of each linear layer is denoted as δ i+1 , the gradient of the model loss L with respect to each weight parameter w i is denoted as The model loss L defaults to using the negative log-likelihood loss function L NLL as the loss function, L = L NLL (θ, y|x) = -log(P θ (y|x)), x is the model input, y is the model output, and P θ (y|x) is the probability that the model output of the target model is y when the model parameters are the overall model parameters θ and the model input is x;
[0007] Configure a weight parameter editing model for each linear layer of the target model, denoted as the editor g i ; All the editors g i have the same model structure and are all implemented based on a two-layer non-linear network. Each layer of the non-linear network consists of a fully connected layer, an activation layer, and a residual connection module; All the editors g i are used to fine-tune the vector u i and the gradient δ i+1 input to the current editor through a two-layer non-linear network and output the adjusted vector u new,i and the gradient δ new,i+1 ;
[0008] Receive the input-output transformation dataset input by the user, denoted as the dataset D set ; The dataset D set includes N D groups of transformation data N D is a configurable positive integer, with a minimum value of 1; 1 ≤ data index j ≤ N D , x j is the model input, is the latest output label corresponding to the input x j ;
[0009] Select any group of transformation data from the dataset D set as the base sample; And set a specified number N K of related samples and an unrelated sample for the base sample, and the training dataset D K is composed of the base sample, N tr equivalent samples and the unrelated sample; And based on the training dataset D tr , the target model trains all the editors g i ; The specified number N K is a preset positive integer, default less than 10; The base sample is denoted as The equivalent sample is denoted as (x k , y k ), 1 ≤ sample index k ≤ N K , The unrelated sample is denoted as (x un , y un ); The input x k of each equivalent sample is related to the input x s of the base sample, and the output label y k is related to the output label of the base sample; The input x un of the unrelated sample is related to the input x sIrrelevant, output label y un is also irrelevant to the output label of the basic sample ;
[0010] After the training of all editors is completed, based on the dataset D set and all the editors g i perform parameter editing on all the weight parameters w of the target model i .
[0011] Preferably, each of the editors g i is composed of an input splicing module, a first fully connected layer, a first activation layer, a first residual connection module, a second fully connected layer, a second activation layer, a second residual connection module, and an output splitting module;
[0012] The input end of the input splicing module is connected to the input end of the current editor, and the output end is respectively connected to the input end of the first fully connected layer and the first input end of the first residual connection module; the output end of the first fully connected layer is connected to the input end of the first activation layer; the output end of the first activation layer is connected to the second input end of the first residual connection module; the output end of the first residual connection module is respectively connected to the input end of the second fully connected layer and the first input end of the second residual connection module; the output end of the second fully connected layer is connected to the input end of the second activation layer; the output end of the second activation layer is connected to the second input end of the second residual connection module; the output end of the second residual connection module is connected to the input end of the output splitting module; the output end of the output splitting module is connected to the output end of the current editor;
[0013] The input splicing module is used to receive the vector u input by the current editor i and the gradient δ i+1 ; and splice the vector u i and the gradient δ i+1 to obtain the corresponding initial vector z 0,i and send it to the first fully connected layer and the first residual connection module;
[0014] The model parameters of the first fully connected layer include a scale vector parameter p 1,i , an offset vector parameter q 1,i , a low-rank matrix parameter A1, and a low-rank matrix parameter B1; the scale vector parameter p 1,i and the offset vector parameter q 1,i are private parameters of each of the editors g i ; the low-rank matrix parameters A1 and B1 are shared parameters of all the editors g i ;
[0015] The first fully-connected layer is used to calculate a corresponding hidden vector h through fully-connected calculation on the initial vector z according to the scale vector parameter p 1,i , the offset vector parameter q 1,i , the low-rank matrix parameter A1, and the low-rank matrix parameter B1 0,i and send it to the first activation layer; 1,i
[0016] wherein, the hidden vector h 1,i is: h 1,i = p 1,i ⊙(A1B1z 0,i + b)+ q 1,i ; b is a preset bias parameter, and the bias parameter b of the first fully-connected layer for all the editors g i is the same; ⊙ is the Hadamard product operator;
[0017] The first activation layer is implemented based on a class of non-linear activation functions σ1; the non-linear activation function σ1 includes at least the ReLU activation function;
[0018] The first activation layer is used to perform activation operation on the hidden vector h 1,i according to the non-linear activation function σ1 to obtain a corresponding activation vector a 1,i and send it to the first residual connection module;
[0019] wherein, the activation vector a 1,i is: a 1,i = σ1(h 1,i );
[0020] The first residual connection module is used to perform element-wise addition on the initial vector z 0,i and the activation vector a 1,i according to the residual connection method to obtain a corresponding first vector z 1,i and send it to the second fully-connected layer;
[0021] wherein, the first vector z 1,i is: z 1,i = z 0,i + a 1,i ;
[0022] The model parameters of the second fully-connected layer include the scale vector parameter p 2,i , the offset vector parameter q 2,i , the low-rank matrix parameter A2, and the low-rank matrix parameter B2; the scale vector parameter p 2,i and the offset vector parameter q 2,i for each of the editors g i private parameters; the low-rank matrix parameters A2 and B2 are shared parameters for all the editors g i ;
[0023] The second fully connected layer is used to perform a fully connected calculation on the first vector z according to the scale vector parameter p 2,i , the offset vector parameter q 2,i , the low-rank matrix parameter A2, and the low-rank matrix parameter B2 to obtain a corresponding hidden vector h 1,i and send it to the second activation layer; 2,i ;
[0024] wherein, the hidden vector h 2,i is: h 2,i = p 2,i ⊙ (A2B2z 1,i ) + q 2,i ;
[0025] The second activation layer is implemented based on a class of non-linear activation functions σ2; the non-linear activation function σ2 includes at least the ReLU activation function;
[0026] The second activation layer is used to perform an activation operation on the hidden vector h according to the non-linear activation function σ2 to obtain a corresponding activation vector a 2,i and send it to the second residual connection module; 2,i ;
[0027] wherein, the activation vector a 2,i is: a 2,i = σ2(h 2,i );
[0028] The second residual connection module is used to perform element-wise addition on the first vector z 1,i and the activation vector a 2,i to obtain a corresponding second vector z 2,i and send it to the output splitting module;
[0029] wherein, the second vector z 2,i is: z 2,i = z 1,i + a 2,i ;
[0030] The output splitting module is used to extract the corresponding vector u 0,i and the gradient δ i from the second vector z i+1 with reference to the splicing position of the vector u 2,i and the gradient δ new,i in the initial vector z new,i+1 and output them.
[0031] Preferably, based on the training dataset D tr and the target model for all the editors g i are trained, specifically including:
[0032] Step 31, extract the corresponding basic samples tr from the training dataset D N K equivalent samples (x k , y k ) and the irrelevant samples (x un , y un ); and make a parameter backup of the current overall model parameters θ of the target model and save them as the corresponding backup parameters θ0; and use the current overall model parameters θ as the corresponding current model parameters θ now ; and form the corresponding editor parameter set φ from the editor parameters of all the editors g i ; and initialize all the parameters of the editor parameter set φ to obtain the current parameter set φ now ; and set a first counter initialized to 0;
[0033] Among them, the editor parameter set φ consists of the private scale vector parameters p i p 1,i / p 2,i , the offset vector parameters q 1,i / q 2,i and the shared low-rank matrix parameters A1 / A2, B1 / B2;
[0034] Step 32, set the editor parameters of all the editors g now based on the current parameter set φ i ;
[0035] Step 33, input the x of the basic sample s into the target model for forward prediction processing, cache each vector u i generated during the processing, and calculate the corresponding model loss
[0036] Among them, is the probability that the output of the target model is now when its model parameters are the current model parameters θ s and the model input is x ;
[0037] Step 34, according to the current model parameters θnow Each of the weight parameters w in i and its corresponding bias parameter b i , the vector u i Calculate the corresponding linear layer output vector u i+1 ; and calculate the model loss L for each vector u i+1 The gradient of the corresponding gradient δ i+1 ; and each of the gradients δ i+1 and its corresponding vector u i Enter the corresponding editor g i Fine-tune to obtain the corresponding vector u new,i and the gradient δ new,i+1 ; and based on each of the vectors u new,i and its corresponding gradient δ new,i+1 Calculate new weight gradients And based on each of the weight parameters w i and its corresponding weight gradient Calculate the corresponding weight parameters And the current model parameter θ now Each of the weight parameters w in i Reset to the corresponding weight parameter
[0038] in,
[0039] The vector u i+1 The calculation method is:
[0040] u i+1 =w i u i +b i ;
[0041] The gradient δ i+1 The calculation method is:
[0042]
[0043] The vector u new,i and the gradient δ new,i+1 The calculation method is:
[0044] g i (u i ,δ i+1 )→(u new,i ,δ new,i+1 ), g i (u i ,δ i+1 ) is the editor g i Expression of
[0045] The weight gradient is calculated as follows:
[0046] T is the transpose symbol;
[0047] The weight parameter is calculated as follows:
[0048]
[0049] Step 35, based on the latest current model parameter θ now reset the model parameters of the target model; and input the x of each equivalent sample (x k , y k ) into the target model for forward prediction processing, and calculate the corresponding first model loss L k ; 1,k ;
[0050] wherein, the first model loss L 1,k defaults to the negative log-likelihood loss function L NLL as the loss function, specifically:
[0051]
[0052] is the probability that the output of the target model is the corresponding y now when its model parameter is the latest current model parameter θ k and the model input is x k ;
[0053] Step 36, input the x of the uncorrelated sample (x un , y un ) into the target model for forward prediction processing, and calculate the corresponding edited probability un Then reset the model parameters of the target model based on the backup parameter θ0; and input the x of the uncorrelated sample (x un , y un , y un ) into the target model for forward prediction processing, and calculate the corresponding pre-edited probability pre and calculate the corresponding second model loss L2 based on the pre-edited probability r pre , the edited probability r post and the divergence loss function L KL ; ;
[0054] wherein, is the probability that the output of the target model is the corresponding y now when its model parameter is the latest current model parameter θnow and the model input is x un in which case its model output is y un probability;
[0055] is the probability that the model output is y when the model parameters of the target model are the backup parameters θ0 and the model input is x un in which case its model output is y un probability;
[0056] The second model loss L2 defaults to using the divergence loss function L KL as the loss function, specifically:
[0057]
[0058] Step 37, from N K of the first model losses L 1,k and the second model loss L2 to form the corresponding editor loss L G ;
[0059] wherein, the editor loss L G is:
[0060]
[0061] λ is a preset balance parameter;
[0062] Step 38, in the direction of minimizing the editor loss L G based on a preset first model optimizer, perform a round of parameter modulation on all the current parameter sets φ i corresponding to the editors g now ; and use the modulated editor parameter set as the new current parameter set φ now ;
[0063] wherein, the first model optimizer at least includes the Adam optimizer;
[0064] Step 39, increment the first counter by 1; and identify whether the incremented first counter exceeds a preset counter threshold; if not, return to Step 32 to continue training; if it has exceeded, stop training, and based on the latest current parameter set φ now solidify the editor parameters of all the editors g i and confirm that the training of all editors is completed.
[0065] Preferably, all the weight parameters w of the target model based on the dataset D set and all the editors g i i Perform parameter editing, specifically including:
[0066] Set the model parameters of the target model based on the backup parameter θ0; and for the dataset D set All the transformed data Perform a round of traversal; and during this round of traversal, use the currently traversed transformed data As the corresponding current transformed data (x c , y c ); and input the x of the current transformed data (x c , y c ) into the target model for forward prediction processing, and cache each vector u c Generated during the processing, and calculate the corresponding model loss i And calculate the linear layer output vector u corresponding to each weight parameter w In the backup parameter θ0 and its corresponding bias parameter b i And the vector u i ; and calculate the gradient of the model loss L with respect to each vector u i To obtain the corresponding gradient δ i+1 ; and input each gradient δ i+1 And its corresponding vector u i+1 Into the corresponding editor g i+1 For fine-tuning to obtain the corresponding vector u i And the gradient δ i ; and calculate the new weight gradient based on each vector u new,i And its corresponding gradient δ new,i+1 ; and calculate the corresponding weight parameter based on each weight parameter w new,i And its corresponding weight gradient new,i+1 ; and reset each weight parameter w in the current backup parameter θ0 To the corresponding weight parameter i ; and perform a parameter reset on the model parameters of the target model once based on the latest backup parameter θ0; Among them, the calculation method of the vector u i Is:
[0067]
[0068] u i+1 = w i+1 i u i + b i ;
[0069] The gradient δ i+1 is calculated as follows:
[0070]
[0071] The vector u new,i and the gradient δ new,i+1 are calculated as follows:
[0072] g i (u i , δ i+1 ) → (u new,i , δ new,i+1 ), where g i (u i , δ i+1 ) is the expression of the editor g i ;
[0073] The weight gradient is calculated as follows:
[0074] T is the transpose symbol;
[0075] The weight parameter is calculated as follows:
[0076]
[0077] In the second aspect of the embodiments of the present invention, there is provided an apparatus for implementing the processing method for parameter editing of the pre-trained large model described in the first aspect above. The apparatus includes: a first preprocessing module, a second preprocessing module, a data receiving module, an editor training module, and a model parameter editing module;
[0078] The first preprocessing module is configured to record the pre-trained large model that needs to be parameter - edited as a target model; the target model includes N w layers of linear layers, and N w is a positive integer greater than 1; the overall model parameters of the target model are denoted as θ, and the weight parameters and bias parameters of each linear layer are denoted as w i , b i , where 1 ≤ layer index i ≤ N w , w i , b i ∈θ; during the forward prediction process of the target model, the input vector of each linear layer is denoted as u i , the gradient of the model loss L with respect to the output vector of each linear layer is denoted as δ i+1 , and the gradient of the model loss L with respect to each weight parameter w i is denoted as The model loss L is defaulted to the negative log - likelihood loss function LNLL is the loss function, L = L NLL (θ, y|x) = -log(P θ (y|x)), where x is the model input, y is the model output, and P θ (y|x) is the probability that the model output of the target model is y when the model parameters are the overall model parameters θ and the model input is x;
[0079] The second preprocessing module is used to configure a weight parameter editing model for each linear layer of the target model, denoted as the editor g i ; All the editors g i have the same model structure, which is implemented based on a two-layer non-linear network. Each layer of the non-linear network consists of a fully connected layer, an activation layer, and a residual connection module; All the editors g i are used to fine-tune the vector u i input to the current editor, the gradient δ i+1 and output the adjusted vector u new,i and gradient δ new,i+1 ;
[0080] The data receiving module is used to receive the input-output transformation data set input by the user, denoted as the data set D set ; The data set D set includes N D groups of transformation data N D is a configurable positive integer, and the minimum can be set to 1; 1 ≤ data index j ≤ N D , x j is the model input, is the latest output label corresponding to the input x j ;
[0081] The editor training module is used to select a group of transformation data from the data set D set as the base sample; And set a specified number N K of related samples and an unrelated sample for the base sample, and the base sample, N K equivalent samples and the unrelated sample form the training data set D tr ; And based on the training data set D tr , the target model trains all the editors g i ; The specified number N K is a preset positive integer, default less than 10; The base sample is denoted as The equivalent sample is denoted as (x k , y k ), 1 ≤ sample index k ≤ NK , the irrelevant samples are denoted as (x un , y un ); the input x of each of the equivalent samples k is related to the input x of the base sample s , and the output label y k is related to the output label of the base sample ; the input x of the irrelevant sample un is not related to the input x of the base sample s , and the output label y un is also not related to the output label of the base sample ;
[0082] The model parameter editing module is used to, after the training of all the editors is completed, based on the dataset D set and all the editors g i perform parameter editing on all the weight parameters w i of the target model.
[0083] A third aspect of the embodiments of the present invention provides an electronic device, including: a memory, a processor, and a transceiver;
[0084] The processor is used to be coupled with the memory, read and execute the instructions in the memory to implement the method steps described in the first aspect above;
[0085] The transceiver is coupled with the processor, and the processor controls the transceiver to perform message sending and receiving.
[0086] A fourth aspect of the embodiments of the present invention provides a computer-readable storage medium, and the computer-readable storage medium stores computer instructions, and when the computer instructions are executed by a computer, the computer is caused to execute the instructions of the method described in the first aspect above.
[0087] The embodiments of the present invention provide a processing method, device, electronic device, and computer-readable storage medium for parameter editing of a pre-trained large model. As can be seen from the above content, the embodiments of the present invention denote the pre-trained large model that needs to perform parameter editing as the target model, and configure a weight parameter editing model for each linear layer of the target model, denoted as the editor g i ; and after receiving the input-output transformation dataset input by the user, select a set of transformation data set from the current dataset D as the base sample, and set a small number of related samples (x k , y k ) and an irrelevant sample (x un , yun ) and form a training data set D with all samples tr ; and based on the training data set D tr and the target model, train all editors g i ; and after the training ends, based on the data set D set and all editors g i perform parameter editing on all weight parameters w of the target model i . The data set D set may include at least one set of transformation data at least. The editor g given in the embodiment of the present invention i is a simple editing model implemented based on a two-layer non-linear network (each layer of the non-linear network is composed of a fully connected layer, an activation layer, and a residual connection module). It has a small number of model parameters, high training efficiency, and low training cost. The technical solution for parameter editing of the target model using the trained editor g i in the embodiment of the present invention is a parameter editing scheme based on reverse gradient calculation. This scheme has a small amount of calculation, high editing flexibility, high editing efficiency, and low editing cost. When processing a parameter update task caused by a small number of input-output transformations (x, y → y * ), if the scheme in the embodiment of the present invention is used, technical effects of improving update flexibility, improving update efficiency, and reducing update cost can be achieved. Description of the Drawings
[0088] Figure 1 is a schematic diagram of a processing method for parameter editing of a pre-trained large model provided in Embodiment 1 of the present invention;
[0089] Figure 2 is a module structure diagram of the editor g provided in Embodiment 1 of the present invention i ;
[0090] Figure 3 is a module structure diagram of a processing device for parameter editing of a pre-trained large model provided in Embodiment 2 of the present invention;
[0091] Figure 4 is a schematic diagram of the structure of an electronic device provided in Embodiment 3 of the present invention. Detailed Embodiment
[0092] In order to make the objectives, technical solutions, and advantages of the present invention clearer, the present invention will be further described in detail below with reference to the drawings. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all of the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.
[0093] Embodiment 1 of the present invention provides a processing method for parameter editing of a pre-trained large model, as follows Figure 1 As shown in the schematic diagram of the processing method for parameter editing of a pre-trained large model provided by Embodiment 1 of the present invention, the method mainly includes the following steps:
[0094] Step 1, denote the pre-trained large model that needs parameter editing as the target model.
[0095] Here, the pre-trained large model in the embodiment of the present invention can be any type of large language model implemented based on the Transformer model structure, or any type of large language model implemented based on the LSTM / Bi-LSTM model structure, or any type of large language model implemented based on the convolutional neural network. It should be noted that regardless of which type of structure the pre-trained large model is implemented based on, the embodiment of the present invention requires that there be a relatively large number of linear layers inside the pre-trained large model, and corresponding linear layers should be configured inside or on the output side of each key component (such as the embedding encoding module, encoder, decoder, etc.) inside the model, and a corresponding linear layer should be equipped at the model output end.
[0096] Some key parameters of the pre-trained large model, that is, the target model, in the embodiment of the present invention are described below:
[0097] 1) The target model includes N w layers of linear layers, where N w is a positive integer greater than 1;
[0098] 2) Denote the overall model parameters of the target model as θ, and denote the weight parameters and bias parameters of each linear layer as w i , b i ; where 1 ≤ layer index i ≤ N w , w i , b i ∈θ;
[0099] 3) Denote the input vector of each linear layer in the forward prediction process of the target model as u i ; denote the gradient of the model loss L with respect to the output vector u i+1 of each linear layer as δ i+1 ; denote the gradient of the model loss L with respect to the weight parameter w i of each linear layer as This gradient can be further expressed as: T is the transpose symbol;
[0100] Here, the model loss L of the target model in the embodiment of the present invention defaults to using the negative log-likelihood loss function L NLL as the loss function, and its calculation process is specifically as follows:
[0101] L = LNLL (θ, y|x) = -log(P θ (y|x)),
[0102] where x is the model input, y is the model output, and P θ (y|x) is the probability that the model output of the target model is y when the model parameters are the overall model parameters θ and the model input is x.
[0103] Step 2, configure a weight parameter editing model for each linear layer of the target model, denoted as editor g i .
[0104] Here, all editors g of the embodiments of the present invention i have the same model structure and are all implemented based on a two-layer non-linear network. Each layer of the non-linear network consists of a fully connected layer, an activation layer, and a residual connection module. All editors g of the embodiments of the present invention i are used to fine-tune the vector u i input to the current editor and the gradient δ i+1 through a two-layer non-linear network and output the adjusted vector u new,i and the gradient δ new,i+1 .
[0105] The editor g of the embodiments of the present invention will be described in detail below. i
[0106] As Figure 2 shown in the module structure diagram of editor g provided in Embodiment 1 of the present invention i , each editor g of the embodiments of the present invention i consists of an input splicing module, a first fully connected layer, a first activation layer, a first residual connection module, a second fully connected layer, a second activation layer, a second residual connection module, and an output splitting module.
[0107] As Figure 2 shown, the connection relationship of each component in editor g i is as follows: the input end of the input splicing module is connected to the input end of the current editor, and the output end is respectively connected to the input end of the first fully connected layer and the first input end of the first residual connection module; the output end of the first fully connected layer is connected to the input end of the first activation layer; the output end of the first activation layer is connected to the second input end of the first residual connection module; the output end of the first residual connection module is respectively connected to the input end of the second fully connected layer and the first input end of the second residual connection module; the output end of the second fully connected layer is connected to the input end of the second activation layer; the output end of the second activation layer is connected to the second input end of the second residual connection module; the output end of the second residual connection module is connected to the input end of the output splitting module; the output end of the output splitting module is connected to the output end of the current editor.
[0108] Editor g i The component functions of each component inside are as follows.
[0109] 1) Input splicing module:
[0110] The input splicing module of the embodiment of the present invention is used to receive the vector u input by the current editor i and the gradient δ i+1 ; and splice the vector u i and the gradient δ i+1 to obtain the corresponding initial vector z 0,i and send it to the first fully connected layer and the first residual connection module.
[0111] 2) First fully connected layer:
[0112] The model parameters of the first fully connected layer of the embodiment of the present invention include the scale vector parameter p 1,i 、the offset vector parameter q 1,i 、the low-rank matrix parameter A1, and the low-rank matrix parameter B1; among them, the scale vector parameter p 1,i and the offset vector parameter q 1,i are the private parameters of each editor g i ; the low-rank matrix parameters A1 and B1 are the shared parameters of all editors g i .
[0113] This first fully connected layer is used to perform a fully connected calculation on the initial vector z 1,i according to the scale vector parameter p 1,i 、the offset vector parameter q 0,i 、the low-rank matrix parameter A1 and the low-rank matrix parameter B1 to obtain the corresponding hidden vector h 1,i and send it to the first activation layer.
[0114] Here, the specific calculation formula of the hidden vector h 1,i is:
[0115] h 1,i = p 1,i ⊙(A1B1z 0,i + b) + q 1,i ;
[0116] where b is a preset bias parameter, and the bias parameter b of all editors g i is the same; ⊙ is the Hadamard product operator.
[0117] 3) First activation layer:
[0118] The first activation layer in the embodiment of the present invention is implemented based on a class of non-linear activation functions σ1. Here, the non-linear activation function σ1 includes at least the ReLU activation function.
[0119] This first activation layer is used to perform an activation operation on the hidden vector h 1,i to obtain the corresponding activation vector a 1,i and send it to the first residual connection module.
[0120] Here, the specific calculation formula for the activation vector a 1,i is:
[0121] a 1,i = σ1(h 1,i ).
[0122] 4) First residual connection module:
[0123] The first residual connection module in the embodiment of the present invention is used to perform element-wise addition on the initial vector z 0,i and the activation vector a 1,i to obtain the corresponding first vector z 1,i and send it to the second fully connected layer.
[0124] Here, the specific calculation formula for the first vector z 1,i is:
[0125] z 1,i = z 0,i + a 1,i .
[0126] 5) Second fully connected layer:
[0127] The model parameters of the second fully connected layer include the scale vector parameter p 2,i , the offset vector parameter q 2,i , the low-rank matrix parameter A2, and the low-rank matrix parameter B2. Among them, the scale vector parameter p 2,i and the offset vector parameter q 2,i are the private parameters of each editor g i ; the low-rank matrix parameters A2 and B2 are the shared parameters of all editors g i .
[0128] This second fully connected layer is used to perform a fully connected calculation on the first vector z 2,i according to the scale vector parameter p 2,i , the offset vector parameter q 1,i , the low-rank matrix parameter A2, and the low-rank matrix parameter B2 to obtain the corresponding hidden vector h 2,i and send it to the second activation layer.
[0129] Here, the hidden vector h2,i The specific calculation formula is:
[0130] h 2,i = p 2,i ⊙(A2B2z 1,i ) + q 2,i .
[0131] 6) Second activation layer:
[0132] The second activation layer of the embodiment of the present invention is implemented based on a class of non - linear activation functions σ2. Here, the non - linear activation function σ2 includes at least the ReLU activation function.
[0133] This second activation layer is used to perform an activation operation on the hidden vector h 2,i to obtain the corresponding activation vector a 2,i and send it to the second residual connection module.
[0134] Here, the specific calculation formula of the activation vector a 2,i is:
[0135] a 2,i = σ2(h 2,i ).
[0136] 7) Second residual connection module:
[0137] The second residual connection module of the embodiment of the present invention is used to perform element - by - element addition on the first vector z 1,i and the activation vector a 2,i to obtain the corresponding second vector z 2,i and send it to the output splitting module.
[0138] Here, the specific calculation formula of the second vector z 2,i is:
[0139] z 2,i = z 1,i + a 2,i .
[0140] 8) Output splitting module:
[0141] The output splitting module of the embodiment of the present invention is used to extract the corresponding vector u 0,i and the gradient δ i from the second vector z i+1 with reference to the splicing position of the vector u 2,i and the gradient δ new,i in the initial vector z new,i+1 and output them.
[0142] Step 3, receive the input - output transformation data set input by the user, denoted as data set Dset 。
[0143] Here, the data set D of the embodiment of the present invention set includes N D groups of transformation data N D is a configurable positive integer, and the minimum can be set to 1; 1 ≤ data index j ≤ N D , x j is the model input,[[]] is the input x j corresponding to the latest output label.[[]]
[0144] Step 4, select a group of transformation data from the data set D set as the base sample; and set a specified number N K of relevant samples and an irrelevant sample for the base sample, and the base sample, N K equivalent samples and the irrelevant sample form the training data set D tr ; and based on the training data set D tr , the target model trains all editors g i ;[[]]
[0145] Specifically, it includes: Step 41, select a group of transformation data from the data set D set as the base sample; and set a specified number N K of relevant samples and an irrelevant sample for the base sample, and the base sample, N K equivalent samples and the irrelevant sample form the training data set D tr ;[[]]
[0146] Here, the specified number N of the embodiment of the present invention K is a preset positive integer, and by default it is less than 10; the base sample of the embodiment of the present invention is denoted as The equivalent sample of the embodiment of the present invention is denoted as (x k , y k ), 1 ≤ sample index k ≤ N K , the irrelevant sample of the embodiment of the present invention is denoted as (x un , y un );[[]]
[0147] It should be noted that the inputs x k of each equivalent sample (x k , y k ) of the embodiment of the present invention are all related to the input x s of the base sample, and the output labels y k are all related to the output label of the base sample; it should also be noted that the irrelevant sample (xun , y un ) input x un is independent of the input x of the base sample s , and the output label y un is also independent of the output label of the base sample ;
[0148] How to understand the correlation and non - correlation between equivalent samples (x k , y k ), non - related samples (x un , y un ) and the base sample is illustrated by a simple example below:
[0149] For example, it is known that the target model learned Knowledge 1 "The father of Person A is Person B and the mother is Person C", Knowledge 2 "The address of Person B is Address B and the address of Person C is Address C", and Knowledge 3 "The address of Person D, who has no relation with Persons A, B, and C, is Address D" after the last pre - training or parameter update; but through fact - checking, we find that the father of Person A should actually be Person E and the address of Person E is Address E; therefore, the model parameters of the target model need to be updated based on the new knowledge content "The father of Person A is Person E and the address of Person E is Address E";
[0150] Assume that the x of the base sample s is the question "Who is the father of Person A?", then, when the parameters of the target model have not been updated, the corresponding output label y s should be the answer "The father of Person A is Person B"; after the parameters of the target model have been updated, the corresponding output label should be the answer "The father of Person A is Person E";
[0151] When x s is the question "Who is the father of Person A?" and is the answer "The father of Person A is Person E", under the condition that the question x k "What is the address of the father of Person A?" and the answer y k "The address of the father of Person A is Address E" are a set of equivalent samples (x , y k , y k ) related to the base sample; while the question x un "What is the address of Person D?" and the answer y un "The address of Person D is Address D" are a set of non - related samples (x , y un , y un ) that are not related to the base sample;
[0152] In summary: 1) For equivalent samples (x k ,y k ) For example, the problem x k The answers before and after the model parameters are updated are different, and the changes in the model output answers are different from the answers ; 2) For irrelevant samples (x un ,y un ) For example, the problem x un The answer to the question x is the same before and after the model parameter update. un The corresponding model output answer will not change;
[0153] Step 42, and based on the training data set D tr , target model for all editors g i Conduct training;
[0154] Specifically include: Step 421, from the training data set D tr Extract the corresponding basic samples N K Equivalent samples (x k ,y k ) and irrelevant samples (x un ,y un );and make a parameter backup of the current overall model parameter θ of the target model and save it as the corresponding backup parameter θ0;and use the current overall model parameter θ as the corresponding current model parameter θ now ; and by all editors g i The editor parameters of the corresponding editor parameter set φ are formed; and all parameters of the editor parameter set φ are initialized to obtain the current parameter set φ now ; and set a first counter initialized to 0;
[0155] The editor parameter set φ consists of all editors g i Private scale vector parameter p 1,i , offset vector parameter q 1,i , scale vector parameter p 2,i , offset vector parameter q 2,i and shared low-rank matrix parameter A1, low-rank matrix parameter B1, low-rank matrix parameter A2 and low-rank matrix parameter B2;
[0156] Step 422, based on the current parameter set φ now For all editors i Set the editor parameters;
[0157] Step 423: The basic sample xs Input the target model for forward prediction processing, and cache each vector u generated during the processing, and calculate the corresponding model loss L; i Cache each vector u generated during the processing, and calculate the corresponding model loss L;
[0158] Here, the model loss L calculated in the current step is:
[0159]
[0160] Wherein, is the probability that the output of the target model is now when its model parameters are the current model parameters θ s and the model input is x ;
[0161] Step 424, according to each weight parameter w now in the current model parameters θ i and its corresponding bias parameter b i , vector u i calculate the corresponding linear layer output vector u i+1 ; and calculate the gradient of the model loss L with respect to each vector u i+1 to obtain the corresponding gradient δ i+1 ; and input each gradient δ i+1 and its corresponding vector u i into the corresponding editor g i for fine-tuning to obtain the corresponding vector u new,i and gradient δ new,i+1 ; and calculate a new weight gradient based on each vector u new,i and its corresponding gradient δ new,i+1 and calculate the corresponding weight parameter based on each weight parameter w i and its corresponding weight gradient and reset each weight parameter w in the current model parameters θ now to the corresponding weight parameter i
[0162] Here:
[0163] a. The calculation method of vector u i+1 is: u i+1 = w i u i + b i ;
[0164] b. The calculation method of gradient δ i+1 is:
[0165] c. Vector u new,i and gradient δ new,i+1 are calculated as follows:
[0166] g i (u i , δ i+1 ) → (u new,i , δ new,i+1 ), g i (u i , δ i+1 ) is the expression of editor g i ;
[0167] d. Weight gradient is calculated as follows:
[0168] e. Weight parameter is calculated as follows:
[0169] Step 425: Reset the model parameters of the target model based on the latest current model parameters θ now ; and input x of each equivalent sample (x k , y k ) into the target model for forward prediction processing, and calculate the corresponding first model loss L k ; 1,k ;
[0170] Here, each first model loss L of the embodiments of the present invention 1,k is defaulted to use the negative log-likelihood loss function L NLL as the loss function, specifically:
[0171]
[0172] Among them, is the probability that the output of the target model is the corresponding y now when the model parameters of the target model are the latest current model parameters θ k and the model input is x k ;
[0173] Step 426: Input x of the irrelevant sample (x un , y un ) into the target model for forward prediction processing, and calculate the corresponding edited probability un Then reset the model parameters of the target model based on the backup parameter θ0; and input x of the irrelevant sample (x un , y, y un ) into the target model for forward prediction processing, and calculate the corresponding probability before editing un and based on the probability r before editing pre , the probability r after editing post and the divergence loss function L KL calculate the corresponding second model loss L2;
[0174] Here, in the embodiments of the present invention is the probability that the output of the target model is y when its model parameters are the latest current model parameters θ now and the model input is x un ; in the embodiments of the present invention un is the probability that the output of the target model is y when its model parameters are the backup parameters θ0 and the model input is x ; in the embodiments of the present invention, the second model loss L2 of the present invention defaults to use the divergence loss function L un as the loss function, specifically: un KL K 1,k
[0175]
[0176] Step 427, compose the corresponding editor loss L K from N 1,k first model losses L G and the second model loss L2;
[0177] Here, the editor loss L in the embodiments of the present invention G is:
[0178]
[0179] where λ is a preset balance parameter;
[0180] Step 428, in the direction of minimizing the editor loss L G , perform a round of parameter modulation on the current parameter set φ i corresponding to all editors g now ; and use the modulated editor parameter set as the new current parameter set φ now ;
[0181] Here, the first model optimizer in the embodiments of the present invention includes at least the Adam optimizer;
[0182] Step 429, increment the first counter by 1; and identify whether the incremented first counter exceeds a preset counter threshold; if not, return to step 422 to continue training; if it has exceeded, stop training, and based on the latest current parameter set φ now for all editors g iSolidify the editor parameters and confirm that the training of all editors is completed.
[0183] Here, the counter threshold is a preset threshold parameter.
[0184] Step 5, after the training of all editors is completed, based on the dataset D set and all editors g i Edit all the weight parameters w i of the target model;
[0185] Specifically, it includes: setting the model parameters of the target model based on the backup parameters θ0; and traversing all the transformed data set of the dataset D once; and during this traversal, take the currently traversed transformed data as the corresponding current transformed data (x c , y c ); and input the x c of the current transformed data (x c , y c ) into the target model for forward prediction processing, cache each vector u i generated during the processing, and calculate the corresponding model loss And calculate the corresponding linear layer output vector u i based on each weight parameter w i in the backup parameter θ0 and its corresponding bias parameter b i and the vector u i+1 ; and calculate the gradient of the model loss L with respect to each vector u i+1 to obtain the corresponding gradient δ i+1 ; and input each gradient δ i+1 and its corresponding vector u i into the corresponding editor g i for fine-tuning to obtain the corresponding vector u new,i and gradient δ new,i+1 ; and calculate the new weight gradient new,i based on each vector u new,i+1 and its corresponding gradient δ And calculate the corresponding weight parameter i based on each weight parameter w and its corresponding weight gradient And reset each weight parameter w i in the current backup parameter θ0 to the corresponding weight parameter And perform a parameter reset on the model parameters of the target model based on the latest backup parameter θ0.
[0186] Here, in the current step is the probability that the model output of the target model is y when its model parameters are the backup parameters θ0 and the model input is x c ; the calculation methods of the parameters in the current step are similar to those in step 424, that is: c
[0187] a. The calculation method of the vector u i+1 is: u i+1 = w i u i + b i ;
[0188] b. The calculation method of the gradient δ i+1 is:
[0189] c. The calculation methods of the vector u new,i and the gradient δ new,i+1 are:
[0190] g i (u i , δ i+1 ) → (u new,i , δ new,i+1 ), g i (u i , δ i+1 ) is the expression of the editor g i ;
[0191] d. The calculation method of the weight gradient is:
[0192] e. The calculation method of the weight parameter is:
[0193] Figure 3 FIG. Figure 3 shows a module structure diagram of a processing device for parameter editing of a pre-trained large model provided in the second embodiment of the present invention. The device is a terminal device or a server for implementing the foregoing method embodiment, or may be a device that enables the foregoing terminal device or server to implement the foregoing method embodiment. For example, the device may be a device or a chip system of the foregoing terminal device or server. As
[0194] shown, the device includes: a first preprocessing module 201, a second preprocessing module 202, a data receiving module 203, an editor training module 204, and a model parameter editing module 205. w The first preprocessing module 201 is used to record the pre-trained large model that needs to be parameter edited as the target model; the target model includes N w layers of linear layers, Nis a positive integer greater than 1; the overall model parameters of the target model are denoted as θ, and the weight parameters and bias parameters of each linear layer are denoted as w i , b i , 1 ≤ layer index i ≤ N w , w i , b i ∈ θ; during the forward prediction process of the target model, the input vectors of each linear layer are denoted as u i , the gradient of the model loss L with respect to the output vector of each linear layer is denoted as δ i+1 , the gradient of the model loss L with respect to each weight parameter w i is denoted as The model loss L defaults to using the negative log-likelihood loss function L NLL as the loss function, L = L NLL (θ, y|x) = -log(P θ (y|x)), x is the model input, y is the model output, and P θ (y|x) is the probability that the model output of the target model is y when the model parameters are the overall model parameters θ and the model input is x.
[0195] The second preprocessing module 202 is used to configure a weight parameter editing model for each linear layer of the target model, denoted as the editor g i ; all editors g i have the same model structure and are all implemented based on a two-layer non-linear network. Each layer of the non-linear network consists of a fully connected layer, an activation layer, and a residual connection module; all editors g i are used to fine-tune the input vector u i , the gradient δ i+1 of the current editor through a two-layer non-linear network and output the adjusted vector u new,i , the gradient δ new,i+1 .
[0196] The data receiving module 203 is used to receive the input-output transformation data set input by the user, denoted as the data set D set ; the data set D set includes N D groups of transformation data N D is a configurable positive integer, and the minimum can be set to 1; 1 ≤ data index j ≤ N D , x j is the model input, is the latest output label corresponding to the input x j .
[0197] The editor training module 204 is used to select a group of transformation data from the data set D set as the base sample; and set a specified number N for the base sampleK one related sample and one unrelated sample, and the training dataset D is composed of a base sample, N K equivalent samples and unrelated samples tr ; and based on the training dataset D tr , the target model trains all editors g i ; the specified quantity N K is a preset positive integer, by default less than 10; the base sample is denoted as the equivalent sample is denoted as (x k , y k ), 1 ≤ sample index k ≤ N K , the unrelated sample is denoted as (x un , y un ); the input x k of each equivalent sample is related to the input x s of the base sample, and the output label y k is related to the output label of the base sample; the input x un of the unrelated sample is irrelevant to the input x s of the base sample, and the output label y un is also irrelevant to the output label of the base sample.
[0198] The model parameter editing module 205 is used to, after the training of all editors is completed, based on the dataset D set and all editors g i perform parameter editing on all weight parameters w i of the target model.
[0199] A processing device for parameter editing of a pre-trained large model provided by an embodiment of the present invention can execute the method steps in the above method embodiment, and its implementation principle and technical effect are similar, which will not be elaborated here.
[0200] It should be noted that it should be understood that the division of each module of the above device is only a division of logical functions. In actual implementation, it can be fully or partially integrated into a physical entity, or physically separated. And these modules can all be implemented in the form of software called by a processing element; they can also all be implemented in the form of hardware; or some modules can be implemented in the form of software called by a processing element, and some modules can be implemented in the form of hardware. For example, the first preprocessing module can be a separately established processing element, or can be integrated in a certain chip of the above device. In addition, it can also be stored in the memory of the above device in the form of program code, and the function of the above determined module can be called and executed by a certain processing element of the above device. The implementation of other modules is similar. In addition, all or part of these modules can be integrated together or can be independently implemented. The processing element described here can be an integrated circuit with signal processing capabilities. In the implementation process, each step of the above method or each of the above modules can be completed by the integrated logic circuit in the processor element or the instruction in the form of software.
[0201] For example, the above modules can be one or more integrated circuits configured to implement the above method, such as: one or more Application Specific Integrated Circuits (ASICs), or one or more Digital Signal Processors (DSPs), or one or more Field Programmable Gate Arrays (FPGAs), etc. Again, when a certain module above is implemented in the form of a processing element scheduling program code, the processing element can be a general-purpose processor, such as a Central Processing Unit (CPU) or other processors that can call program code. Again, these modules can be integrated together and implemented in the form of a System-on-a-chip (SOC).
[0202] In the above embodiments, it can be implemented in whole or in part by software, hardware, firmware, or any combination thereof. When implemented using software, it can be implemented in whole or in part in the form of a computer program product. The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, the processes or functions described in the foregoing method embodiments are generated in whole or in part. The foregoing computer may be a general-purpose computer, a special-purpose computer, a computer network, or other programmable devices. The foregoing computer instructions may be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another. For example, the foregoing computer instructions may be transmitted from one website, computer, server, or data center to another website, computer, server, or data center by wire (such as coaxial cable, optical fiber, Digital Subscriber Line (DSL)) or wireless (such as infrared, wireless, Bluetooth, microwave, etc.). The foregoing computer-readable storage medium may be any available medium that can be accessed by a computer or a data storage device such as a server or data center that includes one or more integrated available media. The available media may be magnetic media (such as floppy disks, hard disks, magnetic tapes), optical media (such as DVDs), or semiconductor media (such as solid state disks (SSDs)), etc.
[0203] Figure 4 FIG. 4 is a schematic structural diagram of an electronic device provided in Embodiment 3 of the present invention. The electronic device may be a terminal device or a server for implementing the method of the foregoing embodiments, or may be a terminal device or a server connected to the foregoing terminal device or server for implementing the method of the foregoing embodiments. As Figure 4 shown, the electronic device may include: a processor 301 (such as a CPU), a memory 302, and a transceiver 303; the transceiver 303 is coupled to the processor 301, and the processor 301 controls the transceiver operations of the transceiver 303. Various instructions may be stored in the memory 302 for completing various processing functions and implementing the processing steps described in the foregoing method embodiments. Preferably, the electronic device according to the embodiments of the present invention further includes: a power supply 304, a system bus 305, and a communication port 306. The system bus 305 is used to implement communication connections between components. The foregoing communication port 306 is used for the electronic device to connect and communicate with other peripherals.
[0204] In Figure 4The system bus 305 mentioned above can be a Peripheral Component Interconnect (PCI) bus, an Extended Industry Standard Architecture (EISA) bus, or the like. The system bus can be divided into an address bus, a data bus, a control bus, etc. For the sake of simplicity, only a thick line is used to represent it in the figure, but it does not mean that there is only one bus or one type of bus. The communication interface is used to implement communication between the database access device and other devices (such as clients, read-write libraries, and read-only libraries). The memory may include Random Access Memory (RAM), and may also include non-volatile memory, such as at least one disk memory.
[0205] The above-mentioned processor can be a general-purpose processor, including a Central Processing Unit (CPU), a Network Processor (NP), a Graphics Processing Unit (GPU), etc.; it can also be a Digital Signal Processor (DSP), an Application Specific Integrated Circuit (ASIC), a Field Programmable Gate Array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components.
[0206] It should be noted that the embodiment of the present invention also provides a computer-readable storage medium, in which instructions are stored, and when they run on a computer, the computer is enabled to execute the methods and processing procedures provided in the above embodiments.
[0207] The embodiment of the present invention provides a processing method, device, electronic device, and computer-readable storage medium for parameter editing of a pre-trained large model. As can be seen from the above content, the embodiment of the present invention records the pre-trained large model that needs to be parameter edited as the target model, and configures a weight parameter editing model for each linear layer of the target model, denoted as the editor g i ; and after receiving the input-output transformation data set input by the user , select a set of transformation data from the current data set D set as the basic sample, and set a small number of related samples (x , y k ) and an unrelated sample (x k , y un ) for the basic sample, and form a training data set D un from all samples; and based on the training data set D tr and the target model for all editors g tr i Train; and after the training is completed, based on the dataset D set and all editors g i perform parameter editing on all the weight parameters w of the target model i , and the dataset D set may include at least one set of transformation data. The editor g given in the embodiments of the present invention i is a simple editing model implemented based on a two-layer non-linear network (each layer of the non-linear network consists of a fully connected layer, an activation layer, and a residual connection module), and it has a small number of model parameters, high training efficiency, and low training cost. The technical solution for performing parameter editing on the target model by using the trained editor g in the embodiments of the present invention i is a parameter editing solution based on reverse gradient calculation. This solution has a small amount of calculation, high editing flexibility, high editing efficiency, and low editing cost. When dealing with the parameter update task caused by a small number of input-output transformations (x, y → y * ), if the solution of the embodiments of the present invention is used, the technical effects of improving the update flexibility, improving the update efficiency, and reducing the update cost can be achieved.
[0208] The steps of the methods or algorithms described in combination with the embodiments disclosed herein can be implemented by hardware, software modules executed by a processor, or a combination of both. The software modules can be placed in a random access memory (RAM), memory, read-only memory (ROM), electrically programmable ROM, electrically erasable programmable ROM, registers, hard disk, removable disk, CD-ROM, or any other form of storage medium well-known in the technical field.
[0209] The specific embodiments described above further elaborate the purpose, technical solutions, and beneficial effects of the present invention. It should be understood that the above are only the specific embodiments of the present invention and are not used to limit the protection scope of the present invention. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present invention shall be included in the protection scope of the present invention.
Claims
1. A processing method for parameter editing of a pre-trained large model, characterized in that, The method includes: Denote the pre-trained large model that needs to perform parameter editing as the target model; the target model contains N w layers of linear layers, where N w is a positive integer greater than 1; denote the overall model parameters of the target model as θ, and the weight parameters and bias parameters of each linear layer as w i , b i , 1 ≤ layer index i ≤ N w , w i , b i ∈θ; during the forward prediction process of the target model, the input vectors of each linear layer are denoted as u i , the gradient of the model loss L with respect to the output vectors of each linear layer is denoted as δ i+1 , the gradient of the model loss L with respect to each weight parameter w i is denoted as The model loss L is defaultly the negative log-likelihood loss function L NLL as the loss function, L = L NLL (θ, y|x) = -log(P θ (y|x)), where x is the model input, y is the model output, and P θ (y|x) is the probability that the model output is y when the model parameters of the target model are the overall model parameters θ and the model input is x; Configure a weight parameter editing model for each linear layer of the target model, denoted as the editor g i ; All the editors g i have the same model structure and are all implemented based on a two-layer non-linear network. Each layer of the non-linear network consists of a fully connected layer, an activation layer, and a residual connection module; All the editors g i are used to fine-tune the vector u i and the gradient δ i+1 input to the current editor through a two-layer non-linear network and output the adjusted vector u new,i and the gradient δ new,i+1 ; The input-output transformation data set received from the user is denoted as data set D set ; The data set D set includes N D groups of transformation data N D is a configurable positive integer, and the minimum can be set to 1; 1 ≤ data index j ≤ N D , x j is the model input is the latest output label corresponding to the input x j ; Select a set of transformed data from the dataset D set as the base sample; and set a specified number N K of related samples and one unrelated sample for the base sample, and form a training dataset D K from the base sample, N tr of the equivalent samples and the unrelated sample; and based on the training dataset D tr , train all the editors g i using the target model; the specified number N K is a preset positive integer, defaulting to less than 10; the base sample is denoted as the equivalent samples are denoted as (x k , y k ), 1 ≤ sample index k ≤ N K , and the unrelated sample is denoted as (x un , y un ); the input x k of each of the equivalent samples is related to the input x s of the base sample, and the output label y k is related to the output label of the base sample; the input x un of the unrelated sample is unrelated to the input x s of the base sample, and the output label y un is also unrelated to the output label of the base sample; After the training of all editors is completed, based on the dataset D set and all the editors g i perform parameter editing on all the weight parameters w of the target model i 2. The method for processing parameter editing of a pre-trained large model according to claim 1, characterized in that: Each of the editors g i is composed of an input splicing module, a first fully-connected layer, a first activation layer, a first residual connection module, a second fully-connected layer, a second activation layer, a second residual connection module, and an output splitting module; The input end of the input splicing module is connected to the input end of the current editor, and the output end is respectively connected to the input end of the first fully connected layer and the first input end of the first residual connection module; the output end of the first fully connected layer is connected to the input end of the first activation layer; the output end of the first activation layer is connected to the second input end of the first residual connection module; The output end of the first residual connection module is respectively connected to the input end of the second fully connected layer and the first input end of the second residual connection module; the output end of the second fully connected layer is connected to the input end of the second activation layer; the output end of the second activation layer is connected to the second input end of the second residual connection module; the output end of the second residual connection module is connected to the input end of the output splitting module; the output end of the output splitting module is connected to the output end of the current editor; The input splicing module is used to receive the vector u input by the current editor i and the gradient δ i+1 ; and splice the vector u i and the gradient δ i+1 to obtain the corresponding initial vector z 0,i and send it to the first fully connected layer and the first residual connection module; The model parameters of the first fully connected layer include a scale vector parameter p 1,i , an offset vector parameter q 1,i , a low-rank matrix parameter A1, and a low-rank matrix parameter B1; the scale vector parameter p 1,i and the offset vector parameter q 1,i are private parameters of each of the editors g i ; the low-rank matrix parameters A1 and B1 are shared parameters of all the editors g i ; The first fully-connected layer is used to perform a fully-connected calculation on the initial vector z 1,i according to the scale vector parameter p 1,i and the offset vector parameter q 0,i and the low-rank matrix parameters A1 and B1 to obtain the corresponding hidden vector h 1,i and send it to the first activation layer; Among them, the hidden vector h 1,i is: h 1,i = p 1,i ⊙(A1B1z 0,i + b) + q 1,i ; b is a preset bias parameter, and the bias parameters b of the first fully connected layers of all the editors g i are the same; ⊙ is the Hadamard product operator; The first activation layer is implemented based on a class of non-linear activation functions σ1; the non-linear activation function σ1 includes at least the ReLU activation function; The first activation layer is used to perform an activation operation on the hidden vector h according to the non-linear activation function σ1 1,i to obtain a corresponding activation vector a 1,i and send it to the first residual connection module; Among them, the activation vector a 1,i is: a 1,i = σ1(h 1,i ); The first residual connection module is used to perform element-wise addition on the initial vector z 0,i and the activation vector a 1,i to obtain the corresponding first vector z 1,i and send it to the second fully-connected layer; wherein, the first vector z 1,i is: z 1,i = z 0,i + a 1,i ; The model parameters of the second fully connected layer include a scale vector parameter p 2,i , an offset vector parameter q 2,i , a low-rank matrix parameter A2, and a low-rank matrix parameter B2; the scale vector parameter p 2,i and the offset vector parameter q 2,i are private parameters of each of the editors g i ; the low-rank matrix parameters A2 and B2 are shared parameters of all the editors g i ; The second fully connected layer is used to perform a fully connected calculation on the first vector z 2,i according to the scale vector parameter p 2,i and the offset vector parameter q 1,i and the low-rank matrix parameters A2 and B2 to obtain the corresponding hidden vector h 2,i and send it to the second activation layer; Among them, the hidden vector h 2,i is: h 2,i = p 2,i ⊙(A2B2z 1,i ) + q 2,i ; The second activation layer is implemented based on a class of non-linear activation functions σ2; the non-linear activation function σ2 includes at least the ReLU activation function; The second activation layer is used to perform an activation operation on the hidden vector h according to the non-linear activation function σ2 2,i to obtain a corresponding activation vector a 2,i and send it to the second residual connection module; Among them, the activation vector a 2,i is: a 2,i = σ2(h 2,i ); The second residual connection module is used to perform element-wise addition on the first vector z 1,i and the activation vector a 2,i to obtain a corresponding second vector z 2,i and send it to the output splitting module; Among them, the second vector z 2,i is: z 2,i = z 1,i + a 2,i ; The output splitting module is used to use the initial vector z 0,i the vector u in i and the gradient δ i+1 as a reference for the splicing position, and extract the corresponding vector u 2,i from the second vector z new,i and the gradient δ new,i+1 and output them.
3. The processing method for parameter editing of a pre-trained large model according to claim 1, wherein, Based on the training dataset D tr and the target model trains all the editors g i The specific process is as follows: Step 31, from the training data set D tr Extract the corresponding basic sample N K The equivalent samples (x k ,y k ) and the irrelevant samples (x un ,y un ); and make a parameter backup of the current overall model parameter θ of the target model as the corresponding backup parameter θ0; and use the current overall model parameter θ as the corresponding current model parameter θ now ; and by all said editors g i The editor parameters of the corresponding editor parameter set φ are formed; and all parameters of the editor parameter set φ are initialized to obtain the current parameter set φ now ; and set a first counter initialized to 0; Among them, the editor parameter set φ consists of all the editors g i private scale vector parameter p 1,i / p 2,i , offset vector parameter q 1,i / q 2,i and shared low-rank matrix parameters A1 / A2, B1 / B2; Step 32, based on the current parameter set φ now set the editor parameters for all the editors g i ; Step 33, input the basic sample of x s into the target model for forward prediction processing, cache each of the vectors u i generated during the processing, and calculate the corresponding model loss Among them, is the probability that the output of the target model is now when its model parameters are the current model parameters θ s and the model input is x ; Step 34, according to the current model parameters θ now for each of the weight parameters w i and its corresponding bias parameter b i , the vector u i calculate the corresponding output vector u of the linear layer i+1 ; and calculate the gradient of the model loss L with respect to each vector u i+1 to obtain the corresponding gradient δ i+1 ; and input each of the gradients δ i+1 and its corresponding vector u i into the corresponding editor g i for fine-tuning to obtain the corresponding vector u new,i and the gradient δ new,i+1 ; and based on each of the vectors u new,i and its corresponding gradient δ new,i+1 calculate the new weight gradient and based on each of the weight parameters w i and its corresponding weight gradient calculate the corresponding weight parameter and reset each of the weight parameters w now in the current model parameters θ i to the corresponding weight parameter Wherein, The vector u i+1 is calculated as follows: u i+1 = w i u i + b i ; The gradient δ i+1 is calculated as follows: The vector u new,i and the gradient δ new,i+1 are calculated as follows: g i (u i ,δ i+1 ) → (u new,i ,δ new,i+1 ), g i (u i ,δ i+1 )(u, δ) is the expression of the editor g i ; The weight gradient is calculated as follows: T is the transpose symbol; The weight parameter is calculated as follows: Step 35, based on the latest current model parameter θ now reset the model parameters of the target model; and input the x of each of the equivalent samples (x k , y k ) into the target model for forward prediction processing, and calculate the corresponding first model loss L k ; 1,k ; Among them, the first model loss L 1,k defaults to the negative log-likelihood loss function L NLL as the loss function, specifically: is the probability that the model output corresponding to y for the target model when its model parameters are the latest current model parameters θ now and the model input is x k ; k Step 36: convert the irrelevant samples (x un ,y un ) un Input the target model for forward prediction processing and calculate the corresponding edited probability Then, the model parameters of the target model are reset based on the backup parameter θ0; and the irrelevant samples (x un ,y un ) un Input the target model for forward prediction processing and calculate the corresponding pre-edit probability And based on the pre-editing probability r pre , edited probability r post And the divergence loss function L KL Calculate the corresponding second model loss L2; Among them, is the probability that the model output of the target model is y when its model parameters are the latest current model parameters θ now and the model input is x un ; un Probability; is the probability that, when the model parameters of the target model are the backup parameters θ0 and the model input is x un the model output is y un ; The second model loss L2 defaults to the divergence loss function L KL as the loss function, specifically: Step 37, from N K of the first model losses L 1,k and the second model loss L2 to form the corresponding editor loss L G ; wherein, the editor loss L G is: λ is a preset balance parameter; Step 38, towards minimizing the loss L of the editor G in the direction of minimizing the loss, based on a preset first model optimizer, optimize all the editors g i for the corresponding current parameter set φ now to perform one round of parameter modulation; and use the modulated editor parameter set as the new current parameter set φ now ; Wherein, the first model optimizer includes at least the Adam optimizer; Step 39: increment the first counter by 1; identify whether the incremented first counter exceeds a preset counter threshold; if not, return to Step 32 to continue training; if so, stop training and based on the latest current parameter set φ now for all the editors g i solidify the editor parameters and confirm that the training of all editors is completed.
4. The method for processing parameter editing of a pre-trained large model according to claim 1, wherein Based on the dataset D set and all the editors g i perform parameter editing on all the weight parameters w of the target model i Specifically, it includes: Set the model parameters of the target model based on the backup parameter θ0; and for the dataset D set All the transformed data Perform a round of traversal; and during this round of traversal, take the currently traversed transformed data As the corresponding current transformed data (x c , y c ); and input the x of the current transformed data (x c , y c ) into the target model for forward prediction processing, cache each vector u c generated during the processing, and calculate the corresponding model loss i And calculate the corresponding linear layer output vector u based on each weight parameter w in the backup parameter θ0 i and its corresponding bias parameter b i , the vector u i ; and calculate the gradient of the model loss L with respect to each vector u i+1 to obtain the corresponding gradient δ i+1 ; and input each gradient δ i+1 and its corresponding vector u i+1 into the corresponding editor g i for fine-tuning to obtain the corresponding vector u i and the gradient δ new,i ; and calculate the new weight gradient based on each vector u new,i+1 and its corresponding gradient δ new,i new,i+1 i ; and calculate the corresponding weight parameter based on each weight parameter w and its corresponding weight gradient i ; and reset each weight parameter w i in the current backup parameter θ0 to the corresponding weight parameter i ; and perform a parameter reset on the model parameters of the target model based on the latest backup parameter θ0 once; Among them, the vector u i+1 is calculated as follows: u i+1 = w i u i + b i ; The gradient δ i+1 is calculated as follows: The vector u new,i and the gradient δ new,i+1 are calculated as follows: g i (u i ,δ i+1 )→(u new,i ,δ new,i+1 ), g i (u i ,δ i+1 ) is the expression of the editor g i ; The weight gradient is calculated as follows: T is the transpose symbol; The weight parameter is calculated as follows:
5. An apparatus for performing the processing method of parameter editing on a pre-trained large model according to any one of claims 1-4, characterized in that, The device includes: a first preprocessing module, a second preprocessing module, a data receiving module, an editor training module, and a model parameter editing module; The first preprocessing module is used to record the pre-trained large model that needs to perform parameter editing as the target model; the target model includes N w linear layers, where N w is a positive integer greater than 1; the overall model parameters of the target model are denoted as θ, and the weight parameters and bias parameters of each linear layer are denoted as w i , b i , 1 ≤ layer index i ≤ N w , w i , b i ∈θ; during the forward prediction process of the target model, the input vectors of each linear layer are denoted as u i , the gradients of the model loss L with respect to the output vectors of each linear layer are denoted as δ i+1 , the gradients of the model loss L with respect to each weight parameter w i are denoted as The model loss L is defaultly the negative log-likelihood loss function L NLL as the loss function, L = L NLL (θ, y|x) = -log(P θ (y|x)), where x is the model input, y is the model output, and P θ (y|x) is the probability that the model output is y when the model parameters of the target model are the overall model parameters θ and the model input is x; The second preprocessing module is used to configure a weight parameter editing model for each linear layer of the target model, denoted as the editor g i ; all the editors g i have the same model structure, which is implemented based on a two-layer non-linear network. Each layer of the non-linear network consists of a fully connected layer, an activation layer, and a residual connection module; all the editors g i are used to fine-tune the vector u i and the gradient δ i+1 input to the current editor through a two-layer non-linear network and output the adjusted vector u new,i and the gradient δ new,i+1 ; The data receiving module is used to receive the input-output transformation data set input by the user, denoted as data set D set ; the data set D set includes N D groups of transformation data N D is a configurable positive integer, with the minimum value set to 1; 1 ≤ data index j ≤ N D , x j is the model input is the latest output label corresponding to the input x j ; The editor training module is used to select a set of transformed data from the dataset D set as a base sample; and set a specified number N K of relevant samples and an irrelevant sample for the base sample, and the training dataset D is composed of the base sample, N K equivalent samples and the irrelevant sample tr ; and based on the training dataset D tr , the target model trains all the editors g i ; the specified number N K is a preset positive integer, default less than 10; the base sample is denoted as The equivalent sample is denoted as (x k , y k ), 1 ≤ sample index k ≤ N K , and the irrelevant sample is denoted as (x un , y un ); the input x k of each equivalent sample is related to the input x s of the base sample, and the output label y k is related to the output label of the base sample; the input x un of the irrelevant sample is irrelevant to the input x s of the base sample, and the output label y un is also irrelevant to the output label of the base sample; The model parameter editing module is used to, after the training of all editors is completed, based on the dataset D set and all the editors g i perform parameter editing on all the weight parameters w i of the target model.
6. An electronic device, characterized in that, It includes: A memory, a processor, and a transceiver; The processor is used to be coupled with the memory, read and execute the instructions in the memory to implement the method according to any one of claims 1-4; The transceiver is coupled with the processor, and the processor controls the transceiver to perform message sending and receiving.
7. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores computer instructions, and when the computer instructions are executed by a computer, the computer is caused to execute the method according to any one of claims 1-4.
Citation Information
Patent Citations
Face adversarial attack sample generation method and system based on attribute editing
CN115775406A
Editing method and device for pre-training language model
CN117851613A
Method and device for training text embedding module of large language model
CN118504714A
Intelligent tile transformation recommendation model training method, system, medium and program
CN118839387A
Parameter synchronization method and device for large model, equipment and medium
CN119211260A