Method and apparatus for using noise perturbation to train pre-trained language model
By adding noise perturbation to the pretrained language model and conducting multi-stage training, the problem of low overfitting and generalization capabilities of large-scale language models is solved, and more efficient model optimization and generalization capabilities are achieved.
Patent Information
- Application Number
- PCT/CN2024/079410
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2024-02-29
- Publication Date
- 2025-09-04
AI Technical Summary
In the prior art, after further training of pre-trained large-scale language models, the model is prone to overfitting and low generalization capabilities.
By calculating the noise perturbation corresponding to each parameter matrix in the pre-trained language model, and updating the parameter matrix according to the noise perturbation, combining the training data set to optimize the bias term and the updated parameter matrix, a multi-stage training method is used to optimize the model.
It effectively avoids model overfitting, improves the generalization ability of the model, and improves training efficiency and model accuracy.
Smart Images

Figure CN2024079410_04092025_PF_FP_ABST
Abstract
Description
Method and device for training pre-trained language model using noise perturbation
[0001] This application is based on the Chinese patent application with application number 2023106147795 and application date May 29, 2023, and claims the priority of the Chinese patent application. The entire content of the Chinese patent application is hereby introduced into this application as a reference. Technical Field
[0002] The present application relates to the field of machine learning technology, and in particular to a method and device for training a pre-trained language model using noise perturbation. Background Art
[0003] In recent years, with the advancement of machine learning technology, an increasing number of large-scale models have been applied to the language field. To ensure that these large-scale models meet requirements and improve training efficiency, it is common to further train pre-trained large-scale models. However, this type of further training often results in models that suffer from overfitting and poor generalization.
[0004] Summary of the Invention
[0005] In view of this, the embodiments of the present application provide a method, device, electronic device and computer-readable storage medium for training a pre-trained language model using noise perturbation to solve the problem in the prior art that further training of a pre-trained large-scale model often results in the final model suffering from overfitting and low generalization ability.
[0006] In a first aspect of an embodiment of the present application, a method for training a pre-trained language model using noise perturbation is provided, comprising: obtaining a training dataset and a pre-trained language model corresponding to a target task; calculating the noise perturbation corresponding to each parameter matrix in the pre-trained language model, and updating the parameter matrix according to the noise perturbation corresponding to each parameter matrix; and optimizing the bias term and the updated parameter matrix in the pre-trained language model based on the target task using the training dataset.
[0007] According to a second aspect of an embodiment of the present application, a device for training a pre-trained language model using noise perturbation is provided, comprising: an acquisition module configured to acquire a training data set and a pre-trained language model corresponding to a target task; a calculation module configured to calculate the noise perturbation corresponding to each parameter matrix in the pre-trained language model, and update the parameter matrix according to the noise perturbation corresponding to each parameter matrix; and a training module configured to optimize the bias term and the updated parameter matrix in the pre-trained language model using the training data set based on the target task.
[0008] According to a third aspect of an embodiment of the present application, an electronic device is provided, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor implements the steps of the above method when executing the computer program.
[0009] According to a fourth aspect of the embodiments of the present application, a computer-readable storage medium is provided, which stores a computer program. When the computer program is executed by a processor, the steps of the above method are implemented.
[0010] The beneficial effects of the embodiments of the present application compared with the prior art are: because the embodiments of the present application obtain a training data set and a pre-trained language model corresponding to the target task; calculate the noise perturbation corresponding to each parameter matrix in the pre-trained language model, and update the parameter matrix according to the noise perturbation corresponding to each parameter matrix; based on the target task, the training data set is used to optimize the bias term and the updated parameter matrix in the pre-trained language model. Therefore, the above-mentioned technical means can solve the problem in the prior art that further training of a pre-trained large-scale model often results in the final model having overfitting and low generalization ability, thereby avoiding model overfitting and improving the model generalization ability. BRIEF DESCRIPTION OF THE DRAWINGS
[0011] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the following briefly introduces the drawings required for use in the embodiments or descriptions of the prior art. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without any creative work.
[0012] FIG1 is a flow chart of a method for training a pre-trained language model using noise perturbation provided in an embodiment of the present application;
[0013] FIG2 is a flow chart of another method for training a pre-trained language model using noise perturbation provided in an embodiment of the present application;
[0014] FIG3 is a schematic diagram of the structure of a pre-trained language model provided in an embodiment of the present application;
[0015] FIG4 is a schematic diagram of the structure of an apparatus for training a pre-trained language model using noise perturbation according to an embodiment of the present application;
[0016] FIG5 is a schematic structural diagram of an electronic device provided in an embodiment of the present application. DETAILED DESCRIPTION
[0017] In the following description, specific details such as specific system structures and techniques are provided for purposes of illustration rather than limitation to facilitate a thorough understanding of the embodiments of the present application. However, it will be apparent to those skilled in the art that the present application may be implemented in other embodiments without these specific details. In other cases, detailed descriptions of well-known systems, devices, circuits, and methods are omitted to avoid obscuring the description of the present application with unnecessary detail.
[0018] FIG1 is a flow chart of a method for training a pre-trained language model using noise perturbation provided by an embodiment of the present application. The method for training a pre-trained language model using noise perturbation in FIG1 can be executed by a computer or server, or software on a computer or server. As shown in FIG1 , the method for training a pre-trained language model using noise perturbation includes:
[0019] S101, obtaining a training dataset and a pre-trained language model corresponding to the target task;
[0020] S102, calculating the noise perturbation corresponding to each parameter matrix in the pre-trained language model, and updating the parameter matrix according to the noise perturbation corresponding to each parameter matrix;
[0021] S103, based on the target task, using the training dataset to optimize the bias term and the updated parameter matrix in the pre-trained language model.
[0022] Pretrained language models contain a large number of bias terms and parameter matrices. The optimization of the parameter matrices in pretrained language models described later in this article refers to the optimization of the updated parameter matrices in the pretrained language model.
[0023] The bias term is the bias unit, bias term, or intercept term, and has the same meaning as b in the linear equation y=wx+b. In the linear equation y=wx+b, b represents the intercept of the function on the y-axis, which controls the distance the function deviates from the origin. The neural network model (the pre-trained language model is a pre-trained model, and the pre-trained model is a neural network model that has been pre-trained) can also be represented by y=Wx+b. Unlike the linear equation, W and b in the neural network model represent matrices, and the trainable parameters of the neural network model can also be expressed as: (W, b), where W represents the parameter matrix and b represents the bias term. The parameters of the neural network model are divided into fixed parameters and trainable parameters. The trainable parameters include: parameter matrix and bias term. Training the neural network model is the process of optimizing the trainable parameters.
[0024] This application can be used in any scenario in the field of language, such as text translation, word order prediction, next sentence prediction, question-answering tasks, named entity recognition tasks, text classification, etc. For example, in the text translation scenario, the target task is the text translation task; the training dataset is the annotated corpus for text translation; the pre-trained language model is a model obtained by pre-training the language model based on the text translation task; based on the text translation task, the training dataset is used to optimize the parameter matrix and bias terms in the pre-trained language model; and the final trained model is used for text translation. Other scenarios are similar to the text translation scenario.
[0025] According to the technical solution provided in the embodiment of the present application, a training data set and a pre-trained language model corresponding to the target task are obtained; the noise perturbation corresponding to each parameter matrix in the pre-trained language model is calculated, and the parameter matrix is updated according to the noise perturbation corresponding to each parameter matrix; based on the target task, the bias term and the updated parameter matrix in the pre-trained language model are optimized using the training data set. The embodiment of the present application reduces the influence of pre-training on the overfitting and generalization ability of the language model by adding noise perturbation to the parameter matrix. Therefore, the above-mentioned technical means can solve the problem in the prior art that further training of a pre-trained large-scale model often results in overfitting and low generalization ability of the resulting model, thereby avoiding model overfitting and improving the generalization ability of the model.
[0026] Furthermore, the noise perturbation corresponding to each parameter matrix in the pre-trained language model is calculated by the following formula: W' i =U(-λ / 2,λ / 2)*std(W i );
[0027] Among them, W' i is the noise perturbation corresponding to the i-th parameter matrix, U(-λ / 2,λ / 2) is the uniformly distributed noise in the range from -λ / 2 to λ / 2, λ is the hyperparameter controlling the noise intensity in the pre-trained language model, std(W i ) is the standard deviation of the data within the i-th parameter matrix.
[0028] Furthermore, each parameter matrix is updated by the following formula:
[0029] W” i =W i +W' i ;
[0030] Among them, W' i is the noise perturbation corresponding to the i-th parameter matrix, W i is the i-th parameter matrix before update, W” i is the updated i-th parameter matrix.
[0031] W'i If and W i The dimensions of W' i Fill so that W' i and W i The dimensions are consistent.
[0032] Furthermore, based on the target task, the training data set is used to optimize the bias items and the updated parameter matrix in the pre-trained language model, including: dividing the training data set into a first training data set and a second training data set according to a first preset ratio, and performing multi-stage training on the pre-trained language model: freezing the parameter matrix in the pre-trained language model, and based on the target task, using the first training data set to optimize the bias items in the pre-trained language model to complete the first stage training of the pre-trained language model; after completing the first stage training, unfreezing the parameter matrix in the pre-trained language model, and based on the target task, using the second training data set to optimize the bias items and parameter matrix in the pre-trained language model to complete the second stage training of the pre-trained language model.
[0033] The first preset ratio is related to the ratio of the parameter matrix to the trainable parameters in the pre-trained language model, and the ratio of the bias term to the trainable parameters. For example, the first preset ratio can be 1:9, and the ratio of the data volume of the first training dataset to the second training dataset is 1:9.
[0034] In this embodiment, the first stage of training is to freeze the parameter matrix and train only the bias item using the first training data set; after the first stage of training is completed, the parameter matrix is unfrozen; the second stage of training is to train the bias item and parameter matrix using the second training data set. The second stage of training is to train the pre-trained language model as a whole.
[0035] Furthermore, based on the target task, the training data set is used to optimize the bias items and the updated parameter matrix in the pre-trained language model, including: dividing the training data set into a first training data set and a second training data set according to a third preset ratio, and performing multi-stage training on the pre-trained language model: freezing the bias items in the pre-trained language model, and based on the target task, using the first training data set to optimize the parameter matrix in the pre-trained language model to complete the first stage training of the pre-trained language model; after completing the first stage training, unfreezing the bias items in the pre-trained language model, and based on the target task, using the second training data set to optimize the bias items and parameter matrix in the pre-trained language model to complete the second stage training of the pre-trained language model.
[0036] In this embodiment, the first stage of training is to freeze the bias item and train only the parameter matrix using the first training data set. After the first stage of training is completed, the bias item is unfrozen. The second stage of training is to train the bias item and the parameter matrix using the second training data set. The second stage of training is to train the pre-trained language model as a whole.
[0037] Furthermore, based on the target task, the bias items and the updated parameter matrix in the pre-trained language model are optimized using the training data set, including: determining the data amount of the training data set; training the pre-trained language model according to the data amount: when the data amount is less than a first preset size, freezing the updated parameter matrix in the pre-trained language model, and based on the target task, optimizing the bias items in the pre-trained language model using the training data set; when the data amount is not less than the first preset size, optimizing the bias items and the updated parameter matrix in the pre-trained language model using the training data set based on the target task.
[0038] The parameter matrix accounts for more than 99% of the trainable parameters in the pre-trained language model, and the bias term accounts for less than 1%. In the embodiment of the present application, when the amount of data is less than the first preset size, the training data set is used to optimize only the bias term in the pre-trained language model (this method is applied to small sample scenarios, and the small sample scenario is a situation where the number of training samples is small). This can greatly reduce the amount of optimized parameters and training time, and at the same time avoid model overfitting when the number of training samples is small. It has been found in practice that only optimizing the bias term in the pre-trained language model can also achieve good results.
[0039] Furthermore, based on the target task, the bias items and the updated parameter matrix in the pre-trained language model are optimized using the training data set, including: determining the data amount of the training data set; training the pre-trained language model according to the data amount: when the data amount is less than a first preset size, freezing the updated parameter matrix in the pre-trained language model, and based on the target task, optimizing the bias items in the pre-trained language model using the training data set; when the data amount is greater than or equal to the first preset size but less than a second preset size, freezing the bias items in the pre-trained language model, and based on the target task, optimizing the updated parameter matrix in the pre-trained language model using the training data set; when the data amount is greater than or equal to the second preset size, optimizing the bias items and the updated parameter matrix in the pre-trained language model using the training data set based on the target task.
[0040] The embodiment of the present application selects a corresponding training method according to the data volume of the training data set to improve the efficiency of training.
[0041] Furthermore, before obtaining the pre-trained language model corresponding to the target task, the method also includes: sequentially connecting multiple linear layers and nonlinear activation functions to obtain a feedforward layer; sequentially connecting an embedding layer, a multi-head attention network, a residual layer, a normalization layer, a feedforward layer, a residual layer, a normalization layer, a fully connected layer and a classification layer to obtain a language model; and pre-training the language model based on the target task to obtain a pre-trained language model.
[0042] Multiple linear layers are serially connected and then connected with a nonlinear activation function as a feedforward layer. The residual layer after the multi-head attention network is used to add the output of the multi-head attention network to the input of the multi-head attention network; the residual layer after the feedforward layer is used to add the output of the feedforward layer to the input of the feedforward layer.
[0043] FIG2 is a flow chart of another method for training a pre-trained language model using noise perturbation provided by an embodiment of the present application. As shown in FIG2 , the method includes:
[0044] S201: Divide the training data set into a first training data set, a second training data set, and a third training data set according to a second preset ratio, and perform multi-stage training on the pre-trained language model:
[0045] S202, freezing the parameter matrix in the pre-trained language model, and optimizing the bias term in the pre-trained language model using the first training dataset based on the target task to complete the first stage training of the pre-trained language model;
[0046] S203, after completing the first stage of training, unfreeze the parameter matrix in the pre-trained language model, freeze the bias term in the pre-trained language model, and optimize the parameter matrix in the pre-trained language model using the second training dataset based on the target task to complete the second stage of training the pre-trained language model;
[0047] S204, after completing the second stage of training, unfreeze the bias items in the pre-trained language model, and based on the target task, use the third training data set to optimize the bias items and parameter matrix in the pre-trained language model to complete the third stage of training of the pre-trained language model.
[0048] The second preset ratio is related to the ratio of the parameter matrix to the trainable parameters in the pre-trained language model, as well as the ratio of the bias term to the trainable parameters. For example, the first preset ratio may be 1:6:3, where the ratio of the data size of the first training dataset, the second training dataset, and the third training dataset is 1:6:3.
[0049] The first stage of training: freeze the parameter matrix and use the first training data set to train only the bias item; after the first stage of training is completed, unfreeze the parameter matrix; the second stage of training: freeze the bias item and use the second training data set to train the parameter matrix; after the second stage of training is completed, unfreeze the bias item; the third stage of training: use the third training data set to train the parameter matrix and bias item. The third stage of training is the training of the pre-trained language model as a whole.
[0050] The embodiment of the present application can greatly improve the accuracy of the final model through multi-stage training of the pre-trained language model.
[0051] In an optional embodiment, multiple linear layers and nonlinear activation functions are connected in sequence to obtain a feedforward layer; an embedding layer, a multi-head attention network, a residual layer, a normalization layer, a feedforward layer, a residual layer, a normalization layer, a fully connected layer, and a classification layer are connected in sequence to obtain a language network, and multiple language networks are serially connected to obtain a language model; a training data set corresponding to a target task is obtained; a noise perturbation corresponding to each parameter matrix in the language model is calculated, and the parameter matrix is updated according to the noise perturbation corresponding to each parameter matrix; based on the target task, the training data set is used to optimize the bias term and the updated parameter matrix in the language model.
[0052] This embodiment does not pre-train the language model, but directly performs formal training on the language model. The above technical means can solve the problems of overfitting and low generalization ability of the trained model in the existing technology, thereby avoiding model overfitting and improving model generalization ability.
[0053] Figure 3 is a schematic diagram of the structure of a language model provided by an embodiment of the present application. As shown in Figure 3, the language model includes, from the input end to the output end, an embedding layer, a multi-head attention network, a residual layer, a normalization layer, a feedforward layer, a residual layer, a normalization layer, a fully connected layer, and a classification layer.
[0054] The residual layer after the feedforward layer is used to add the output of the feedforward layer to the input of the feedforward layer and output the result of the addition; the residual layer after the multi-head attention network is used to add the output of the multi-head attention network to the input of the multi-head attention network and output the result of the addition.
[0055] Figure 3 is also a schematic diagram of the structure of the pre-trained language model. The pre-trained language model is a pre-trained language model.
[0056] The language model can also be a BERT model, an XLNET model, a RoBERTa model, or an ELECTRA model. During model training, the optimizer used can be the Adam optimizer, the AdamW optimizer, the AdaGrad optimizer, or the RMSProp optimizer.
[0057] In an optional embodiment, a training data set and a pre-trained language model corresponding to a target task are obtained; the noise perturbation corresponding to the network parameters of each network layer in the pre-trained language model is calculated, and the network parameters are updated according to the noise perturbation corresponding to each network parameter; based on the target task, the updated network parameters in the pre-trained language model are optimized using the training data set to complete the training of the pre-trained language model.
[0058] Furthermore, the noise perturbation corresponding to each network parameter in the pre-trained language model is calculated by the following formula: W' i =U(-λ / 2,λ / 2)*std(W i );
[0059] Among them, W' i is the noise perturbation corresponding to the i-th network parameter, U(-λ / 2,λ / 2) is the uniformly distributed noise in the range from -λ / 2 to λ / 2, λ is the hyperparameter controlling the noise intensity in the pre-trained language model, std(W i ) is the standard deviation of the internal data of the i-th network parameter.
[0060] Furthermore, each network parameter is updated by the following formula: W” i =W i +W' i ;
[0061] Among them, W' i is the noise perturbation corresponding to the i-th network parameter, W i is the i-th network parameter before update, W” i is the updated i-th network parameter.
[0062] W' i If and W i The dimensions of W' i Fill so that W' i and W i The dimensions are consistent.
[0063] In an optional embodiment, multiple linear layers and nonlinear activation functions are connected in sequence to obtain a feedforward layer; an embedding layer, a multi-head attention network, a residual layer, a normalization layer, a feedforward layer, a residual layer, a normalization layer, a fully connected layer, and a classification layer are connected in sequence to obtain a language network, and multiple language networks are serially connected to obtain a language model; a training data set corresponding to a target task is obtained; a noise perturbation corresponding to each network parameter in the language model is calculated, and the network parameter is updated according to the noise perturbation corresponding to each network parameter; based on the target task, the updated network parameters in the language model are optimized using the training data set.
[0064] In an optional embodiment, a training data set and a pre-trained language model corresponding to a target task are obtained; the noise perturbation corresponding to the network parameters of each network layer in the pre-trained language model is calculated, and the network parameters are updated according to the noise perturbation corresponding to each network parameter; based on the target task, the updated network parameters in the pre-trained language model are optimized using the training data set to complete the training of the pre-trained language model.
[0065] In an optional embodiment, a training data set and a pre-trained language model corresponding to a target task are obtained; first network parameters and second network parameters corresponding to the bias term and parameter matrix in the pre-trained language model are determined respectively; a noise perturbation corresponding to each second network parameter in the pre-trained language model is calculated, and the second network parameter is updated according to the noise perturbation corresponding to each parameter matrix; based on the target task, the first network parameters and the updated second network parameters in the pre-trained language model are optimized using the training data set.
[0066] In an optional embodiment, based on the target task, the bias term and the updated parameter matrix in the language model are optimized using a training data set, including: obtaining a trained target language model; inputting multiple training samples in the training data set into the language model and the target language model, and outputting the first processing result and the second processing result corresponding to each training sample respectively; calculating the contrast loss using a triplet loss function according to the first processing result and the second processing result corresponding to each training sample and the second processing result corresponding to another training sample with different semantics from the training sample; calculating the classification loss using a cross-entropy loss function according to the first processing result and the label corresponding to each training sample; and updating the network parameters of the language model based on the contrast loss and the classification loss to complete the training of the language model.
[0067] The triplet loss function is triplet(). The first processing result and the second processing result corresponding to a certain training sample are A1 and A2 respectively, and the second processing result corresponding to another training sample with different semantics from the training sample is A3 (the other training sample with different semantics from the training sample is randomly determined in the training data set). The loss value corresponding to the first language corpus is equal to triplet (A1, A2, A3), and the loss values corresponding to all training samples are added together to form the contrast loss. The contrast loss and the classification loss are weighted and summed according to the preset weights to form the total loss, and the model parameters of the language model are updated according to the total loss. The embodiment of the present application can solve the problem of overfitting of the translation model in the prior art by introducing the contrast loss into the model training, thereby improving the generalization performance of the model.
[0068] All of the above optional technical solutions can be combined in any way to form optional embodiments of the present application, and will not be described in detail here.
[0069] The following are device embodiments of the present application, which can be used to implement the method embodiments of the present application. For details not disclosed in the device embodiments of the present application, please refer to the method embodiments of the present application.
[0070] FIG4 is a schematic diagram of an apparatus for training a pre-trained language model using noise perturbation provided by an embodiment of the present application. As shown in FIG4 , the apparatus for training a pre-trained language model using noise perturbation includes:
[0071] Acquisition module 401 is configured to acquire a training dataset and a pre-trained language model corresponding to a target task;
[0072] A calculation module 402 is configured to calculate the noise perturbation corresponding to each parameter matrix in the pre-trained language model, and update the parameter matrix according to the noise perturbation corresponding to each parameter matrix;
[0073] The training module 403 is configured to optimize the bias term and the updated parameter matrix in the pre-trained language model using the training dataset based on the target task.
[0074] Pretrained language models contain a large number of bias terms and parameter matrices. The optimization of the parameter matrices in pretrained language models described later in this article refers to the optimization of the updated parameter matrices in the pretrained language model.
[0075] The bias term is the bias unit, bias term, or intercept term, and has the same meaning as b in the linear equation y=wx+b. In the linear equation y=wx+b, b represents the intercept of the function on the y-axis, which controls the distance the function deviates from the origin. The neural network model (the pre-trained language model is a pre-trained model, and the pre-trained model is a neural network model that has been pre-trained) can also be represented by y=Wx+b. Unlike the linear equation, the W in the neural network model represents a matrix, and the trainable parameters of the neural network model can also be expressed as: (W, b), where W represents the parameter matrix and b represents the bias term. The parameters of the neural network model are divided into fixed parameters and trainable parameters. The trainable parameters include: parameter matrix and bias term. Training the neural network model is the process of optimizing the trainable parameters.
[0076] This application can be used in any scenario in the field of language, such as text translation, word order prediction, next sentence prediction, question-answering tasks, named entity recognition tasks, text classification, etc. For example, in the text translation scenario, the target task is the text translation task; the training dataset is the annotated corpus for text translation; the pre-trained language model is a model obtained by pre-training the language model based on the text translation task; based on the text translation task, the training dataset is used to optimize the parameter matrix and bias terms in the pre-trained language model; and the final trained model is used for text translation. Other scenarios are similar to the text translation scenario.
[0077] According to the technical solution provided in the embodiment of the present application, a training data set and a pre-trained language model corresponding to the target task are obtained; the noise perturbation corresponding to each parameter matrix in the pre-trained language model is calculated, and the parameter matrix is updated according to the noise perturbation corresponding to each parameter matrix; based on the target task, the bias term and the updated parameter matrix in the pre-trained language model are optimized using the training data set. The embodiment of the present application reduces the influence of pre-training on the overfitting and generalization ability of the language model by adding noise perturbation to the parameter matrix. Therefore, the above-mentioned technical means can solve the problem in the prior art that further training of a pre-trained large-scale model often results in overfitting and low generalization ability of the resulting model, thereby avoiding model overfitting and improving the generalization ability of the model.
[0078] Optionally, the calculation module 402 is further configured to calculate the noise disturbance corresponding to each parameter matrix in the pre-trained language model by the following formula: i =U(-λ / 2,λ / 2)*std(W i );
[0079] Among them, W' i is the noise perturbation corresponding to the i-th parameter matrix, U(-λ / 2,λ / 2) is the uniformly distributed noise in the range from -λ / 2 to λ / 2, λ is the hyperparameter controlling the noise intensity in the pre-trained language model, std(W i ) is the standard deviation of the data within the i-th parameter matrix.
[0080] Optionally, the calculation module 402 is further configured to update each parameter matrix by the following formula: i =W i +W' i ;
[0081] Among them, W' i is the noise perturbation corresponding to the i-th parameter matrix, W i is the i-th parameter matrix before update, W” i is the updated i-th parameter matrix.
[0082] W'i If and W i The dimensions of W' i Fill so that W' i and W i The dimensions are consistent.
[0083] Optionally, the training module 403 is further configured to divide the training data set into a first training data set and a second training data set according to a first preset ratio, and perform multi-stage training on the pre-trained language model: freeze the parameter matrix in the pre-trained language model, and based on the target task, use the first training data set to optimize the bias item in the pre-trained language model to complete the first stage training of the pre-trained language model; after completing the first stage training, unfreeze the parameter matrix in the pre-trained language model, and based on the target task, use the second training data set to optimize the bias item and parameter matrix in the pre-trained language model to complete the second stage training of the pre-trained language model.
[0084] The first preset ratio is related to the ratio of the parameter matrix to the trainable parameters in the pre-trained language model, and the ratio of the bias term to the trainable parameters. For example, the first preset ratio can be 1:9, and the ratio of the data volume of the first training dataset to the second training dataset is 1:9.
[0085] In this embodiment, the first stage of training is to freeze the parameter matrix and train only the bias item using the first training data set; after the first stage of training is completed, the parameter matrix is unfrozen; the second stage of training is to train the bias item and parameter matrix using the second training data set. The second stage of training is to train the pre-trained language model as a whole.
[0086] Optionally, the training module 403 is further configured to divide the training data set into a first training data set and a second training data set according to a third preset ratio, and perform multi-stage training on the pre-trained language model: freeze the bias item in the pre-trained language model, and based on the target task, use the first training data set to optimize the parameter matrix in the pre-trained language model to complete the first stage training of the pre-trained language model; after completing the first stage training, unfreeze the bias item in the pre-trained language model, and based on the target task, use the second training data set to optimize the bias item and parameter matrix in the pre-trained language model to complete the second stage training of the pre-trained language model.
[0087] In this embodiment, the first stage of training is to freeze the bias item and train only the parameter matrix using the first training data set. After the first stage of training is completed, the bias item is unfrozen. The second stage of training is to train the bias item and the parameter matrix using the second training data set. The second stage of training is to train the pre-trained language model as a whole.
[0088] Optionally, the training module 403 is further configured to determine the data size of the training data set; train the pre-trained language model according to the data size: when the data size is less than a first preset size, freeze the updated parameter matrix in the pre-trained language model, and optimize the bias item in the pre-trained language model based on the target task using the training data set; when the data size is not less than the first preset size, optimize the bias item and the updated parameter matrix in the pre-trained language model based on the target task using the training data set.
[0089] The parameter matrix accounts for more than 99% of the trainable parameters in the pre-trained language model, and the bias term accounts for less than 1%. In the embodiment of the present application, when the amount of data is less than the first preset size, the training data set is used to optimize only the bias term in the pre-trained language model (this method is applied to small sample scenarios, and the small sample scenario is a situation where the number of training samples is small). This can greatly reduce the amount of optimized parameters and training time, and at the same time avoid model overfitting when the number of training samples is small. It has been found in practice that only optimizing the bias term in the pre-trained language model can also achieve good results.
[0090] Optionally, the training module 403 is further configured to determine the data volume of the training data set; train the pre-trained language model according to the data volume: when the data volume is less than a first preset size, freeze the updated parameter matrix in the pre-trained language model, and optimize the bias item in the pre-trained language model based on the target task using the training data set; when the data volume is greater than or equal to the first preset size but less than a second preset size, freeze the bias item in the pre-trained language model, and optimize the updated parameter matrix in the pre-trained language model based on the target task using the training data set; when the data volume is greater than or equal to the second preset size, optimize the bias item and the updated parameter matrix in the pre-trained language model based on the target task using the training data set.
[0091] The embodiment of the present application selects a corresponding training method according to the data volume of the training data set to improve the efficiency of training.
[0092] Optionally, the acquisition module 401 is also configured to sequentially connect multiple linear layers and nonlinear activation functions to obtain a feedforward layer; sequentially connect an embedding layer, a multi-head attention network, a residual layer, a normalization layer, a feedforward layer, a residual layer, a normalization layer, a fully connected layer and a classification layer to obtain a language model; and pre-train the language model based on the target task to obtain a pre-trained language model.
[0093] Multiple linear layers are serially connected and then connected with a nonlinear activation function as a feedforward layer. The residual layer after the multi-head attention network is used to add the output of the multi-head attention network to the input of the multi-head attention network; the residual layer after the feedforward layer is used to add the output of the feedforward layer to the input of the feedforward layer.
[0094] Optionally, the training module 403 is further configured to divide the training data set into a first training data set, a second training data set, and a third training data set according to a second preset ratio, and perform multi-stage training on the pre-trained language model: freeze the parameter matrix in the pre-trained language model, and optimize the bias item in the pre-trained language model using the first training data set based on the target task to complete the first stage training of the pre-trained language model; after completing the first stage training, unfreeze the parameter matrix in the pre-trained language model, freeze the bias item in the pre-trained language model, and optimize the parameter matrix in the pre-trained language model using the second training data set based on the target task to complete the second stage training of the pre-trained language model; after completing the second stage training, unfreeze the bias item in the pre-trained language model, and optimize the bias item and parameter matrix in the pre-trained language model using the third training data set based on the target task to complete the third stage training of the pre-trained language model.
[0095] The second preset ratio is related to the ratio of the parameter matrix to the trainable parameters in the pre-trained language model, as well as the ratio of the bias term to the trainable parameters. For example, the first preset ratio may be 1:6:3, where the ratio of the data size of the first training dataset, the second training dataset, and the third training dataset is 1:6:3.
[0096] The first stage of training: freeze the parameter matrix and use the first training data set to train only the bias item; after the first stage of training is completed, unfreeze the parameter matrix; the second stage of training: freeze the bias item and use the second training data set to train the parameter matrix; after the second stage of training is completed, unfreeze the bias item; the third stage of training: use the third training data set to train the parameter matrix and bias item. The third stage of training is the training of the pre-trained language model as a whole.
[0097] The embodiment of the present application can greatly improve the accuracy of the final model through multi-stage training of the pre-trained language model.
[0098] Optionally, the training module 403 is further configured to sequentially connect multiple linear layers and nonlinear activation functions to obtain a feedforward layer; sequentially connect an embedding layer, a multi-head attention network, a residual layer, a normalization layer, a feedforward layer, a residual layer, a normalization layer, a fully connected layer, and a classification layer to obtain a language network, and serially connect multiple language networks to obtain a language model; obtain a training data set corresponding to a target task; calculate the noise perturbation corresponding to each parameter matrix in the language model, and update the parameter matrix according to the noise perturbation corresponding to each parameter matrix; based on the target task, use the training data set to optimize the bias term and the updated parameter matrix in the language model.
[0099] This embodiment does not pre-train the language model, but directly performs formal training on the language model. The above technical means can solve the problems of overfitting and low generalization ability of the trained model in the existing technology, thereby avoiding model overfitting and improving model generalization ability.
[0100] Optionally, the training module 403 is also configured to obtain a training data set and a pre-trained language model corresponding to the target task; calculate the noise perturbation corresponding to the network parameters of each network layer in the pre-trained language model, and update the network parameters according to the noise perturbation corresponding to each network parameter; based on the target task, use the training data set to optimize the updated network parameters in the pre-trained language model to complete the training of the pre-trained language model.
[0101] Optionally, the calculation module 402 is further configured to calculate the noise disturbance corresponding to each network parameter in the pre-trained language model by the following formula: i =U(-λ / 2,λ / 2)*std(W i );
[0102] Among them, W' i is the noise perturbation corresponding to the i-th network parameter, U(-λ / 2,λ / 2) is the uniformly distributed noise in the range from -λ / 2 to λ / 2, λ is the hyperparameter controlling the noise intensity in the pre-trained language model, std(W i ) is the standard deviation of the internal data of the i-th network parameter.
[0103] Optionally, the calculation module 402 is further configured to update each network parameter using the following formula: i =W i +W' i ;
[0104] Among them, W' i is the noise perturbation corresponding to the i-th network parameter, W i is the i-th network parameter before update, W” i is the updated i-th network parameter.
[0105] W' i If and W i The dimensions of W' i Fill so that W' i and W i The dimensions are consistent.
[0106] Optionally, the training module 403 is further configured to sequentially connect multiple linear layers and nonlinear activation functions to obtain a feedforward layer; sequentially connect an embedding layer, a multi-head attention network, a residual layer, a normalization layer, a feedforward layer, a residual layer, a normalization layer, a fully connected layer, and a classification layer to obtain a language network, and serially connect multiple language networks to obtain a language model; obtain a training data set corresponding to a target task; calculate the noise perturbation corresponding to each network parameter in the language model, and update the network parameter according to the noise perturbation corresponding to each network parameter; based on the target task, use the training data set to optimize the updated network parameters in the language model.
[0107] Optionally, the training module 403 is also configured to obtain a training data set and a pre-trained language model corresponding to the target task; calculate the noise perturbation corresponding to the network parameters of each network layer in the pre-trained language model, and update the network parameters according to the noise perturbation corresponding to each network parameter; based on the target task, use the training data set to optimize the updated network parameters in the pre-trained language model to complete the training of the pre-trained language model.
[0108] Optionally, the training module 403 is further configured to obtain a training data set and a pre-trained language model corresponding to the target task; determine the first network parameters and the second network parameters corresponding to the bias term and the parameter matrix in the pre-trained language model respectively; calculate the noise perturbation corresponding to each second network parameter in the pre-trained language model, and update the second network parameter according to the noise perturbation corresponding to each parameter matrix; based on the target task, use the training data set to optimize the first network parameters and the updated second network parameters in the pre-trained language model.
[0109] Optionally, the training module 403 is also configured to obtain a trained target language model; input multiple training samples in the training data set into the language model and the target language model, and output the first processing result and the second processing result corresponding to each training sample respectively; calculate the contrast loss using the triplet loss function according to the first processing result and the second processing result corresponding to each training sample and the second processing result corresponding to another training sample with different semantics from the training sample; calculate the classification loss using the cross-entropy loss function according to the first processing result and the label corresponding to each training sample; update the network parameters of the language model based on the contrast loss and the classification loss to complete the training of the language model.
[0110] The triplet loss function is triplet(). The first processing result and the second processing result corresponding to a certain training sample are A1 and A2 respectively, and the second processing result corresponding to another training sample with different semantics from the training sample is A3 (the other training sample with different semantics from the training sample is randomly determined in the training data set). The loss value corresponding to the first language corpus is equal to triplet (A1, A2, A3), and the loss values corresponding to all training samples are added together to form the contrast loss. The contrast loss and the classification loss are weighted and summed according to the preset weights to form the total loss, and the model parameters of the language model are updated according to the total loss. The embodiment of the present application can solve the problem of overfitting of the translation model in the prior art by introducing the contrast loss into the model training, thereby improving the generalization performance of the model.
[0111] It should be understood that the size of the serial numbers of the steps in the above embodiments does not mean the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of this application.
[0112] Figure 5 is a schematic diagram of an electronic device 5 provided in an embodiment of the present application. As shown in Figure 5, the electronic device 5 of this embodiment includes: a processor 501, a memory 502, and a computer program 503 stored in the memory 502 and executable by the processor 501. When the processor 501 executes the computer program 503, the steps of the above-described method embodiments are implemented. Alternatively, when the processor 501 executes the computer program 503, the functions of the modules / units in the above-described device embodiments are implemented.
[0113] Electronic device 5 can be a desktop computer, laptop, PDA, cloud server, or other electronic device. Electronic device 5 can include, but is not limited to, a processor 501 and a memory 502. Those skilled in the art will appreciate that FIG5 is merely an example of electronic device 5 and does not limit the electronic device 5. The electronic device 5 may include more or fewer components than shown, or different components.
[0114] The processor 501 may be a central processing unit (CPU), or other general-purpose processors, digital signal processors (DSP), application-specific integrated circuits (ASIC), field-programmable gate arrays (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc.
[0115] The memory 502 can be an internal storage unit of the electronic device 5, such as a hard disk or memory of the electronic device 5. The memory 502 can also be an external storage device of the electronic device 5, such as a plug-in hard disk, a Smart Media Card (SMC), a Secure Digital (SD) card, a flash memory card, etc. The memory 502 can also include both an internal storage unit of the electronic device 5 and an external storage device. The memory 502 is used to store computer programs and other programs and data required by the electronic device.
[0116] Those skilled in the art will clearly understand that for the sake of convenience and brevity of description, only the division of the above-mentioned functional units and modules is used as an example for illustration. In actual applications, the above-mentioned functions can be distributed and completed by different functional units and modules as needed, that is, the internal structure of the device can be divided into different functional units or modules to complete all or part of the functions described above. The functional units and modules in the embodiments can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The above-mentioned integrated units can be implemented in the form of hardware or in the form of software functional units.
[0117] If the integrated module / unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the present application implements all or part of the process in the above-mentioned embodiment method, and can also be completed by instructing the relevant hardware through a computer program. The computer program can be stored in a computer-readable storage medium, and when the computer program is executed by the processor, it can implement the steps of the above-mentioned various method embodiments. The computer program may include computer program code, which may be in source code form, object code form, executable file or some intermediate form. The computer-readable medium may include: any entity or device capable of carrying computer program code, recording medium, USB flash drive, mobile hard disk, magnetic disk, optical disk, computer memory, read-only memory (ROM), random access memory (RAM), electric carrier signal, telecommunication signal and software distribution medium. It should be noted that the content contained in the computer-readable medium can be appropriately increased or decreased according to the requirements of legislation and patent practice in the jurisdiction. For example, in some jurisdictions, according to legislation and patent practice, computer-readable media do not include electric carrier signals and telecommunication signals.
[0118] The above embodiments are only used to illustrate the technical solutions of the present application, rather than to limit them. Although the present application has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or make equivalent replacements for some of the technical features therein. These modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the various embodiments of the present application, and should all be included in the scope of protection of the present application.
Claims
1. A method for pre-training a language model using noise perturbation training, characterized in that: include: Obtain the training dataset and pre-trained language model corresponding to the target task; Calculating the noise perturbation corresponding to each parameter matrix in the pre-trained language model, and updating the parameter matrix according to the noise perturbation corresponding to each parameter matrix; Based on the target task, the bias term and the updated parameter matrix in the pre-trained language model are optimized using the training dataset.
2. The method according to claim 1, characterized in that The noise perturbation corresponding to each parameter matrix in the pre-trained language model is calculated by the following formula: W' i =U(-λ / 2,λ / 2)*std(W i ); Among them, W' i is the noise perturbation corresponding to the i-th parameter matrix, U(-λ / 2,λ / 2) is the uniformly distributed noise in the range from -λ / 2 to λ / 2, λ is the hyperparameter controlling the noise intensity in the pre-trained language model, std(W i ) is the standard deviation of the data within the i-th parameter matrix.
3. The method according to claim 1, characterized in that Update each parameter matrix by the following formula: W” i =W i +W' i ; Among them, W' i is the noise perturbation corresponding to the i-th parameter matrix, W i is the i-th parameter matrix before update, W” i is the updated i-th parameter matrix.
4. The method according to claim 1, wherein Based on the target task, optimizing the bias term and the updated parameter matrix in the pre-trained language model using the training dataset includes: Dividing the training data set into a first training data set and a second training data set according to a first preset ratio, and performing multi-stage training on the pre-trained language model: Freezing a parameter matrix in the pre-trained language model, and optimizing a bias term in the pre-trained language model using the first training dataset based on the target task to complete a first phase of training of the pre-trained language model; After completing the first stage of training, the parameter matrix in the pre-trained language model is unfrozen, and based on the target task, the bias term and parameter matrix in the pre-trained language model are optimized using the second training dataset to complete the second stage of training of the pre-trained language model.
5. The method according to claim 1, wherein Based on the target task, optimizing the bias term and the updated parameter matrix in the pre-trained language model using the training dataset includes: Dividing the training data set into a first training data set, a second training data set, and a third training data set according to a second preset ratio, and performing multi-stage training on the pre-trained language model: Freezing a parameter matrix in the pre-trained language model, and optimizing a bias term in the pre-trained language model using the first training dataset based on the target task to complete a first phase of training of the pre-trained language model; After completing the first stage of training, unfreezing the parameter matrix in the pre-trained language model, freezing the bias term in the pre-trained language model, and optimizing the parameter matrix in the pre-trained language model using the second training dataset based on the target task to complete the second stage of training the pre-trained language model; After completing the second stage of training, the bias items in the pre-trained language model are unfrozen, and based on the target task, the bias items and parameter matrices in the pre-trained language model are optimized using the third training data set to complete the third stage of training of the pre-trained language model.
6. The method according to claim 1, characterized in that Based on the target task, optimizing the bias term and the updated parameter matrix in the pre-trained language model using the training dataset includes: Determining the data volume of the training data set; The pre-trained language model is trained according to the amount of data: When the data amount is less than a first preset size, freezing the updated parameter matrix in the pre-trained language model, and optimizing the bias item in the pre-trained language model using the training data set based on the target task; When the data amount is not less than the first preset size, based on the target task, the bias term and the updated parameter matrix in the pre-trained language model are optimized using the training data set.
7. The method according to claim 1, characterized in that Before obtaining the pre-trained language model corresponding to the target task, the method further includes: Connect multiple linear layers and nonlinear activation functions in sequence to obtain a feedforward layer; Connect the embedding layer, multi-head attention network, residual layer, normalization layer, feedforward layer, residual layer, normalization layer, fully connected layer and classification layer in sequence to obtain the language model; The language model is pre-trained based on the target task to obtain the pre-trained language model.
8. The method according to claim 1, characterized in that The method further comprises: Connect multiple linear layers and nonlinear activation functions in sequence to obtain a feedforward layer; Connect the embedding layer, multi-head attention network, residual layer, normalization layer, feedforward layer, residual layer, normalization layer, fully connected layer and classification layer in sequence to obtain a language network. Connect multiple language networks in series to obtain a language model. Obtain the training dataset corresponding to the target task; Calculate the noise perturbation corresponding to each parameter matrix in the language model, and update the noise perturbation corresponding to each parameter matrix. New parameter matrix; Based on the target task, the training dataset is used to optimize the bias terms and updated parameter matrix in the language model.
9. The method according to claim 1, characterized in that The method further comprises: Obtain the training dataset and pre-trained language model corresponding to the target task; Calculating the noise perturbation corresponding to the network parameters of each network layer in the pre-trained language model, and updating the network parameters according to the noise perturbation corresponding to each network parameter; Based on the target task, the updated network parameters in the pre-trained language model are optimized using the training data set to complete the training of the pre-trained language model.
10. A device for training a pre-trained language model using noise perturbation, characterized in that: include: An acquisition module is configured to acquire a training dataset and a pre-trained language model corresponding to a target task; A calculation module is configured to calculate the noise disturbance corresponding to each parameter matrix in the pre-trained language model, and update the parameter matrix according to the noise disturbance corresponding to each parameter matrix; A training module is configured to optimize the bias term and the updated parameter matrix in the pre-trained language model using the training dataset based on the target task.
Citation Information
Patent Citations
User abnormal mode recognition method, device and equipment
CN113052324A
Knowledge distillation-based text classification method and system
CN114818902A
Terminal suitability judgment method and device, electronic equipment and storage medium
CN115734029A
Method and device for training pre-training language model by using noise disturbance
CN116362351A
Efficient automatic punctuation with robust inference
US20210319176A1
Cited By
Text reasoning method, product, equipment and storage medium
CN120851223A