Dual consistency regularization method and system for Transformer supervised learning for sequence tasks
By adding perturbations to the training input sequence and constructing base and mean models, and using consistency loss to adjust parameters, the overfitting problem of the Transformer model under small-scale data is solved, and the robustness and performance of the sequence generation model are improved.
Patent Information
- Application Number
- CN202310629724.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-05-31
- Publication Date
- 2025-10-03
- Estimated Expiration
- 2043-05-31
AI Technical Summary
The Transformer model is prone to overfitting in sequence tasks with small-scale training data, resulting in a significant decrease in the model's prediction results for small perturbation inputs during the prediction phase and insufficient robustness.
By adding perturbations to the training input sequence, building a base model and a mean model, the exponential moving average is used to migrate the base model parameters, and the model parameters are adjusted to improve robustness by combining the output feature probability distribution and prediction probability consistency loss of the encoder and decoder.
It effectively alleviates the overfitting problem under few-sample training conditions, improves the robustness of the sequence generation model, and improves performance in tasks such as machine translation and text summarization.
Smart Images

Figure CN116611473B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of artificial intelligence application technology, and in particular to a dual consistency regularization method and system for Transformer supervised learning for sequence tasks. Background Art
[0002] Common sequence tasks include machine translation, text summarization, and intelligent dialogue answer generation. A common characteristic of sequence tasks is that each text sample presents an exponential or infinite number of candidate label sequence combinations. For part-of-speech tagging, assuming there are n parts of speech, then for a text of length L, there are nL possible label sequence combinations. For sequence generation tasks like machine translation, regardless of model constraints, the possible label sequences—that is, target language word sequences—are infinite. Based on these characteristics, the problem can be structured, enabling more efficient model learning and training.
[0003] The Transformer model has achieved excellent performance across various natural language processing (NLP) tasks and is the core network for pre-trained language models. Given a sentence or paragraph as input, each word in the input sequence is first converted into its corresponding word vector, along with a position vector representing each word's position in the sequence. These word vectors are then fed into the Transformer network, where a self-attention mechanism is used to learn relationships between words and encode their context. A feed-forward network then undergoes nonlinear transformations to output vector representations of each word that incorporate contextual features. The Transformer can be used for both encoding and decoding. Decoding involves obtaining a desired outcome based on a sentence input, such as machine translation (inputting a source language sentence and outputting a target language sentence) or reading comprehension (inputting a document and question and outputting an answer). During decoding, the decoded words undergo a self-attention mechanism, followed by another attention mechanism with the encoded hidden state sequence. A linear layer then maps these vectors to a vector of the size of the vocabulary. Each vector represents the output probability of a word in the vocabulary, and a softmax layer then generates the output probability for each word. Deep neural networks based on the Transformer model, which are large in scale and have many layers, are very powerful in sequence tasks. However, they are prone to overfitting when training on small amounts of data. During the prediction phase, the model's prediction results for slightly perturbed inputs are significantly lower than when the input is not perturbed. Summary of the Invention
[0004] To this end, the present invention provides a dual consistency regularization method and system for Transformer supervised learning for sequence tasks, which solves the overfitting problem of model training under the condition of scarce labeled data in sequence tasks, improves the robustness of sequence generation models, and facilitates their application in sequence tasks such as machine translation and text summarization.
[0005] According to the design scheme provided by the present invention, a dual consistency regularization method for Transformer supervised learning for sequence tasks is provided, including:
[0006] Add perturbations to the training input sequence to obtain perturbation sequence data for model training;
[0007] Based on the perturbed sequence data, the training loss of the base model and the consistency loss between the base model and the mean model are determined. The base model is an end-to-end model for sequence tasks modeled using the Transformer structure, and the mean model is a model structure obtained by migrating the backpropagation update parameters of the base model using the exponential moving average.
[0008] Obtain the overall training loss of the base model based on the base model training loss and the consistency loss between the base model and the mean model;
[0009] The base model parameters are adjusted based on the overall training loss to obtain the target sequence task end-to-end model for performing the sequence task.
[0010] As a dual consistency regularization method for Transformer supervised learning for sequence tasks of the present invention, further, perturbations are added to the training input sequence, including: adding different perturbations to the training input sequence respectively to obtain first perturbation sequence data for basic model training and second perturbation sequence data for mean model training.
[0011] As the dual consistency regularization method of Transformer supervised learning for sequence tasks in the present invention, further, in the mean model structure, the mean model parameter update rule is expressed as: θ m ←λθ m +(1-λ)θ b , where θ b is the basic model parameter, and the basic model parameter is updated by back propagation, θ m is the mean model parameter, and λ is the model parameter balance weight during the training phase.
[0012] As a dual consistency regularization method for Transformer supervised learning for sequence tasks in the present invention, further, the consistency loss between the base model and the mean model is determined based on the perturbed sequence data, including:
[0013] First, the perturbation sequence data is input into the base model and the mean model respectively, the encoder output features are obtained based on the base model encoder and the mean model encoder, and the probability distribution of the corresponding encoder output features is obtained based on the softmax layer in the model;
[0014] Next, the encoder output features are input to the base model decoder and the mean model decoder to obtain the decoder predicted output probability;
[0015] Then, the encoder output feature probability distribution distance similarity and the decoder prediction output probability distance similarity between the base model and the mean model are measured, and the consistency loss of the base model and the mean model is constructed based on the distance similarity.
[0016] As the dual consistency regularization method of Transformer supervised learning for sequence tasks in the present invention, the consistency loss of the base model and the mean model is further expressed as:
[0017] Among them, p m,i (x)‖、p b,i (x) are the probability distribution of the base model encoder output feature and the mean model encoder output feature probability distribution for the training input sequence x, respectively, x b and x m are the disturbance sequence data of the basic model input and the disturbance sequence data of the mean model input, p b (y i ∣x b ,y <i ), p m (y i ∣x m ,y <i ) are the predicted output probability of the basic model decoder and the predicted output probability of the mean model decoder for the corresponding input perturbation sequence data, respectively. KL () represents the KL divergence function, i represents the current iteration round, y i Represents the predicted token generated by the decoder in the current iteration round, and V represents all vocabulary.
[0018] As a dual consistency regularization method for Transformer supervised learning for sequence tasks in the present invention, further, when determining the basic model training loss based on the perturbed sequence data, the cross entropy function is used to represent the basic model training loss.
[0019] As the dual consistency regularization method for Transformer supervised learning for sequence tasks in the present invention, the overall training loss of the basic model is expressed as: L = L CE +αLcon , where L CE Represents the basic model training loss, L con represents the consistency loss of the base model and the mean model, and a represents the balance weight parameter.
[0020] Furthermore, the present invention also provides a Transformer supervised learning dual consistency regularization system for sequence tasks, comprising: a training data construction module, a training loss determination module and a model parameter adjustment module, wherein,
[0021] The training data construction module is used to add perturbations to the training input sequence to obtain perturbation sequence data for model training;
[0022] A training loss determination module is used to determine the base model training loss and the consistency loss between the base model and the mean model based on the perturbation sequence data. The base model is an end-to-end model for sequence tasks modeled using a Transformer structure, and the mean model is a model structure obtained by migrating the backpropagation update parameters of the base model based on the base model using an exponential moving average. The overall training loss of the base model is obtained based on the base model training loss and the consistency loss between the base model and the mean model.
[0023] The model parameter adjustment module is used to adjust the basic model parameters based on the overall training loss to obtain the target sequence task end-to-end model for performing the sequence task.
[0024] Beneficial effects of the present invention:
[0025] During the training process, the present invention constrains the outputs of the encoder and decoder respectively to alleviate the overfitting problem under the condition of few-sample training. It does not use unlabeled data, and makes high-confidence and high-consistency feature extraction and prediction for inputs with similar distances in the feature space. By measuring the consistency between the features extracted by the encoders of the two models and the predicted probabilities of the decoders, the solution space of the model is doubly constrained, thereby achieving the purpose of regularizing the internal parameters of the model. It can be used for supervised learning under conditions of few samples, can improve the robustness of the sequence generation model finally obtained by training, and is convenient for application in sequence task scenarios such as machine translation and text summary generation. BRIEF DESCRIPTION OF THE DRAWINGS
[0026] Figure 1 Schematic diagram of the dual consistency regularization process for Transformer supervised learning for sequence tasks in the embodiment;
[0027] Figure 2 This is a schematic diagram of the principle of the end-to-end model for sequence tasks in the embodiment. DETAILED DESCRIPTION
[0028] In order to make the objectives, technical solutions and advantages of the present invention clearer and more understandable, the present invention is further described in detail below with reference to the accompanying drawings and technical solutions.
[0029] Supervised learning is the process of adjusting the parameters of a classifier using a set of samples of known categories to achieve the required performance; it is a machine learning task that infers a function from labeled training data. The training data includes a set of training examples. In supervised learning, each instance consists of an input object and a desired output value. The supervised learning algorithm analyzes the training data and generates an inferred function that can be used to map new instances, allowing the algorithm to correctly determine the class labels of unseen instances. To address the overfitting problem under few-sample training conditions, in the embodiments of the present invention, see Figure 1 As shown in the figure, a dual consistency regularization method for Transformer supervised learning for sequence tasks is provided, including:
[0030] S101, adding disturbance to the training input sequence to obtain disturbance sequence data for model training;
[0031] S102. Determine the training loss of the base model and the consistency loss between the base model and the mean model based on the perturbation sequence data, wherein the base model is an end-to-end model for sequence tasks using a Transformer structure, and the mean model is a model structure obtained by migrating the back-propagation update parameters of the base model based on the base model using an exponential moving average.
[0032] S103. Obtaining the overall training loss of the base model based on the base model training loss and the consistency loss between the base model and the mean model;
[0033] S104. Adjust the basic model parameters based on the overall training loss to obtain a target sequence task end-to-end model for performing the sequence task.
[0034] During training, the encoder and decoder outputs are constrained to mitigate overfitting in low-sample training conditions. Based on the target model loss function, a mean model is constructed by integrating the model parameters from previous rounds. By calculating the consistency between the mean model and the base model, regularization constraints are applied to the model's solution space to improve the training performance of sequential task models.
[0035] As a preferred embodiment, further, adding disturbance to the training input sequence includes: adding different disturbances to the training input sequence respectively to obtain first disturbance sequence data for basic model training and second disturbance sequence data for mean model training respectively.
[0036] In a specific implementation, the consistency loss between the base model and the mean model is determined based on the perturbation sequence data, which can be designed to include the following contents: first, the perturbation sequence data is respectively input into the base model and the mean model, the encoder output features are obtained based on the base model encoder and the mean model encoder, and the probability distribution of the corresponding encoder output features is obtained based on the softmax layer in the model; then, the encoder output features are input into the base model decoder and the mean model decoder to obtain the decoder predicted output probability; then, the encoder output feature probability distribution distance similarity and the decoder predicted output probability distance similarity between the base model and the mean model are measured, and the consistency loss of the base model and the mean model is constructed based on the distance similarity.
[0037] EMA, or Exponential Moving Average (EXPMA or EMA), is a trend indicator. In deep learning, EMA can be used to average model parameters to improve test metrics and increase model accuracy and robustness.
[0038] See also Figure 2 As shown, the basic model of the previous training round can be used to build the mean model. The two models have the same structure. The parameters of the mean model are updated using the exponential moving average (EMA). Instead of using the optimized parameters from the final training iteration as the mean model parameters, the exponential moving average of the parameters in all training iterations is used to reduce the noise and fluctuation of the time series data, reduce the fluctuation noise of the model parameters, and make the model parameters more likely to approach the local minimum. In this embodiment, the update rule of the mean model parameters can be set as
[0039] θ m ←λθ m +(1-λ)θ b
[0040] Among them, θ b As the basic model θb The parameters of θ are updated by back propagation. m is the mean model g θm Since the model performance is poor in the early stages of training, if the initial model weight is too high, the training effect will be quite poor. Therefore, during training, λ can be adjusted to follow a cosine schedule from 0.996 to 1 to balance the model parameter weights at different training stages.
[0041] In the training phase of the sequence task model of machine translation, it is assumed that the input source sentence x is (x1, x2, ..., x T), the given input x is perturbed differently and then input into the mean model and the basic model, denoted as x b and x m , the feature outputs of the two model encoders are recorded as f b (x) and f m (x), the predicted output of the decoder is recorded as p b (y i ∣x b ,y <i ) and p m (y i ∣x m ,y <i ). f b (x) and f m (x) is converted into a probability distribution p through the softmax layer b (x) and p m (x). For similar inputs, the model encoder and decoder should have consistent outputs. The KL divergence can be used to measure the distance between the distributions of the encoder output and decoder output of two models and define the consistency loss, which can be expressed as:
[0042]
[0043] The total objective loss for model training can be defined as:
[0044] L=L CE +αL con
[0045] Among them, L CE is the basic model training loss, and α is the weight to balance the two losses.
[0046] Furthermore, based on the above method, an embodiment of the present invention also provides a Transformer supervised learning dual consistency regularization system for sequence tasks, comprising: a training data construction module, a training loss determination module and a model parameter adjustment module, wherein:
[0047] The training data construction module is used to add perturbations to the training input sequence to obtain perturbation sequence data for model training;
[0048] A training loss determination module is used to determine the base model training loss and the consistency loss between the base model and the mean model based on the perturbation sequence data. The base model is an end-to-end model for sequence tasks modeled using a Transformer structure, and the mean model is a model structure obtained by migrating the backpropagation update parameters of the base model based on the base model using an exponential moving average. The overall training loss of the base model is obtained based on the base model training loss and the consistency loss between the base model and the mean model.
[0049] The model parameter adjustment module is used to adjust the basic model parameters based on the overall training loss to obtain the target sequence task end-to-end model for performing the sequence task.
[0050] To verify the effectiveness of this solution, the following is a further explanation based on experimental data:
[0051] Experiments were conducted on small-sample machine translation tasks. In the German-English experiment, the sizes of the training set, validation set, and test set were 160k, 7.3k, and 6.5k, respectively; in the English-Vietnamese experiment, the sizes of the training set, validation set, and test set were 133k, 1.5k, and 1.3k, respectively.
[0052] The end-to-end model baseline system was built using the fairseq open-source deep learning library, and the non-parametric knowledge distillation framework was implemented using the faiss deep learning package. Both models adopted the Transformer architecture. The encoder and decoder layers were 6, the embedding size was 256, the number of attention heads was 4, and the filter size of the feedforward neural network was 1024. The hyperparameter α was set to 0.75. The experimental results are shown in Table 1.
[0053] Table 1 Experimental results
[0054]
[0055] The experimental results above demonstrate that our DOCR solution can effectively improve model performance, reduce overfitting, and demonstrate complementary effects with other regularization methods. Compared to the basic Transformer model, our DOCR solution achieves improvements of +2.57 BLEU scores and +2.63 BLEU scores on the IWSLT'14 German-English and IWSLT'15 English-Vietnamese datasets, respectively. When combined with other regularization methods, model performance improves by an average of +5.41 BLEU scores. Furthermore, experiments using the kNN-KD Sota model for knowledge distillation demonstrate that our regularization method can significantly improve strong baseline models.
[0056] Unless otherwise specifically stated, the relative steps, numerical expressions and values of the components and steps set forth in these embodiments do not limit the scope of the present invention.
[0057] The various embodiments in this specification are described in a progressive manner, with each embodiment focusing on the differences from other embodiments. Reference can be made to the common and similar parts between the various embodiments. For the systems disclosed in the embodiments, since they correspond to the methods disclosed in the embodiments, the description is relatively simple, and the relevant parts can be referred to the method description.
[0058] The units and method steps of each example described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, computer software, or a combination of the two. In order to clearly illustrate the interchangeability of hardware and software, the components and steps of each example have been generally described in terms of function in the above description. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. A person of ordinary skill in the art may use different methods to implement the described functions for each specific application, but such implementation is not considered to be beyond the scope of the present invention.
[0059] Those skilled in the art will appreciate that all or part of the steps in the above method can be performed by a program instructing related hardware. The program can be stored in a computer-readable storage medium, such as a read-only memory, a magnetic disk, or an optical disk. Alternatively, all or part of the steps in the above embodiment can be implemented using one or more integrated circuits. Accordingly, each module / unit in the above embodiment can be implemented in the form of hardware or software functional modules. The present invention is not limited to any specific combination of hardware and software.
[0060] Finally, it should be noted that the above-described embodiments are only specific implementation methods of the present invention, which are used to illustrate the technical solutions of the present invention, rather than to limit them. The scope of protection of the present invention is not limited thereto. Although the present invention has been described in detail with reference to the above-described embodiments, those skilled in the art should understand that any person skilled in the art can modify or easily conceive of changes to the technical solutions described in the above-described embodiments within the technical scope disclosed by the present invention, or replace some of the technical features therein with equivalents. Such modifications, changes, or replacements do not deviate from the spirit and scope of the technical solutions of the embodiments of the present invention, and should be included in the scope of protection of the present invention. Therefore, the scope of protection of the present invention shall be subject to the scope of protection of the claims.
Claims
1. A dual consistency regularization method for Transformer supervised learning for sequence tasks, applied to machine translation, characterized by: Include: Adding different perturbations to the training input sequence to obtain first perturbation sequence data for basic model training and second perturbation sequence data for mean model training, respectively, wherein the input sequence is a source sentence; The perturbation sequence data is input into the base model and the mean model respectively, the encoder output features are obtained based on the base model encoder and the mean model encoder, and the probability distribution of the corresponding encoder output features is obtained based on the softmax layer in the model; the encoder output features are input into the base model decoder and the mean model decoder respectively, and the decoder predicted output probability is obtained; the encoder output feature probability distribution distance similarity and the decoder predicted output probability distance similarity between the base model and the mean model are measured, and the consistency loss of the base model and the mean model is constructed based on the distance similarity. Among them, the base model is an end-to-end model for sequence tasks modeled using the Transformer structure, and the mean model is a model structure obtained by migrating the back-propagation update parameters of the base model based on the base model using the exponential moving average; Obtain the overall training loss of the base model based on the base model training loss and the consistency loss between the base model and the mean model; The base model parameters are adjusted based on the overall training loss to obtain the target sequence task end-to-end model for performing the sequence task.
2. The dual consistency regularization method for Transformer supervised learning for sequence tasks according to claim 1, characterized in that In obtaining the mean model structure, the mean model parameter update rule is expressed as: θ m ←λθ m +(1-λ)θ b , where θ b is the basic model parameter, and the basic model parameter is updated by back propagation, θ m is the mean model parameter, and λ is the model parameter balance weight during the training phase.
3. The dual consistency regularization method for Transformer supervised learning for sequence tasks according to claim 1, characterized in that The consistency loss of the base model and the mean model is expressed as: Among them, p m,i (x)|、p b,i (x) are the probability distribution of the base model encoder output feature and the mean model encoder output feature probability distribution for the training input sequence x, respectively, x b and x m are the disturbance sequence data of the basic model input and the disturbance sequence data of the mean model input, p b (y i ∣x b ,y <i ), p m (y i ∣x m ,y <i ) are the predicted output probability of the basic model decoder and the predicted output probability of the mean model decoder for the corresponding input perturbation sequence data, respectively. KL () represents the KL divergence function, i represents the current iteration round, y i Represents the predicted token generated by the decoder in the current iteration round, and V represents all vocabulary.
4. The dual consistency regularization method for Transformer supervised learning for sequence tasks according to claim 1, characterized in that When determining the base model training loss based on perturbation sequence data, the cross entropy function is used to represent the base model training loss.
5. The dual consistency regularization method for Transformer supervised learning for sequence tasks according to claim 1 or 4, characterized in that The overall training loss of the base model is expressed as: Among them, L CE Represents the basic model training loss, L con represents the consistency loss of the base model and the mean model, and a represents the balance weight parameter.
6. A dual consistency regularization system for Transformer supervised learning for sequence tasks, characterized by: The method according to claim 1 is implemented, comprising: a training data construction module, a training loss determination module and a model parameter adjustment module, wherein: The training data construction module is used to add perturbations to the training input sequence to obtain perturbation sequence data for model training; A training loss determination module is used to determine the base model training loss and the consistency loss between the base model and the mean model based on the perturbation sequence data. The base model is an end-to-end model for sequence tasks modeled using a Transformer structure, and the mean model is a model structure obtained by migrating the backpropagation update parameters of the base model based on the base model using an exponential moving average. The overall training loss of the base model is obtained based on the base model training loss and the consistency loss between the base model and the mean model. The model parameter adjustment module is used to adjust the basic model parameters based on the overall training loss to obtain the target sequence task end-to-end model for performing the sequence task.
7. An electronic device, characterized in that: The method comprises a memory and a processor, wherein the processor and the memory communicate with each other via a bus; the memory stores program instructions executable by the processor, and the processor calls the program instructions to execute the method according to any one of claims 1 to 5.
8. A computer-readable storage medium, characterized in that The computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the method according to any one of claims 1 to 5 is implemented.
Citation Information
Patent Citations
Domain generalization and domain adaptive learning method based on data expansion consistency
CN111382871A
Training method and device of grammar error correction model, equipment and storage medium
CN115062611A