Variable Importance Learning for Multi-Task Models
By assigning scaling factors and mixed weights to tasks in multi-task learning, and optimizing the model parameter set, the problem of limited task weight estimation is solved, achieving more efficient model training and performance improvement.
Patent Information
- Application Number
- CN202080103533.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2020-10-02
- Publication Date
- 2025-07-04
- Estimated Expiration
- 2040-10-02
AI Technical Summary
Existing multitasking learning methods are difficult to effectively utilize mutually beneficial information between tasks during training. Especially in deep learning, task weight estimation is limited by gradient estimation and cannot utilize negative migration, resulting in limited model training efficiency and performance.
By assigning candidate scale factors to multiple tasks, performing optimization steps to form a refined model parameter set, and determining a mixed weight set through predetermined evaluation criteria, allowing positive, zero and negative values to optimize target task performance.
It improves the training efficiency and performance of the target task, can effectively utilize the information of auxiliary tasks, is suitable for complex optimizers, allows negative migration, and improves the overall performance of the model.
Smart Images

Figure CN116057537B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to multi-task models, and more particularly to improving the training of such models. Background Art
[0002] Multi-task learning can be used to train models applicable to various fields. For example, these models can be used for Natural Language Understanding (NLU) problems. These problems include, but are not limited to, Natural Language Inference (NLI), entailment or Word Sense Disambiguation (WSD). All problems come with their own datasets for training, development, and testing purposes, typically consisting of natural language input pairs (e.g., sentences or sentence pairs) and associated discrete labels (e.g., "yes" or "no"), where the label is one of a predefined closed set of possible labels (classification problems), or an associated continuous score, e.g., in the range from 0 to 5 (regression problems). The goal of all NLU systems is to learn from the provided data to predict the correct output for new inputs.
[0003] In multi-task learning, many such datasets from various usually related tasks are combined together, and based on a shared underlying model, classifiers are trained for each task on the premise that the information obtained from the training of different tasks is mutually beneficial.
[0004] The following challenges pose a problem: either determining tasks associated in a mutually beneficial way with their training signals before training, or learning the relevance and benefits of tasks during model training.
[0005] In some implementations, multi-task learning can train a model for a target task, where one of the available tasks is considered the task whose performance is to be optimized, and all other tasks are only considered auxiliary tasks, i.e., their role in model updates is only to improve the performance of the target task.
[0006] One way to address this problem previously was to manually select suitable tasks by human experts (usually system developers). The experts would select those tasks from the range of all available tasks that should provide mutually beneficial information based on their expertise and / or intuition. One drawback of this method is that the experts need to manually analyze the available data for all tasks to reach an informed conclusion, which is a laborious and costly task. Another drawback is that the experts may make mistakes, or their intuition may be wrong when training the model. This is especially true when training larger deep neural network models, as their decision-making process is opaque and may be very different from human judgment.
[0007] In classical machine learning (ML) and deep learning (DL), the technical solutions for multi-task learning problems usually automatically infer task weights based on the model performance during training. In DL, there are various methods to estimate the impact of the model update of task A on the performance of another task B or the performance of all other tasks, for example, measured by task classification accuracy. Then, the task-specific weights can be evaluated and adjusted based on this performance. Existing technologies usually use the gradients of the model to estimate task weights and scale them to the range of 0 to 1, where 0 means completely ignoring the task and 1 means fully utilizing its model update.
[0008] This technique has several drawbacks. First, existing methods usually estimate task weights based on the gradients obtained from training input-output pairs and apply them to these gradients. They are limited to using very simple model optimization methods, usually only Stochastic Gradient Descent (SGD), because manipulating gradients before the optimization step easily interferes with the internal state of more complex optimizers such as ADAM. This means that ultimately, any such method simply learns the task-specific learning rate to be used with SGD and is not easily (or not at all) applicable to more advanced optimizers.
[0009] Another drawback of existing technology methods is that they limit task weights to the range of 0 to 1. Assume that there is only positive transfer (i.e., a positive impact on task performance) between the tasks being trained, or a given task has an adverse effect on performance. In the former case, the weight will be greater than 0; in the latter case, the weight will be 0. This excludes the possibility of leveraging negative transfer, for example, by flipping the sign of the task model update, in which case this flip makes the target task perform better.
[0010] A method to overcome these problems needs to be developed. Summary of the Invention
[0011] According to one aspect, the present invention provides an apparatus including one or more processors for training a machine learning model with an initial set of model parameters, the apparatus being configured to train the model using a respective training data set for each of a plurality of tasks, the plurality of tasks including a target task for which the model is to be trained and one or more auxiliary tasks, the apparatus being configured to train the model by performing the following steps: assigning at least one candidate scaling factor to each of the plurality of tasks; for each of one or more of the plurality of tasks, performing one or more optimization steps on the respective task according to the respective candidate scaling factor to form a respective refined set of model parameters; performing a machine learning operation through one or more predetermined evaluation criteria and based on one or more predetermined constraints to evaluate the performance of the refined set of model parameters of the one or more tasks of the plurality of tasks in the target task, thereby determining a set of mixing weights; updating the set of model parameters of the machine learning model according to the refined model parameters weighted by the set of mixing weights; and updating the candidate scaling factor according to the set of mixing weights. In this way, the apparatus can use the model updates for the auxiliary tasks in cases where the model updates contribute to driving the performance of the target task.
[0012] Using the one or more auxiliary tasks during training can optimize the performance of the target task. In this way, the training process of the target task can obtain mutually beneficial information from the one or more auxiliary tasks.
[0013] The set of mixing weights may include one or more of positive values, zero values, and negative values. This can allow the use of negatively correlated tasks to improve the performance of the target task. In cases where the model update of an auxiliary task is detrimental to the target task, it is still possible to utilize its data by setting the mixing ratio of that task to a negative number, thereby generating an update that may be beneficial to the target task instead. In cases where an auxiliary task cannot be utilized at all, the method can still set the weight to 0 and ignore the task.
[0014] The apparatus may be configured to assign a plurality of candidate scaling factors to each of the plurality of tasks. In this way, it is possible to assign different candidate scaling factors to different layers of each task.
[0015] The apparatus may also be configured to determine, for each of the plurality of tasks, an optimal set of mixing weights related to the performance of the target task. In this way, the apparatus can train a model with good performance for the target task.
[0016] The apparatus may also be configured to adjust the at least one scaling factor based on the optimal set of mixing weights for each of the plurality of tasks. In this way, the apparatus can use the best scaling factor for each task.
[0017] The model architecture may include multiple layers, and the device may be used to apply different scaling factors to each of at least some of the multiple layers for at least one of the multiple tasks. The model architecture may include multiple layers, and the device may also be used to apply a single scaling factor to all layers of the model for each of the multiple tasks. The α variable assigned to the task may be a "global" variable (applying exactly one mixing ratio for each task to all layers of the model), or a "local" variable (where each task for each model layer uses one variable). In either case, the number of additional parameters introduced into the model is minimal and is generally negligible compared to the parameters of the original multi-task model.
[0018] The device may be used to assign a candidate scaling factor with a value of 1 to each of the multiple tasks at the start of the training. This can provide a simple initial value that can be optimized in successive iterations.
[0019] The model architecture may include an embedder and at least one encoder block, the embedder including multiple layers, each encoder block including a self-attention mechanism and having at least one linear layer and at least one normalization layer. For example, a pre-trained RoBERTa model may be used as the input encoder, the encoder having encoder layers and a task-specific classification head, and including a simple feed-forward layer at the top.
[0020] One or more of the multiple tasks may be randomly sampled from the multiple tasks. This can effectively evaluate the impact of different tasks on the training of the target task.
[0021] The target task may be a predetermined task to be optimized through the model training. This enables the use of datasets from various related tasks to train the model, where classifiers are trained for each of the tasks, based on the assumption that the information obtained from the training of different tasks is mutually beneficial and improves the performance related to the target task.
[0022] The model may be used to perform Natural Language Understanding (NLU) tasks. The model may allow a computer to interpret and use human language input. These problems include but are not limited to Natural Language Inference (NLI), entailment or Word Sense Disambiguation (WSD).
[0023] The multiple tasks may correspond to relevant problems in the field of human language understanding. Preferably, the training data set for each task includes natural language input pairs (e.g., sentences or sentence pairs) and associated discrete labels (e.g., "yes" or "no"). The labels may be one of a predefined closed set of possible labels (classification problem), or associated continuous scores, e.g., within the range of 0 to 5 (regression problem). The goal of such a natural language understanding system may be to learn from the provided data to predict the correct output for new inputs.
[0024] Natural language understanding tasks may include BoolQ, CommitBank, CoPA, RTE, WiC, MRPC, WNLI, CNLI, CoLA. Using any auxiliary task can improve the performance of the target task in NLU tasks.
[0025] The updated candidate scale factor may indicate the benefit of the corresponding task to the training performance of the target task. Thus, the tasks can be weighted according to the impact of the tasks on the target task training.
[0026] The step of performing machine learning operations to evaluate the performance of the refined model parameter set may include: evaluating the performance of combinations of the refined model parameter sets. Thereby, determining the mixing weights can be more accurate than by heuristic methods.
[0027] According to another aspect, the present invention provides a method of training a machine learning model having an initial model parameter set using the respective training data sets of each of a plurality of tasks, the plurality of tasks including a target task for training the model and one or more auxiliary tasks, the method comprising: assigning at least one candidate scale factor to each of the plurality of tasks; for each of the one or more tasks of the plurality of tasks, performing one or more optimization steps on the corresponding task according to the corresponding scale factor to form a corresponding refined model parameter set; performing machine learning operations through one or more predefined evaluation criteria and based on one or more predefined constraints to evaluate the performance of the refined model parameter sets of the one or more tasks of the plurality of tasks in the target task through the one or more predefined evaluation criteria, thereby determining a set of mixing weights; updating the model parameter set of the machine learning model according to the refined model parameter set weighted by the set of mixing weights; and updating the candidate scale factor according to the set of mixing weights. Thus, in cases where model updates help to boost the performance of the target task, the method can use the model updates for auxiliary tasks.
[0028] According to another aspect, the present invention provides an apparatus for implementing a model trained by the above method. The apparatus can be used to implement models related to tasks including natural language processing. BRIEF DESCRIPTION OF THE DRAWINGS
[0029] The present invention will now be described by way of example with reference to the accompanying drawings.
[0030] In the drawings:
[0031] Figure 1 A multi-task model is schematically shown;
[0032] Figure 2 A single transformation block is schematically shown (provided by "The Illustrated Transformer", http: / / jalammar.github.io / illustrated-transformer / );
[0033] FIG. 3(a) schematically shows the model training in the training phase;
[0034] FIG. 3(b) schematically shows the update of the α variable in the training phase;
[0035] Figure 4 An exemplary algorithm for training a model is shown;
[0036] Figure 5 An apparatus for training a machine learning model is schematically shown, which includes a processor and a memory;
[0037] Figure 6 An example of a method for training a machine learning model is shown. DETAILED DESCRIPTION
[0038] The method described herein uses multi-task learning to train a model for a target task. The target task is one of the available tasks, which is regarded as the task whose performance needs to be optimized, and all other tasks are only regarded as auxiliary tasks, that is, their role in model update is only to improve the performance of the target task. Therefore, using any auxiliary task during training can provide the performance of the target task.
[0039] Given a set of classification tasks T = {t1, t2,..., t n}, each classification task is associated with training and validation data, a target task t * ∈ T and W = {ω1, ω2,..., ω n}, the parameters θ of a deep neural network are trained based on the data of all tasks in T, while minimizing the objective loss of the model on the target task development set, where, is the prediction of the model for the target task data instance:
[0040]
[0041] The cross - entropy loss of model training is used for all tasks in the task set T, and task - specific weights are used to augment the standard loss function:
[0042]
[0043] where y t is a one - hot vector encoding the correct label for an instance of task t ∈ T, is the classifier prediction for that instance, · denotes the inner product of the true label and the prediction, and ω t is the current weight associated with the task. While the ω weights in Equation (2) can take any value (preferably positive), it should be clear that in the case where ω t = 1, this loss is the same as standard unweighted multi - task training and the same as for single - task training (where ω t = 1^ω t′ = 0, ). In any case, using this means that the loss is not weighted.
[0044] Although in multi - task training each task can only naturally reduce its own loss actively, it is clear from Equation (1) that ultimately only the performance of the target task matters. Since the model parameters are shared among tasks, this objective is affected by the updates of the target task as well as the auxiliary tasks, so it is necessary to find the optimal interpolation of task - specific model updates related to the performance of the target task. Generally speaking, there are two possible outcomes when using auxiliary tasks in training. When the auxiliary task t′ is beneficial to the target task t * i.e., training with t′ also exactly pushes the model towards a lower dev - set loss for t * , it is very likely that the associated ω t′ needs to be maintained or increased. On the other hand, if it is proven that training with the auxiliary task is detrimental to the target task performance, its weight needs to be reduced or even set to zero, effectively eliminating t′ from training.
[0045] The method described in this paper provides a way to estimate the impact of tasks on the performance of the target task and to adjust the weights accordingly during model training.
[0046] The method uses a method based on updating the model to replace the traditional task weight estimation, where the model update is represented by the difference in the model parameters before and after one or more (preferably multiple) optimization steps (or optimization processes) (hereinafter referred to as the model "increment").
[0047] Rather than iteratively estimating the scaling weights through a heuristic loop of weight application, target task evaluation, and weight rescaling, the method described herein uses a meta-optimizer to estimate the optimal mixing ratio of task-specific model increments related to the target task performance, and uses the mixing ratio to scale the increments during training. However, based on the determined optimal mixing ratio of the increments, the task-specific weights are updated.
[0048] Furthermore, while the weights of the tasks can be restricted to positive weights, the mixing ratio for any given task can be negative, allowing the weights to become arbitrarily large or small, including negative values.
[0049] Additional parameters (hereinafter referred to as α variables) are introduced into the model training process, which are used to scale the model updates of individual tasks by finding the optimal mixing ratio related to the target task performance metric. These variables are associated with the task-specific model updates actually implemented.
[0050] These α variables can be "global" variables (applying exactly one mixing ratio for each task to all layers of the model), or "local" variables (where one variable is used for each task of each model layer).
[0051] In either case, the number of additional parameters introduced into the model is minimized and is typically negligible compared to the parameters of the original multi-task model. For example, in one implementation, the original model had a total of over 124 million parameters, while the method introduced an additional 8 (global, 8 tasks) or 1773 (local, 9 tasks x 197 weights and biases) α variables.
[0052] Although in some implementations, some additional data of the target task can be set aside to estimate the α variables, in practice, it is sufficient to estimate the α variables on a random subset of the training data or the development data, so there are no additional data requirements, such as having to specifically collect a new tuning set for the α variables.
[0053] As Figure 1As shown, one implementation of the training method described herein may use a pre-trained RoBERTa model as the input encoder 101, which has encoder layers 1 to n shown at 102 to 104 and task-specific classification heads 1 to t shown at 105 to 107 (i.e., each task has its own classification head), and includes a simple feed-forward layer at the top.
[0054] Roberta is based on BERT, which is a well-established Natural Language Processing (NLP) model built on a Transformer architecture, and the blocks of this model are as Figure 2 shown.
[0055] In Figure 2 it, the encoder block 201 includes a self-attention mechanism 202, a normalization layer 203, feed-forward layers 204, 205, and a second normalization layer 206.
[0056] In one example, the freely available encoder model Huggingface (https: / / huggingface.co / ), which is based on pre-trained RoBERTa, can be used. The RoBERTa-based model includes an embedder (3 embedding layers + 1 normalization layer) and 12 encoder blocks, and each encoder block includes a self-attention mechanism (3 linear layers), a self-output layer (1 linear layer + 1 normalization layer), an intermediate layer (1 linear layer), and an output module (1 linear layer + 1 normalization layer). Each linear layer and normalization layer comes with an additional bias term, and there are a total of 197 weights and biases.
[0057] At each embedding layer, linear layer, normalization layer, and each bias, a scale factor for each task (referred to here as the α variable) is injected into the model. When globally optimizing the task mixture using the α variable, all α variables for a given task are optimized to have the same value, while local optimization is for all individual layer and bias variables.
[0058] Thus, in some implementations, for each task, a single scale factor is applied to all layers of the model.
[0059] In other implementations, for at least one of the multiple tasks, different scale factors are applied to each of at least some of the multiple layers.
[0060] At the start of training, the values of all task weights are initialized to 1.
[0061] The training of the model includes two main phases that are alternated, as shown in FIGS. 3(a) and 3(b).
[0062] Figure 3(a) schematically shows the first stage of the training: the model training stage. Any number of tasks are sampled from the plurality of available tasks (i.e., the target task and at least one auxiliary task). Each task in one or more of the sampled tasks is used to independently update the underlying model using the current weights of the task. Thereafter, the differences in the module parameters (increments) are collected. Then, the model is reset.
[0063] Thus, in the first stage, after assigning at least one candidate scaling factor to each of the plurality of tasks (the target task and at least one auxiliary task) and adopting the model with its initial set of model parameters, one or more optimization steps are performed on one or more of the plurality of tasks using the corresponding training data set of the task according to the corresponding candidate scaling factor of the task to form a corresponding refined set of model parameters. The one or more optimization steps are performed independently, i.e., multiple tasks can be sampled, but the respective increments are based on the optimization steps for each individual task. Then, the model increment Δ = {δ1 + δ2 + δ3 …} can be determined as the difference between the refined set of model parameters and the initial set of model parameters (θ).
[0064] In Figure 3(a), there are five available tasks 301 to 305. Task 3 (303) is the target task, and the other tasks are auxiliary tasks. One or more optimization steps are performed independently on Task 1 (301) and Task 4 (304). For each, the optimization step 306 will produce updated model parameters. At 307, the difference between the previous model parameters and the new model parameters is saved in the form of the increment of the task. The model is reset and trained using the remaining tasks 306. Additionally, the increment of this task is also saved at 307.
[0065] In the second stage, as shown in Figure 3(b) (showing the α variable update stage), the increments previously collected from the first stage are used to find the best mixing ratio related to the model performance of the target task. Initially, at the start of step 2, all α variables are set to 1, indicating the use of all updates of a single task.
[0066] An optimization step is performed for the α variables and their impact on the performance of the target task (Task 3 (303)) using a "meta-optimizer", as shown at 308. After the update, the optimized α variables represent the best "mixing ratio" of the increments collected in the first stage. Each α can have any value. A value greater than 1.0 indicates that the corresponding increment should be increased, i.e., the corresponding task is more beneficial to the target task. A value below 1.0 indicates that the benefit of a certain task is smaller and its increment should be reduced. A negative α value indicates that a certain task is detrimental to the target task and its update should be performed in the opposite direction. These mixing weights are used to insert the collected model increments, which are then applied to the model. Based on the newly discovered best mixing ratio, the individual task weights are updated.
[0067] Therefore, in the second stage, machine learning operations are performed through one or more predetermined evaluation criteria and based on one or more predetermined constraints to evaluate the performance of the refined model parameter sets of one or more of the multiple tasks (for which optimization steps have been performed in the first stage) in the target task, thereby determining a set of mixing weights.
[0068] Then, the model parameter set of the machine learning model is updated according to the refined model parameter set weighted by the set of mixing weights (α variables), and the candidate scale factor is updated according to the set of mixing weights.
[0069] Then, the training will return to the first stage (Figure 3(a)), and this loop continues to iterate a predetermined number of times on the training data, or until the model performance of the target task converges. Then, the optimal model parameter set is selected and used as the final model parameter after training for inference.
[0070] In summary, if the auxiliary update helps improve the performance of the target task, it is best to use the auxiliary update; if the auxiliary update cannot help improve the performance of the target task, it is ignored. During training, the method starts with the current model parameters of each individual task and performs task-specific weighted model updates on a portion of the available training data. The method collects the generated model increments, i.e., the difference between the model parameters before and after the individual task updates, and resets the model. After this increment collection stage, the α variables are used as additional parameters that are optimized through gradient descent to find a good interpolation of the individual task model updates for the loss of the data for the target task.
[0071] The best increment interpolation found is used to update the model parameters, and the new interpolation parameters are used to update the task-specific weights.
[0072] Figure 4An exemplary algorithm for training the model is shown. As data, the algorithm employs a model M with parameters θ, a task set T = {t1, t2, ……, t n}, a target task t * , training data for each task t development data for the target task the maximum number of training epochs ε, the ratio ρ of task training data to inner-loop samples, and the number s of α-tuning steps. The result is an updated set of model parameters optimized for the performance of the target task t * .
[0073] Figure 5 FIG. 500 is a schematic diagram of a device 500 for performing the methods described herein. The device 500 may be implemented on devices such as laptop computers, tablet computers, smart phones, or televisions.
[0074] The device 500 includes a processor 501 that is configured to process the data set in the manner described herein. For example, the processor 501 may be implemented as a computer program running on a programmable device such as a GPU or a central processing unit (CPU). The device 500 includes a memory 502 that is configured to communicate with the processor 501. The processor 502 may be a non-volatile memory. The processor 501 may also include a cache ( Figure 5 not shown in FIG. 500) that may be used to temporarily store data from the memory 502. The device may include multiple processors and multiple memories. The memory may store data executable by the processor. The processor may be configured to operate according to a computer program stored on a machine-readable storage medium in a non-transitory form. The computer program may store instructions for causing the processor to perform its method in the manner described herein.
[0075] Devices such as 500 may also implement a model trained by the methods described herein.
[0076] Figure 6A flowchart is shown that summarizes an example method of training a machine learning model with an initial set of model parameters using a respective training dataset for each of a plurality of tasks, the plurality of tasks including a target task for which the model is to be trained and one or more auxiliary tasks. The method includes steps 601 to 605 as follows. In step 601, the method includes: assigning at least one candidate scaling factor for each of the plurality of tasks. In step 602, the method includes: for each of one or more of the plurality of tasks, performing one or more optimization steps on the respective task according to the respective candidate scaling factor to form a respective refined set of model parameters. In step 603, the method includes: performing a machine learning operation by one or more predetermined evaluation criteria and based on one or more predetermined constraints to evaluate the performance of the refined set of model parameters of the one or more of the plurality of tasks in the target task, thereby determining a set of mixing weights. In step 604, the method includes: updating the set of model parameters of the machine learning model according to the refined model parameters weighted by the set of mixing weights. In step 605, the method includes: updating the candidate scaling factor according to the set of mixing weights.
[0077] Thus, the machine learning system described herein directly utilizes the auxiliary training data and uses the scaling factor for each task in the form of the α variable assigned to the model layer. The model update is performed based on the task data, and the α parameter is used to scale the model update (the difference between the model parameters θ before and after the update). The increment of the model parameters is collected between Phase 1 and Phase 2, as shown in FIG. 3(a). Then, the α update is estimated in Phase 2.
[0078] Thus, the model parameter update (increment) is used to estimate the weights. In previous methods, the indirect weights were usually obtained through a gating mechanism or the like. Direct learning is usually performed using gradient weights, while the method described herein uses parameter updates.
[0079] As described above, injecting task-based α variables (scaling factors) into the model on a task-by-task or layer-by-layer basis only results in a small number of additional parameters for the model. For example, a RoBERTa-based model has approximately 124 million parameters. The number of additional parameters required using the method described herein is equal to the number of tasks, or |number of tasks| * (|weights| + |biases|).
[0080] In addition, the available training data can also be optimally used by taking random samples of the training set or development set as adjustment data.
[0081] One advantage of this solution is that it allows the use of complex optimizers during task training, rather than being limited to using stochastic gradient descent (SGD). Optimizers such as Adam or AdamW (see https: / / arxiv.org / abs / 1711.05101) have been shown to enable faster optimization and higher final model performance than SGD. Therefore, the multi-task training algorithm for the target task can use other optimizers in addition to the standard SGD, thereby improving the performance of the model.
[0082] Another advantage is that the method allows the use of negatively correlated tasks to improve the performance of the target task. In cases where the model update of the auxiliary task is detrimental to the target task, it is still possible to utilize its data by setting the mixing ratio of the task to a negative number, thereby generating updates that may be beneficial to the target task instead. In cases where an auxiliary task cannot be utilized at all, the method can still set the weight to 0 and ignore the task.
[0083] One implementation uses AdamW for model and head optimization and SGD with momentum for α variable optimization. In this implementation, model deltas are collected for 25% of any training dataset, and then a variable number of α optimization steps are performed using the collected model variables. These settings, as well as the optimizers actually used, are hyperparameters that need to be adjusted empirically. The training scheme itself makes no assumptions about these hyperparameters. Therefore, for example, we can use optimizers that are completely different from those recommended for the base model, head, or α optimization.
[0084] In some implementations, the methods described herein can be model-agnostic, task-agnostic, and optimizer-agnostic.
[0085] The method can be conveniently implemented as target-task-oriented multi-task learning for natural language understanding problems. The model can be used to allow a computer to interpret and use human language input. These problems include, but are not limited to, Natural Language Inference (NLI), entailment, or Word Sense Disambiguation (WSD). All of these problems come with their own datasets for training, development, and testing purposes, typically consisting of natural language input pairs (e.g., sentences or sentence pairs) and associated discrete labels (e.g., "yes" or "no"), where the label is one of a predefined closed set of possible labels (classification problem), or an associated continuous score, e.g., in the range from 0 to 5 (regression problem). The goal of all NLU systems is to learn from the provided data to predict the correct output for new inputs.
[0086] For example, natural language understanding tasks can include BoolQ, CommitBank, CoPA, RTE, WiC, MRPC, WNLI, CNLI, CoLA, which are all recognized tasks in the natural language understanding research community and are also represented in the GLUE and / or SuperGLUE NLU benchmarks. Using any auxiliary task can improve the performance of the target task in the NLU task.
[0087] The model can also be used for classification tasks in such fields. For example, among predefined possible categories, the model can be used to determine the correct category for a given natural language input.
[0088] For example, the model can be used in entailment or contradiction scenarios as follows:
[0089] Student in lecture hall / Student indoors → Entailment
[0090] Student in lecture hall / Student outdoors → Contradiction
[0091] Student in lecture hall / Student having a science class → Neutral
[0092] The model can also be used to determine reasonable continuations as follows:
[0093] The stain on the shirt is removed, [because] I bleached the shirt.
[0094] The stain on the shirt is removed, [because] I mended the shirt.
[0095] Therefore, the method can be used to train various tasks.
[0096] The applicant hereby separately discloses each individual feature described herein and any combination of two or more such features. With the ordinary knowledge of those skilled in the art, such features or combinations can be implemented as a whole based on this specification, regardless of whether such features or combinations of features can solve any problems disclosed herein, and are not limited to the scope of the claims. This application shows that various aspects of the present invention can be constituted by any such individual feature or combination of features. Given the foregoing description, various modifications within the scope of the present invention will be apparent to those skilled in the art.
Claims
1. An apparatus comprising one or more processors for training a machine learning model with an initial set of model parameters, characterized in that, The device is used to train the model using the respective training datasets of each of the multiple tasks, the multiple tasks including a target task for which the model is to be trained and one or more auxiliary tasks, and the device is used to train the model by performing the following steps: The model is used to perform natural language understanding tasks, and its input includes: a computer interpreting and using human language; Assign at least one candidate scaling factor to each of the multiple tasks; For each of the one or more tasks among the multiple tasks, perform one or more optimization steps on the respective task according to the respective candidate scaling factor to form a respective refined model parameter set; Perform machine learning operations through one or more predetermined evaluation criteria and based on one or more predetermined constraints to evaluate the performance of the refined model parameter sets of the one or more tasks among the multiple tasks in the target task, thereby determining a set of mixing weights; the set of mixing weights includes one or more of positive values, zero values, and negative values; Update the model parameter set of the machine learning model according to the refined model parameters weighted by the set of mixing weights; Update the candidate scaling factor according to the set of mixing weights.
2. The device according to claim 1, characterized in that Using the one or more auxiliary tasks during training is for optimizing the performance of the target task.
3. The device according to claim 1 or 2, characterized in that, The device is used to assign multiple candidate scaling factors to each of the multiple tasks.
4. The device according to claim 1 or 2, characterized in that, The device is also used to determine, for each of the multiple tasks, an optimal set of mixing weights related to the performance of the target task.
5. The device according to claim 4, characterized in that, The device is also used to adjust the at least one scaling factor based on the optimal set of mixing weights for each of the multiple tasks.
6. The device according to claim 1 or 2, characterized in that, The model architecture includes multiple layers, and the device is used to apply different scaling factors to each layer in at least some of the multiple layers for at least one of the multiple tasks.
7. The device according to claim 1 or 2, characterized in that, The model architecture includes multiple layers, and the device is used to apply a single scaling factor to all layers of the model for each of the multiple tasks.
8. The device according to claim 1 or 2, characterized in that, The device is used to assign a candidate scaling factor with a value of 1 to each of the multiple tasks at the start of the training.
9. The device according to claim 1 or 2, characterized in that, The model architecture includes an embedder and at least one encoder block, the embedder includes multiple layers, and each encoder block includes a self-attention mechanism and has at least one linear layer and at least one normalization layer.
10. The device according to claim 1 or 2, characterized in that, Randomly sample one or more of the multiple tasks from the multiple tasks.
11. The device according to claim 1 or 2, characterized in that, The target task is a predetermined task to be optimized through model training.
12. The device according to claim 1, wherein, The multiple tasks correspond to related problems in the field of human language understanding.
13. The device according to claim 1 or 2, characterized in that, The updated candidate scaling factor indicates the benefit of the respective task to the training performance of the target task.
14. The device according to claim 1 or 2, characterized in that, The step of performing machine learning operations to evaluate the performance of the refined model parameter sets includes: evaluating the combined performance of the refined model parameter sets.
15. A method for training a machine learning model with an initial set of model parameters using a respective training data set for each of a plurality of tasks, characterized in that, The multiple tasks include a target task for training the model and one or more auxiliary tasks, the model being for performing natural language understanding tasks, the input of which includes: a computer interpreting and using human language; the method includes: Assigning at least one candidate scaling factor for each of the multiple tasks; For each of one or more of the multiple tasks, performing one or more optimization steps on the corresponding task according to the corresponding scaling factor to form a corresponding refined model parameter set; Performing machine learning operations through one or more predetermined evaluation criteria and based on one or more predetermined constraints to evaluate the performance of the refined model parameter set of the one or more tasks in the multiple tasks in the target task through the one or more predetermined evaluation criteria, thereby determining a set of mixing weights; Updating the model parameter set of the machine learning model according to the refined model parameter set weighted by the set of mixing weights; Updating the candidate scaling factor according to the set of mixing weights.
16. An apparatus for training a machine learning model, characterized in that, It is for implementing a model trained by the method according to claim 15 above.
Citation Information
Patent Citations
Gradient normalization systems and methods for adaptive loss balancing in deep multitask networks
US20190130275A1