Multi-model joint denoising training
Through the multi-model joint denoising training method, the problem of insufficient training data in resource-scarce languages is solved, and the model performance of natural language comprehension tasks is improved.
Patent Information
- Application Number
- CN202110338761.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-03-30
- Publication Date
- 2025-05-06
- Estimated Expiration
- 2041-03-30
AI Technical Summary
In resource-scarce languages, the lack of reliable training data limits the performance of machine learning models in natural language understanding (NLU) tasks.
Through the multi-model joint denoising training method, multiple models are used to denoise the training samples, and these models are trained using the denoised training samples to improve the performance of the model.
This method effectively improves the quality of training samples, improves the NLU task performance of the model in resource scarce languages, and promotes the training of more robust models.
Smart Images

Figure CN115146654B_ABST
Abstract
Description
Background Art
[0001] Natural Language Understanding (NLU) is a technology that uses natural language to communicate with computers. It aims to enable computers to understand and use natural language to achieve communication between humans and computers, thereby replacing humans to perform various tasks related to natural language, such as spoken language understanding (SLU) tasks, machine reading comprehension (MRC) tasks, question answering (QA) tasks, etc. NLU tasks can be performed by trained machine learning models. The performance of machine learning models in performing NLU tasks depends on a large amount of reliable training data. For resource-rich languages such as English, there is large-scale human-annotated training data for some NLU tasks. Therefore, these NLU tasks have excellent performance on resource-rich languages. Summary of the invention
[0002] This Summary is provided to introduce a group of concepts that will be further described in the following Detailed Description. This Summary is not intended to identify key features or essential features of the claimed subject matter, nor is it intended to be used to limit the scope of the claimed subject matter.
[0003] The embodiments of the present disclosure provide a method and apparatus for multi-model joint denoising training. A plurality of models may be obtained. A set of training samples may be denoised by the plurality of models. The plurality of models may be trained using a set of denoised training samples.
[0004] It should be noted that one or more of the above aspects include the features specifically pointed out in the following detailed description and claims. The following description and drawings set forth in detail certain illustrative features of the one or more aspects. These features are merely indicative of the various ways in which the principles of various aspects may be implemented, and the present disclosure is intended to include all such aspects and their equivalents. BRIEF DESCRIPTION OF THE DRAWINGS
[0005] The disclosed aspects will be described below in conjunction with the accompanying drawings, which are provided to illustrate rather than limit the disclosed aspects.
[0006] Figure 1 An exemplary process for synthesizing and translating training samples and generating training samples according to an embodiment of the present disclosure is shown.
[0007] Figure 2 An exemplary process for obtaining multiple training corpora according to an embodiment of the present disclosure is shown.
[0008] Figure 3 An exemplary process for multi-model joint denoising training according to an embodiment of the present disclosure is shown.
[0009] Figure 4 Another exemplary process for multi-model joint denoising training according to an embodiment of the present disclosure is shown.
[0010] Figure 5 An exemplary process for performing denoising and training according to an embodiment of the present disclosure is shown.
[0011] Figure 6 Detailed description is a flowchart of an exemplary method for multi-model joint denoising training according to an embodiment of the present disclosure.
[0012] Figure 7 An exemplary device for multi-model joint denoising training according to an embodiment of the present disclosure is shown.
[0013] Figure 8 An exemplary device for multi-model joint denoising training according to an embodiment of the present disclosure is shown. DETAILED DESCRIPTION
[0014] The present disclosure will now be discussed with reference to several exemplary embodiments. It should be understood that the discussion of these embodiments is only used to enable those skilled in the art to better understand and thereby implement the embodiments of the present disclosure, and does not teach any limitation on the scope of the present disclosure.
[0015] It is desirable to extend NLU tasks such as SLU tasks, MRC tasks, QA tasks, etc. to resource-scarce languages, such as German, Spanish, French, etc. However, for resource-scarce languages, there are only a few or even no reliable training data, which restricts the performance of machine learning models when performing NLU tasks for resource-scarce languages. For specific NLU tasks, the training data of resource-rich languages can be used to augment the training data of resource-scarce languages. In this article, the resource-rich language can be referred to as the source language, and the resource-scarce language can be referred to as the target language. The training data may also be referred to as a training data set in this article, which may be composed of multiple training samples. A training sample may refer to a single training instance contained in a training data set. A training sample of a target language for a specific NLU task may be synthesized in a variety of ways, thereby augmenting the training data set of the target language for the NLU task. For example, the text in the training sample of the source language may be translated into the text of the target language by machine translation technology, and the annotations in the training sample of the source language may be mapped to the annotations on the target language side by alignment technology, so that the training sample of the target language may be obtained. This method for synthesizing training samples can be called a translation method. Alternatively, a training sample of the target language can be generated by a large-scale neural network, such as a generative adversarial network, a variational autoencoder, a pre-trained language model, etc. This method for synthesizing training samples can be called a generation method. However, whether the training samples synthesized by the translation method or the generation method often contain some wrong or inaccurate annotations, which results in poor quality of the synthesized training samples. The wrong or inaccurate annotations in the training samples can be considered as noise in the training samples. Training samples containing wrong annotations or inaccurate annotations can be considered as noisy training samples.
[0016] Some methods can be used to improve the quality of synthesized training samples. Improving the quality of training samples can also be considered as denoising the training samples. However, these methods only consider training samples synthesized in a single way, that is, either only consider training samples synthesized by translation, or only consider training samples synthesized by generation. For example, an attention mechanism can be used to achieve annotation alignment and recognition between the target language and the source language, thereby improving the quality of the translated training samples of the target language. This method only considers training samples synthesized by translation. In addition, before generating training samples of the target language through a language model, the language model can be optimized using a training data set of the target language obtained through machine translation technology to improve the ability of the language model to generate training samples, thereby improving the quality of the generated training samples of the target language. This method only considers training samples synthesized by generation.
[0017] The embodiment of the present disclosure proposes to denoise a set of training samples through multiple models, and the denoised set of training samples can be used to train the multiple models, and the trained multiple models can be further used to perform NLU tasks corresponding to the set of training samples. Since the denoising and training process of the embodiment of the present disclosure is jointly performed by multiple models, the method can also be called a multi-model joint denoising training method.
[0018] In one aspect, embodiments of the present disclosure propose a series of mechanisms for denoising a set of training samples, which can be executed during the training of a model. The denoising mechanism according to the embodiments of the present disclosure may include, for example, a collaborative training mechanism for selecting training samples for a current model from a set of training samples through other models in a plurality of models, a weight determination mechanism for determining the weights of training samples for calculating training losses through a plurality of models, an annotation update mechanism for updating the annotations of training samples through a plurality of models, and the like. The above mechanisms can effectively improve the quality of training samples, and thereby improve the performance of a plurality of models trained using such training samples.
[0019] On the other hand, the various denoising mechanisms proposed in the embodiments of the present disclosure may be applicable to training samples synthesized in a variety of ways, such as training samples synthesized by translation, training samples synthesized by generation, etc. In this article, a training sample whose language is a source language and contains reliable annotations may be referred to as a source training sample. In addition, a training sample synthesized by translation may be referred to as a translation training sample, and a training sample synthesized by generation may be referred to as a generation training sample. The various denoising mechanisms proposed in the embodiments of the present disclosure may be applicable to denoising a group of training samples including source training samples, translation training samples, generation training samples, etc. A group of denoised training samples may be used to jointly train multiple models. Training a model using a training data set including multiple training samples may help train a more robust model.
[0020] In another aspect, the various denoising mechanisms proposed in the embodiments of the present disclosure can be applied not only to a group of cross-language training samples, but also to a group of single-language training samples. A group of training samples composed of training samples in different languages can be referred to as a group of cross-language training samples, and a group of training samples composed of training samples in the same language can be referred to as a group of single-language training samples. For example, for a certain NLU task, the number of training samples in the source language may be insufficient. In this case, the training samples of the source language for the NLU task can also be synthesized by translation or generation, such as translation training samples of the source language, generation training samples of the source language, etc. Various denoising mechanisms according to the embodiments of the present disclosure can be applied to denoising a group of single-language training samples including source training samples, translation training samples of the source language, generation training samples of the source language, etc. It should be understood that although the foregoing discussion and the following discussion may involve an example of denoising a group of cross-language training samples, the embodiments of the present disclosure are not limited thereto, but can denoise a group of single-language training samples in a similar manner.
[0021] Figure 1 An exemplary process 100 for synthesizing translation training samples and generating training samples according to an embodiment of the present disclosure is shown. In process 100, translation training samples in a target language can be synthesized and training samples can be generated from source training samples. The process 100 is described below using an example in which the source language is English and the target language is Spanish.
[0022] The source training sample 102 may be, for example, a training sample for an SLU task. SLU is a key part in a task-oriented dialog system, which aims to parse user utterances into predefined semantic representations, such as intents, slot-value pairs, etc. Therefore, the training sample for the SLU task may include utterances and corresponding intent annotations and slot-value pair annotations. For example, the source training sample 102 may include an utterance "When will my next alarm start" and an intent annotation "get_alarm". In addition, the term "next" in the utterance is marked as "B-ordinal", which may indicate that the slot type of the text segment where the term "next" is located is "ordinal", where "B-" indicates that the term "next" is located at the beginning of the text segment. Accordingly, the slot-value pair annotation corresponding to the utterance in the source training sample 102 may be "ordinal=next", where the slot type is "ordinal", and the value corresponding to the slot type "ordinal" is the term "next".
[0023] At 110, the source training sample 102 may be translated into a translation training sample 112 through operations such as translation and alignment. In one embodiment, the text in the source training sample may be translated into the text in the target language through a known machine translation technology, and the annotations in the source training sample may be mapped to the annotations on the target language side through known alignment technologies such as attention weight, fastalign, GIZA++, etc., so that the translation training sample in the target language may be obtained. The translation training sample X corresponding to the utterance x may be defined by the following formula:
[0024] X=[I;(s 1 ,v 1 ),…,(s p ,v p );(x 1 ,…x L )] (1)
[0025] Among them, I is the intention label, is a slot-value pair annotation, and (x 1 ,…x L ) is the token sequence of utterance x, where s i is the slot type, v i Is with slot type s i The corresponding value, and v i is a term in the target language discourse, which can be obtained by aligning terms between the source language discourse and the target language discourse.
[0026] For example, the utterance "When will my next alarm start" in the source training sample 102 can be translated into the Spanish utterance "Cuando va a empezar mi siguientealarma" by known machine translation techniques, and the annotations in the source training sample 102, namely the intent annotation "get_alarm" and the slot-value pair annotation "ordinal=next", can be mapped into annotations on the Spanish side, such as the intent annotation "get_alarm" and the slot-value pair annotation "ordinal=empezar" by known alignment techniques. The utterance "Cuando va a empezar mi siguientealarma", the intent annotation "get_alarm" and the slot-value pair annotation "ordinal=empezar" can form the translation training sample 112.
[0027] In order to further enhance the diversity of the training samples of the target language, additional training samples of the target language can be synthesized by generative means. For example, the generated training samples of the target language can be generated based on the translation training samples of the target language. In one embodiment, the generated training samples of the target language can be generated by a pre-trained generative model. Preferably, before the training samples are generated by the generative model, the generative model can be optimized, such as fine-tuning. For example, a group of translation training samples, such as a group of translation training samples obtained by the above process, can be used to optimize the generative model. The translation training samples used to optimize the generative model can be defined as in formula (1) above. Preferably, when optimizing the generative model using the translation training samples, noise can also be injected into the translation training samples by applying text infilling to the translation training samples. For example, for each translation training sample, 35% of the words in the translation training sample can be shielded by randomly sampling a segment length according to a Poisson distribution.
[0028] The optimized generative model can generate generative training samples in the target language based on the translated training samples in the target language. Similar to the optimized generative model, when generating training samples through the generative model, noise can be injected into the translated training samples by applying text padding to the translated training samples. The translated training samples injected with noise can be provided as input data to the generative model to generate generative training samples in the target language. For example, at 120, noise can be injected into the translated training sample 112 by applying text padding to obtain input data 122. The input data 122 can include the intent annotation "get_alarm", the slot value pair annotation "ordinal=empezar", and the term sequence " <mask>start mynext <mask> <es>",in" <mask>" replaces the blocked term, " " is the sentence ending term, and " <es>" is the corresponding language identifier symbol.
[0029] The input data 122 may be provided to the generative model 130. The generative model 130 may be, for example, a model obtained by performing the above optimization process on a pre-trained language model, such as multilingual Bidirectional and Auto-Regressive Transformers (mBART). The generative model 130 may generate a plurality of candidate Spanish generative training samples based on the input data 122. For example, the generative training sample 132 may include an intent annotation "get_alarm", a slot value pair annotation "ordinal=empezar", and an utterance "Cuando va aempezar mi siguiente <es>". Training samples that do not contain accurate annotations can be filtered out from the generated multiple candidate generated training samples. For example, training samples that do not contain the intent annotation "get_alarm" and / or the slot value pair annotation "ordinal=empezar" included in the input data 122 can be filtered out from the multiple candidate generated training samples.
[0030] It should be understood that Figure 1 The process 100 in the figure is only an example of the process for synthesizing translation training samples and generating training samples. According to actual application requirements, the process for synthesizing translation training samples and generating training samples may include any other steps, and may include more or fewer steps. In addition, although the training samples of the target language are synthesized by the translation mode and the generation mode in the process 100, when the training samples of the source language are insufficient, the training samples of the source language can be synthesized by a similar process. For example, the training samples of the source language can be translated into training samples of other languages, and then the training samples of other languages can be reversely translated back to the training samples of the source language. Since different expressions are generated during the translation process, different training samples of the source language can be constructed. In addition, a small amount of training samples of the source language can be provided to the generation model. The generation model can generate additional training samples of the source language.
[0031] Through process 100, a translation training sample and a generation training sample are synthesized. The synthesized translation training sample and the generation training sample can be used to construct a translation training corpus and a generation training corpus, respectively. In this article, a training corpus constructed by a set of translation training samples can be referred to as a translation training corpus, and a training corpus constructed by a set of generation training samples can be referred to as a generation training corpus. Figure 2 An exemplary process 200 for obtaining multiple training corpora according to an embodiment of the present disclosure is shown. In process 200, at least one translation training corpus and at least one generation training corpus can be obtained through source training corpora. Herein, a corpus including a set of source training samples can be referred to as a source training corpus.
[0032] The source training corpus 210 may include a set of source training samples 210-1 to 210-F (F≥1). Each source training sample 210-f (1≤f≤F) may be translated into a corresponding translation training sample of the target language. For example, Figure 1 The translation and alignment operation 110 in the process is used to translate the source training samples into translation training samples.
[0033] For each source training sample, multiple machine translation technologies or multiple alignment technologies can be used, and accordingly, multiple translation training samples corresponding to the source training sample can be obtained. A group of translation training samples for the source training corpus 210 obtained using the same machine translation technology and the same alignment technology can be combined in a translation training corpus. As an example, M translation training corpora 220-1 to 220-M (M≥1) can be obtained, and each translation training corpus 220-m (1≤m≤M) can, for example, correspond to a specific machine translation technology and alignment technology. In addition, a group of translation training samples 220-m-1 to 220-mF included in the translation training corpus 220-m can correspond to a group of source training samples 210-1 to 210-F, respectively.
[0034] In addition, for each translation training corpus 220-m in the group of translation training corpus 220-1 to 220-M, a group of generated training corpus can be generated based on the translation training corpus. For example, a group of generated training corpus 230-1 to 230-N (N≥1) can be generated based on the translation training corpus 220-1. For each translation training sample, a generated training sample corresponding to the translation training sample can be generated by the generative model. For example, Figure 1 The injection noise operation 120 in the embodiment of the present invention and the use of the generative model 130 to generate generated training samples corresponding to the translated training samples.
[0035] For each translation training sample, a variety of generation models can be used, and accordingly, a plurality of generation training samples corresponding to the translation training sample can be obtained. A group of translation training samples obtained using a generation model for the same translation training corpus can be combined in a generation training corpus. As an example, for the translation training corpus 220-1, N generation training corpora 230-1 to 230-N (N≥1) can be obtained, and each generation training corpus 230-n (1≤n≤N) can, for example, correspond to a specific generation model. In addition, a group of generation training samples 230-n-1 to 230-nF included in the generation training corpus 230-n can correspond to a group of translation training samples 220-1-1 to 220-1-F, respectively.
[0036] It should be understood that Figure 2 The process 200 in is merely an example of a process for obtaining multiple training corpora. Depending on actual application requirements, the process for obtaining multiple training corpora may include any other steps, and may include more or fewer steps. In addition, although in process 200, for the sake of simplicity, the translation training samples in the translation training corpus and the generated training samples in the generated training corpus correspond one-to-one to the source training samples in the source training corpus, that is, using a specific translation technology and a specific alignment technology, a source training sample can be translated into a translation training sample, and a specific generation model can generate a generated training sample based on a translation training sample, but the embodiments of the present disclosure are not limited to this. According to actual application requirements, using a specific translation technology and a specific alignment technology, a source training sample can also be translated into multiple translation training samples. In addition, a specific generation model can generate multiple generated training samples based on a translation training sample. For example, as described above in combination Figure 1 As described, a generative model can generate multiple candidate generative training samples. After filtering out the training samples that do not contain accurate annotations from the multiple candidate generative training samples, several training samples can be randomly sampled from the remaining generative training samples to construct a generative training corpus for the generative model. In addition, it should be understood that the generative model can also generate one or more generative training samples based on the source training samples. Accordingly, the generative training corpus can be constructed directly from the source training corpus.
[0037] The following will take the SLU task as an example to illustrate the process of multi-model joint denoising training. The training samples for the SLU task can include utterances and corresponding intent annotations and slot-value pair annotations. Where L is the length of the term sequence, which can have a corresponding intent tag y I and slot type annotation sequence in is for the i-th term x i The slot type annotation of . The utterance x can be provided to the encoder To obtain the contextual hidden representation H of the utterance x. The above process can be expressed by the following formula:
[0038]
[0039] Where Θ is the encoder Parameters, h 0 is a sentence-level representation of the intent classification for utterance x, and h i (1≤i≤L) is the term-level representation of the slot filling for utterance x. The intent probability distribution p of utterance x can be obtained by applying a linear transformation and a softmax operation I (x; Θ) and slot type probability distribution The above process can be expressed by the following formula:
[0040] p I (x; Θ) = softmax(W I ·h 0 +b I ) (3)
[0041]
[0042] in, C I is the set of intention annotations, C S is a set of slot type annotations based on the BIO tag mode. and is the output matrix, and b I and b S It is bias.
[0043] Figure 3 An exemplary process 300 for multi-model joint denoising training according to an embodiment of the present disclosure is shown. In process 300, multiple models can be initialized, a set of training samples can be jointly denoised by the initialized multiple models, and the multiple models can be jointly trained using the denoised set of training samples.
[0044] At 302, multiple training corpora may be obtained. For example, two or more of a source training corpus, at least one translation training corpus, and at least one generated training corpus may be obtained. The training samples in the multiple training corpora may be based on the same language or different languages. For example, Figure 2 The process 200 in the embodiment is used to obtain a plurality of training corpora.
[0045] At 304, multiple models may be obtained. For example, K (K ≥ 1) models may be obtained. The multiple models may be multiple models having the same structure for performing a specific NLU task.
[0046] Preferably, before denoising a set of training samples through the multiple models, the multiple models may be preliminarily trained. In this article, the preliminary training process may also be referred to as an initialization process. At 306, multiple initialization data groups may be formed based on multiple training corpora. In this article, the training data used to initialize the model may be referred to as an initialization data group. For example, multiple initialization data groups may be formed based on the source training corpus obtained at 302, at least one translation training corpus, and at least one generated training corpus. The number of initialization data groups may be consistent with the number of models. For example, in a case where there are 2 models, namely, model and Model In the case of, two initialization data groups can be formed based on multiple training corpora. In addition, preferably, since the source training corpus is a training corpus containing reliable annotations, the source training corpus can be included in each initialization data group. That is, each initialization data group can include at least the source training corpus. As an example, an initialization data group can be formed Initialize data group etc., among which It can be the source training corpus, can be one of at least one translation training corpus, and It can be one of at least one generated training corpus. In the case of having multiple translation training corpora or multiple generated training corpora, different translation training corpora or different generated training corpora can be used to form multiple different initialization data sets.
[0047] At 308, the plurality of models may be respectively initialized using the plurality of initialization data sets. For example, the plurality of models may be initialized using the initialization data sets. To initialize the model As an example, you can use the initialization data set To initialize the model And use the initialization data set To initialize the model Assume that each training corpus is a training corpus for the SLU task. In one embodiment, the model can be initialized by minimizing the cross entropy loss as shown in the following formula:
[0048]
[0049] Among them, x is the utterance in the training sample, y I is the intention annotation of utterance x, is the slot type annotation of the i-th token in utterance x, and p I (x; Θ k )and The models are Get the predicted probability distribution of intent and the predicted probability distribution of slot type.
[0050] Optionally, multiple rounds of initialization can be performed on multiple models by performing steps 306 and 308 multiple times. At 310, it can be determined whether a predetermined number of initialization rounds has been reached. If it is determined at 310 that the predetermined number of initialization rounds has not been reached, the process 300 can return to step 306, where multiple initialization data groups can be re-formed based on multiple training corpora, and at 308, the multiple models obtained in the previous round are initialized using the re-formed multiple initialization data groups. The re-formed multiple initialization data groups can be the same as or different from the multiple initialization data groups formed in the previous round.
[0051] If it is determined at step 310 that the predetermined number of initialization rounds has been reached, process 300 may proceed to step 312, where the plurality of training corpora obtained at step 302 may be combined into a training data set. Translation training corpus Generate training corpus Combined into a training data set
[0052] The training data set can be used to jointly denoise multiple models. For example, a set of training samples can be denoised by multiple models, and the denoised set of training samples can be used to train multiple models. A set of training samples can be further denoised by the trained multiple models. In one embodiment, the training data set can be provided to multiple models as a whole. In this case, a set of training samples can correspond to the training data set. Steps 314 to 318 in process 300 illustrate an exemplary process using this embodiment.
[0053] At step 314 in process 300, the training data set may be denoised by multiple models. For example, for each model in the multiple models, a training sample for the model may be selected from the training data set by other models in the multiple models. Alternatively or additionally, for each training sample in the training data set, a weight for calculating the training loss may be determined by multiple models. Figure 5 An exemplary process for selecting training samples and determining weights is described below.
[0054] At 316, a plurality of models may be trained using the denoised training data set. For example, for each of the plurality of models, the model may be trained using the training sample selected at 314. For example, for each of the selected training samples, the weight for calculating the training loss of the training sample determined at 314 may be applied to the model training. Figure 5 To illustrate an example process for training multiple models.
[0055] At 318, the training data set may be further denoised using the trained multiple models. For example, the labels of one or more training samples in the training data set may be updated using the trained multiple models. The updated labels may be used for the next round of denoising and training. Figure 5 An exemplary process for further denoising the training data set is illustrated.
[0056] According to an embodiment of the present disclosure, optionally, multiple rounds of joint denoising training can be performed on multiple models by performing steps 314 to 318 multiple times. At 320, it can be determined whether a predetermined number of training rounds has been reached. If it is determined at 320 that the predetermined number of training rounds has not been reached, the process 300 can return to step 314 and perform steps 314 to 318 again. If it is determined at 320 that the predetermined number of training rounds has been reached, the process 300 can proceed to step 322, and at 322, the process 300 can end.
[0057] Figure 4 Another exemplary process 400 for multi-model joint denoising training according to an embodiment of the present disclosure is shown. Steps 402 to 412 in process 400 may correspond to Figure 3 Steps 302 to 312 in process 300 in . Through steps 402 to 412, multiple models can be initialized, and a training data set composed of multiple training corpora can be obtained. The training data set can be used to perform joint denoising training on multiple models. For example, a set of training samples can be denoised by multiple models, and a set of denoised training samples can be used to train multiple models. In one embodiment, the training data set can be divided into multiple training data subsets. In this case, a set of training samples can correspond to one training data subset among multiple training data subsets in the training data set, and the multiple training data subsets can be used to iteratively perform denoising and training processes.
[0058] At 414, the training data set can be partitioned into a plurality of training data subsets.
[0059] The denoising and training process can be iteratively performed for multiple training data subsets in the training data set. In each iteration, a training data subset can be denoised by multiple models, and multiple models can be trained using the denoised training data subsets.
[0060] At 416, a training data subset may be denoised by multiple models. For example, for each model in the multiple models, a training sample for the model may be selected from the training data subset by other models in the multiple models. Alternatively or additionally, for each training sample in the training data subset, a weight for calculating the training loss may be determined by multiple models. Figure 5 An exemplary process for selecting training samples and determining weights is described below.
[0061] At 418, a plurality of models may be trained using the denoised subset of training data. For example, for each of the plurality of models, the model may be trained using the training sample selected at 416. For example, for each of the selected training samples, the weight for calculating the training loss for the training sample determined at 416 may be applied to the model training. Figure 5 To illustrate an example process for training multiple models.
[0062] At 420, the training data subset may be further denoised using the trained multiple models. For example, the labels of one or more training samples in the training data subset may be updated using the trained multiple models. The updated labels may be used for the next round of denoising and training. Figure 5 2 to illustrate an exemplary process for further denoising a subset of the training data.
[0063] Optionally, steps 416 to 420 may be iteratively performed for multiple training data subsets in the training data set. At 422, it may be determined whether all training data subsets in the training data set have been traversed. If it is determined at 422 that all training data subsets in the training data set have not been traversed, the process 400 returns to step 416 and performs steps 416 to 420 for the next training data subset. If it is determined at 422 that all training data subsets in the training data set have been traversed, the process 400 may proceed to step 424.
[0064] According to an embodiment of the present disclosure, optionally, multiple rounds of joint denoising training can be performed on multiple models by performing steps 414 to 422 multiple times. At 424, it can be determined whether a predetermined number of training rounds has been reached. If it is determined at 424 that the predetermined number of training rounds has not been reached, the process 400 can return to step 414 and perform steps 414 to 422 again. In this case, at 414, the training data set can be re-divided into multiple training data subsets, and the denoising and training process at steps 416 to 420 can be re-iteratively performed based on the re-divided multiple training data subsets.
[0065] If it is determined at 424 that the predetermined number of training rounds has been reached, process 400 may proceed to step 426 where process 400 may end.
[0066] pass Figure 3 Process 300 or Figure 4 The process 400 in the embodiment can perform denoising training on multiple models. The multiple models that have been trained for denoising can be used to perform the same NLU task as the NLU task for which the training data set is targeted. In one embodiment, one model can be selected from the multiple models to perform the NLU task. In another embodiment, the multiple models can be combined into a model set to perform the NLU task. For example, multiple prediction results obtained by the multiple models based on the input data can be obtained, and the prediction result with the largest number of occurrences can be selected as the final prediction result for the input data.
[0067] It should be understood that Figure 3 Process 300 and Figure 4 The process 400 in is merely an example of a process for multi-model joint denoising training. Depending on the actual application requirements, the process for multi-model joint denoising training may include any other steps, and may include more or fewer steps. For example, the process of initializing multiple models can be omitted from process 300 or 400, so that a set of training samples can be denoised directly through multiple models. In addition, it should be understood that although the foregoing discussion and the following discussion may involve denoising training samples for SLU tasks, and accordingly, multiple models trained using such training samples can be used to perform SLU tasks, the embodiments of the present disclosure are not limited thereto, but can be denoised in a similar manner for training samples for other NLU tasks, such as MRC tasks, QA tasks, etc., and the trained multiple models can be used for corresponding other NLU tasks.
[0068] Figure 5 An exemplary process 500 for performing denoising and training according to an embodiment of the present disclosure is shown. In process 500, a set of training samples may be denoised by multiple models, and multiple models may be trained using the denoised set of training samples. Process 500 may correspond to Figure 3 Steps 314 to 318 of , or Figure 4 In step 416 to 420 of the process 500 corresponds to Figure 3 In the case of steps 314 to 318 in the process 500, the process 500 can be performed using a training data set, that is, a set of training samples in the process 500 can correspond to the training data set. Figure 4 In the case of steps 416 to 420 in the process 500, one of the multiple training data subsets in the training data set may be used to perform the process 500, that is, a group of training samples in the process 500 may correspond to one training data subset.
[0069] First, for each model in the plurality of models, preferably, the training samples for the model can be selected from a set of training samples by other models in the plurality of models. The noise in the training corpus obtained by one method is usually independent of the noise in the training corpus obtained by another method. For example, the noise in the training corpus obtained by translation is usually independent of the noise in the training corpus obtained by generation. In addition, as mentioned above, Figure 3 or Figure 4 As described, multiple initialization data sets can be formed based on multiple training corpora, and multiple models can be initialized using the multiple initialization data sets. Through the initialization process, each model can learn corresponding knowledge from the corresponding training corpora, so that more accurate prediction results can be obtained when making predictions based on the corresponding training samples. The training corpora contained in the multiple initialization training data sets used to initialize the multiple models are different from each other, so the initialized multiple models can have better performance in different training corpora. For example, assuming that the model Is to use the initial training data set To initialize, the model Based on the source training corpus Or translation training corpus When predicting the training samples in the model, more accurate prediction results are obtained; Is to use the initial training data set To initialize, the model Based on the source training corpus Translation training corpus Or generate training corpus The more accurate the prediction results are, the more accurate the prediction results are. Each model can make predictions based on each training sample in a set of training samples. The training samples with more accurate prediction results can be provided to other models.
[0070] For example, at 510, for one of the multiple models, a predetermined proportion of training samples can be filtered out from a set of training samples by other models, and at 520, the remaining training samples in the set of training samples can be determined as training samples for the model. Generally speaking, training samples with smaller prediction losses may have more accurate annotations. Therefore, for the model, a predetermined proportion of training samples with larger prediction losses obtained at other models can be filtered out from a set of training samples. The training sample selection process can be expressed, for example, by the following formula:
[0071]
[0072] in, represents a set of training samples, Represents a set of training samples Select the model for training samples, δ represents the predetermined filtering ratio, and Represents other models and a set of training samples The corresponding prediction loss.
[0073] Training samples for each model can be selected through steps 510 and 520. Since the training samples for each model are selected through other models, this mechanism can also be called a collaborative training mechanism.
[0074] After selecting a training sample for each model from a set of training samples, preferably, for each training sample in a set of training samples, a weight for calculating the training loss of the training sample can be determined by multiple models. The disclosed embodiment proposes a weight determination mechanism. In one embodiment, the weight of the training sample can be determined according to the consistency of multiple prediction results obtained by multiple models based on the training sample. If the multiple prediction results obtained by multiple models based on the training sample are inconsistent, the training sample is likely to be noisy. For example, at 530, multiple prediction results obtained by multiple models based on the training sample can be obtained, and at 540, the weight of the training sample for calculating the training loss can be determined according to the consistency of the multiple prediction results. The consistency of multiple prediction results can be associated with the uncertainty of multiple prediction results. The lower the consistency of multiple prediction results for a specific training sample, the greater the disagreement between the multiple prediction results, the higher the uncertainty, and the more likely the training sample is to be noisy, so its weight should be lower.
[0075] The uncertainty of the prediction result can be expressed as u, which can be defined, for example, by the following formula:
[0076]
[0077] in:
[0078]
[0079] At 550, the plurality of models may be trained using the corresponding training samples. For example, for each of the plurality of models, the model may be trained using the training samples selected at 510 and 520. For example, for each of the selected training samples, the weights for calculating the training loss of the training sample determined at 530 and 540 may be applied to the model training. The training loss corresponding to the training sample including the utterance x can be calculated as follows:
[0080]
[0081] Where w = e -u , is the current intention annotation of utterance x, is the current slot type annotation of the i-th token in utterance x, and p I (x; Θ k )and The models are The predicted probability distribution of intent and slot type obtained based on utterance x.
[0082] For each model, the model can be trained by minimizing the total training loss for the model corresponding to a set of training samples.
[0083] After training multiple models, preferably, a set of training samples can be further denoised by multiple models. For example, at 560, the annotations of one or more training samples in the set of training samples can be updated by the trained multiple models. is a training corpus containing reliable annotations, so it comes from the source training corpus The annotations of the training samples can remain unchanged, and only the annotations from the translation training corpus are updated. and / or generate training corpus In one embodiment, for each training sample in one or more training samples whose annotations are to be updated, multiple prediction results obtained by multiple trained models based on the training sample can be obtained, and the annotation of the training sample can be updated based on the multiple prediction results. The above process can be expressed by the following formula:
[0084]
[0085] The updated annotations can be used for the next round of denoising and training. In addition, annotations can be updated in a variety of ways, such as modifying the intent annotation of the utterance, modifying the slot type annotation of the text segment, modifying the BIO boundary of the slot, modifying both the slot type annotation and the BIO boundary, etc.
[0086] The above-mentioned collaborative training mechanism, weight determination mechanism and annotation update mechanism can effectively improve the quality of training samples, and thus improve the performance of multiple models trained using such training samples. Figure 5 The process 500 in is merely an example of a process for performing denoising and training. Depending on the actual application requirements, the process for performing denoising and training may include any other steps, and may include more or fewer steps. For example, although the above description uses a collaborative training mechanism, a weight determination mechanism, and a label update mechanism, in some embodiments, any one or two of these three mechanisms may be omitted from process 500. For example, without performing a collaborative training mechanism, all training samples in a set of training samples may be used to train each model. Alternatively, without performing a weight determination mechanism, all training samples may have equal weights when calculating training losses. In addition, without adopting a label update mechanism, each round of denoising and training process may be performed based on the same label.
[0087] Figure 6 is a flowchart of an exemplary method 600 for multi-model joint denoising training according to an embodiment of the present disclosure.
[0088] At 610, a plurality of models can be obtained.
[0089] At 620, a set of training samples may be denoised by the plurality of models.
[0090] At 630, the plurality of models may be trained using the denoised set of training samples.
[0091] In one embodiment, the set of training samples may be from a training data set. The training data set may include two or more of a source training corpus, at least one translation training corpus, and at least one generated training corpus.
[0092] In one embodiment, the set of training samples may be based on the same language or different languages.
[0093] In one embodiment, the method 600 may further include: initializing the multiple models using multiple initialization data sets. The multiple initialization data sets may be formed by at least one training corpus in the training data set including the set of training samples.
[0094] In one implementation, the set of training samples may correspond to a training data set.
[0095] In one embodiment, the set of training samples may correspond to one of a plurality of training data subsets in a training data set. The plurality of training data subsets may be used to iteratively perform the denoising and the training.
[0096] In each iteration, the denoising may include denoising a subset of the training data through the multiple models, and the training may include training the multiple models using the denoised subset of the training data.
[0097] In one embodiment, denoising a set of training samples may include: for each model in the multiple models, selecting training samples for the model from the set of training samples through other models in the multiple models.
[0098] The selecting the training samples for the model may include: filtering out a predetermined proportion of training samples from the set of training samples by using the other model; and determining the remaining training samples in the set of training samples as the training samples for the model.
[0099] In one implementation, denoising a set of training samples may include: for each training sample in the set of training samples, determining, by using the multiple models, a weight of the training sample for calculating a training loss.
[0100] Determining the weight of the training sample may include: determining the weight according to consistency of multiple prediction results obtained by the multiple models based on the training sample.
[0101] In one embodiment, the method 600 may further include: further denoising the set of training samples using multiple trained models.
[0102] Further denoising the set of training samples may include: updating the annotations of one or more training samples in the set of training samples by using multiple trained models.
[0103] The updating of the annotation may include, for each of the one or more training samples: obtaining a plurality of prediction results obtained by a plurality of trained models based on the training sample; and updating the annotation of the training sample based on the plurality of prediction results.
[0104] The one or more training samples may come from translation training corpus and / or generation training corpus.
[0105] In one embodiment, the set of training samples may correspond to a training data set. The training data set may be used to perform the denoising and the training in multiple rounds.
[0106] In one embodiment, the set of training samples may be from a training data set. The training data set may be used to perform the denoising and the training in multiple rounds. In each round, the set of training samples may correspond to one of multiple training data subsets in the training data set, and the multiple training data subsets may be used to iteratively perform the denoising and the training.
[0107] It should be understood that method 600 may also include any steps / processes for multi-model joint denoising training according to the above-mentioned embodiments of the present disclosure.
[0108] Figure 7 An exemplary apparatus 700 for multi-model joint denoising training according to an embodiment of the present disclosure is shown.
[0109] The apparatus 700 may include: a model acquisition module 710 for obtaining a plurality of models; a training sample denoising module 720 for denoising a set of training samples by using the plurality of models; and a model training module 730 for training the plurality of models using the denoised set of training samples. In addition, the apparatus 700 may also include any other module configured for multi-model joint denoising training according to the above-mentioned embodiments of the present disclosure.
[0110] Figure 8 An exemplary apparatus 800 for multi-model joint denoising training according to an embodiment of the present disclosure is shown. The apparatus 800 may include: at least one processor 810; and a memory 820 storing computer executable instructions. When the computer executable instructions are executed, the at least one processor 810 may: obtain multiple models, denoise a set of training samples using the multiple models, and train the multiple models using the denoised set of training samples.
[0111] In one embodiment, when the computer executable instructions are executed, they may also cause the at least one processor 810 to initialize the multiple models using multiple initialization data groups, where the multiple initialization data groups are formed by at least one training corpus in a training data set that includes the set of training samples.
[0112] It should be understood that the processor 810 may also execute any other steps / processes of the method for multi-model joint denoising training according to the above-mentioned embodiments of the present disclosure.
[0113] The embodiments of the present disclosure provide a computer program product for multi-model joint denoising training, including a computer program, which is executed by at least one processor to: obtain multiple models; denoise a set of training samples by the multiple models; and train the multiple models using the denoised set of training samples. In addition, the computer program can also be executed to implement any other steps / processes of the method for multi-model joint denoising training according to the above-mentioned embodiments of the present disclosure.
[0114] Embodiments of the present disclosure may be embodied in a non-transitory computer-readable medium. The non-transitory computer-readable medium may include instructions that, when executed, cause one or more processors to perform any operation of the method for multi-model joint denoising training according to the embodiments of the present disclosure as described above.
[0115] It should be understood that all operations in the above-described methods are merely exemplary, and the present disclosure is not limited to any operations or the order of these operations in the methods, but should cover all other equivalent transformations under the same or similar concept. In addition, unless otherwise specified or it is clear from the context that it is directed to a singular form, the articles "a" and "an" as used in this specification and the appended claims should generally be interpreted as meaning "one" or "one or more".
[0116] It should also be appreciated that all modules in the above described device can be implemented in various ways. These modules can be implemented as hardware, software, or a combination thereof. In addition, any module in these modules can be further divided into submodules or combined together in function.
[0117] Processors have been described in conjunction with various devices and methods. These processors can be implemented using electronic hardware, computer software or any combination thereof. Whether these processors are implemented as hardware or software will depend on specific application and the overall design constraints imposed on the system. As an example, the processor provided in the present disclosure, any part of the processor or any combination of processors can be realized using a microprocessor, a microcontroller, a digital signal processor (DSP), a field programmable gate array (FPGA), a programmable logic device (PLD), a state machine, a gated logic unit, a discrete hardware circuit, and other suitable processing components configured to perform the various functions described in the present disclosure. The function of the processor provided in the present disclosure, any part of the processor or any combination of processors can be realized using software executed by a microprocessor, a microcontroller, a DSP or other suitable platforms.
[0118] Software should be broadly considered to mean instructions, instruction sets, codes, code segments, program codes, programs, subroutines, software modules, applications, software applications, software packages, routines, subroutines, objects, running threads, processes, functions, etc. Software can reside in a computer-readable medium. A computer-readable medium can include, for example, a memory, which can be, for example, a magnetic storage device (e.g., a hard disk, a floppy disk, a magnetic stripe), an optical disk, a smart card, a flash memory device, a random access memory (RAM), a read-only memory (ROM), a programmable ROM (PROM), an erasable PROM (EPROM), an electrically erasable PROM (EEPROM), a register, or a removable disk. Although the memory is shown as being separated from the processor in the multiple aspects provided in the present disclosure, the memory can also be located inside the processor, such as a cache or a register.
[0119] The above description is provided to enable any person skilled in the art to practice the various aspects described herein. Various modifications to these aspects will be apparent to those skilled in the art, and the general principles defined herein may be applied to other aspects. Therefore, the claims are not intended to be limited to the aspects shown herein. All structural and functional equivalents to the elements of the various aspects described in this disclosure that are known or to be known to those of ordinary skill in the art are expressly incorporated herein and covered by the claims.< / es> < / es> < / mask> < / es> < / mask> < / mask>
Claims
1. A method for multi-model joint denoising training, comprising: Obtaining a plurality of training corpora, wherein the plurality of training corpora include two or more of a source training corpus, at least one translation training corpus, and at least one generated training corpus; Get multiple models; Forming a plurality of initialization data sets based on the plurality of training corpora; Initializing the multiple models using the multiple initialization data groups, wherein the multiple initialization data groups are formed by at least one training corpus, and the training corpora included in the multiple initialization data groups used to initialize the multiple models are different from each other; Determining whether a predetermined number of initialization rounds has been reached; In response to determining that the predetermined number of initialization rounds has been reached, combining the plurality of training corpora into a training data set; Denoising a set of training samples by the multiple models, wherein the set of training samples is from the training data set, and wherein the set of training samples includes utterances, intent annotations corresponding to the utterances, and slot-value pair annotations corresponding to terms in the utterances; and Training the multiple models using a set of denoised training samples; The training data set is further denoised by the trained multiple models, The denoising of a set of training samples comprises, for each training sample in the set of training samples: Determining weights of the training samples for calculating training losses by the multiple models, and wherein determining the weights of the training samples comprises: determining the weights according to the consistency of multiple prediction results obtained by the multiple models based on the training samples; and The labels of one or more training samples in the set of training samples are updated by the trained multiple models, and wherein the updating of the labels comprises: obtaining multiple prediction results obtained by the trained multiple models based on the training samples; and updating the labels of the training samples based on the multiple prediction results.
2. The method according to claim 1, wherein: The set of training samples are based on the same language or different languages.
3. The method according to claim 1, wherein: The set of training samples corresponds to a training data set.
4. The method according to claim 1, wherein: The set of training samples corresponds to one training data subset among a plurality of training data subsets in a training data set, and the plurality of training data subsets are used to iteratively perform the denoising and the training.
5. The method according to claim 4, wherein: In each iteration: The denoising comprises: denoising a training data subset using the multiple models, and The training includes training the multiple models using a denoised subset of training data.
6. The method according to any one of claims 1, 3 and 4, wherein: Denoising a set of training samples includes: For each model in the plurality of models, training samples for the model are selected from the set of training samples by other models in the plurality of models.
7. The method according to claim 6, wherein: The selecting of training samples for the model comprises: Filtering a predetermined proportion of training samples from the set of training samples by using the other model; and The remaining training samples in the set of training samples are determined as training samples for the model.
8. The method according to claim 1, wherein: The one or more training samples are from translation training corpus and / or generation training corpus.
9. The method according to claim 1, wherein: The set of training samples corresponds to a training data set, and the training data set is used to perform the denoising and the training in multiple rounds.
10. The method according to claim 1, wherein: The training data set is used to perform the denoising and the training in multiple rounds, and In each round, the set of training samples corresponds to one training data subset among a plurality of training data subsets in the training data set, and the plurality of training data subsets are used to iteratively perform the denoising and the training.
11. A device for multi-model joint denoising training, comprising: at least one processor; as well as a memory storing computer executable instructions that, when executed, cause the at least one processor to: Obtaining a plurality of training corpora, wherein the plurality of training corpora include two or more of a source training corpus, at least one translation training corpus, and at least one generated training corpus; Get multiple models, Forming a plurality of initialization data sets based on the plurality of training corpora; Initializing the multiple models using the multiple initialization data groups, wherein the multiple initialization data groups are formed by at least one training corpus, and the training corpora included in the multiple initialization data groups used to initialize the multiple models are different from each other; Determining whether a predetermined number of initialization rounds has been reached; In response to determining that the predetermined number of initialization rounds has been reached, combining the plurality of training corpora into a training data set; Denoising a set of training samples by using the multiple models, wherein the set of training samples is from the training data set, and wherein the set of training samples includes utterances, intent annotations corresponding to the utterances, and slot-value pair annotations corresponding to terms in the utterances, and Training the multiple models using a set of denoised training samples; The training data set is further denoised by the trained multiple models, The denoising of a set of training samples comprises, for each training sample in the set of training samples: Determining weights of the training samples for calculating training losses by the multiple models, and wherein determining the weights of the training samples comprises: determining the weights according to the consistency of multiple prediction results obtained by the multiple models based on the training samples; and The labels of one or more training samples in the set of training samples are updated by the trained multiple models, and wherein the updating of the labels comprises: obtaining multiple prediction results obtained by the trained multiple models based on the training samples; and updating the labels of the training samples based on the multiple prediction results.
12. A computer program product for multi-model joint denoising training, comprising a computer program, the computer program being executed by at least one processor for: Get multiple training corpora, among which, The plurality of training corpora include two or more of a source training corpus, at least one translation training corpus, and at least one generated training corpus; Get multiple models; Forming a plurality of initialization data sets based on the plurality of training corpora; Initializing the multiple models using the multiple initialization data groups, wherein the multiple initialization data groups are formed by at least one training corpus, and the training corpora included in the multiple initialization data groups used to initialize the multiple models are different from each other; Determining whether a predetermined number of initialization rounds has been reached; In response to determining that the predetermined number of initialization rounds has been reached, combining the plurality of training corpora into a training data set; Denoising a set of training samples by the multiple models, wherein the set of training samples is from the training data set, and wherein the set of training samples includes utterances, intent annotations corresponding to the utterances, and slot-value pair annotations corresponding to terms in the utterances; and Training the multiple models using a set of denoised training samples; The training data set is further denoised by the trained multiple models, The denoising of a set of training samples comprises, for each training sample in the set of training samples: Determining weights of the training samples for calculating training losses by the multiple models, and wherein determining the weights of the training samples comprises: determining the weights according to the consistency of multiple prediction results obtained by the multiple models based on the training samples; and The labels of one or more training samples in the set of training samples are updated by the trained multiple models, and wherein the updating of the labels comprises: obtaining multiple prediction results obtained by the trained multiple models based on the training samples; and updating the labels of the training samples based on the multiple prediction results.
Citation Information
Patent Citations
Face expression recognition method based on sample weight distribution and deep learning
CN110276248A
Machine learning model training method and device and sample processing method and device
CN111340233A
Machine translation model obtaining method and device, text translation method and device and storage medium
CN111859994A