Non-autoregressive machine translation method based on source end-target end common mask

By introducing cross-attention-based uncovered source-side word mask-prediction task and semantic unit mask based on mutual information in the machine translation model, combined with the random mask strategy, the problem of missing translation in the mask-prediction non-autoregressive machine translation method is solved, and more accurate and efficient translation results are achieved.

CN120012787APending Publication Date: 2025-05-16SHANGHAI JIAOTONG UNIV
View PDF 0 Cites 3 Cited by

Patent Information

Application Number
CN202311533492.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2023-11-16
Publication Date
2025-05-16

AI Technical Summary

Technical Problem

The existing mask-predicted non-autoregressive machine translation methods have problems with missing translation, resulting from insufficient modeling of word information at the source and insufficient random mask at the target.

Method used

A non-autoregressive machine translation method based on source-target-side common mask is adopted. By introducing a cross-attention-based uncovered source-side word mask-prediction task in the encoder module, and using a semantic unit mask based on mutual information in the decoder module, combined with a random mask strategy, a hybrid mask strategy is formed.

Benefits of technology

It effectively reduces the situation of translation missing, enhances the pertinence of source-end word representation, and alleviates the problem of translation missing by learning complete semantic unit information.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120012787A_ABST
    Figure CN120012787A_ABST
Patent Text Reader

Abstract

The invention provides a non-autoregressive machine translation method and system based on a source end-target end common mask. The method comprises the steps that mask-prediction is conducted on uncovered source end words based on encoder-decoder attention, and semantic unit masks based on mutual information are added to target end words on the basis of random token masks. According to the method, the under-translated source end words are masked, so that expression of enhanced word vectors is more targeted, and the problem of translation missing is relieved. According to the method, the untranslated source end words are positioned and masked according to the encoder-decoder attention matrix and the prediction result of the Transform, additional external supervision information and tools do not need to be introduced, and the method is simple and direct. Besides, for the problem that a random mask method of a target end is insufficient, a semantic unit mask based on mutual information is added on the basis of a random token mask, and a mixed mask strategy is formed.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of machine translation, and in particular to a non-autoregressive machine translation method and system based on a source-target common mask. Background Art

[0002] Machine translation refers to the translation of sentences in the source language into sentences in the target language through computers. It is an important branch of artificial intelligence and natural language processing, and has important significance for social and economic development and cultural exchanges.

[0003] Most traditional neural machine translation models follow the autoregressive generation method, that is, generating sentences word by word from left to right. The advantage of the autoregressive generation method is that the model can refer to the previously generated translation content when inferring the token at the current moment, thereby maintaining the fluency and coherence of the text generation. However, the disadvantage of the autoregressive machine translation method is that its generation speed is slow and it cannot fully utilize the parallelism of GPU computing. Therefore, the non-autoregressive machine translation model was proposed. The biggest difference between it and the autoregressive machine translation is that it breaks the left-to-right generation method that depends on the previous text, greatly improves the speed of machine translation, and makes more full use of the computing resources of the GPU. Non-autoregressive machine translation can play an important role in scenarios such as timely translation and online translation that have high requirements for translation timeliness.

[0004] Existing non-autoregressive machine translation models can be divided into two categories according to the generation method: full non-autoregressive machine translation and iterative improvement non-autoregressive machine translation. Among them, the characteristic of full non-autoregressive machine translation is that it can generate all translation content at the same time. In full non-autoregressive machine translation, in order to better grasp the dependencies between translation content, researchers have proposed various methods to enhance full non-autoregressive machine translation, including introducing reordering modules, introducing syntactic structures, introducing latent variable models, etc. The characteristic of iterative improvement non-autoregressive machine translation is to use multiple rounds of iterations to continuously improve the translation content. Representative works include mask-predict model, Levenshtein Transformer, Insert Transformer, etc.

[0005] The method of the present invention belongs to the mask-predict translation method in iteratively improving non-autoregressive machine translation. The mask-predict translation method includes two processes: training and reasoning. During the training process, the input of the target end is randomly masked, and the model predicts the mask content based on the source end input and the visible content of the target end. As the training proceeds, the proportion of random masks will gradually increase. In the reasoning stage, the model predicts based on the fully masked target end input, and masks the words with the lowest probability of predicted output, and predicts again as the next round of input until the model's prediction no longer changes, or the maximum number of iterations is reached and stops.

[0006] Although the existing mask-prediction non-autoregressive machine translation has achieved good performance in translation performance (accuracy and speed), it still has the problem of missing translation. The missing translation problem refers to the fact that some semantic information of the source sentence is not translated in the translation result. The missing translation problem in traditional mask-prediction non-autoregressive machine translation is caused by two reasons: insufficient modeling of source word information and insufficient random masking of the target.

[0007] First, in the end-to-end machine translation model, the source input provides a lot of information for the target translation, but the traditional mask-prediction non-autoregressive machine translation method lacks direct supervision measures for the learning of source word representations during the training process, resulting in the inability to learn effective representations of source words, and thus unable to provide effective information for the translation process of the target. Existing iterative improvement non-autoregressive machine translation methods lack attention to source word modeling. Although some methods mask source words (representative works such as JM-NAT and AMOM), they all use random masking measures, which lacks specificity.

[0008] Secondly, the existing iterative improvement of non-autoregressive machine translation uses a method of randomly masking the target words during training. However, random masking may cover up subwords of words, or parts of a phrase, making it impossible for the model to model and understand a complete semantic unit, which in turn leads to translation loss during reasoning. Most of the existing iterative improvement of non-autoregressive machine translation methods use random masking measures to mask the target end. Although random masking is simple and direct, it makes the model pay less attention to the prediction and modeling of semantic units containing multiple tokens, which leads to translation loss problems. At the same time, using only a masking strategy based on semantic units will cause the model to ignore the correction of single incorrect words predicted.

[0009] Therefore, there is a need in the market for a non-autoregressive machine translation method and system based on a source-target common mask that can avoid the problem of missing translation during reasoning. Summary of the invention

[0010] In view of the defects in the prior art, the object of the present invention is to provide a non-autoregressive machine translation method and system based on source-target common mask.

[0011] A non-autoregressive machine translation method based on source-target common mask provided by the present invention comprises:

[0012] Step S1: building a machine translation model;

[0013] Step S2: Obtain bilingual corpus and learning rate and input them into the machine translation model, train the machine translation model to obtain model parameters that meet the requirements, and obtain the final translation result based on the trained machine translation model.

[0014] Preferably, the machine translation model comprises an encoder module, a non-autoregressive decoder module and a length prediction module;

[0015] The encoder module is used to map the input sequence to a hidden layer vector in the source language space through a series of encoding layers;

[0016] The non-autoregressive decoder module is used to decode the source information provided by the encoder and output the generated sequence in parallel;

[0017] The length prediction module is used to determine the length to be translated.

[0018] Preferably, the encoder module includes an auxiliary task of cross-attention based mask-prediction of uncovered source words;

[0019] The auxiliary tasks include selection of uncovered source words, masking of uncovered source words and training.

[0020] Preferably, the selection of the uncovered source words comprises: extracting a word alignment matrix A according to the attention allocated to all source words for each target word in the cross-attention matrix in the decoder in the machine translation model, and finding the covered source words and the uncovered source words in the sentence;

[0021] The covered source words are allocated to each target word the source word with the greatest attention;

[0022] The uncovered source-end words are the remaining source-end words that are not aligned by any target-end words;

[0023] For the extracted word alignment matrix A, the formula is as follows:

[0024]

[0025] Among them, S i,j Represents the i-th word y on the target sidei and the jth source word x j The alignment score between .

[0026] Preferably, the masking and training of the uncovered source-end words comprises: randomly selecting a token from the uncovered source-end word set for masking to obtain a masked sentence;

[0027] The encoder module needs to maximize the following conditional probability:

[0028] P(x M |x R )=softmax(FFN(R E ))

[0029] R E =Encoder(x R )

[0030] Among them, x represents the original source sentence, x M Indicates the uncovered source term, x R Indicates that the uncovered words are replaced with <mask>The masked source sentence, R E Represents the vector representation of the source sentence after masking obtained by the encoder;

[0031] The training loss of the encoder is:

[0032]

[0033] Among them, L encoder represents the loss function of the encoder.

[0034] Preferably, the non-autoregressive decoder module includes a semantic unit mask based on mutual information;

[0035] The mutual information-based semantic unit masking includes selecting continuous semantic units in a sentence through pointwise mutual information PMI, and masking the continuous semantic units.

[0036] Preferably, the pointwise mutual information PMI is defined as follows: the probability of occurrence of a certain n-gram is defined as the ratio between the number of occurrences of the n-gram in the corpus and the number of occurrences of all n-grams. Based on the probability P of the n-gram, the PMI measures the amount of information brought by the common occurrence of the n-grams.

[0037] For n-gramw 1 …w n , "w 1 …w n The PMI of ” is as follows:

[0038]

[0039] Among them, w j represents the jth token in the n-gram, p(w j ) represents w j The probability of occurrence.

[0040] Preferably, the modeling process of the decoder is represented as follows:

[0041] P(y M |y R ,x)=softmax(FFN(R D ))

[0042] R D =Decdoer(y R ,R E )

[0043] Among them, y M Represents the valid correct data ground truth tokens of the masked token, R D represents the top hidden state of the decoder, R E represents the top-level output of the encoder, y R represents the target sequence after random masking;

[0044] The loss function calculation formula of the decoder is as follows:

[0045] L decoder =-∑P(y M |y R ,x)

[0046] Among them, L decoder represents the loss function of the non-autoregressive decoder.

[0047] Preferably, a length marker is added to the front end of the encoder module input, and the output of the encoder length marker is used as the length of the translated sentence. The loss function of the length prediction is as follows:

[0048] L len =-∑P(l|x)

[0049] Among them, L len represents the loss function for length prediction, l represents the length of the target sentence, and x represents the source input sentence.

[0050] A non-autoregressive machine translation system based on source-target common mask provided by the present invention comprises:

[0051] Module M1: Building a machine translation model;

[0052] Module M2: Obtain bilingual corpus and learning rate and input them into the machine translation model, train the machine translation model to obtain model parameters that meet the requirements, and obtain the final translation result based on the trained machine translation model.

[0053] Compared with the prior art, the present invention has the following beneficial effects:

[0054] 1. The present invention masks the under-translated source words, thereby making the expression of the enhanced word vector more targeted and alleviating the problem of missing translations. In addition, the present invention locates and masks the untranslated source words based on the Transformer's encoder-decoder attention matrix and prediction results, without the need to introduce additional external supervision information and tools. The method is simple and direct.

[0055] 2. The present invention proposes a hybrid masking strategy, which uses a combination of mutual information-based semantic unit masking and random masking strategies on the target side during training, so that the model can pay attention to the modeling of complete semantic units and complete the task of correcting single words, thereby reducing the missing translation. Among them, the present invention proposes to use simulated annealing to balance the use of the two masking strategies. BRIEF DESCRIPTION OF THE DRAWINGS

[0056] Other features, objects and advantages of the present invention will become more apparent from the detailed description of non-limiting embodiments made with reference to the following drawings:

[0057] Figure 1 It is a schematic diagram of the working method of the present invention. DETAILED DESCRIPTION

[0058] The present invention is described in detail below in conjunction with specific embodiments. The following embodiments will help those skilled in the art to further understand the present invention, but are not intended to limit the present invention in any form. It should be noted that, for those of ordinary skill in the art, several changes and improvements can also be made without departing from the concept of the present invention. These all belong to the protection scope of the present invention.

[0059] The English explanations of the present invention are as follows: NMT, Neural Machine Translation, neural machine translation; AT, Autoregressive Translation, autoregressive translation; NAT, Non-autoregressive Translation, non-autoregressive machine translation; IRM, Iterative Refinement Model, iterative improvement model.

[0060] The present invention specifically includes applying mask prediction based on the attention mechanism to uncovered words on the source side, and implementing a hybrid masking strategy based on word mutual information on the target side words, so as to respectively solve the problems of insufficient modeling of source side word representation and insufficient random masking on the target side in iteratively improving non-autoregressive machine translation. First, in view of the defect of insufficient learning of source side word representation, the present invention provides an uncovered source side word mask-prediction task based on the attention mechanism as an auxiliary learning task for the machine translation task. Secondly, in view of the problem of insufficient random masking method on the target side, the present invention adds a semantic unit mask based on mutual information on the basis of random token masking to form a hybrid masking strategy. Among them, the semantic unit masking method based on mutual information masks complete semantic units (words or phrases), thereby helping the model to learn complete semantic information and alleviate the problem of missing translation.

[0061] According to the non-autoregressive machine translation method based on source-target common mask provided by the present invention, Figure 1 As shown, including:

[0062] Step S1: Construct a translation model. The translation model is a sequence-to-sequence Transformer translation model, which includes three modules, namely, an encoder module, a non-autoregressive decoder module, and a length prediction module. The structure and training objectives of each of the above modules are described in detail as follows:

[0063] Encoder module: The function of this module is to map the input sequence to a hidden layer vector in the source language space through a series of encoding layers. In order to alleviate the problem of missing translation on the target side, the present invention adds an auxiliary task of mask-prediction of uncovered source words based on cross-attention to the encoder module, and enhances the modeling of source word vectors by predicting uncovered source words. The task is divided into two steps, namely the selection of uncovered source words and the masking and training of uncovered source words.

[0064] The first step is to select uncovered source words based on cross-attention. According to the attention assigned by each target word to all source words in the cross-attention matrix in the decoder of the machine translation model, the word alignment matrix A is extracted, and the covered and uncovered source words in the sentence are found. Among them, the covered words are the source words that are assigned the maximum attention to each target word, and the uncovered words are the remaining source words that are not aligned by any target word.

[0065] Specifically, for a source input of length x and a target input of length y, at the i-th layer of the decoder, the attention head-averaged encoder-decoder cross-attention matrix is ​​defined as W i ∈R y×x , where the matrix elements Measures the correlation between the hidden state of the decoder and the encoder output of the i-th layer. The alignment score matrix is ​​S∈R y×x , is the attention matrix between the top-level (lth layer) decoder and encoder, that is, The element S in the matrix S i,j is the i-th word y on the target side i and the source word x j The alignment score between .

[0066] The word alignment matrix A is defined as:

[0067]

[0068] Then randomly select a token from the uncovered source word set for masking to obtain the masked sentence. The original source sentence is denoted as x, and the uncovered source word is denoted as x. M , use the uncovered words <mask>The masked source sentence is denoted as x R The encoder module needs to maximize the following conditional probability:

[0069] P(x M |x R )=softmax(FFN(R E ))

[0070] R E =Encoder(x R )

[0071] Among them, R E Represents the vector representation of the source sentence after masking obtained by the encoder.

[0072] Finally, the training loss of the encoder is expressed as:

[0073]

[0074] Among them, L encoder represents the loss function of the encoder.

[0075] Non-autoregressive decoder module: The function of the decoder module is to decode the source information given by the encoder and output the generated sequence in parallel.

[0076] During training, the traditional non-autoregressive decoder based on the mask-predict paradigm needs to predict the masked tokens in the target input, and the masking strategy is derived from the pre-training method of BERT, namely Random-token Masking. The Random-token Masking method is simple and fast, and can predict the missing content based on the neighboring tokens of the masked token. However, its disadvantage is that the masked token may be a subword in a complete word, or a word in a common phrase, which makes the model unable to learn the semantics of the complete phrase or word, and thus leads to missing and wrong word predictions. On the other hand, current research proposes a series of methods to mask complete phrases or words to promote the translation model to learn complete semantic unit information and alleviate the problem of missing translation. Representative works include Whole-word Masking, Entity-phrase Masking, etc. However, using only semantic unit masking cannot effectively deal with the situation of correcting the predicted single wrong token.

[0077] In order to promote the model to learn complete semantic fragment information and effectively deal with the prediction of single words in continuous semantic fragments, the present invention proposes a masking method that integrates semantic unit masking and random-token masking. Among them, Pointwise Mutual Information (PMI) is used based on semantic unit masking to select continuous semantic units in sentences and mask them.

[0078] First, let’s define pointwise mutual information (PMI). First, define the probability of occurrence of a certain n-gram as the ratio between the number of times the n-gram appears in the corpus and the number of times all n-grams appear. Based on the probability P of the n-gram, PMI measures the amount of information brought by the co-occurrence of the n-grams. For a given two tokens w 1 and w 2 , "w 1 w 2 The PMI of ” is:

[0079]

[0080] Similarly, for n-gramw 1 …w n , "w 1 …w n The PMI of ” is:

[0081]

[0082] Among them, w j represents the jth token in the n-gram, p(w j ) represents w j The lower the mutual information, the lower the relationship between the words in the n-gram; the higher the mutual information, the higher the correlation between the words in the n-gram, and the more likely it is a word composed of a string of subwords, or a phrase composed of a string of words. The n-grams in the corpus are sorted from high to low according to the mutual information, and then a mask vocabulary is constructed, and n-grams with a length of 2-5 and more than 10 occurrences in the corpus are selected. In this embodiment, the size of the mask vocabulary is set to 800k.

[0083] Next, a hybrid masking strategy is applied during the training of the translation model, that is, a hybrid masking method of using PMI-based semantic unit masking and random-token masking. The random-token masking method is used in the early stages of training because the model is not powerful enough at the beginning of training and the difficulty of predicting individual words is relatively low. As the training progresses and the model's capabilities improve, the PMI-based semantic unit masking method is used so that the model can grasp more complete semantic information. Specifically, given the target input y and the ratio of the number of masks to the sentence length k, the masking method is defined as:

[0084] s~Bernoulli(p)

[0085]

[0086] in, is a Bernoulli distribution with parameter p, PMI MASK represents the semantic unit masking operation based on PMI, RANDOM MASK represents the random masking operation based on token, and y R Represents the target sequence after random masking. In order to balance the two operations and make the increase in training difficulty and the improvement in model capabilities more matched, p is gradually annealed from 1 to 0 during the training process.

[0087] During training, the modeling process of the decoder can be expressed as:

[0088] P(y M |y R ,x)=softmax(FFN(R D ))

[0089] R D =Decoder(y R ,R e )

[0090] Among them, y M Represents the valid correct data ground truth tokens of the masked token, R D represents the top hidden state of the decoder, R E represents the top-level output of the encoder. Finally, the loss function of the decoder is calculated as follows:

[0091] L decoder =-∑P(y M |y R ,x)

[0092] Among them, L decoder Loss function for the non-autoregressive decoder.

[0093] Length prediction module: Since non-autoregressive machine translation completes the generation of all translations at the same time, unlike the autoregressive machine translation model which uses the [EOS] marker to represent the end of the translation, it is necessary to determine the length of the translation in advance. The method of the present invention follows the mask-prediction non-autoregressive translation length determination method, that is, adding a length marker [LENGTH] at the very front of the encoder input. The output of the encoder length marker is used as the length of the translated sentence, and the loss function of the length prediction is as follows:

[0094] L len =-∑P(l|x)

[0095] Among them, L len represents the loss function of length prediction, l represents the length of the target sentence, and x represents the source input sentence. Combining the above encoder module, non-autoregressive translation decoder module, and length prediction module, the overall loss function is:

[0096] L total =L encoder +L decoder +L len

[0097] Where L total Represents the overall loss function obtained by the three modules: encoder module, non-autoregressive translation decoder module, and length prediction module.

[0098] Step S2: input bilingual corpus (X, Y) and learning rate γ, train the translation model to output trained model parameters θ, and obtain the final translation result based on the trained translation model.

[0099] Furthermore, the specific description of the training and reasoning process of the method of the present invention is as follows:

[0100] During the training process, the input is the bilingual corpus (X, Y) and the learning rate γ, and the expected output is the trained model parameter θ. Repeat the following steps until the model converges: for the input sentence x in the bilingual corpus, input it into the encoder to obtain the source language vector; for the target translation y in the bilingual corpus, mask it using the hybrid masking strategy mentioned above to obtain y R , and y R The non-autoregressive encoder is input to predict the masked target words. Then, the cross-attention layer on the decoder side obtains a one-to-one correspondence between the target sentence and the source sentence, and finds the uncovered source words to mask and obtain x R , predict the masked source words. Finally, calculate the overall loss in the above process, i.e., L total =L encoder +L decoder +L len , and then use the optimizer to backpropagate the gradients and update the parameter values.

[0101] During the inference process, first, for the sentence to be translated x, the encoder maps it to the source vector space and predicts the length L of the translated sentence. Note that the encoder no longer performs the mask prediction process. Then, the target uses a vector of length L to predict the length of the translated sentence. <mask>The sequence initializes the target input. The decoder outputs a possible translation result based on the input and the information given by the encoder, and masks the tokens in the translation result whose confidence is less than the threshold. Finally, the masked translation is input to the decoder for prediction again. Iterates continuously until the two outputs no longer change or the maximum number of iterations is reached.

[0102] The present invention performs masking-prediction on uncovered source words based on encoder-decoder attention, and adds semantic unit masking based on mutual information to target words on the basis of random token masking. Among them, source word masking-prediction based on cross-attention refers to corresponding to the untranslated source words according to the prediction results based on the target end of cross-attention, and then masking and predicting these untranslated source words, thereby enhancing the representation modeling of source words and alleviating the problem of missing translation. Target semantic unit masking based on mutual information refers to a masking strategy for multi-words / phrases based on point mutual information. During training, the target input is masked according to the mutual information, thereby helping the model learn the information of the complete semantic unit, thereby alleviating the problem of missing translation.

[0103] The present invention aims to improve the masking method in the mask-predictive translation process to promote the source and target words to better grasp the context and better learn the feature representation, thereby alleviating the problem of missing translation. The technical problem to be solved by the present invention is to alleviate the missing translation problem existing in the existing mask-predictive non-autoregressive machine translation method.

[0104] The present invention also provides a non-autoregressive machine translation system based on a source-target common mask. The non-autoregressive machine translation system based on a source-target common mask can be implemented by executing the process steps of the non-autoregressive machine translation method based on a source-target common mask, that is, those skilled in the art can understand the non-autoregressive machine translation method based on a source-target common mask as a preferred implementation of the non-autoregressive machine translation system based on a source-target common mask.

[0105] A non-autoregressive machine translation system based on source-target common mask provided by the present invention comprises:

[0106] Module M1: Build a machine translation model. The machine translation model includes an encoder module, a non-autoregressive decoder module, and a length prediction module.

[0107] The encoder module is used to map the input sequence to a hidden layer vector in the source language space through a series of encoding layers. The encoder module includes an auxiliary task of mask-prediction of uncovered source words based on cross-attention. The auxiliary tasks include the selection of uncovered source words, masking of uncovered source words, and training. The selection of uncovered source words includes: according to the attention assigned to all source words in the cross-attention matrix in the decoder of the machine translation model, the word alignment matrix A is extracted, and the covered source words and uncovered source words in the sentence are found. The covered source words are the source words that are assigned the maximum attention to each target word. The uncovered source words are the remaining source words that are not aligned by any target word.

[0108] For the extracted word alignment matrix A, the formula is as follows:

[0109]

[0110] Among them, S i,j Represents the i-th word y on the target side i and the jth source word x j The alignment score between .

[0111] The masking and training of uncovered source words includes: randomly selecting a token from the uncovered source word set for masking, and obtaining the masked sentence. The encoder module needs to maximize the following conditional probability:

[0112] P(x M |x R )=softmax(FFN(R E ))

[0113] R E =Encoder(x R )

[0114] Among them, x represents the original source sentence, x M Indicates the uncovered source term, x R Indicates that the uncovered words are replaced with <mask>The masked source sentence, R E It represents the vector representation of the source sentence after masking obtained by the encoder. The training loss of the encoder is:

[0115] L encoder =-∑P(x M |x R )

[0116] Among them, L encoder represents the loss function of the encoder.

[0117] The non-autoregressive decoder module is used to decode the source information given by the encoder and output the generated sequence in parallel. The non-autoregressive decoder module includes a semantic unit mask based on mutual information. The semantic unit mask based on mutual information includes selecting continuous semantic units in a sentence through point-wise mutual information PMI and masking the continuous semantic units. The point-wise mutual information PMI is defined as: the probability of occurrence of a certain n-gram is defined as the ratio between the number of occurrences of the n-gram in the corpus and the number of occurrences of all n-grams. Based on the probability P of the n-gram, PMI measures the amount of information brought by the co-occurrence of n-grams. For n-gramw 1 …w n , "w 1 …w n The PMI of ” is as follows:

[0118]

[0119] Among them, w j represents the jth token in the n-gram, p(w j ) represents w j The probability of occurrence of . The modeling process of the decoder is expressed as follows:

[0120] P(y M |y R ,x)=softmax(FFN(R D ))

[0121] R D =Decoder(y R ,R E )

[0122] Among them, y M Represents the valid correct data ground truth tokens of the masked token, R D represents the top hidden state of the decoder, R E represents the top-level output of the encoder, y R Represents the target sequence after random masking.

[0123] The loss function calculation formula of the decoder is as follows:

[0124] L decoder =-∑P(y M |y R ,x)

[0125] Among them, L decoder represents the loss function of the non-autoregressive decoder.

[0126] The length prediction module is used to determine the length of the translation. A length marker is added to the front end of the encoder module input, and the output of the encoder length marker is used as the length of the translated sentence. The loss function of the length prediction is as follows:

[0127] L len =-∑P(l|x)

[0128] Among them, L len represents the loss function for length prediction, l represents the length of the target sentence, and x represents the source input sentence.

[0129] Module M2: Obtain bilingual corpus and learning rate and input them into the machine translation model, train the machine translation model and output the trained model parameters, and obtain the final translation result based on the trained machine translation model.

[0130] Those skilled in the art know that, in addition to realizing the system and its various devices, modules, and units provided by the present invention in a purely computer-readable program code, it is entirely possible to realize the same functions in the form of logic gates, switches, application-specific integrated circuits, programmable logic controllers, and embedded microcontrollers by logically programming the method steps. Therefore, the system and its various devices, modules, and units provided by the present invention can be considered as a hardware component, and the devices, modules, and units included therein for realizing various functions can also be regarded as structures within the hardware component; the devices, modules, and units for realizing various functions can also be regarded as both software modules for realizing the method and structures within the hardware component.

[0131] The above describes the specific embodiments of the present invention. It should be understood that the present invention is not limited to the above specific embodiments, and those skilled in the art can make various changes or modifications within the scope of the claims, which does not affect the essence of the present invention. In the absence of conflict, the embodiments of the present application and the features in the embodiments can be combined with each other arbitrarily.< / mask> < / mask> < / mask> < / mask>

Claims

1. A non-autoregressive machine translation method based on source-target joint masking, characterized in that: include: Step S1: building a machine translation model; Step S2: Obtain bilingual corpus and learning rate and input them into the machine translation model, train the machine translation model to obtain model parameters that meet the requirements, and obtain the final translation result based on the trained machine translation model.

2. The non-autoregressive machine translation method based on source-target common mask according to claim 1, characterized in that: The machine translation model includes an encoder module, a non-autoregressive decoder module and a length prediction module; The encoder module is used to map the input sequence to a hidden layer vector in the source language space through a series of encoding layers; The non-autoregressive decoder module is used to decode the source information provided by the encoder and output the generated sequence in parallel; The length prediction module is used to determine the length to be translated.

3. The non-autoregressive machine translation method based on source-target common mask according to claim 2, characterized in that: The encoder module includes an auxiliary task of cross-attention based mask-prediction of uncovered source words; The auxiliary tasks include selection of uncovered source words, masking of uncovered source words and training.

4. The non-autoregressive machine translation method based on source-target common mask according to claim 3, characterized in that: The selection of the uncovered source words includes: extracting a word alignment matrix A according to the attention allocated to all source words for each target word in the cross-attention matrix in the decoder of the machine translation model, and finding the covered source words and the uncovered source words in the sentence; The covered source words are allocated to each target word the source word with the greatest attention; The uncovered source-end words are the remaining source-end words that are not aligned by any target-end words; For the extracted word alignment matrix A, the formula is as follows: Among them, S i,j Represents the target word y i and source word x j The alignment score between i,j Represents the i-th word y on the target side i and the jth source word x j The alignment score between .

5. The non-autoregressive machine translation method based on source-target common mask according to claim 3, characterized in that: The masking and training of uncovered source words includes: randomly selecting a token from the uncovered source word set for masking to obtain a masked sentence; The encoder module needs to maximize the following conditional probability: P(x M |x R )=softmax(FFN(R E )) R E =Encoder(x R ) Among them, x represents the original source sentence, x M Indicates the uncovered source term, x R Indicates that the uncovered words are replaced with <mask>The masked source sentence, R E Represents the vector representation of the source sentence after masking obtained by the encoder;< / mask> The training loss of the encoder is: L encoder =-∑P(x M |x R ) Among them, L encoder represents the loss function of the encoder.

6. The non-autoregressive machine translation method based on source-target common mask according to claim 2, characterized in that: The non-autoregressive decoder module includes a semantic unit mask based on mutual information; The mutual information-based semantic unit masking includes selecting continuous semantic units in a sentence through pointwise mutual information PMI, and masking the continuous semantic units.

7. The non-autoregressive machine translation method based on source-target common mask according to claim 6, characterized in that: The pointwise mutual information PMI is defined as follows: the probability of occurrence of a certain n-gram is defined as the ratio between the number of occurrences of the n-gram in the corpus and the number of occurrences of all n-grams. Based on the probability P of the n-gram, the PMI measures the amount of information brought by the common occurrence of the n-grams. For n-grams w1…w n , "w1…w n The PMI of ” is as follows: Among them, w j represents the jth token in the n-gram, p(w j ) represents w j The probability of occurrence.

8. The non-autoregressive machine translation method based on source-target common mask according to claim 6, characterized in that: The modeling process of the decoder is represented as follows: P(and M |and R ,x)=softmax(FFN(R D )) R D =Decoder(y R ,R E ) Among them, y M Represents the valid correct data ground truth tokens of the masked token, R D represents the top hidden state of the decoder, R E represents the top-level output of the encoder, y R represents the target sequence after random masking; The loss function calculation formula of the decoder is as follows: L decoder =-∑P(and M |and R ,x) Among them, L decoder represents the loss function of the non-autoregressive decoder.

9. The non-autoregressive machine translation method based on source-target common mask according to claim 2, characterized in that: A length marker is added to the front end of the encoder module input. The output of the encoder length marker is used as the length of the translated sentence. The loss function of the length prediction is as follows: L len =-∑P(l|x) Among them, L len represents the loss function for length prediction, l represents the length of the target sentence, and x represents the source input sentence.

10. A non-autoregressive machine translation system based on source-target joint masking, characterized in that: include: Module M1: Building a machine translation model; Module M2: Obtain bilingual corpus and learning rate and input them into the machine translation model, train the machine translation model to obtain model parameters that meet the requirements, and obtain the final translation result based on the trained machine translation model.

Citation Information

Cited By

  • Voice recognition system based on an intended action of an occupant

    US12567406B2

  • Context based translation and ordering of webpage text elements

    US20240427834A1

  • Voice recognition system based on an intended action of an occupant

    US20250218432A1