A machine translation method capable of learning future information
By adding a future information network module to the decoder end of the neural machine translation model, using the gated loop unit and attention mechanism to learn future information of sentences, the problem that the neural machine translation model cannot learn future information is solved, and the corpus utilization rate and translation performance are improved.
Patent Information
- Application Number
- CN202210016283.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-01-07
- Publication Date
- 2025-06-24
- Estimated Expiration
- 2042-01-07
AI Technical Summary
Neural machine translation models cannot learn future information during training, resulting in low corpus utilization.
The future information network module is added to the decoder end of the Transformer model, and the future information of the sentence is learned using the gated loop unit and attention mechanism, and incorporate it into the decoding process.
The neural machine translation model's information capture ability of parallel corpus is improved, the translation performance is improved, and the model's learning ability is improved.
Smart Images

Figure CN114528853B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to a neural machine translation technology, specifically a machine translation method capable of learning future information. Background Art
[0002] Machine translation is to use a computer to translate one natural language into another natural language. In the early stage, rule-based methods were used for machine translation. Scientists used the linguistic rules summarized by linguists to write programs, attempting to make the computer translate according to the rules. Subsequently, example-based machine translation developed. This method relies on a sentence example library. The more sentences the example library has, the better the translation quality. With the research of IBM, machine translation entered the era of statistical machine translation. This paradigm uses statistical knowledge to statistically analyze the corresponding relationship between the source language and the target language. During decoding, the sentence with the highest value obtained by multiplying the translation probability by the language model probability is regarded as the decoded sentence.
[0003] Nowadays, machine translation has entered the era of neural machine translation. The earliest neural machine translation paradigm uses an encoder-decoder structure, which runs through all neural machine translation models. The encoder of the first traceable neural machine translation model uses a convolutional neural network to capture the features in the source language sentence, and the decoder uses a recurrent neural network. However, the convolutional neural network cannot capture the dependency information of long sentences, and due to the weight sharing at each time step, the recurrent neural network is prone to gradient explosion. In addition, when dealing with long sentence sequences, the recurrent neural network is prone to gradient vanishing problems. Subsequently, LSTM was introduced into the neural machine translation model to alleviate the long sequence dependency problem and the gradient vanishing problem.
[0004] With the exploration of the attention mechanism in machine learning, it is found to be very suitable for the machine translation task. In machine translation, the relationship between each word and its context words is not uniform. By introducing the attention mechanism, it is possible to naturally learn the relationship between each word and its context words in the sentence, assign a greater weight to the words with higher correlation, and improve the efficiency of model feature learning. In addition, in the attention mechanism, in each calculation, the current word can simultaneously capture its relationship with any word in the context. That is to say, the distance between each word and other words is 1, while in the recurrent neural network, the distance between the nth word and the first word is n - 1. In other words, the attention mechanism is very suitable for capturing long-distance dependencies in sentences.
[0005] In current neural machine translation models, future information is masked during both the training and inference decoding processes. In this way, the information that the model can learn during the training process will have a certain loss. Summary of the Invention
[0006] In view of the problem that in the training process of neural machine translation, there is no ability to learn future information and the utilization rate of the corpus is relatively low, the present invention proposes a machine translation method capable of learning future information, introducing future information into the neural machine translation model to improve the model training paradigm and enhance the utilization rate of the model for the corpus.
[0007] To solve the above technical problems, the technical solution adopted by the present invention is:
[0008] The present invention provides a machine translation method capable of learning future information, including the following steps:
[0009] 1) Use the Transformer model based on the self-attention mechanism as the basic framework, and add a future information network module at the decoder end to learn the future information of the sentence, and construct a machine translation model capable of learning future information;
[0010] 2) Process the training data, perform data cleaning and word segmentation on the source language and target language parallel sentence pairs, and use the existing word embedding model to convert the obtained corresponding words into their corresponding word embedding representations;
[0011] 3) Initialize the parameters of the machine translation model capable of learning future information using a uniform distribution, and use an adaptive learning rate algorithm to optimize the model training;
[0012] 4) In the encoder, first use the self-attention mechanism to calculate the word embeddings, and then send the calculated values into the feed-forward neural network in the machine translation model capable of learning future information to obtain more information in the word embedding vectors. After performing this operation n times, the model learns the feature information of the sentence;
[0013] 5) The machine translation model capable of learning future information uses the encoder-decoder attention mechanism to learn the correlation information between the source language and the target language. At the same time, the information of the encoder is also sent into the network structure capable of learning future information to assist the decoder in learning the future information of the sentence; finally, the future information network sends the learned information back to the decoder to assist the decoder in decoding;
[0014] 6) Use the trained machine translation model capable of learning future information for machine translation. During the translation process, the model only uses the basic Transformer decoder for decoding to implement the machine translation method capable of learning future information.
[0015] In step 1), a future information network module is added at the decoder end to learn the future information of the sentence. The algorithm flow of the future information network is as follows:
[0016] 101) Perform a linear transformation on the future information:
[0017] z i=GRU(z i-1 , h i-1 )
[0018] where z i is the future sentence information of the i-th sentence, z i-1 is the future sentence information of the (i - 1)-th sentence, h i-1 is the recurrent unit state of the (i - 1)-th sentence, and GRU is the gated recurrent unit;
[0019] 102) Fuse the future information with the current state information:
[0020] o i = sigmoid(z i W + b)
[0021]
[0022] where o i is the gating information, used to learn what proportion of the future sentence information z i to retain, sigmoid(·) is the activation function, W and b are the learnable parameters of the model, y i is the current state, is the current state after fusing the future information;
[0023] 103) Fuse the current state with the encoder information to learn the correlation information from the source language to the target language:
[0024]
[0025] where is the output value of the future information network, attention(·) is the attention mechanism algorithm, is the current state after fusing the future information, x i is the feature value captured by the encoder.
[0026] In step 101), the gated recurrent unit GRU is used to perform a linear transformation on the information. This structure consists of a reset gate and an update gate. The reset gate is used to control the proportion of the information of the previous hidden layer state. The formula is as follows:
[0027] r t = sigmoid([h t-1 , x t W r )
[0028] where r t represents the proportion of the previous hidden layer information to be retained, h t-1 is the previous hidden layer information, x t is the current input, Wr are trainable weight parameters;
[0029] The update gate is used to update the memory information, and its calculation formula is as follows:
[0030] u t = sigmoid([h t-1 , x t W u )
[0031] where u t is the weight value obtained by the update gate, and W u are trainable weight parameters;
[0032] After that, the calculated r t and u t are used to update the hidden state at the current time step, and its specific calculation formula is as follows:
[0033]
[0034]
[0035] where is the value containing the hidden state of the previous moment and the information of the current moment, W h are trainable weight parameters, tanh(·) is the activation function, and h t is the hidden state information at the current moment.
[0036] In step 103), the attention(·) attention mechanism algorithm is used to capture the relationship between future sentence information and current sentence information, and the calculation formula of this algorithm is as follows:
[0037] attention(q, k, v) = softmax(α) · v
[0038]
[0039] where q, k, and v are the inputs of the algorithm, representing query, key, and value respectively, softmax(·) is the normalization function, α is the attention weight matrix, which is used to represent the correlation size of each position, and d k is the hidden layer dimension of k.
[0040] In step 3), the model is initialized with a uniform distribution, and the adaptive learning rate algorithm is used to optimize the model training. The specific process is as follows:
[0041] 301) Randomly sample the model parameters from a uniform distribution. The specific implementation of this method is:
[0042]
[0043] where \(w\) is the weight matrix, \(d\) in and \(d\) out represent the input and output dimension sizes of \(w\) respectively, and \(U(\cdot)\) represents the uniform distribution;
[0044] 302) The model training is optimized using the adaptive learning rate algorithm, and the specific formula of this optimization algorithm is:
[0045] \(v\) t =\(\beta_1v\) t-1 +(1 - \(\beta_1\))\(g\) t
[0046] \(s\) t =\(\beta_2s\) t-1 +(1 - \(\beta_2\))\(g\) t 2
[0047]
[0048] where \(v\) t is the momentum of the mini - batch stochastic gradient at the \(t\) - th moment, \(\beta_1\) and \(\beta_2\) are settable parameters, \(g\) t is the mini - batch stochastic gradient at the \(t\) - th moment, \(x\) t is the model parameter value at the \(t\) - th moment, \(\alpha\) is the learning rate, its value is greater than 0, \(\epsilon\) is a constant added to avoid the denominator being 0, and \(s\) t is the momentum at the \(t\) - th moment.
[0049] The present invention has the following beneficial effects and advantages:
[0050] 1. The present invention introduces future information into neural machine translation to assist the decoder in training, explores the learning ability of neural machine translation, improves the deficiencies of the existing neural machine translation paradigm, and enhances the information capture ability of the neural machine translation model for parallel corpora, and also improves the translation performance of the model.
[0051] 2. Although the present invention is an improvement in the Transformer model, the future information network module is independent of the Transformer model and can be applied to any model structure. BRIEF DESCRIPTION OF THE DRAWINGS
[0052] Figure 1 is the encoder structure diagram;
[0053] Figure 2 is the decoder structure diagram;
[0054] Figure 3 is the future information network structure diagram in the method of the present invention;
[0055] Figure 4 It is the structure diagram of GRU;
[0056] Figure 5 It is a schematic diagram of the Transformer model structure for introducing future information. Specific implementation manners
[0057] The present invention will be further described below in conjunction with the accompanying drawings of the specification.
[0058] The present invention is a machine translation method for a network that can learn future information, specifically including the following steps:
[0059] 1) Use the Transformer model based on the self-attention mechanism as the basic framework, and add a future information network module at the decoder end to learn the future information of the sentence, and construct a machine translation model that can learn future information;
[0060] The Transformer model uses the encoder (as shown in Figure 1 )-decoder (as shown in Figure 2 ) framework as the basic framework, and add a future information network module (as shown in Figure 3 ) at the decoder end to learn the future information of the sentence. The future information network module mainly uses GRU to capture future information (as shown in Figure 4 ), and finally constitutes a machine translation model that can learn future information (as shown in Figure 5 );
[0061] In step 1), add a future information network module at the decoder end to learn the future information of the sentence. The algorithm process of the future information network is as follows:
[0062] 101) Perform a linear transformation on the future information:
[0063] z i = GRU(z i-1 , h i-1 )
[0064] where z i is the future sentence information of the i-th sentence, z i-1 is the future sentence information of the (i - 1)-th sentence, h i-1 is the recurrent unit state of the (i - 1)-th sentence, and GRU is the gated recurrent unit;
[0065] 102) Fuse the future information with the current state information:
[0066] o i = sigmoid(z i W + b)
[0067]
[0068] where o i is the gating information, used to learn what proportion of future sentence information needs to be retained. i , sigmoid(·) is the activation function, W and b are the parameters that the model can learn, and y i is the current state, is the current state after fusing future information;
[0069] 103) Fuse the current state with the encoder information to learn the correlation information from the source language to the target language:
[0070]
[0071] where is the output value of the future information network, attention(·) is the attention mechanism algorithm, is the current state after fusing future information, and x i is the eigenvalue captured by the encoder.
[0072] In step 101), a gated recurrent unit GRU is used to perform a linear transformation on the information. This structure consists of a reset gate and an update gate. The reset gate is used to control the proportion of information of the previous hidden layer state. The formula is as follows:
[0073] r t = sigmoid([h t-1 , x t W r )
[0074] where r t represents the proportion of the previous hidden layer information that needs to be retained, h t-1 is the previous hidden layer information, x t is the current input, and W r is the trainable weight parameter;
[0075] The update gate is used to update the memory information. Its calculation formula is as follows:
[0076] u t = sigmoid([h t-1 , x t W u )
[0077] where u t is the weight value obtained by the update gate, and W u is the trainable weight parameter;
[0078] After that, use the calculated r t and ut Update the hidden state of the current time step, and its specific calculation formula is as follows:
[0079]
[0080]
[0081] Where is the value containing the hidden state of the previous moment and the information of the current moment, and W h is the trainable weight parameter, tanh(·) is the activation function, and h t is the hidden state information of the current moment.
[0082] In step 103), the attention(·) attention mechanism algorithm is used to capture the relationship between the future sentence information and the current sentence information. The calculation formula of this algorithm is as follows:
[0083] attention(q,k,v) = softmax(α)·v
[0084]
[0085] Where q, k, and v are the inputs of the algorithm, representing query, key, and value respectively. softmax(·) is the normalization function, and α is the attention weight matrix, which is used to represent the correlation size of each position. d k is the hidden layer dimension of k.
[0086] 2) Process the training data, perform data cleaning, word segmentation on the source language and target language parallel sentence pairs, and use the existing word embedding model to convert the obtained corresponding words into their corresponding word embedding representations;
[0087] In step 2), process the training data, perform data cleaning, word segmentation on the source language and target language parallel sentence pairs, and use the word embedding model to convert the obtained corresponding words into their corresponding word embedding representations. The specific process is as follows:
[0088] 201) Use the general cleaning process to clean and filter the data set, delete the sentences containing garbled characters, html tags or website addresses in the sentences, and delete the sentences with too large or too small length ratios of the source language to the target language;
[0089] 202) Use the word segmentation tool to perform word segmentation on the data set, and use bpe to perform sub-word segmentation on the words in the sentences;
[0090] 203) Use the word2vec tool to obtain the word embedding vector information corresponding to the words in the data set.
[0091] 3) Initialize the parameters of the machine translation model that can learn future information using a uniform distribution, and use an adaptive learning rate algorithm to optimize the model training;
[0092] Step 3) Initialize the parameters of the model using a uniform distribution and optimize the model training using an adaptive learning rate algorithm. The specific process is as follows:
[0093] 301) Randomly sample the model parameters from a uniform distribution. The specific implementation of this method is:
[0094]
[0095] where w is the weight matrix, d in and d out represent the input and output dimension sizes of w respectively, and U(·) represents the uniform distribution;
[0096] 302) Use an adaptive learning rate algorithm to optimize the model training. The specific formula of this optimization algorithm is:
[0097] v t = β1v t-1 + (1 - β1)g t
[0098] s t = β2s t-1 + (1 - β2)g t 2
[0099]
[0100] where v t is the momentum of the mini-batch stochastic gradient at the t-th moment, β1 and β2 are adjustable parameters, g t is the mini-batch stochastic gradient at the t-th moment, x t is the value of the model parameters at the t-th moment, α is the learning rate, its value is greater than 0, ∈ is a constant added to avoid the denominator being zero, and s t is the momentum at the t-th moment.
[0101] 4) In the encoder, first use the self-attention mechanism to calculate the word embeddings, and then send the calculated values into the feed-forward neural network in the machine translation model that can learn future information to obtain more information in the word embedding vectors. After performing this operation n times, the model learns the feature information of the sentence;
[0102] 5) The machine translation model that can learn future information uses the encoder-decoder attention mechanism to learn the correlation information between the source language and the target language. At the same time, the information of the encoder is also sent into the network structure that can learn future information to assist the decoder in learning the future information of the sentence. Finally, the future information network sends the learned information back to the decoder to assist the decoder in decoding;
[0103] 6) Use the trained machine translation model that can learn future information for machine translation. During the translation process, the model only uses the basic Transformer decoder for decoding to implement the machine translation method that can learn future information.
[0104] Taking the IWSLT German-English translation as an example, Table 1 shows three example sentences, indicating the effectiveness of the present invention in improving the quality of the model translation. Among them, the base model is a neural machine translation model that does not use the future information module, and our model is the model of the present invention. Table 2 shows the effectiveness of the method of the present invention in the translation evaluation BLEU compared with the neural machine translation model that does not use the future information module.
[0105] Table 1 Comparison of IWSLT German-English translation example sentences
[0106]
[0107] Table 2 Comparison of IWSLT German-English translation BLEU
[0108] Model BLEU Base model 34.8 Our model 35.5
[0109] The present invention verifies the effectiveness of the model in the IWSLT English-German task and uses the BLEU value as the translation performance evaluation index. The present invention introduces the future information of the sentence through the future information network module. During the learning process at the decoding end, the future information is integrated into the auxiliary decoding end for decoding, alleviating the problems such as insufficient corpus learning and imperfect model structure caused by the previous neural machine translation model's inability to utilize future information, and bringing better performance and possibilities to the machine translation model.
Claims
1. A machine translation method capable of learning future information, characterized in that It includes the following steps: 1) Use the Transformer model based on the self-attention mechanism as the basic framework, and add a future information network module at the decoder end to learn the future information of the sentence, and construct a machine translation model that can learn future information; 2) Process the training data, perform data cleaning and word segmentation on the source language and target language parallel sentence pairs, and use the existing word embedding model to convert the corresponding words obtained into their corresponding word embedding representations; 3) Initialize the parameters of the machine translation model that can learn future information using a uniform distribution, and use an adaptive learning rate algorithm to optimize the model training; 4) In the encoder, first use the self-attention mechanism to calculate the word embeddings, and then send the calculated values into the feed-forward neural network in the machine translation model that can learn future information to obtain more information in the word embedding vector. After performing this operation n times, the model has learned the feature information of the sentence; 5) The machine translation model that can learn future information uses the encoder-decoder attention mechanism to learn the correlation information between the source language and the target language. At the same time, the information of the encoder is also sent into the network structure that can learn future information to assist the decoder in learning the future information of the sentence; finally, the future information network sends the learned information back to the decoder to assist the decoder in decoding; 6) Use the trained machine translation model that can learn future information for machine translation. During the translation process, the model only uses the basic Transformer decoder for decoding to implement the machine translation method that can learn future information.
2. The machine translation method capable of learning future information according to claim 1, wherein: In step 1), a future information network module is added at the decoder end to learn the future information of the sentence. The algorithm process of the future information network is as follows: 101) Perform a linear transformation on the future information: z i = GRU(z i-1 , h i-1 ) where z i is the future sentence information of the i-th sentence, and z i-1 is the future sentence information of the (i - 1)-th sentence, h i-1 is the recurrent unit state of the (i - 1)-th sentence, and GRU is the gated recurrent unit; 102) Fuse the future information with the current state information: o i = sigmoid(z i W + b) where o i is gating information, used to learn what proportion of future sentence information needs to be retained z i , sigmoid(·) is the activation function, W and b are parameters that can be learned by the model, y i is the current state, is the current state after fusing future information; 103) Fuse the current state with the encoder information to learn the correlation information from the source language to the target language: wherein is the output value of the future information network, and attention(·) is the attention mechanism algorithm, is the current state after integrating future information, and x i is the eigenvalue captured by the encoder.
3. The machine translation method capable of learning future information according to claim 2, wherein: In step 101), use the gated recurrent unit GRU to perform a linear transformation on the information. This structure consists of a reset gate and an update gate. The reset gate is used to control the information ratio of the previous hidden layer state, and the formula is as follows: r t = sigmoid([h t-1 , x t W r ) where r t represents the proportion of the previous hidden layer information to be retained, h t-1 is the hidden layer information at the previous moment, x t is the input at the current moment, and W r is the trainable weight parameter; The update gate is used to update the memory information, and its calculation formula is as follows: u t = sigmoid([h t-1 , x t W u ) where u t is the weight value obtained by the update gate, and W u is a trainable weight parameter; After that, use the calculated r t and u t to update the hidden state at the current time step. The specific calculation formula is as follows: Among them is the value containing the hidden layer state at the previous moment and the information at the current moment, W h is a trainable weight parameter, tanh(·) is the activation function, h t is the hidden layer state information at the current moment.
4. The machine translation method capable of learning future information according to claim 2, characterized in that: In step 103), the attention(·) attention mechanism algorithm is used to capture the relationship between the future sentence information and the current sentence information. The calculation formula of this algorithm is as follows: attention(q,k,v)=softmax(α)·v Among them, q, k, and v are the inputs of the algorithm, representing query, key, and value respectively. softmax(·) is a normalization function, and α is the attention weight matrix, which is used to represent the correlation size of each position, and d k is the hidden layer dimension of k.
5. A machine translation method capable of learning future information according to claim 1, characterized in that: In step 3), initialize the model parameters using a uniform distribution, and use an adaptive learning rate algorithm to optimize the model training. The specific process is as follows: 301) Randomly sample the model parameters from a uniform distribution. The specific implementation of this method is: where w is the weight matrix, d in and d out represent the input and output dimension sizes of w respectively, and U(·) represents the uniform distribution; 302) Use an adaptive learning rate algorithm to optimize the model training. The specific formula of this optimization algorithm is: v t = β1v t-1 + (1 - β1)g t s t = β2s t-1 + (1 - β2)g t 2 where v t is the momentum of the mini-batch stochastic gradient at the t-th moment, β1 and β2 are adjustable parameters, g t is the mini-batch stochastic gradient at the t-th moment, x t is the model parameter value at the t-th moment, α is the learning rate, whose value is greater than 0, ∈ is a constant added to avoid the denominator being zero, s t is the momentum at the t-th moment.
Citation Information
Patent Citations
Text-level neural machine translation method based on context memory network
CN111160050A
Neural network machine translation method, model and model forming method
CN111401081A