A method for improving the performance of a neural machine translation system
By adjusting the absolute position coding generation rules of the neural machine translation system to product-form sine cosine coding, the system's shortcomings in capturing position information ability and robustness are solved, and performance improvement is achieved, which is specifically manifested as an improvement of BLEU value.
Patent Information
- Application Number
- CN202210090738.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-01-26
- Publication Date
- 2025-06-20
- Estimated Expiration
- 2042-01-26
AI Technical Summary
The existing neural machine translation system has shortcomings in capturing position information ability and robustness, which leads to insensitive to the input language word order and is difficult to play an auxiliary role in important scenarios.
By adjusting the absolute position encoding generation rule of the neural machine translation system to product cosine cosine encoding, and using this rule to generate absolute position encoding during training, improving the system's position information capture ability.
Without increasing the number of parameters and calculations, the performance of the neural machine translation system is improved, which is specifically manifested in the improvement of 0.5BLEU values and 0.4BLEU values on WMT14 and WMT16 tasks respectively.
Smart Images

Figure CN114528855B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to a neural machine translation technology, in particular to a method for improving the performance of a neural machine translation system. Background Art
[0002] Machine Translation (MT for short) is a discipline that uses computers to translate between natural languages. It is a branch of natural language processing research and also one of the ultimate goals of artificial intelligence. Compared with human translation, although there is a certain gap in translation quality, the benefits brought by the high efficiency and low cost of machine translation are extremely significant, and it is of great significance for promoting human cultural exchanges.
[0003] Early machine translation research mainly relied on rule-based methods. Especially in the 1970s, expert systems represented by rule-based methods were the most representative research fields in artificial intelligence. Its main idea is to use dictionaries and manually written rule bases as translation knowledge and complete translation with a combination of a series of rules. The drawback of this method is that a large number of linguistics experts are required to construct rules, and the rules formulated are difficult to unify. Even conflicts may occur between manually defined rules, resulting in poor scalability and maintainability of rule-based translation systems.
[0004] Until the 1990s, statistical machine translation gradually emerged. It uses statistical models to automatically learn translation knowledge from monolingual or bilingual corpora. Statistical machine translation uses monolingual corpora to learn language models, uses bilingual parallel corpora to learn translation models, and uses these statistical models to model the translation process. The entire process does not require manual writing of rules, nor does it require constructing translation templates from examples. Whether it is words, phrases, or even syntactic structures, statistical machine translation systems can automatically learn. However, statistical machine translation requires statistical analysis of a large amount of bilingual parallel corpora to construct statistical translation models to complete translation. Since 2005, statistical machine translation has entered a decade of golden age. During this period, various statistical machine translation models emerged in an endless stream, and classic phrase-based models and syntax-based models were also proposed one after another.
[0005] Since 2014, with the development of machine learning technology, neural machine translation based on deep learning has gradually emerged. It has achieved obvious advantages in most tasks in just a few years. In neural machine translation, word strings are represented as real-valued vectors. The translation process is not carried out on discrete words and phrases, but is calculated in the real-valued vector space, so it has essentially changed the way of representing word sequences. In neural machine translation, the sequence-to-sequence conversion process can be implemented by an encoder-decoder framework. The encoder encodes the input source language to form a dense semantic vector, and then the decoder performs autoregressive decoding in combination with the semantic vector to generate the final translation result of the target language. This method does not require additional artificial feature engineering and directly uses neural networks for modeling. Similarly, a large amount of bilingual corpus is required for training.
[0006] Currently, neural machine translation systems have achieved good results. If the neural machine translation model is trained with optimal parameters to achieve a strong enough representation ability, then compared with traditional rule-based machine translation methods and statistical machine translation methods, it has greater advantages in terms of translation speed and quality. However, compared with professional human translation, there is still a significant gap. Therefore, further optimizing the performance of neural machine translation has become a difficult problem to be solved.
[0007] Neural machine translation systems have been widely used in real life, but there are still problems such as poor ability to capture position information and poor robustness. With the increasing popularity of international communication, although neural machine translation has excellent performance. However, in important scenarios, neural machine translation systems can only play an auxiliary role. Therefore, improving the performance of neural machine translation systems has become a key issue in the application of machine translation technology. In traditional neural network system performance methods, a series of complex operations need to be carried out on the system to achieve better performance improvement effects, which are time-consuming and difficult to reproduce, greatly limiting the practical application of machine translation technology. Summary of the Invention
[0008] Aiming at the problems in the existing neural machine translation technology, such as insufficient ability to capture position information, resulting in the neural machine translation system being insensitive to the word order of the input language and poor robustness, the technical problem to be solved by the present invention is to provide an absolute position encoding generation rule, which can improve the performance of the neural machine translation system without increasing the number of parameters and the amount of calculation of the neural machine translation system.
[0009] To solve the above technical problems, the technical solution adopted by the present invention is:
[0010] The present invention provides a method for improving the performance of a neural machine translation system, including the following steps:
[0011] 1) Process the training data and initialize the parameters of the neural machine translation system, and the parameter initialization rule follows the Xavier parameter initialization rule;
[0012] 2) Adjust the absolute position encoding generation rule in the neural machine translation system to a multiplicative sine-cosine encoding generation rule;
[0013] 3) Input the training data, read the absolute position encoding generated in step 2) into the neural machine translation system, add it to the word vector of the input source sentence to obtain a word vector fused with position information, and send it into the neural machine translation model;
[0014] 4) Use the gradient descent method to train the neural machine translation system until convergence, and the training process is consistent with the training process of the existing neural machine translation system;
[0015] 5) During the decoding process, for the absolute position encoding generation rule in the neural machine translation system, it should be consistent with the multiplicative sine-cosine encoding generation rule in step 2).
[0016] In step 1), process the training data and initialize the parameters of the neural machine translation system, and the parameter initialization rule follows the Xavier parameter initialization rule. Specifically:
[0017] 101) Select the parameters to be trained, including each word vector in the vocabulary, the parameters of each layer in the encoder and decoder, and the parameters of the decoder output layer;
[0018] 102) Initialize the parameters in step 101) using the Xavier parameter initialization rule. The specific formula is as follows:
[0019]
[0020] where w represents the parameter to be trained, U represents the uniform distribution, n in and n out represent the input and output dimensions of the parameter to be trained respectively.
[0021] In step 2), adjust the absolute position encoding generation rule in the neural machine translation system to a multiplicative sine-cosine encoding. Specifically:
[0022]
[0023]
[0024] where pos represents the position dimension index of the position encoding, 2j represents the hidden layer dimension index of the position encoding, and d represents the size of the hidden layer dimension of the position encoding.
[0025] In step 4), the neural machine translation system is trained until convergence, and the training process is consistent with that of the existing neural machine translation system;
[0026] 401) Input the training data into the neural machine translation system after modifying the position encoding generation rule, and calculate the objective function L for the training data. The calculation formula is as follows:
[0027]
[0028] where (x, y) represents the input and target of the training, w represents the trainable parameters of the model, and l(·) represents the loss function of the neural machine translation system;
[0029] 402) Backpropagate the gradient of the loss, calculate the gradient of the trainable parameters in the neural machine translation model, and update the parameter formula as follows:
[0030]
[0031] where t represents the number of update steps, α is the learning rate, representing the size of the update step, which needs to be continuously updated and adjusted as the training progresses;
[0032] 403) Continuously update the trainable parameters of the model using the formulas in steps 401) and 402) until the loss of the neural machine translation model for the training data converges.
[0033] The present invention has the following beneficial effects and advantages:
[0034] 1. The method of the present invention improves the absolute position encoding generation rule of the neural machine translation system in machine translation. By expanding the difference in encoding between different positions, the ability of the neural machine translation system to capture position information is ultimately improved to achieve the purpose of improving the performance of the neural machine translation system.
[0035] 2. The method for improving the performance of the neural machine translation system proposed by the present invention achieves a performance improvement effect of 0.5 BLEU value on the WMT14 English-German task, and at the same time, the number of parameters and the amount of calculation of the model do not increase. BRIEF DESCRIPTION OF THE DRAWINGS
[0036] Figure 1 is a flowchart of a method for improving the performance of a neural machine translation system according to the present invention;
[0037] Figure 2 is a schematic diagram of the present invention applied to an existing neural machine translation system. DETAILED DESCRIPTION OF THE INVENTION
[0038] The present invention will be further described below with reference to the accompanying drawings of the specification.
[0039] AsFigure 1 As shown in the figure, a method for improving the performance of a neural machine translation system according to the present invention includes the following steps:
[0040] 1) Initialize the parameters of the neural machine translation system, and the parameter initialization rule follows the Xavier parameter initialization rule;
[0041] 2) Adjust the absolute position encoding generation rule in the neural machine translation system to a multiplicative sine-cosine encoding to generate the absolute position encoding of the neural machine translation system;
[0042] 3) Integrate the adjusted neural network system and read in the training data;
[0043] 4) Use the gradient descent method to train the neural machine translation system until convergence, and the training process is consistent with the training process of the existing neural machine translation system;
[0044] 5) Use the neural machine translation system trained by the present invention for machine translation. As Figure 2 shown, during the decoding process, use the rule in step 2) to generate the absolute position encoding of the neural machine translation system, send the sentence into the neural machine translation system, and decode word by word in an autoregressive manner to obtain the translation result.
[0045] In step 1), process the training data and initialize the parameters of the neural machine translation system. The parameter initialization rule follows the Xavier parameter initialization rule, specifically:
[0046] 101) Select the parameters to be trained, including each word vector in the vocabulary, the parameters of each layer in the encoder and decoder, and the parameters of the decoder output layer;
[0047] 102) Initialize the parameters in step 101) using the Xavier parameter initialization rule. The specific formula is as follows:
[0048]
[0049] where w represents the parameter to be trained in 101), U represents a uniform distribution, n in and n out respectively represent the input and output dimensions of the parameter to be trained.
[0050] This step uses the aligned source language and target language bilingual data as the training data.
[0051] In step 2), adjust the absolute position encoding generation rule in the neural machine translation system to a multiplicative sine-cosine encoding, specifically:
[0052]
[0053]
[0054] Among them, pos represents the position dimension index of the position encoding, 2j represents the hidden layer dimension index of the position encoding, and d represents the size of the hidden layer dimension of the position encoding.
[0055] In step 4), the neural machine translation system is trained until convergence, and the training process is consistent with that of the existing neural machine translation system.
[0056] 401) Input the training data into the neural machine translation system after modifying the absolute position encoding generation rule, and calculate the objective function L for the training data. The calculation formula is as follows:
[0057]
[0058] Among them, (x, y) represents the input and target of the training, w represents the trainable parameters of the model, and l(·) represents the loss function of the neural machine translation system.
[0059] 402) Backpropagate the loss, calculate the gradients of the trainable parameters in the neural machine translation model, and update the parameters with the following formula:
[0060]
[0061] Among them, t represents the number of update steps, α is the learning rate, representing the size of the update step, which needs to be continuously updated and adjusted as the training progresses.
[0062] 403) Continuously update the trainable parameters of the model using the formulas in steps 401) and 402) until the loss of the neural machine translation model for the training data converges.
[0063] In step 5), during the decoding process, for the absolute position encoding generation rule in the neural machine translation system, it should be consistent with step 2).
[0064] The specific process of this method is as Figure 1 shown. The specific application process is as follows: After segmenting the source language text to be translated, it is sent into the model. Parameter calculations are performed in the encoder, and then it is sent into the decoder for word-by-word decoding in an autoregressive manner. Finally, the text of the target language is calculated and output, as Figure 2 shown.
[0065] Taking the calf translator as an example of the application effect, by applying the method for improving the performance of the neural machine translation system, without increasing any number of parameters and computational amount, the baseline system has improved the neural machine translation system by 0.5 BLEU on the WMT14 English-to-German task; on the WMT16 English-to-Romanian task, the baseline system has improved the neural machine translation system by 0.4 BLEU. The present invention defines a new generation rule for absolute position encoding for the neural machine translation system: changing the absolute position encoding rule of the neural machine translation system to a sine-cosine encoding generation rule; during the model training process, using this rule to generate absolute position encoding, so as to improve the ability of the neural machine translation system to capture position information; finally, a neural translation system that is more sensitive to word order information and has better performance is trained.
Claims
1. A method for improving the performance of a neural machine translation system, characterized in that Including the following steps: 1) Process the training data and initialize the parameters of the neural machine translation system, and the parameter initialization rule follows the Xavier parameter initialization rule; 2) Adjust the absolute position encoding generation rule in the neural machine translation system to a product - type sine - cosine encoding generation rule; 3) Input the training data, read the absolute position encoding generated in step 2) into the neural machine translation system, add it to the word vector of the input source sentence to obtain a word vector fused with position information, and send it into the neural machine translation model; 4) Use the gradient descent method to train the neural machine translation system until convergence, and the training process is consistent with the training process of the existing neural machine translation system; 5) During the decoding process, for the absolute position encoding generation rule in the neural machine translation system, it should be consistent with the product - type sine - cosine encoding generation rule in step 2); In step 2), adjust the absolute position encoding generation rule in the neural machine translation system to a product - type sine - cosine encoding, specifically: where pos represents the position dimension index of the position encoding, 2j represents the hidden layer dimension index of the position encoding, and d represents the size of the hidden layer dimension of the position encoding.
2. The method for improving the performance of a neural machine translation system according to claim 1, characterized in that: In step 1), process the training data and initialize the parameters of the neural machine translation system, and the parameter initialization rule follows the Xavier parameter initialization rule, specifically: 101) Select the parameters to be trained, including each word vector in the vocabulary, the parameters of each layer in the encoder and decoder, and the parameters of the decoder output layer; 102) Initialize the parameters in step 101) using the Xavier parameter initialization rule, and the specific formula is as follows: wherein represents the parameter to be trained, represents a uniform distribution, and respectively represent the input and output dimensions of the parameter to be trained.
3. The method for improving the performance of a neural machine translation system according to claim 1, characterized in that: In step 4), train the neural machine translation system until convergence, and the training process is consistent with the training process of the existing neural machine translation system; 401) Input the training data into the neural machine translation system after modifying the positional encoding generation rule, and calculate the objective function with respect to the training data , and the calculation formula is as follows: wherein represents the input and target of training, represents the trainable parameters of the model, represents the loss function of the neural machine translation system; 402) Back - propagate the gradient of the loss, calculate the gradient of the parameters to be trained in the neural machine translation model, and update the parameter formula as follows: Among them represents the number of update steps is the learning rate, which represents the size of the update step and needs to be continuously updated and adjusted as the training process progresses; 403) Continuously update the parameters to be trained of the model using the formulas in step 401) and step 402) until the loss of the neural machine translation model for the training data converges.
Citation Information
Patent Citations
Encoder-decoder framework pre-training method for neural machine translation
CN111382580A
Chinese-Vietnamese unsupervised neural machine translation method based on shared encoder
CN112287694A