A machine translation method applied to a stable deep machine translation model
By adopting gradient standardization methods in deep machine translation models, the problems of gradient disappearance and gradient explosion in deep models during training are solved, and a more stable training process and higher computing efficiency are achieved, especially suitable for semi-precision training scenarios.
Patent Information
- Application Number
- CN202210126056.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-02-10
- Publication Date
- 2025-06-20
- Estimated Expiration
- 2042-02-10
AI Technical Summary
Deep machine translation models are prone to gradient vanishing and gradient explosion problems during training, especially when using semi-precision training, which is more serious, resulting in instability in training and high computing resources.
The gradient standardization method is used to calculate the mean and standard deviation of the gradient matrix, and rescaling the gradient values that are too small or too large are rescaling to ensure that the gradient remains within a reasonable range during the backpropagation process.
This method makes the training process of deep models more stable, reduces the risk of gradient vanishing and gradient explosion, improves the training speed and computing efficiency of the model, and performs well in semi-precision training scenarios.
Smart Images

Figure CN114528856B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to a method for training a machine translation model, specifically a machine translation method applied to a stable deep machine translation model. Background Art
[0002] Machine Translation is a classic task in the field of natural language processing, and its goal is to translate the source language (the language to be translated) into the target language (the translated language) by a computer. In today's world of interconnectedness, the significance of language is particularly important. In fields such as politics, business, education, and healthcare, language is an irreplaceable medium for information transmission. Faced with the vast amounts of data on the Internet today, it is difficult for humans to analyze this information on their own. Therefore, machine translation technology, as a powerful translation means, has been widely accepted by people.
[0003] Although scholars put forward the idea of using machines to complete the translation process as early as the 17th century, due to the complexity of human languages and language phenomena such as polysemy and ellipsis, there is still no method that can perfectly translate one language into another. However, machine translation has been developed for a century, and today's machine translation has achieved some results. Looking at the development history of machine translation, the development of machine translation methods has roughly gone through three stages, namely rule-based machine translation methods, statistic-based machine translation methods, and neural network-based machine translation methods. Rule-based machine translation methods originated around the 1930s to 1950s of the last century. This method uses manually defined translation rules to complete the translation process. However, these translation rules need to be defined by linguists, and when there are more and more rules, conflicts will occur between the rules, which makes it difficult to construct a rule-based machine translation system and very difficult to maintain. Therefore, rule-based machine translation methods were quickly replaced by statistic-based machine translation methods. Statistic-based machine translation methods define features manually and use statistical translation models to complete translation based on these features. Although this method does not require manually defining translation rules, the definition of features is still subjective. A good feature may make the results of machine translation very excellent. And the design of features depends on the experience of researchers, which makes the effect of statistic-based machine translation methods unstable. Around 2014, neural networks were introduced into the machine translation task. In a very short time, neural network-based machine translation methods replaced the previous methods and became the most concerned machine translation methods. This machine translation method neither requires manually defining features nor manually defining translation rules. It only needs to convert the text into input data and input the data into the neural network to complete the translation. It greatly improves the performance of the machine translation system and the cost of building the machine translation system. It is the best machine translation method so far.
[0004] Although machine translation technology has been developed for a long time, due to the complexity of human languages, such as language phenomena like polysemy and ellipsis, there is still no method that can perfectly translate one language into another. In response, some researchers' solution is to increase the number of parameters in the model to enhance the representation ability of the machine translation model. One of the most intuitive manifestations is to build a deeper machine translation system with more layers and a more complex structure. However, simply increasing the number of layers of the neural network still faces many problems. For example, when the number of layers of the neural network deepens, the backpropagation process of the neural network is more likely to have problems such as gradient vanishing or gradient explosion, and problems such as unstable training, large consumption of computing time and computing resources will also occur.
[0005] Regarding the problem that deep models are difficult to train, researchers have proposed many technical methods, such as network initialization strategies, learning strategies during training, simplifying the calculations of some layers in the network, adjusting the model structure, etc. to assist in the training of deep networks. On this basis, the experimental results of deep networks have also proven their powerful performance and potential. Although the deep models used in existing machine translation have remarkable effects, they face problems such as a large number of parameters, high training consumption, and difficulty in deployment. Facing the problem of high computational resource consumption of deep networks, an effective solution is to adopt the training method of mixed-precision training. Since this method uses 16-bit floating-point numbers (half-precision) to calculate network parameters instead of the original 32-bit floating-point numbers (full-precision), this method is also called half-precision training. At the same time, in half-precision training, in order to prevent errors from accumulating during the update of network parameters, the storage of parameters in the network still selects full-precision numbers. Compared with the calculation using full-precision numbers in the original network, using half-precision numbers for calculation can greatly accelerate the training speed of the network. However, using half-precision numbers for calculation will inevitably affect the calculation accuracy to a certain extent, and due to the calculation errors during half-precision number calculation, when using half-precision numbers for training, it will also exacerbate the problems of gradient vanishing and gradient explosion that occur in the backpropagation process of neural networks. At this time, in the half-precision scenario, it is very necessary to solve the problems of gradient vanishing and gradient explosion of deep models. Summary of the Invention
[0006] On the premise of using the half-precision method to train deep models, it can be found that for deep models, the occurrence of gradient vanishing and gradient explosion on the one hand stems from the fact that as the number of neural network layers increases, during the training process, the error contained in the gradient information will gradually increase. On the other hand, due to the parameters represented by half-precision having a smaller representation range, it further exacerbates the problems of gradient vanishing and gradient explosion existing in deep networks.
[0007] The present invention provides a training method for a deep machine translation model of machine translation applied to a stable deep machine translation model, that is, a gradient normalization method, which uses the normalization method to rescale too small or too large gradient values, so that the numerical value of the gradient remains within a reasonable range during the backpropagation process without losing gradient information. The gradient normalization method can ensure that when using half-precision to train a deep model, the gradient information of the model is more stable, and the deep model can be trained faster and better.
[0008] To solve the above technical problems, the technical solution adopted by the present invention is:
[0009] The present invention provides a machine translation method applied to a stable deep machine translation model, including the following steps:
[0010] 1) During the training process of the deep model, when performing each backpropagation calculation based on the input source language sentence X, two matrices μ and σ are set for the gradient matrix G of the encoder parameters in the deep model;
[0011] 2) The matrices μ and σ are used to represent the mean and standard deviation of the gradient matrix G respectively, and the values of the matrices μ and σ are calculated based on the gradient matrix G;
[0012] 3) During each backpropagation process during model training, the gradient matrix G is normalized using the matrices μ and σ;
[0013] 4) By means of a preset strategy and prior knowledge, the usage stage and position of this normalization are controlled to perform fast and accurate training of the machine translation model, and finally improve the quality of the output translation during decoding of the translation.
[0014] Step 1) During the training process of the deep model, when performing each backpropagation calculation according to a source language sentence X of length n = {x1, x2,... x n}, two matrices μ and σ are set for the gradient matrix G of the parameters of some layers in the encoder part of the deep model, specifically:
[0015] G ∈ G Encoder_layer
[0016] where the gradient matrix G is the parameter gradient matrix of a certain layer of the encoder. The matrices μ, σ and the gradient matrix G have the same dimensions in the batch dimension and the length dimension. Here, it is assumed that the dimension of the gradient matrix G is (dim Batch_size , dim Len , dim Embeding_size ), then the dimensions of the matrices μ and σ are (dim Batch_size , dim Len ).
[0017] In step 2), the values of the matrices μ and σ are calculated based on the gradient matrix G, specifically:
[0018] The matrices μ and σ are the mean and standard deviation of the gradient matrix G, and the calculation methods of the mean and standard deviation are as follows:
[0019]
[0020] where x k is a parameter in the gradient matrix G; μ i is a parameter in the matrix μ, which represents the mean of the parameters in the gradient matrix G. The parameters are represented by the set S i indicates, m is the number of parameters in the set S i , and the set S ivaries according to the different standardization methods adopted; similarly, σ i is an element in the matrix σ, and ∈ is a non-zero minimum value;
[0021] Step 3) During the backpropagation process of model training, use the matrices μ and σ to perform a standardization operation on the gradient matrix G, and the result LN of the standardization operation LN G can be expressed as:
[0022]
[0023] At this time, when the model performs backpropagation, the gradient information after standardization for the l-th layer can be expressed as:
[0024]
[0025] where Loss is the final loss of the model, L represents the total number of layers of the model, x l represents the input matrix of the l-th layer, F(·) represents the repeatedly used function in the model, and θ k represents the parameter of the function F in the k-th layer of the model.
[0026] Step 4) Control the usage stage and position of this standardization through preset strategies and prior knowledge. The use of gradient standardization is continuous in terms of the number of training rounds and the model structure, ensuring that gradient standardization is applied to consecutive layers of the model, and ensuring that the gradient standardization method is used continuously within a continuous period during the training stage. During the training process, optimize the parameters in the model according to the standardized gradients, and obtain a better optimal translation than the original model output during the decoding process according to the model parameters better optimal translation
[0027] The present invention has the following beneficial effects and advantages:
[0028] 1. The present invention provides a machine translation method applied to a stable deep machine translation model, which can make the training process more stable, is not limited to specific models and tasks, and has good versatility.
[0029] 2. The method of the present invention is orthogonal to other methods for stable model training, such as some initialization methods for deep models and structural adjustments of neural networks, etc.
[0030] 3. Since the method of the present invention makes the gradient information in backpropagation more stable, this method reduces the uncertainty during model training, making the effects of the model easier to reproduce on different devices.
[0031] 4. The method of the present invention can ensure the stable training of the model in the half-precision training scenario without reducing the model performance, making the training process of the model both fast and accurate. BRIEF DESCRIPTION OF THE DRAWINGS
[0032] Figure 1 It is a schematic diagram of the Transformer model structure related to the method of the present invention;
[0033] Figure 2 It is a schematic diagram of the backpropagation calculation related to the method of the present invention;
[0034] Figure 3 It is a position diagram of the gradient normalization in the sublayer of the Transformer model related to the method of the present invention. DETAILED DESCRIPTION OF THE INVENTION
[0035] The present invention will be further described below in conjunction with the drawings of the specification.
[0036] A machine translation method applied to a stable deep machine translation model of the present invention includes the following steps:
[0037] 1) During the training process of the deep model, when performing each backpropagation calculation according to the input source language sentence X, set two matrices μ and σ for the gradient matrix G of the encoder parameters in the deep model;
[0038] 2) Use the matrices μ and σ to represent the mean and standard deviation of the gradient matrix G respectively, and calculate the values of the matrices μ and σ according to the gradient matrix G;
[0039] 3) During each backpropagation process during model training, use the matrices μ and σ to normalize the gradient matrix G;
[0040] 4) Control the usage stage and position of this normalization through preset strategies and prior knowledge to achieve the fast and accurate training of the machine translation model, and finally improve the quality of the output translation when decoding.
[0041] In the present invention, during the training process of the deep model, the matrices μ and σ are used to represent the mean and variance of the gradient matrix G of the parameters in the model, and the gradients of the deep model are normalized through the matrices μ and σ. Using the gradient normalization method during the training process of the deep machine translation model can improve the performance of the model, and this gradient normalization method can make the training process of the deep model more stable, especially under the premise of half-precision training. This method takes the Transformer, a commonly used translation model in current machine translation, as an example to explain this method, and its model structure is as Figure 1 shown. It should be noted that this method is not restricted by the model structure and tasks and has good versatility.
[0042] Step 1) During the training process of the deep model, when performing a backpropagation calculation based on a source sentence X = {x1, x2, …, x n} of length n, for the gradient matrix G of the layer parameters in the encoder part of the deep model, two matrices μ and σ are set, specifically:
[0043] G ∈ G Encoder_layer
[0044] where the gradient matrix G is the parameter gradient matrix of a certain layer of the encoder. The matrices μ, σ and the gradient matrix G have the same dimensions in the batch and length one-dimensional dimensions. Taking the Transformer model as an example, in this model, the dimension of the gradient matrix G can be expressed as (dim Batch_size , dim Len , dim Embeding_size ), and the three dimensions of this vector are determined by the batch size (Batch_size), the sentence length (Len), and the feature dimension (Embeding_size) respectively. At this time, the dimensions of the matrices μ and σ can be expressed as (dim Batch_size , dim Len ).
[0045] Step 2) Use the matrices μ and σ to represent the mean and standard deviation of the gradient matrix G respectively, and calculate the values of the matrices μ and σ according to the gradient matrix G, specifically:
[0046] The matrices μ and σ are the mean and standard deviation of the gradient matrix G, and the calculation methods of the mean and standard deviation are as follows:
[0047]
[0048] where x k is a parameter in the gradient matrix G; μ i is a parameter in the matrix μ, which represents the mean of the parameters in the gradient matrix G, and the parameters are represented by the set S i , m is the number of parameters in the set S i , and the set S i varies according to the different normalization methods adopted; similarly, σ i is an element in the matrix σ, and ∈ is a non-zero minimum value;
[0049] Although there are many commonly used normalization methods in current models, the batch normalization method, which is widely used today, is adopted in this method. Therefore, S i can be expressed by the following formula:
[0050] S i = {k|k E = iE}
[0051] Among them, i E represents an integer value whose range is [0, dim Embeding_size ), and k E represents the index value of the current element k in the feature dimension. This formula means that in batch normalization, according to i E construct dim Embeding_size sets S. According to the set S, all elements with the same feature dimension are normalized once, so batch normalization calculates the mean and variance of the parameters along the batch dimension and the sentence length dimension;
[0052] Step 3) During the backpropagation process of model training, use the matrices μ and σ to normalize the gradient matrix G, and the result LN of the normalization operation G can be expressed as:
[0053]
[0054] When the model performs backpropagation, the gradient of the input of the l-th layer can be expressed as:
[0055]
[0056] Among them, Loss is the final loss of the model, L represents the total number of layers of the model, x l represents the input matrix of the l-th layer, F represents the function reused in the model, and θ represents the parameters of the function F in the model.
[0057] It can be found from the above formula that the right side of the equation consists of two terms. The first term is the derivative of the model loss function with respect to the input of the last layer of the model, and the second term contains gradient information related to the total number of layers of the model. It can be seen that when the number of layers of the model is more, the effect of this term will be exponentially amplified, which is the source of gradient explosion and gradient disappearance. Therefore, normalization can be used in this term to constrain the gradient information.
[0058] Based on the assumption that a more stable training comes from a descent process with low variance and low noise, this method normalizes the gradient, which can be expressed as:
[0059]
[0060] Although, after using the gradient normalization method, only the gradient information of the current layer is changed, and there is still a quantity related to the depth of the model in the gradient information, as shown by the last term on the right side of the above equation. However, in fact, different from other previous methods that intervene in the gradient information, this method performs gradient normalization during the backpropagation of the gradient information. That is, the previous methods always intervene in the gradient after the backpropagation ends, such as in the way of gradient clipping. But in this method, gradient normalization is applied during the reverse calculation process, that is, the normalized gradient information will continue to be passed to the shallow layer of the model, further intervening in the gradient solving process of the shallow layer of the model. This is an important difference between this method and the previous methods, as Figure 2 shown;
[0061] Step 4) Control the usage stage of this normalization through a preset strategy and prior knowledge. For example, use this normalization method in the later stage of the training process. At the same time, it can also be adjusted according to prior knowledge or dynamic strategies in the early stage or middle stage of training, etc. During the training process, optimize the parameters in the model according to the normalized gradient, and a better model parameter can be obtained. During the decoding process, a better optimal translation can be obtained based on this model parameter than the original model output translation a better optimal translation It should be noted that the use of normalization is continuous. This is because using normalization on the gradient at the beginning will cause large fluctuations in the gradient, but as the training progresses, the change of the gradient will tend to be stable. Therefore, during the training stage, it is generally ensured to use this normalization method within a continuous period of time. In addition, from the perspective of each layer of the model, the application of model gradient normalization is also continuous, generally applied to several consecutive layers. And after starting to use this method during the training process, it should be ensured to use this method until the end of the model training. The specific number of application layers and which round of the training stage to start applying this method should be based on the prior knowledge related to the task and the model. When the Transformer model is trained using the iwslt14de-en dataset, this method is recommended to be applied to the shallow layer of the encoder and used in the later stage of the training stage. The application position on the sublayer is as Figure 3 shown.
[0062] Step 5) Using the gradient normalization method during the training process of the deep neural machine translation model can improve the performance of the model. And this gradient normalization method can make the training process of the deep model more stable. Especially under the premise of half-precision training, the deep model is less likely to have the phenomena of gradient vanishing and gradient explosion.
[0063] In this embodiment, the method of the present invention is applied to a deep machine translation model based on half-precision training, which can significantly reduce the problems of gradient vanishing and gradient explosion during the training of the deep model. For example, on an 8-GPU server with 1080Ti GPUs, a 30-layer deep Transformer model can be trained using half-precision. In some cases, this model is extremely prone to gradient vanishing and gradient explosion, resulting in training failure. After using this method, the situation of training failure hardly occurs. This makes the training process of the deep model more stable, which is beneficial for large models such as deep models to provide more stable services on devices with large resources such as cloud platforms. On the iwslt14de-en dataset, an improvement of 0.5 BLEU points can be achieved. The following table compares the models using this method and those not using this method. When translating the same sentence, different translation results can be found. After using this method, the translation of the model is closer to the standard answer.
[0064]
[0065] In addition, this method is not limited to specific hardware and systems, and this method can be used to train a machine translation model more stably on various devices and tasks.
Claims
1. A machine translation method applied to a stable deep machine translation model, characterized in that Including the following steps: 1) During the training process of the deep model, when performing each backpropagation calculation based on the input source language sentence X, set two matrices μ and σ for the gradient matrix G of the encoder parameters in the deep model; 2) Use the matrices μ and σ to represent the mean and standard deviation of the gradient matrix G respectively, and calculate the values of the matrices μ and σ according to the gradient matrix G; 3) During each backpropagation process during model training, use the matrices μ and σ to standardize the gradient matrix G; 4) Control the usage stage and location of the normalization through preset strategies and prior knowledge to quickly and accurately train the machine translation model, and finally improve the quality of the output translation during decoding. The quality of the translation; The values of the matrices μ and σ calculated according to the gradient matrix G in step 2) are specifically: The matrices μ and σ are the mean and standard deviation of the gradient matrix G, and the calculation methods of the mean and standard deviation are as follows: where x k is a parameter in the gradient matrix G; μ i is a parameter in the matrix μ, which represents the mean of the parameters in the gradient matrix G, and the parameters are represented by the set S i where m is the number of parameters in the set S i and the set S i varies according to the different normalization methods adopted; similarly, σ i is an element in the matrix σ, and ∈ is a non-zero minimum value; Step 3) During the backpropagation of model training, use matrices μ and σ to perform a normalization operation on the gradient matrix G, and the result LN of the normalization operation LN G can be expressed as: At this time, when the model performs backpropagation, the normalized gradient information of the l-th layer can be expressed as: Among them, Loss is the final loss of the model, L represents the total number of layers of the model, and x l represents the input matrix of the l-th layer, F(·) represents the function reused in the model, and θ k represents the parameter of the function F in the k-th layer of the model.
2. The machine translation method according to claim 1, applied to a stable deep machine translation model, characterized in that: Step 1) During the training process of the deep model, for each backpropagation calculation based on a source language sentence X = {x1, x2, … x n}, when performing each backpropagation calculation, two matrices μ and σ are set for the gradient matrix G of the encoder layer parameters in the deep model, specifically: G belongs to G Encoder_layer Among them, the gradient matrix G is the parameter gradient matrix of a certain layer of the encoder. The matrices μ and σ have the same dimensions as the gradient matrix G in the batch dimension and the length dimension. Here, it is assumed that the dimension of the gradient matrix G is (dim Batch_size ,dim Len ,dim Embeding_size ), then the dimensions of the matrices μ and σ are (dim Batch_size ,dim Len ).
3. The machine translation method according to claim 1, applied to a stable deep machine translation model, characterized in that: Step 4) Control the usage stage and position of the normalization through preset strategies and prior knowledge. The use of gradient normalization is continuous in terms of the number of training rounds and the model structure, ensuring that gradient normalization is applied to consecutive layers of the model and that the gradient normalization method is used continuously during the training stage. During the training process, optimize the parameters in the model according to the normalized gradient, and obtain a better optimal translation than the output translation of the original model based on the model parameters during the decoding process. A better optimal translation
Citation Information
Patent Citations
Neural machine translation-oriented position coding method and computer storage medium
CN110399619A
Compression method for deep neural machine translation model of small mobile equipment
CN112257469A