A Mongolian-Chinese neural machine translation method based on multiple constraint items

By adding semantic constraints, parameter constraints and vocabulary constraints to the Mongolian-Chinese neural machine translation model, and using dynamic integration strategies of cross-entropy training and reinforcement training, the problems of training sample distribution deviation and semantic loss are solved, and efficient translation of translations and efficient training of models are achieved.

CN114818743BActive Publication Date: 2025-05-27INNER MONGOLIA UNIV OF TECH
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202210277518.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-03-21
Publication Date
2025-05-27
Estimated Expiration
2042-03-21

AI Technical Summary

Technical Problem

The existing Mongolian and Chinese neural machine translation models are prone to training sample distribution deviations and semantic losses during the training process, resulting in low translation quality and low training efficiency.

Method used

The Mongolian and Han neural machine translation method based on multi-constraint terms is adopted. By adding semantic constraints, parameter constraints and vocabulary constraints to the translation model, and combining dynamic integration strategies of cross-entropy training and reinforcement training, the training efficiency of the model and the readability and fluency of the translation are improved.

Benefits of technology

It effectively alleviates the problems of training sample distribution deviation and semantic loss, improves the readability and fluency of the translation, and improves the training efficiency of the model.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114818743B_ABST
    Figure CN114818743B_ABST
Patent Text Reader

Abstract

A Mongolian-Chinese neural machine translation method based on multiple constraint items. First, a model training process based on reinforcement learning is constructed for the Mongolian-Chinese neural machine translation task. Then, on the basis of the reinforcement model, the constraint conditions are further improved for the optimization objective of training, including: adding a semantic constraint module to alleviate the problem of poor translation fluency caused by a single BLEU value evaluation system; performing parameter constraints on the training process to improve the training efficiency of the model; and performing vocabulary constraints on the corpus to reduce the number of out-of-vocabulary words in the translation. The present invention alleviates the problems of poor sequence structure analysis ability and low training efficiency of the model in the low-resource Mongolian-Chinese machine translation task by adjusting the overall constraint method. At the same time, for the problem of large variance caused by reinforcement training, the present invention adopts the methods of mean reward and pruning beam search to effectively alleviate the negative impacts brought by the above reinforcement training.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of machine translation, and particularly relates to a Mongolian-Chinese neural machine translation method based on multiple constraints. Background Art

[0002] When processing Mongolian texts using a deep learning model, it is necessary to perform training constraints on other metrics in addition to the standard BLEU value to ensure the quality of Mongolian translations with scarce resources.

[0003] In the existing translation models, during the training process, due to the scarcity of Mongolian resources, it is easy to generate biases in the distribution of training samples and semantic losses, resulting in low model training efficiency and inconsistencies in the evaluation metrics during the training and inference stages. Summary of the Invention

[0004] In order to overcome the above-mentioned drawbacks of the prior art, the purpose of the present invention is to provide a Mongolian-Chinese neural machine translation method based on multiple constraints to improve the training efficiency of the translation model, enhance the readability and fluency of the translation, and improve the Mongolian-Chinese neural machine translation model.

[0005] In order to achieve the above purpose, the technical solution adopted by the present invention is as follows:

[0006] A Mongolian-Chinese neural machine translation method based on multiple constraints includes the following steps:

[0007] Step 1, on the basis of taking the BLEU value as the training optimization target, add constraint conditions to construct a Mongolian-Chinese neural machine translation model based on reinforcement learning. The constraint conditions include:

[0008] (1) Semantic constraint to alleviate the problem of poor translation fluency caused by a single BLEU value evaluation system;

[0009] (2) Parameter constraint to improve the training efficiency of the model;

[0010] (3) Vocabulary constraint on the corpus to reduce the number of out-of-vocabulary words in the translation;

[0011] Step 2, train the constructed translation model. The training strategies include cross-entropy training, reinforcement training, and the dynamic integration of the two trainings.

[0012] In one embodiment, in step 1, a translation model is constructed, mapping the reinforcement learning algorithm into machine translation. The learning body, environment, parameter optimization method, exploration action, learning body parameter state, and reward in reinforcement learning respectively correspond to the translation module, input at different time steps, model training method, sequence decoding, network parameters, and BLEU value of the sequence in machine translation; the input at different time steps is the Mongolian stem and affix or word representation unit, and the sequence decoding is decoding with one element as a unit.

[0013] In one embodiment, the construction of the mapping regards machine translation as a process of exploring rewards for a single variable by the Markov decision method, that is:

[0014] The learning body continuously changes its parameter state through iterative exploration actions to obtain more rewards, and continuously interacts with the environment based on this iterative training. The process is expressed as M = <S, A, P{s, a}, R>, where s ∈ S, s represents the parameter state of the learning body, S is the finite set of the learning body parameter states, used to represent the network parameters of the translation model at different time steps during training; a ∈ A, a represents the exploration action, A is the finite set of exploration actions, used to record the exploration actions of the learning body in different states; P{s, a} represents that the learning body predicts the next parameter state s′ according to the parameter state s and exploration action a at the current time step; R represents the immediate reward obtained by the learning body after taking the exploration action a.

[0015] Based on the above mapping, the translation model obtains rewards and punishments in the exploration actions by continuously optimizing the network parameters from a random state, and finds the optimal network parameters under the reward and punishment mechanism, then continuously updates and makes the optimal choice.

[0016] In one embodiment, for cross-entropy training, at each time step, the translation model takes the element vector e t ∈ E in the standard translation E and the output h t of the neural network hidden layer as inputs, calculates the output h t+1 of the hidden layer at the next time step until the translation model finally outputs the complete translation vector E′. Among them, E is expressed as E <e 1 , e 2 , …, e t , …, e T >, e t is the element vector encoded at the t-th time step in E, and T is the vector length of E; E′ is expressed as E′ <e′ 1 , e′ 2 , …, e′ t , … e′ T >, e′ t is the vector element decoded at the t-th time step in E′; ht is a floating-point real-valued vector used to encode the results of the computational output before the current time step during training;

[0017] In the reinforcement training, the translation model makes corresponding decoding actions according to the REINFORCE algorithm with the BLEU value as the optimization objective, and observes the reward by comparing the predicted sequence in the candidate set composed of multiple translations output by the current model according to probability with the best sequence. The sequence length is equal to the vector length of the vector E, that is, T.

[0018] In one embodiment, the dynamic integration refers to dynamically adding reinforcement training and cross-entropy training to the overall training process of the translation model within the limited training cycles of Mongolian-Chinese machine translation. The execution algorithm is as follows:

[0019] 1) Perform stem-and-affix-level embedding training on Mongolian to obtain the corresponding word vector space, and given three variables: the cross-entropy training round N CE and the reinforcement training round N RL as well as the sequence length T;

[0020] 2) In each round of training iteration, starting from the sequence length as the starting flag and 1 as the ending flag, use cross-entropy as the loss function to train for N CE rounds;

[0021] 3) With Δ as the difference (Δ is a training hyperparameter, set to 2 in the Mongolian-Chinese machine translation task and set to 3 in Mongolian-Cyrillic Mongolian), gradually reduce the number of cross-entropy training rounds and correspondingly increase the number of reinforcement training rounds N RL ;

[0022] 4) Use N RL for training, use cross-entropy training in the first T - Δ steps, and use reinforcement training for the remaining training steps until the reinforcement training fully participates in the overall training and the translation model converges.

[0023] In one embodiment, in the process of constructing the translation model, Mongolian is subjected to embedding training at the stem-and-affix granularity to obtain the corresponding Mongolian vector space representation, and a suffix-level input reward is established for the Mongolian stem-and-affix representation during the reinforcement training process.

[0024] In one embodiment, in both the cross-entropy training and the reinforcement training, the context vector c t is used as the input, and the context vector c t encodes the context to be used when generating the sequence-level output; the translation model learns a recursive function to calculate h t+1 and outputs the element vector e′ t+1, the element vector constitutes a certain element word in the translated text according to the subsequent model decoding probability:

[0025] h t+1 = f(e t , h t , c t ) (1)

[0026] e′ t+1 ~p θ (e′|e t , h t+1 ) (2)

[0027] All activation functions f() are selected as rectified linear unit functions, ~ represents following a distribution, p θ () represents network parameters, which is a unified representation of the floating-point vectors of the weight matrix, the hidden layer output matrix, and the context vector matrix;

[0028] Among them, in cross-entropy training, the cross-entropy loss is adjusted according to the current parameter state set of the translation model to maximize the probability of sequence decoding. For the standard translation E, the cross-entropy loss loss ~CrossEntrop aims to minimize the following formula:

[0029]

[0030] e′ t+1 = argmaxp θ (e|e t , h t+1 ) (4)

[0031] The overall training objective of the translation model is to maximize p θ (e′|e t , h t+1 );

[0032] In reinforcement training, the reward for the best sequence comes from the standard translation. The objective of the reinforcement training is to obtain the network parameters of the translation model that maximize the expected return reward, and the loss is defined as the negative expectation of the reward:

[0033]

[0034] During the decoding process, since the training of the overall model is based on probability calculation to obtain the result, after a sequence decoding is completed, e′ t can also represent the t-th element in the generated sequence, and r() represents the reward corresponding to the sequence decoded and generated by the translation model;

[0035] In reinforcement training, the negative expectation is approximated by randomly sampling or precisely sampling a single word in the output sequence of the translation model, so that the gradient of the loss can be calculated.

[0036] In one embodiment, the semantic constraint participates in the reward calculation in each iteration by comparing the standard translation with the predicted output. As part of the common reward, the semantic constraint reward Rs sem The sequence-level return reward is obtained by calculating the cosine angle sem(e′, e) between the element vectors e′ and e in two sentences. The calculation process is expressed as:

[0037]

[0038] BLEU reward Rb BLEU It is expressed as:

[0039]

[0040]

[0041] where p T represents the precision term calculated by the standard N-gram grammar, w T represents the weight matrix, and bp is the penalty term for the standard length, which weakens the reward according to the comparison result of the decoded output sequence length c and the standard sequence length l; The training return reward based on reinforcement learning is expressed as a joint formula with μ as the control term:

[0042] Reward = μRs BLEU +(1 - μ)Rb sem (9)

[0043] According to the above joint formula, guide the translation model to generate a stable sequence;

[0044] The parameter constraint means that during the model training, after the loss calculation in each round of iteration, the current training iteration is mapped to the execution range of the problem by using gradient mapping, and then the current gradient descent calculation is performed to complete the error transmission; In Mongolian-Chinese machine translation, the soft function constraint is used to improve the objective function and the penalty coefficient, which plays a role in vector normalization, that is:

[0045] Min: R(x) + λc(x) 2

[0046] where R(x) and c(x) are the objective function and the penalty function respectively. R(x) and c(x) are solved alternately by the gradient mapping method, and finally the minimum solution of the loss function is achieved. In this method, the parameter constraint aims to reduce the redundant parameter input in training and achieve the same or approximate convergence effect in a low-volume model parameter;

[0047] The vocabulary constraint means that when training the corpus for vector representation, low-frequency words with a word frequency lower than 3 in the vocabulary are replaced with the original words or fixed words. At the same time, during the decoding process, when performing Beamsearch column search on the original words, the constructed candidate sequences are trimmed, and the sequences with more low-frequency words and the last 5 sequences with lower probabilities are trimmed.

[0048] In one embodiment, the training reward is implemented in one of the following two ways:

[0049] 1) In each batch-size training, the translation model uses the BLEU reward R for the first batch, that is, the training sequences of batchsize - Δ, BLEU , and uses the common reward Reward = R for the remaining Δ training sequences BLEU +R sem , where batchsize is the batch processing size. In subsequent training, the number of sequences using the cross-entropy loss is reduced by Δ in each batchsize and this process is repeated, thereby obtaining the BLEU reward R BLEU and the semantic constraint reward R sem calculation results. That is, the training cycle of the translation model starts from batchsize and ends until it is finally reduced to 1;

[0050] 2) Set μ as a hyperparameter to adapt to different tasks.

[0051] In one embodiment, after adding the constraint conditions, the decoding process of the translation model is as follows:

[0052] 1) Perform embedding training on the Mongolian-Chinese bilingual corpus to obtain vector representation;

[0053] 2) Pre-train the translation model to obtain the initial state space of reinforcement training;

[0054] 3) Calculate the similarity between elements of the output sequences based on 2) and obtain the semantic constraint reward;

[0055] 4) Calculate the reinforcement loss according to the obtained BLEU reward and semantic constraint reward based on the following algorithm:

[0056]

[0057] 5) Decode the sampled text, p(y t ) = softmax(W s f(y t )(co, y t-1 , h t-1 ) + b z ), where p(y t) represents the probability obtained by the target output word at time t, co represents the compressed representation of the Mongolian vector representation, h t-1 represents the output of the hidden layer state at time t-1, b z represents the calculation of the offset, y t represents the model output result at time t, W represents the weight between neural nodes, and f() represents the activation function of the neural network;

[0058] 6) Based on the hidden layer state decoded at the t-th time step, the context vector, and the target word y at time t-1 t-1 predict the probability of the output word y t .

[0059] Compared with the existing Mongolian-Chinese neural machine translation algorithms, the present invention first adds a reinforcement learning algorithm to the training of the translation model, solving the problems of long-existing training sample distribution deviation and inconsistent evaluation metrics in Mongolian-Chinese machine translation; secondly, for the reward calculation method of reinforcement learning, the Mongolian language is calculated with word stem and affix level character rewards, which can effectively explore the associations between elements in the sequence, and a dynamic integration algorithm is used to make dynamic adjustments between reinforcement training and conventional training, enabling the model to be fully trained and alleviating the problem of insufficient model training caused by data sparsity; finally, by adding a semantic constraint term in the reinforcement training, the problem of semantic loss caused by the model training being solely optimized with the BLEU value is alleviated, solving the problem of poor translation fluency that has long existed in Mongolian-Chinese translation and improving the readability of the translation. Brief Description of the Drawings

[0060] Figure 1 is an end-to-end Mongolian-Chinese neural machine translation architecture based on the attention mechanism.

[0061] Figure 2 shows the Mongolian word stem and affix mapping algorithm.

[0062] Figure 3 is an example of the training sample distribution deviation in Mongolian-Chinese machine translation.

[0063] Figure 4 is a diagram of the model sampling process based on dynamic integration.

[0064] Figure 5 is the problem of inconsistent evaluation metrics in the training and inference stages.

[0065] Figure 6 is the training process of the reinforcement model.

[0066] Figure 7 is the overall structure diagram of the model training based on reinforcement learning. Detailed Embodiments

[0067] The implementation manners of the present invention will be described in detail below with reference to the accompanying drawings and embodiments.

[0068] First, a brief introduction to the Mongolian-Chinese neural machine translation process is given:

[0069] (1) Mongolian-Chinese neural machine translation

[0070] Principle: The Mongolian-Chinese neural machine translation uses a deep learning training model to convert Mongolian (or Cyrillic Mongolian) text into Chinese text. The principle is to preprocess the Mongolian text corpus data through denoising and segmentation so that the corresponding embedded vector representation can be better characterized and feature-extracted in the model. The model establishes a data mapping relationship (model parameters) based on bilingual (Mongolian-Chinese) under the guidance of a selected machine learning algorithm, and then through multiple iterative trainings, this mapping relationship gradually becomes accurate, and finally has good generalization ability for unknown test data.

[0071] Problem: Mongolian machine translation has long been limited by the lack of parallel corpora and scarce resources, resulting in a relatively serious data sparsity problem in the data participating in model training. In specific Mongolian-Chinese translation, the above problems are mainly reflected in that the model cannot converge quickly and has low training efficiency within limited training data and training cycles, resulting in an unsatisfactory BLEU value evaluation of the translation.

[0072] (2) Neural network training process

[0073] Principle: The present invention adopts an end-to-end sequence model architecture based on attention. In this architecture, both bilingual texts appear in the form of discrete vectors. ① The input Mongolian Source: X(x 1 , x 2 , …, x n ) and the output Chinese Target: Y(y 1 , y 2 , …, y n ) are all discrete characters; ② The sentences of unequal lengths in the corpus are converted into sequences of equal length (n) through padding as the model input; ③ The mapping lengths between the two ends of the input and output are not equal, with the identifier <eos>is the sequence end flag, and the specific structure is as Figure 1 shown.

[0074] Problem: There are biases in the distribution of training samples and inconsistencies in the evaluation metrics during the training and inference phases. For the sake of easy understanding, taking Figure 2 Mongolian-Chinese test sentences as an example, during the training phase, in an LSTM decoding process that unfolds in sequence, the hidden layer output of the previous time step and the standard label 'agglutinative' are involved in predicting the 'language' label. However, during the inference phase, since there is no standard label, the hidden layer output of the previous time step and the wrongly predicted label 'adhesive' of the model itself at the previous time step are involved in predicting 'language', and 'adhesive' is the wrong output of the model, which will cause this error to be continuously transmitted in subsequent decoding.

[0075] In addition, during the training phase, the quality of model training is mainly measured by calculating the loss function, while during the inference phase, the BLEU value calculation is used to measure the decoding effect of the model, as Figure 3 shown.

[0076] (3) Reinforcement learning method

[0077] Principle: Reinforcement learning is a method of rewarding exploration of data through a reward and punishment mechanism, and continuously updating the model's records and judgments of rewards and punishments during iterative training. The main idea of applying reinforcement learning in machine translation is to map each module in machine translation to different entities in the reinforcement training process, use reinforcement training to obtain rewards and punishments by continuously trying and exploring from a random state, find corresponding rules under such a reward and punishment mechanism, and then continuously update the algorithm characteristics of making optimal choices to update and train the parameters of the translation model. Reinforcement learning obtains optimal training parameters by predicting future actions in a limited Mongolian-Chinese parallel corpus dataset. It can greatly alleviate the demand for large-scale parallel corpora in the resource-scarce Mongolian-Chinese machine translation task. In addition, optimizing the evaluation metrics of text processing by using the characteristic of targeted optimization of the training objective by reinforcement learning is also one of the features of the present invention.

[0078] Problem: Applying reinforcement learning to the Mongolian-Chinese neural machine translation task is a brand-new attempt for low-resource translation tasks. Its biggest technical feature is that reinforcement training can transform the evaluation metrics of machine translation into the training objective of the overall model. This strategy enables the model to optimize the final evaluation content during continuous training, and can effectively alleviate the drawback of insufficient training in Mongolian-Chinese machine translation.

[0079] In the present invention, the biggest problem encountered in applying the reinforcement learning method is the need to consider the monotonic optimization problem brought about by reinforcement training. So-called monotonic optimization means that after applying reinforcement training, the overall training of the model will monotonically use the BLEU value of the machine translation evaluation metric as the parameter update target for each iteration, while ignoring the fluency of the overall translation. This seemingly contradicts the credibility of the BLEU score, but it is prevalent in most low-resource translation tasks because the calculation of model parameters and the direction of gradient update in reinforcement training will completely follow the N-gram method of BLEU to obtain the corresponding candidate set probabilities, resulting in poor readability of the translation. This is also the main problem to be solved by the present invention, that is, how to apply reinforcement learning to Mongolian-Chinese neural machine translation and, on this basis, improve the reinforcement training process to solve the training sample distribution deviation and semantic loss problems in low-resource tasks.

[0080] Based on this, a Mongolian-Chinese neural machine translation method based on multiple constraint terms of the present invention aims to construct a Mongolian-Chinese neural machine translation model based on reinforcement learning, add multiple constraint modules on the basis of using the BLEU value as the training optimization target, adjust the constraint training strategy, and integrate reinforcement training into conventional training through a dynamic integration algorithm for the Mongolian-Chinese translation task, and finally realize neural machine translation according to the current training strategy.

[0081] Specifically, it includes the following steps:

[0082] Step 1, on the basis of using the BLEU value as the training optimization target, aiming at the agglutinative characteristics of Mongolian and the training optimization target, add constraint conditions to construct a Mongolian-Chinese neural machine translation model based on reinforcement learning. Among them, the constraint conditions of the present invention include:

[0083] (1) Semantic constraint to alleviate the problem of poor translation fluency brought about by a single BLEU value evaluation system;

[0084] (2) Parameter constraint to improve the efficiency of model training;

[0085] (3) Vocabulary constraint on the corpus to reduce the number of out-of-vocabulary words in the translation;

[0086] Step 2, train the constructed translation model, and the training strategies include cross-entropy training (CrossEntropy, CE), reinforcement training (REINFORCE), and the dynamic integration of the two trainings.

[0087] Specifically, the construction of the translation model of the present invention maps the reinforcement learning algorithm to machine translation, and solves the sampling problem of training sample distribution deviation by means of a dynamic sampling algorithm and a reward training mechanism in the way of formulating rewards. The specific mapping relationship is shown in Table 1.

[0088] Table 1 Correspondence between the constructed Markov property of reinforcement learning and each entity in machine translation

[0089]

[0090]

[0091] Among them, the construction of the mapping is based on Markov <m>The decision-making method regards machine translation as a process of exploring rewards with a single variable, that is:

[0092] Learning body <agent>Through iterative exploration actions Continuously change its parameter status <s>To obtain more return rewards <r>, and continuously interact with the environment based on this iterative training. The process is represented as M = <S, A, P{s, a}, R>, where s ∈ S, s represents the parameter state of the learning entity, S is a finite set of the parameter states of the learning entity, used to represent the network parameters of the translation model at different time steps during training; a ∈ A, a represents the exploration action, A is a finite set of exploration actions, used to record the exploration actions of the learning entity in different states; P{s, a} represents that the learning entity predicts the next parameter state s′ according to the parameter state s and the exploration action a at the current time step; R represents the immediate reward obtained by the learning entity after taking the exploration action a, and R = R(s, a).

[0093] Based on the mapping, the translation model continuously optimizes the network parameters from a random state, thereby obtaining rewards and punishments in the exploration actions, and finding the optimal network parameters under the reward and punishment mechanism, then continuously updating and making the optimal choice.

[0094] Another function of the reinforcement learning of the present invention is to integrate the targeted training of the optimization objective into the training process of machine translation, and solve the problem of inconsistent evaluation metrics in the training stage and the inference stage in conventional training. The main idea of the method is to optimize the BLEU value of the evaluation metric of machine translation by using the characteristics of targeted optimization of the training objective by reinforcement learning, and combine the Markov decision algorithm to explore the set of optimal parameters in the unknown state. The biggest feature is to use the reward mechanism to weigh the optimal solution in the whole state in the unknown state.

[0095] In the construction process of the translation model of the present invention, Mongolian can also be embedded and trained at the stem and affix granularity to obtain the corresponding vector space representation of Mongolian, so as to establish an affix-level input reward for the Mongolian stem and affix representation during the reinforcement training process. Compared with other granularities, the stem and affix granularity representation is more accurate in the character reward calculation of reinforcement training and can better reflect the sequence-level reward. However, the training strategy and model architecture provided by the present invention are not limited to the granularity of the corpus or the input representation. The algorithm for obtaining the Mongolian stem and affix mapping is as< / r> < / s> <s> Figure 2 as shown

[0096] Applying reinforcement learning in Mongolian-Chinese machine translation makes the training objective more targeted, but there are also more corresponding training conditions and requirements. For example, general reinforcement training requires a large amount of computing resources. Therefore, in order to obtain the optimal solution faster, a pre-trained model can be pre-trained to ensure that the initial state can be maintained in a better state.

[0097] In the cross-entropy training of the present invention, the translation model takes the element vector e t ∈E in the standard translation E and the output h t of the neural network hidden layer as inputs at each time step, and calculates the output h t+1 of the hidden layer at the next time step until the translation model finally outputs the complete translation vector E'. Among them, E is expressed as E <e 1 , e 2 , …, e t , …, e T >, e t is the element vector encoded at the t-th time step in E, and T is the vector length of E; E' is expressed as E' <e' 1 , e' 2 , …, e' t , … e' T >, e' t is the vector element decoded at the t-th time step in E'; h t is a floating-point real-valued vector used to encode the result of the computational output before the current time step in the training.

[0098] In the reinforcement training, the translation model makes corresponding decoding actions according to the REINFORCE algorithm with the BLEU value as the optimization target, and observes the reward by comparing the predicted sequence with the best sequence in the candidate set composed of multiple translations output by the current model according to probability at the end of each output predicted sequence or after the sequence length. The sequence length is equal to the vector length of E, that is, T.

[0099] In the cross-entropy training and the reinforcement training, the context vector c t can be used as an input. The context vector c t encodes the context to be used when generating the sequence-level output; the translation model learns a recursive function to calculate h t+1 , and outputs the element vector e' t+1 at the (t + 1)-th time step. This element vector constitutes a certain element word in the translation according to the subsequent model decoding probability:

[0100] h t+1 = f(e t , h t ,c t ) (1)

[0101] e′ t+1 ~p θ (e′|e t ,h t+1 ) (2)

[0102] In the present invention, all activation functions f() are selected as the Rectified Linear Unit (Relu). ~ indicates following a distribution, and p θ () represents network parameters, which is a unified representation of the floating-point vectors of the weight matrix, the hidden layer output matrix, and the context vector matrix.

[0103] Among them, in cross-entropy training, the cross-entropy loss is adjusted according to the current parameter state set of the translation model to maximize the probability of sequence decoding. For the standard translation E, the cross-entropy loss Loss ~CrossEntropy aims to minimize the following formula:

[0104]

[0105] e′ t+1 =argmaxp θ (e|e t ,h t+1 ) (4)

[0106] The overall training objective of the translation model is to maximize p θ (e′|e t ,h t+1 ).

[0107] In reinforcement training, the reward for the best sequence comes from the standard translation (in the present invention, the reinforcement rewards of the standard translation are all regarded as 1). The objective of reinforcement training is to obtain the network parameters of the translation model that maximize the expected reward. Therefore, the calculation method of the present invention is to define the loss as the negative expectation of the reward:

[0108]

[0109] During the decoding process, since the training of the overall model is based on probability calculation to obtain the result, after a sequence decoding is completed, e′ t can also represent the t-th element in the generated sequence, and r() represents the reward corresponding to the sequence decoded by the translation model.

[0110] In actual reinforcement training, the negative expectation is approximated by randomly sampling or precisely sampling a single word in the output sequence of the translation model, so as to reasonably calculate the gradient of the loss.

[0111] Furthermore, in the reinforcement training, the more complex the model is, the more likely it is to have the problem of excessive variance. To prevent the training process from being affected, the present invention sets up a reward mean to alleviate the problem of excessive variance during gradient calculation. represents the average reward of the result at the (t + 1) time step.

[0112] The present invention adjusts the overall constrained training method and finally realizes Mongolian-Chinese reinforcement neural machine translation according to the current training strategy. The present invention constructs a corresponding multi-constrained training strategy in the training of the translation model according to the semantic characteristics of Mongolian, so as to alleviate the problems of poor sequence structure analysis ability and low training efficiency of the model in the low-resource Mongolian-Chinese machine translation task, and solves the problems of sampling bias of neural machine translation training sample distribution and inconsistent evaluation indicators by means of a dynamic sampling algorithm and a reward training mechanism. At the same time, for the problem of large variance caused by reinforcement training, the present invention adopts the methods of mean reward and pruning beam search to effectively alleviate the negative impacts brought by the above reinforcement training.

[0113] In the present invention, dynamic integration means that within the limited training cycle of Mongolian-Chinese machine translation, reinforcement training and cross-entropy training are dynamically added to the overall training process of the translation model, so that reinforcement training can guide the training in conventional training. It uses a reinforcement learning algorithm for sequence-level training, directly optimizes the BLEU value through reinforcement training, and then gradually introduces reinforcement training and cross-entropy training into the overall training process during the specific training process. The core idea of the algorithm is to set a hyperparameter to control the participation degree of two optimization methods, namely the CE loss and the Reinforcement Learning (RL) loss, in each round of iterative training, and finally make the reinforcement training fully participate in the entire training process to solve the serious problems of training sample distribution deviation and inconsistent training and evaluation indicators caused by data sparsity in Mongolian-Chinese neural machine translation. Refer to Figure 4 , and the specific implementation of the algorithm is as follows:

[0114] 1) Perform stem and affix level embedding training on Mongolian to obtain the corresponding word vector space, and give three variables: the cross-entropy training round N cE 、the reinforcement training round N RL and the sequence length T;

[0115] 2) In each round of training iteration, starting from the sequence length T and ending with 1, use cross-entropy as the loss function to train for N CE rounds;

[0116] 3) With Δ as the difference (Δ is a training hyperparameter, set to 2 in the Mongolian-Chinese machine translation task and set to 3 in Mongolian-Cyrillic Mongolian), gradually reduce the number of cross-entropy training rounds and correspondingly increase the number of reinforcement training rounds N RL ;

[0117] 4) Use N RL for training. In the first T - Δ steps, use cross - entropy training, and for the remaining training steps, use reinforcement training until the reinforcement training fully participates in the overall training and the translation model converges.

[0118] On this basis, the present invention adds semantic constraints. That is, semantic constraints are added on the basis of the constructed reinforcement training and used as part of the training objective to jointly guide the training of the model with the training objective based on BLEU. The semantic constraint term constructed in the present invention is that when the model decodes to the end of the sequence, a reward will be obtained from the calculation of the BLEU value and the constructed semantic constraint term. The specific method is to jointly use the semantic constraint term and the BLEU value as controllable terms to form the optimization objective of the model. Similar to the sequence - level BLEU reward, the semantic constraint participates in the reward calculation in each iteration by comparing the standard translation with the predicted output and is used as part of the common reward (Reward). The semantic constraint reward R sem The sequence - level reward is obtained by calculating the cosine angle sem(e′, e) between the element vectors e′ and e in two sentences. The calculation process is expressed as:

[0119]

[0120] BLEU reward R BLEU is expressed as:

[0121]

[0122]

[0123] where p T represents the precision term calculated by the standard N - gram grammar, w T represents the weight matrix, and bp is the penalty term for the standard length, which weakens the reward according to the comparison result between the length c of the decoded output sequence and the length l of the standard sequence. In summary, the training reward based on reinforcement learning can be expressed as a joint formula with μ as the control term:

[0124] Reward = μR BLEU +(1 - μ)R sem (9)

[0125] Guide the translation model according to the above joint formula to generate a stable sequence.

[0126] Parameter constraint means that during model training, after the loss calculation in each iteration, the current training iteration is mapped to the execution range of the problem using gradient mapping, and then the current gradient descent calculation is performed to complete the error transmission; in the Mongolian-Chinese machine translation of the present invention, a soft function constraint is used to improve the objective function and the penalty coefficient, which plays a role in vector normalization, that is:

[0127] Min:R(x)+λc(x) 2

[0128] Where R(x) and c(x) are the objective function and the penalty function respectively. The penalty function is used to solve the optimization problem under constraints. In the neural machine translation process of the present invention, the constrained loss function can be transformed into an unconstrained objective function through the penalty function, which can greatly reduce the computational difficulty during training. λ represents the penalty coefficient, which is set as a hyperparameter in model training in the present invention. For the application of the penalty function during training, parameter constraints on the loss function can be performed if and only if the loss function reaches an extreme point. Optionally, the penalty function can be selected from classical interior penalty functions or exterior penalty functions. The only requirement is that the selected penalty function can be approximated by a second-order Taylor expansion to obtain an estimate. By using the gradient mapping method to alternately solve R(x) and c(x), the minimum solution of the loss function is finally achieved. In this method, parameter constraints are aimed at reducing redundant parameter input during training and achieving the same or approximate convergence effect with low-volume model parameters.

[0129] Vocabulary constraint means that when training the corpus to make vector representations, the vocabulary is appropriately trimmed to alleviate the phenomenon of a large number of out-of-vocabulary words during the decoding process. Specifically, it can be: replacing all low-frequency words with a frequency lower than 3 in the vocabulary with the original words or fixed words. At the same time, during the decoding process, when performing Beamsearch beam search on the original words, the constructed candidate sequences are trimmed, and the sequences with more low-frequency words and the last 5 orders of probabilities are trimmed (in Mongolian-Chinese translation, generally ten translation candidate sequences are constructed according to probabilities). This method can reduce the phenomenon of out-of-vocabulary words and unknown symbols in the translation, avoid the computational burden brought by multiple redundant beam search processes to the model, and ensure the translation accuracy.

[0130] Based on the above reward method, for the Mongolian-Chinese translation task, the training return reward of the present invention is implemented in one of the following two ways:

[0131] 1. In each batch-size training, the translation model uses the BLEU reward R for the first batch of training sequences, that is, batchsize - Δ BLEU , and uses the common reward Reward = R for the remaining Δ training sequences BLEU +R sem , where batchsize is the batch processing scale. In the Mongolian and Cyrillic Mongolian translation tasks applied in the present invention, Δ is set to [(10%-15%)*batchsize]. In subsequent training, the number of sequences using cross-entropy loss is reduced by Δ in each batchsize and this process is repeated, thereby obtaining the BLEU reward R BLEU and the semantic constraint reward R sem Calculation result, that is, the training cycle of the translation model starts from batchsize and ends until it is finally reduced to 1;

[0132] 2. Set μ as a hyperparameter to adapt to different tasks.

[0133] 1) During the batch sequence training process, the reinforcement reward R_BLEU is gradually applied to the overall training process. Specifically, in the training of the first batch size (batchsize), the reinforcement reward R_BLEU is used except for the difference Δ, and in the subsequent remaining Δ sequences, the BLEU reward R_BLEU and the semantic constraint reward R_sem, that is, (R_BLEU+R_sem), are used. The value of Δ ranges from [(10%-20%)*batchsize]. The final reward is to reduce (Δ) the number of sequences using cross-entropy loss for each batch and repeat the training until the batch iteration ends.

[0134] 2) Control the ratio of the two rewards by setting hyperparameters to adapt to the Mongolian-Chinese translation tasks with different corpus scales. This method is mainly applied in the translation training involving Cyrillic Mongolian.

[0135] The training strategy of the present invention is similar to that of a multi-class logistic regression classifier. The gradient of the loss is the difference between the predicted value and the 1 / N representation of the actual target word, which can be expressed as:

[0136]

[0137]

[0138] After adding the constraint conditions, the decoding process of the translation model of the present invention is as follows:

[0139] 1). Embed and train the Mongolian-Chinese bilingual corpus to obtain vector representations;

[0140] 2). Pre-train the translation model to obtain the initial state space of reinforcement training;

[0141] 3). Calculate the similarity between elements of the output sequences on the basis of 2) and obtain the semantic constraint reward;

[0142] 4). According to the obtained BLEU reward and semantic constraint reward, perform reinforcement loss calculation according to the following algorithm:

[0143]

[0144] 5), Decode the sampled text, p(y t ) = soft max(W s f(y t )(co, y t-1 , h t-1 ) + b z ), where p(y t ) represents the probability obtained by the target output word at time t, co represents the compressed representation of the Mongolian vector representation, h t-1 represents the output of the hidden layer state at time t - 1, b z represents the calculation offset, y t represents the model output result at time t, W represents the weight between neural nodes, and f() represents the activation function of the neural network;

[0145] 6), Based on the hidden layer state decoded at the t-th time step, the context vector, and the target word y t-1 at time t - 1, predict the probability of the output word y t .

[0146] In the translation model of the present invention, to solve the problem of semantic loss caused by a single BLEU value as the optimization target, the present invention adds a semantic constraint term and a parameter constraint term in the reinforcement training part, and takes them as part of the reinforcement reward to participate in the reward calculation of the overall training, as follows:

[0147] Reward = μR BLEU + (1 - μ)R sem (12)

[0148]

[0149]

[0150] Then, according to the actual model training situation in Mongolian-Chinese neural machine translation, one of the aforementioned two training return reward methods is adopted, so as to guide the model to generate stable sequence results with two rewards.

[0151] Based on the above principle, Figure 5 The more detailed training process of the translation model of the present invention is given, including: (1) constructing a pre-training model to provide a better initial parameter space for reinforcement training; (2) sampling: different from the conventional training sampling method, the input of model training is the element decoding sample at each time step, and the sample of the standard translation does not participate in the sampling here, which is also the biggest feature of the reinforcement training constructed in the present invention. (3) Reward calculation: calculating the reinforcement reward and semantic constraint reward in terms of stem affixes or word granularity; (4) Loss calculation: backpropagating the error according to the reward obtained in (2) to update the network parameters; (5) Decoding prediction: determining the candidate set by combining the loss calculation in step (4) to output the decoding probability, and using the sampling output in this round of prediction as the input of the state at the next time step. The overall training process of the model is as Figure 6 shown.

[0152] The following takes a specific Mongolian-Chinese parallel corpus translation process as an example to illustrate the constructed reinforcement training process:

[0153] (1) Pre-training model

[0154] Input: Mongolian-Chinese bilingual sequence I love machine learning.

[0155] Processing: The LSTM / Transformer translation model encodes and decodes the input.

[0156] Output: Model parameters after iterative training.

[0157] (2) Reward calculation

[0158] Input: Mongolian-Chinese bilingual sequence segmented by stem affix granularity I love machine learning.

[0159] Processing:

[0160] ① Word embedding training → Sparse word embedding vector

[0161] ② Obtain the current sequence reward using formulas (6)-(9).

[0162] Output: Sequence-level training reward Reward.

[0163] (3) Sampling

[0164] Input: Sparse word embedding vector

[0165] Processing: Sampling the output of the previous time step at the current time step instead of the standard accident to solve the 'training sample distribution deviation' problem.

[0166] Output: The sampling result of the sparse vector is used as the input for the next prediction.

[0167] (4) Loss calculation

[0168] Input: Output of step (3) + Reward

[0169] Processing: ① Calculate the loss using formulas (10)-(11).

[0170] ② Calculate the partial derivative of ① to obtain the gradient.

[0171] ③ Backpropagate according to the gradient to update the network.

[0172] Output: Loss partial derivative matrix

[0173] (5) Decoding prediction

[0174] Input: Vocabulary + Prediction probability matrix p

[0175] Processing: Traverse the vocabulary according to p to find the word with the highest probability.

[0176] Output: Select the candidate sequence with the highest probability.

[0177] Figure 7 The figure shows the dynamic integration process in the reinforcement training unfolded in time sequence. Similarly, CE represents the cross-entropy training loss, and RL represents the reinforcement training loss. If the length of the sequence is T, after N CE rounds of cross-entropy pre-training, continue with N RL rounds of reinforcement learning training. In each round of sequence training, the CE loss is used for (T - Δ) steps, and the RL loss is used for the remaining time steps. In the Mongolian-Chinese translation task, when the corpus size is less than 200,000, Δ is set to (2 - 3), and when the corpus size exceeds 200,000, Δ can be set as a hyperparameter. The dynamic is that the model anneals the number of steps using the CE loss for each sequence to (T - 2Δ), and iteratively trains until the entire sequence is trained only using RL.

[0178] System structure constraints: The number of nodes in the hidden layer of the neural network <= D n 、Δ <= 3、The step size for calculating the character-level reward in Mongolian characters <= 3.

[0179] Decision variable: Input the Mongolian embedding representation at the encoder end, and convert the character-level reward corresponding to the representation into the corresponding loss in the reinforcement training to update the network weights.

[0180] Among them, D n is the upper bound of the vocabulary size.< / s> < / agent> < / m> < / eos>

Claims

1. A Mongolian-Chinese neural machine translation method based on multiple constraint items, characterized in that, it includes the following steps: Step 1, on the basis of taking the BLEU value as the training optimization target, add constraint conditions to construct a Mongolian-Chinese neural machine translation model based on reinforcement learning. The constraint conditions include: (1) Semantic constraints are used to alleviate the problem of poor translation fluency caused by a single BLEU value evaluation system; the semantic constraints participate in the reward calculation in each iteration by comparing the standard translation with the predicted output and are used as part of the common reward Reward, and the semantic constraint reward is R sem The sequence-level return reward is obtained by calculating the cosine angle sem(e′, e) between the element vectors e′ and e in two sentences; (2) Parameter constraint to improve the efficiency of model training; the parameter constraint means that during model training, after the loss calculation in each round of iteration, the current training iteration is mapped to the execution range of the problem by gradient mapping, and then the current gradient descent calculation is performed to complete the error transmission; in Mongolian-Chinese machine translation, a soft function constraint is used to improve the objective function and the retrograde penalty coefficient, which plays a role in vector normalization, that is: Min: R(x) + λc(x) 2 where R(x) and c(x) are the objective function and the penalty function respectively, and R(x) and c(x) are solved alternately by the gradient mapping method, and finally the minimum solution of the loss function is achieved. λ represents the penalty coefficient; (3) Vocabulary constraint on the corpus to reduce the number of out-of-vocabulary words in the translation; the vocabulary constraint means that when the corpus is used for vector representation training, low-frequency words with a word frequency lower than 3 in the vocabulary are replaced with the original words or fixed words. At the same time, during the decoding process, when performing Beamsearch column search on the original words, the constructed candidate sequences are cropped, and the sequences with more than the set value of low-frequency words and the last 5 orders of probability are cropped; Step 2, train the constructed translation model. The training strategies include cross-entropy training, reinforcement training, and the dynamic integration of the two trainings.

2. The Mongolian-Chinese neural machine translation method based on multiple constraint items according to claim 1, characterized in that, in the step 1 of constructing the translation model, the reinforcement learning algorithm is mapped into machine translation. The learning body, environment, parameter optimization method, exploration action, learning body parameter state, and reward of reinforcement learning respectively correspond to the translation module, input at different time steps, model training method, sequence decoding, network parameters, and BLEU value of the sequence in machine translation; the input at different time steps is the Mongolian stem affix or word representation unit, and the sequence decoding is decoding with one element as a unit.

3. The Mongolian-Chinese neural machine translation method based on multiple constraint items according to claim 2, characterized in that, the construction of the mapping regards machine translation as a process of exploring rewards for a single variable by the Markov decision method, that is: The learning body continuously changes its parameter state through iterative exploration actions to obtain more rewards, and continuously interacts with the environment on the basis of this iterative training. The process is expressed as M = <S, A, P{s, a}, R>, where s ∈ S, s represents the parameter state of the learning body, S is a finite set of the parameter states of the learning body, which is used to represent the network parameters of the translation model at different time steps during training; a ∈ A, a represents the exploration action, A is a finite set of exploration actions, which is used to record the exploration actions of the learning body in different states; P{s, a} represents that the learning body predicts the next parameter state s′ according to the parameter state s and the exploration action a at the current time step; R represents the immediate reward obtained by the learning body after taking the exploration action a; Based on the mapping, the translation model starts from a random state and continuously optimizes the network parameters, thereby obtaining rewards and punishments in the exploration actions, finding the optimal network parameters under the reward and punishment mechanism, and then continuously updating to make the optimal choice.

4. The Mongolian-Chinese neural machine translation method based on multiple constraint items according to claim 2, characterized in that, In the cross-entropy training, at each time step, the translation model takes the element vector e t ∈E in the standard translation E and the output h t of the neural network hidden layer as inputs, and calculates the output h t+1 of the hidden layer at the next time step until the translation model finally outputs the complete translation vector E′. Among them, E is expressed as E<e 1 ,e 2 ,…,e t ,…,e T >, where e t is the element vector encoded at the t-th time step in E, and T is the vector length of E; E′ is expressed as E′<e′ 1 ,e′ 2 ,…,e′ t ,…e′ T >, where e′ t is the vector element decoded at the t-th time step in E′; h t is a floating-point real-valued vector used to encode the result of the computational output before the current time step in the training; in the reinforcement training, the translation model makes corresponding decoding actions according to the REINFORCE algorithm with the BLEU value as the optimization target, and observes the return reward at the end of each predicted sequence or after the sequence length in the candidate set composed of multiple translations output by the current model according to probability. The sequence length is equal to the vector length of E, that is, T.

5. The Mongolian-Chinese neural machine translation method based on multiple constraint items according to claim 4, characterized in that, the dynamic integration means that within the limited training cycle of the Mongolian-Chinese machine translation, the reinforcement training and the cross-entropy training are dynamically added to the overall training process of the translation model. The execution algorithm is as follows: 1) Embed and train Mongolian words at the stem and affix levels to obtain the corresponding word vector space, and given three variables: the cross-entropy training round N CE , the reinforcement training round N RL and the sequence length T; 2) In each round of training iteration, starting from the sequence length as the starting flag and 1 as the ending flag, use cross-entropy as the loss function to train for N CE rounds; 3) Gradually reduce the number of cross-entropy training rounds with Δ as the difference, and correspondingly increase the number of reinforcement training rounds N RL , where Δ is a training hyperparameter; 4) Use N RL for training. Use cross-entropy training in the first T - Δ steps, and use reinforcement learning for the remaining training steps until the reinforcement learning is fully involved in the overall training and the translation model converges.

6. The Mongolian-Chinese neural machine translation method based on multiple constraint items according to claim 4, characterized in that, in the construction process of the translation model, the Mongolian language is embedded and trained at the stem and affix granularity to obtain the corresponding vector space representation of the Mongolian language, and the affix-level input reward is established for the Mongolian language stem and affix representation during the reinforcement training process.

7. The Mongolian-Chinese neural machine translation method based on multiple constraint items according to claim 4, characterized in that, In both the cross-entropy training and the reinforcement training, the context vector c t is used as an input, and the context vector c t encodes the context to be used when generating a sequence-level output; the translation model learns a recursive function to compute h t+1 , and outputs the element vector e′ t+1 at the (t + 1)-th time step, and this element vector forms a certain element word in the translation according to the subsequent model decoding probability: h t+1 = f(e t , h t , c t ) (1) e′ t+1 ~p θ (e′|e t ,h t+1 ) (2) All activation functions f() are selected as rectified linear unit functions, ~ denotes following a distribution, and p θ () represents network parameters, which is a unified representation of the floating-point vectors of the weight matrix, the hidden layer output matrix, and the context vector matrix; Among them, in cross-entropy training, the cross-entropy loss is adjusted according to the current parameter state set of the translation model to maximize the probability of sequence decoding. For the standard translation E, the goal of the cross-entropy loss Loss ~CrossEntropy is to minimize the following formula: The overall training objective of the translation model is to maximize pθ(e′|e t ,h t+1 ); in the reinforcement training, the reward of the best sequence comes from the standard translation. The goal of the reinforcement training is to obtain the network parameters of the translation model that maximize the expected return reward. The loss is defined as the negative expectation of the reward: r() represents the reward corresponding to the sequence generated by the translation model decoding; in the reinforcement training, the negative expectation is approximated by random sampling or exact sampling of individual words in the output sequence of the translation model, so as to calculate the gradient of the loss.

8. The Mongolian-Chinese neural machine translation method based on multiple constraint items according to claim 1, characterized in that, The semantic constraint reward R sem and the calculation process of the joint reward Reward is expressed as: BLEU Reward R BLEU Expressed as: where p T represents the precision term calculated by the standard N-gram grammar, w T represents the weight matrix, bp is the penalty term of the standard length, and the reward is weakened according to the comparison result of the length c of the decoded output sequence and the length l of the standard sequence; the training return reward based on reinforcement learning is expressed as a joint formula with μ as the control term: Reward=μR BLEU +(1-μ)R sem (9) according to the above combined formula, the translation model is guided to generate a stable sequence.

9. The Mongolian-Chinese neural machine translation method based on multiple constraint items according to claim 8, characterized in that, the training return reward is realized in one of the following two ways: 1) In each batch-scale training, the translation model uses the BLEU reward R for the first batch of training sequences, i.e., batchsize - Δ, and uses the common reward Reward = R BLEU for the remaining Δ training sequences, BLEU +R sem where batchsize is the batch processing scale. In subsequent training, the number of sequences using the cross-entropy loss is reduced by Δ in each batchsize and this process is repeated, thereby obtaining the BLEU reward R BLEU and the semantic constraint reward R sem calculation results. That is, the training cycle of the translation model starts from batchsize and ends until it is finally reduced to 1; 2) Set μ as a hyperparameter to adapt to different tasks.

10. The Mongolian-Chinese neural machine translation method based on multiple constraint items according to claim 1, characterized in that, after adding the constraint conditions, the decoding process of the translation model is as follows: 1) The Mongolian-Chinese bilingual corpus is embedded and trained to obtain a vector representation; 2) The translation model is pre-trained to obtain the initial state space of the reinforcement training; 3) Calculate the similarity between elements in the output sequences on the basis of 2) and obtain the semantic constraint reward; 4) Calculate the reinforcement loss according to the obtained BLEU reward and semantic constraint reward according to the following algorithm: 5) Decode the sampled text, p(y t ) = softmax(W s f(y t )(co, y t-1 , h t-1 ) + b z ), where p(y t ) represents the probability obtained by the target output word at time t, co represents the compressed representation of the Mongolian vector representation, h t-1 represents the output of the hidden layer state at time t - 1, b z represents the calculation offset, y t represents the model output result at time t, W s represents the weight between neural nodes, and f() represents the activation function of the neural network; 6) The probability of predicting the output word y based on the hidden layer state decoded at the t-th time step, the context vector, and the target word y at time t-1 t-1 Predict the output word y t of probability

Citation Information

Patent Citations

  • Mongolian and Chinese inter-translation method based on monolingual corpus training

    CN108829685A

  • Mongolian-Chinese machine translation method based on neural network Turing machine

    CN110619127A