A Convolution-Enhanced Mongolian-Chinese Neural Machine Translation Method
By combining convolutional neural network and Transformer, the convolution-enhanced Mongolian and Chinese neural machine translation method is used to solve the problem of difficulty in feature extraction in Mongolian and Chinese translation, and achieve higher translation accuracy and quality.
Patent Information
- Application Number
- CN202210542568.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-05-18
- Publication Date
- 2025-06-20
- Estimated Expiration
- 2042-05-18
AI Technical Summary
There are problems such as lack of parallel corpus and difficulty in semantic feature extraction in the translation between Mongolian and Chinese, which leads to inaccurate translation, poor local feature extraction ability, and inaccurate word vector representation.
The convolution-enhanced Mongolian and Han neural machine translation method is adopted, combined with convolutional neural network and Transformer, and the local and global dependencies of text sequences are modeled through efficient parameters to improve feature extraction capabilities.
It improves the accuracy of machine translation, enhances feature extraction capabilities, improves translation quality, and shows a higher blue score (BLUE value).
Smart Images

Figure CN115099245B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of artificial intelligence, relates to machine translation, and particularly relates to a convolutional enhanced Mongolian-Chinese neural machine translation method. Background Art
[0002] In the translation between Mongolian and Chinese, problems such as lack of parallel corpus and difficulty in semantic feature extraction lead to many deficiencies in the translation process, including inaccurate translation, poor local feature extraction ability, and inaccurate word vector representation. A better solution is to use Transformer. Currently, Transformer is one of the most commonly used neural network architectures in natural language processing, especially in machine translation. It mainly consists of stacked layers, and each stacked layer is composed of two sub-layers with residual connections: the multi-head self-attention sub-layer and the feedforward neural network (FFN) sub-layer. For a given sentence, the multi-head self-attention sub-layer considers the semantics and dependencies of words in different positions and uses this information to capture the internal sentence structure and expression. The FFN sub-layer is applied to each position separately to encode the context of each position into a higher-level representation in the same way. Although Transformer has proven to be effective in many tasks, its design principles have not been fully understood, and the advantages of the architecture have not been fully utilized.
[0003] The Transformer architecture based on the self-attention mechanism has been widely adopted in sequence modeling because it can capture long-distance interactions and has high training efficiency. In addition, convolutional neural networks can gradually capture local context layer by layer through local receptive fields. However, encoder models with self-attention mechanisms or convolutions have their respective limitations. Although Transformers are good at modeling long-term global context, their ability to extract fine-grained local feature patterns is poor. Therefore, when used in Mongolian-Chinese neural machine translation, there are still problems of low translation quality caused by feature loss. Summary of the Invention
[0004] In order to overcome the above-mentioned shortcomings of the prior art, the purpose of the present invention is to provide a convolutional enhanced Mongolian-Chinese neural machine translation method, which combines convolutional neural networks and Transformers to model the local and global dependencies of text sequences in a parameter-efficient manner, improve the feature extraction ability as much as possible, and then improve the accuracy of machine translation.
[0005] In order to achieve the above purpose, the technical solution adopted by the present invention is:
[0006] A convolutional enhanced Mongolian-Chinese neural machine translation method, comprising:
[0007] Step 1, preprocess the Mongolian data and Chinese data respectively;
[0008] Step 2, construct a translation model based on the Transformer network, and the translation model includes an input module, a Conformer module, a Decoder module, and an output module;
[0009] Step 3, train the translation model;
[0010] Step 4, evaluate the translation model using the BLEU value.
[0011] Compared with the prior art, the present invention makes full use of the structural advantages of the Transformer to improve the capture of local information in the translation model. The present invention modifies the structure of the Transformer encoder, combines the convolutional neural network with the Transformer, performs global and local modeling on sentences, and will achieve the best of global and local (the self-attention mechanism in the Transformer learns global interactions, while the convolution effectively captures local correlations based on relative offsets), improving the BLEU value of the translation. Description of the Drawings
[0012] Figure 1 It is the structure diagram of the macaron.
[0013] Figure 2 It is the structure diagram of the Conformer.
[0014] Figure 3 It is the overall framework diagram of the present invention. Detailed Embodiments
[0015] The following describes the embodiments of the present invention in detail with reference to the drawings and embodiments.
[0016] Aiming at the shortcomings of the Transformer architecture in the machine translation task, the present invention uses a new perspective to understand the architecture, and utilizes the correlation between the Transformer architecture and the multi-particle dynamic system (MPDS) in physics. MPDS is a well-established research field, and its purpose is to use differential equations to simulate how a collection of particles moves in space. In MPDS, the behavior of each particle is usually modeled by two factors respectively. The first factor is convection, which is related to the mechanism of each particle and independent of other particles in the system; the second factor is diffusion, which simulates the movement of particles generated by other particles in the system.
[0017] Inspired by the relationship between ordinary differential equations (ODEs) and neural networks, the Transformer network can be naturally interpreted as a numerical ODE solver for the first-order convection-diffusion equation in MPDS. Specifically, the multi-head self-attention sub-layer corresponds to the diffusion term, which performs semantic transformation at one position and semantic transfer at all other positions; the FFN sub-layer is identically applied to each position separately, corresponding to the convection term. The number of stacked layers in the Transformer corresponds to the time dimension in the ODE. In this way, the stacking of the multi-head self-attention sub-layers and the position-based FFN sub-layers with residual connections can be regarded as numerically solving the ODE problem using the Lie-Trotter splitting scheme and the Euler method. Through this interpretation, a new understanding of using the Transformer to learn the context representation of sentences is obtained: the word sequence can be regarded as the initial positions of a group of particles, and the hidden layers of the stacked Transformer can be regarded as the positions where the particles move at different time points in a high-dimensional space.
[0018] The motion dynamics of multiple particles in space is one of the important problems in fluid mechanics and astrophysics. The behavior of each particle is usually modeled by two factors: the first factor involves its motion mechanism without considering other particles. For example, it is caused by an external force outside the system, usually called convection; the second factor is about the motion caused by other particles, usually called diffusion. Mathematically, assuming there are n particles in a d-dimensional space, being the position of the i-th particle at time t, the dynamics of the i-th particle can be expressed as
[0019]
[0020] The function F(x i (t), [x1(t),..., x n (t)], t) represents the diffusion term characterizing the interaction between particles, and the function G(x, t) is a function that takes the position x and time t as inputs and represents the convection term. There are two coupled terms on the right side of this equation describing different physical phenomena, and the numerical methods for directly solving such ODEs are complex. The Lie-Trotter splitting scheme is the simplest splitting method. It decomposes the right side of the formula into functions F(·) and G(·), and alternately solves the individual dynamics. More precisely, calculating x i (t) to x i (t + γ) using the Lie-Trotter splitting format of the Euler method
[0021]
[0022] From time t to time t+γ, the Lie-Trotter splitting method first solves the ODE regarding F(·) and obtains an intermediate position. Then, starting from it solves the second ODE regarding G(·) to obtain t+γ.
[0023] Reorganize the two sub-layers of the Transformer to make its form match the ODE described above. Let x l =(x l,1 ,...,x l,n ) be the input of the l-th layer of the Transformer, where n is the sequence length, and x l,i is a real-valued vector for any i in . is the output of the multi-head self-attention sub-layer at position i, and the calculation of can be written as
[0024]
[0025]
[0026] is the dot product of the input x l,i and x l,j with the linear projection matrices and , that is, Regarding as the normalized value of the pairwise dot product of , formula (3) can be re-expressed as
[0027]
[0028] where represents all the trainable parameters of the l-th multi-head self-attention sub-layer.
[0029] Re-plan the FFN layer, is put into the FFN sub-layer and outputs x l+1,i . The calculation formula of x l+1,i is as follows:
[0030]
[0031] where represents all the trainable parameters of the l-th layer in the FFN layer.
[0032] Combining formula (5) and formula (6), re-define the Transformer as:
[0033]
[0034]
[0035] It can be seen that the Transformer (Equations (7)-(8)) is similar to a multi-particle ODE solver (Equations (1)-(2)). In fact, a connection can be formally established between an ODE solver with a splitting scheme and a stacked Transformer, whereby the Transformer can be regarded as a numerical ODE solver.
[0036] The above series of equations provides a physical interpretation of natural language processing and offers a new perspective on the Transformer architecture. First, it provides a unified view of the heterogeneous components in the Transformer. The multi-head self-attention sublayer is regarded as a diffusion term representing particle interactions, while the FFN sublayer is regarded as a convection term. Together, these two terms naturally form the convection-diffusion equation in physics. Second, this interpretation advances the understanding of the latent representation of language through the Transformer. Regarding the features of words in a sequence (i.e., embeddings) as the initial positions of particles, the latent representation of a sentence extracted by the Transformer can be interpreted as particles moving in a high-dimensional space.
[0037] However, the Lie-Trotter splitting scheme is the simplest one in ODE solvers and has a relatively high error. In the present invention, the Strang-Marchuk splitting scheme is incorporated into the design of the neural network. To reduce the error, the present invention makes a simple modification to the Lie-Trotter splitting format by dividing the one-step numerical solution procedure of G(·) into two half-steps: one step before F(·) and one step after F(·). This improved splitting scheme is called the Strang-Marchuk splitting scheme. Mathematically, for calculating x i (t) to x i (t + γ), the Strang-Marchuk splitting scheme can be understood as
[0038]
[0039] Mapping the Strang-Marchuk splitting scheme onto the neural network design indicates that the Transformer should also have three sublayers instead of two sublayers. By replacing the functions γF and γG with MultiHeadAtt and FFN, we get
[0040]
[0041] Combined with Figure 1It can be seen that the new layer consists of three sub-layers. The hidden vectors at each different position first pass through the first FFN sub-layer with half-step residual connections, and then the output vectors are fed into a multi-head self-attention sub-layer. In the last step, the vectors output by the multi-head self-attention sub-layer are put into the second FFN sub-layer with half-step residual connections. Since the FFN-attention-FFN structure is similar to "macaron", this layer will be called the macaron layer.
[0042] Models with self-attention or convolution have their limitations. Although Transformers are good at modeling long-term global context, they are poor at extracting fine-grained local feature patterns. On the other hand, Convolutional Neural Networks (CNNs) are good at leveraging local information and are used as practical computational blocks in vision. They learn shared position-based kernel functions on local windows, which maintain translational variance and are able to capture features such as edges and shapes. Recent research has shown that combining CNNs and multi-head self-attention mechanisms is better than using them alone. In machine translation, the global and local interactions are both important for parameter efficiency. In the present invention, how to organically combine convolution and self-attention in a machine translation model is studied. Global and local interactions are both important for parameter efficiency.
[0043] To achieve the capture of global and local information of sentences, the present invention will use a new combination of multi-head self-attention mechanism and convolution, which will achieve the best of global and local (the multi-head self-attention mechanism learns global interactions, while convolution effectively captures local correlations based on relative offsets). Inspired by the macaron structure, a new combination of multi-head self-attention and convolution, sandwiched between a pair of feed-forward modules, is called Conformer (as Figure 2 shown). Like the macaron network, half-step residual weights are used in the FFN. And the second FFN sub-layer is followed by the last layer - the normalization layer, which means that for the input x i of the i-th Conformer sub-module, the output y i of this block is:
[0044]
[0045] where Conv is the convolutional neural network layer and Layernorm is the normalization layer.
[0046] Based on the above ideas, the present invention applies it to machine translation and specifically provides a convolution-enhanced Mongolian-Chinese neural machine translation method, which mainly includes the following steps:
[0047] Step 1, preprocess the Mongolian data and Chinese data respectively.
[0048] In this embodiment, a Mongolian corpus of 1.6 million sentences and a Chinese corpus crawled from the Internet are used as the dataset. Before model training, simple preprocessing is required. Different from languages such as Mongolian and English, there are no spaces between Chinese words. Therefore, Chinese words need to be segmented. The Jieba Chinese word segmentation technology, which is currently mainstream, is adopted in the present invention. Then, BPE is used for segmentation, while for Mongolian, BPE is directly used for segmentation. The segmentation process can, to a certain extent, alleviate the influence of low-frequency words.
[0049] Step 2. Based on the above idea, a translation model is constructed. Referring to Figure 3 , in terms of architecture, the translation model of the present invention mainly includes an input module, a Conformer module, a Decoder module, and an output module. Specific introductions are given one by one below.
[0050] Input module: The words in the Mongolian sentence sequence that have been segmented in Step 1 are sequentially encoded by word vectors and then added with positional encoding (Positional Encoding, PE) to obtain the position of the current word (i.e., the distance between different words in a sentence). The calculation formula is as follows:
[0051]
[0052]
[0053] The positional encoding is a two-dimensional matrix, where the rows represent words and the columns represent word vectors; among them, pos represents the absolute position of the word in the sentence, pos = 0, 1, 2…, d model represents the dimension of the word vector of the word, i represents the position of the word vector, and the meanings of PE(pos, 2i) and PE(pos, 2i + 1) are to add a sin variable to the even positions of the word vector of each word and a cos variable to the odd positions, so as to fill the entire positional encoding matrix.
[0054] The Conformer module is stacked by multiple identical Conformer sub-modules. Each Conformer sub-module consists of a first FFN sub-layer, a first multi-head self-attention sub-layer, a convolutional layer, and a second FFN sub-layer. Both the first FFN sub-layer and the second FFN sub-layer have a half-step residual connection.
[0055] The calculation performed by the first FFN sub-layer is as follows:
[0056] FFN(x) = max(0, xW1 + b1)W2 + b2
[0057] Among them, W1 and W2 are two weight matrices with opposite dimensions, and b1 and b2 are hyperparameters; FFN(x) passes through three weight matrices W Q 、WK and W V Multiply them separately to obtain the Query vector Q1, Keys vector K1, and Values vector V1 required for the calculation of the first multi-head self-attention sub-layer;
[0058] The calculations performed in the first multi-head self-attention sub-layer are as follows:
[0059] In the first step, calculate the correlation score vector score1 between the words in the Mongolian sentence:
[0060] score1 = Q1 · K1 T
[0061] That is, calculate the dot product of each vector in Q1 with each vector in K1 to obtain a matrix form.
[0062] In the second step, normalize the correlation score vector, denoted as score′1. The main purpose of normalization is to make the gradient stable during model training. The calculation is as follows:
[0063]
[0064] where is the dimension of K1;
[0065] In the third step, use the softmax function to convert the normalized correlation score vector score′1 into a probability distribution between [0, 1], and at the same time highlight the relationship between words more prominently, that is, softmax(score′1);
[0066] In the fourth step, multiply the probability distribution by the corresponding Values vector V1. The calculation is as follows:
[0067] Z1 = softmax(score′1)V1
[0068] The convolutional layer, as an actual calculation block, is good at utilizing local information. It takes Z1 as the input. The convolutional layer starts with a gating mechanism, a pointwise convolution (the size of the convolutional kernel is 1×1×n, where n is the number of channels in the previous layer, and there will be as many Feature Maps as there are convolutional kernels), and an activation unit (GLU). Next is a one-dimensional depth convolutional layer, and batchnorm is deployed after convolution to facilitate model training. After convolutional calculation, the obtained result is denoted as Z′1.
[0069] The second FFN sub-layer takes the result Z′1 obtained from convolutional calculation as the input and calculates FFN(Z′1). Its calculation formula is the same as that of the first FFN sub-layer.
[0070] The structures of each Conformer sub-module are the same, and each Conformer sub-module performs cyclic calculations. The number of cycles is set according to actual needs. In the first FFN sub-layer of the first Conformer sub-module, x takes the output of the input module. In the first FFN sub-layers of the remaining Conformer sub-modules, x takes the FFN(Z′1) of the previous Conformer sub-module. In the second FFN sub-layer of the last Conformer sub-module, the FFN(Z′1) is further calculated to obtain the Keys vector K3 and Values vector V3 required for the calculation of the Decoder module.
[0071] As described above, a final layer - the normalization layer - can also be set after the second FFN sub-layer.
[0072] The Decoder module is stacked by multiple identical Decoder sub-modules, and each Decoder sub-module consists of a masked multi-head attention sub-layer, a second multi-head self-attention sub-layer, and a third FNN sub-layer.
[0073] The input of the Decoder module has two types: one is the input during training, and the other is the input during prediction. The input during training is the Chinese translation corresponding to the Mongolian sentence. For example, the input of Conformer The Decoder then inputs the corresponding translation "It will rain tomorrow". The input during prediction is the start symbol for the first time, and the output of the translation model at the previous moment is input each time after that.
[0074] In the masked multi-head attention sublayer, masking means covering certain values so that they have no effect during parameter update. There are two types of masks involved in the Decoder. One is used to handle the variable length of the input sentence sequence, and the other is used to prevent future information from being leaked. For handling the variable length of the input sentence sequence, since the length of the input sequence in each batch is different. That is, the input sequences need to be aligned. Specifically, zeros are padded after the shorter sequences. However, if the input sequence is too long, the content on the left side of the input sentence is intercepted and the excess is simply discarded. The specific approach is to add a very large negative number (negative infinity) to each component value in the vector processed by Conformer. In this way, after softmax, the probabilities at these positions will approach 0, and this progressive operation is actually a tensor where each value is a boolean value, and the places with the value false are where we need to perform the processing. For preventing future information from being leaked, it is to ensure that the Decoder cannot see future information. That is, for a sequence, the decoded output should only depend on the output before the current moment and not on the output after the current moment. Therefore, the information after the current moment needs to be hidden. The specific method is to generate an upper triangular matrix with all values in the upper triangle being 0 and apply this upper triangular matrix to each sequence.
[0075] Specifically, in the masked multi-head attention sublayer. The input is multiplied by three weight matrices W Q 、W K and W V respectively to obtain the Query vector Q2, Keys vector K2, and Values vector V2 required for the calculation of the masked multi-head self-attention sublayer. The subsequent calculations are specifically described as follows:
[0076] First step, calculate the correlation score vector score2 between the input words:
[0077] score2 = Q2 · K2 T
[0078] That is, each vector in Q2 is dotted with each vector in K2 to obtain a matrix form. The upper triangular elements of this matrix are set to 0 to obtain the matrix score′2;
[0079] Second step, normalize score′2, denoted as score″2. The main purpose of normalization is to make the gradient stable during model training. The calculation is as follows:
[0080]
[0081] where, is the dimension of K2;
[0082] In the third step, the normalized correlation score vector score″2 is converted into a probability distribution between [0, 1] through the softmax function, while further highlighting the relationship between words, that is, softmax(score″2);
[0083] In the fourth step, this probability distribution is multiplied by the corresponding Values vector V2, and the calculation is as follows:
[0084] Z2 = softmax(score″2)V2
[0085] In the second multi-head attention sub-layer, the same operations as those in the first multi-head self-attention sub-layer are performed. Specifically, the output Z2 obtained from the masked multi-head attention sub-layer is multiplied by the weight matrix W Q to obtain the Query vector Q3 required for the calculation of the second multi-head self-attention sub-layer. The subsequent calculations are as follows:
[0086] In the first step, calculate the correlation score vector score3:
[0087] score3 = Q3·K3 T
[0088] That is, each vector in Q3 is calculated with each vector in K3 to obtain a matrix form.
[0089] In the second step, normalize the correlation score vector, denoted as score′3. The purpose of normalization is mainly to make the gradient stable during model training. The calculation is as follows:
[0090]
[0091] where, is the dimension of K3;
[0092] In the third step, the normalized correlation score vector score′3 is converted into a probability distribution between [0, 1] through the softmax function, while further highlighting the relationship between words, that is, softmax(score′3);
[0093] In the fourth step, this probability distribution is multiplied by the corresponding Values vector V3, and the calculation is as follows:
[0094] Z3 = softmax(score')V3
[0095] The third FNN sub-layer takes Z3 as the input and calculates FFN(Z3), and its calculation formula is the same as that of the first FFN sub-layer.
[0096] Each sub-module of the Decoder module performs cyclic calculations. In the masked multi-head attention sub-layer of the first sub-module of the Decoder module, the input is the sum of the word vectors of the Chinese translation corresponding to the Mongolian sentence input by the Conformer module and the corresponding position encoding. In the masked multi-head attention sub-layers of the remaining sub-modules of the Decoder module, the input is the FFN(Z3) of the previous sub-module of the Decoder module. The FFN(Z3) of the last sub-module of the Decoder module is the output of the Decoder module.
[0097] The structures of all sub-modules of the Decoder module are the same and are used repeatedly for decoding. The number of cycles is set according to actual needs.
[0098] The output module first performs a linear transformation on the word vectors processed by the Decoder module, then obtains the output probability distribution through the Softmax function, and finally outputs the word corresponding to the maximum probability as the predicted output through the dictionary.
[0099] Step 3: Train the translation model.
[0100] Step 4: Evaluate the translation model using the BLUE value. The model that meets the set standards can be used for Mongolian-Chinese translation.
[0101] In an embodiment of the present invention, taking translation as an example, the source language sentence is segmented into Correspondingly, the parallel corpus (standard translation) "It will rain tomorrow" is segmented into "-、tomorrow、will、rain". After successively passing through word vector encoding and adding position encoding, the obtained matrix of the source language is calculated by the Conformer module to obtain the K and V matrices and sent to the multi-head self-attention sub-layer in the Decoder module. After "-、tomorrow、will、rain" successively pass through word vector encoding and add position encoding, the obtained matrix of the target language is calculated by the Decoder module. A linear transformation is performed on the word vectors processed by the Decoder module, then the output probability distribution is obtained through the Softmax function, and finally the translation "It will rain tomorrow" with the highest probability is output.
Claims
1. A convolution-enhanced Mongolian-Chinese neural machine translation method, characterized in that, Including: Step 1: Preprocess Mongolian data and Chinese data respectively; Step 2: Build a translation model based on the Transformer network, and the translation model includes an input module, a Conformer module, a Decoder module, and an output module; Step 3: Train the translation model; Step 4: Evaluate the translation model using the BLUE value; For the input module, the words in the Mongolian sentence sequence after word segmentation in Step 1 are sequentially added with position encoding after being encoded by word vectors to obtain the position of the current word, and the calculation formula is as follows: The positional encoding is a two-dimensional matrix, where the rows represent words and the columns represent word vectors; among them, pos represents the absolute position of the word in the sentence, pos = 0, 1, 2…, d model represents the dimension of the word vector of the word, i represents the position of the word vector. The meanings of PE(pos, 2i) and PE(pos, 2i + 1) are to add the sin variable to the even positions of the word vector of each word and the cos variable to the odd positions, so as to fill the entire positional encoding matrix; The Conformer module is stacked by multiple identical Conformer sub-modules, and each Conformer sub-module consists of a first FFN sub-layer, a first multi-head self-attention sub-layer, a convolutional layer, and a second FFN sub-layer; The calculation performed by the first FFN sub-layer is as follows: FFN(x) = max(0, xW1 + b1)W2 + b2 Among them, W1 and W2 are two weight matrices with opposite dimensions, and b1 and b2 are hyperparameters; FFN(x) is multiplied by three weight matrices W Q , W K and W V respectively to obtain the Query vector Q1, Keys vector K1, and Values vector V1 required for the calculation of the first multi-head self-attention sublayer; The calculation in the first multi-head self-attention sub-layer is as follows: The first step: Calculate the correlation score vector score1 between words in the Mongolian sentence; The second step: Normalize the correlation score vector, denoted as score′1, and the calculation is as follows: Among them, is the dimension of K1; The third step: Through the softmax function, convert the normalized correlation score vector score′1 into a probability distribution between [0, 1], that is, softmax(score′1); The fourth step: Multiply the probability distribution by the corresponding Values vector V1, and the calculation is as follows: Z1 = softmax(score′1)V1 The convolutional layer takes Z1 as the input and obtains Z′1 through convolutional calculation; The second FFN sub-layer takes Z′1 as the input and calculates FFN(Z′1); Each Conformer sub-module calculates cyclically. In the first FFN sub-layer of the first Conformer sub-module, x takes the output of the input module. In the first FFN sub-layers of the remaining Conformer sub-modules, x takes the FFN(Z′1) of the previous Conformer sub-module. In the second FFN sub-layer of the last Conformer sub-module, FFN(Z′1) is further calculated to obtain the Keys vector K3 and Values vector V3 required for the Decoder module calculation; Both the first FFN sub-layer and the second FFN sub-layer have half-step residual connections.
2. The convolution-enhanced Mongolian-Chinese neural machine translation method according to claim 1, characterized in that, In Step 1, a Mongolian corpus of 1.6 million sentences and a Chinese corpus crawled from the Internet are used as the data set. For the preprocessing, Chinese is subjected to word segmentation and then segmented using BPE; For Mongolian, directly segment using BPE.
3. The convolution-enhanced Mongolian-Chinese neural machine translation method according to claim 1, characterized in that, A normalization layer is also provided behind the second FFN sub-layer.
4. The convolution-enhanced Mongolian-Chinese neural machine translation method according to claim 1, characterized in that, The Decoder module is stacked by multiple identical Decoder sub-modules, and each Decoder sub-module consists of a masked multi-head attention sub-layer, a second multi-head self-attention sub-layer, and a third FNN sub-layer; There are two types of inputs to the Decoder module: one is the input during training, and the other is the input during prediction; the input during training is the Chinese translation corresponding to the Mongolian sentence, and the input during prediction is the start symbol for the first time, and the output of the translation model at the previous moment is the input for each subsequent time; In the masked multi-head attention sub-layer, the input is multiplied by three weight matrices W Q , W K , and W V respectively to obtain the Query vector Q2, Keys vector K2, and Values vector V2 required for the calculation of the masked multi-head self-attention sub-layer. The subsequent calculations are as follows: In the first step, calculate the correlation score vector score2 between the input words: score2 = Q2 · K2 T That is, calculate the dot product of each vector in Q2 and each vector in K2 to obtain a matrix form. Set the upper triangular elements of this matrix to 0 to obtain the matrix score'2; In the second step, normalize score'2, denoted as score''2, and the calculation is as follows: Among them, is the dimension of K2; In the third step, through the softmax function, convert the normalized correlation score vector score''2 into a probability distribution between [0,1], that is, softmax(score''2); In the fourth step, multiply the probability distribution by the corresponding Values vector V2, and the calculation is as follows: Z2 = softmax(score''2)V2 In the second multi-head attention sub-layer, the output Z2 obtained from the masked multi-head attention sub-layer is multiplied by the weight matrix W Q to obtain the Query vector Q3 required for the calculation of the second multi-head self-attention sub-layer. The subsequent calculations are as follows: In the first step, calculate the correlation score vector score3: score3 = Q3 · K3 T In the second step, normalize the correlation score vector, denoted as score'3, and the calculation is as follows: Among them, is the dimension of K3; In the third step, through the softmax function, convert the normalized correlation score vector score'3 into a probability distribution between [0,1], that is, softmax(score'3); In the fourth step, multiply the probability distribution by the corresponding Values vector V3, and the calculation is as follows: Z3 = softmax(score')V3 The third FFN sub-layer takes Z3 as the input and calculates FFN(Z3); Each sub-module of the Decoder module is calculated iteratively. In the masked multi-head attention sub-layer of the first sub-module of the Decoder module, the input is the sum of the word vector corresponding to the Chinese translation of the Mongolian sentence input by the Conformer module and the corresponding position encoding. In the masked multi-head attention sub-layer of the remaining sub-modules of the Decoder module, the input is the FFN(Z3) of the previous Decoder module sub-module. The FFN(Z3) of the last Decoder module sub-module is the output of the Decoder module.
5. The convolutional enhancement-based Mongolian-Chinese neural machine translation method according to claim 1, characterized in that, In the output module, first perform a linear transformation on the word vector processed by the Decoder module, then obtain the output probability distribution through the Softmax function, and finally output the word corresponding to the maximum probability through the dictionary as the prediction output.
Citation Information
Patent Citations
A neural machine translation method based on dynamic configuration decoding
CN109933808A
Machine translation method and machine translation device
CN112699693A