Neural Network Machine Translation Method, Model and Model Formation Method
By introducing future information models into the decoder of neural network machine translation model, integrating historical and future information, the problem that existing models cannot utilize untranslated word sequence information is solved, improving the accuracy of translation and reducing the phenomenon of missed translation.
Patent Information
- Application Number
- CN201811534845.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2018-12-14
- Publication Date
- 2025-07-01
- Estimated Expiration
- 2038-12-14
AI Technical Summary
Existing neural network machine translation models cannot obtain information about untranslated word sequences when predicting current words, resulting in the inability to fully utilize future information.
By setting up a future information model in the decoder of the neural network machine translation model, the model represents a fusion of the current predicted word and the first attention hidden layer representation of the generated word and the second attention hidden layer representation of the current predicted word and the possible future words, and performs decoding training of the source language sequence from left to right and right to left.
This enables the neural network machine translation model to use the historical information of the generated words before the current predicted words and the future information of possible future words after the current predicted words, thereby improving the accuracy of translation and improving the phenomenon of missed translation.
Smart Images

Figure CN111401081B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of machine translation. Specifically, the present invention relates to a method for forming a neural network machine translation model, a neural network machine translation model, a neural network machine translation method, a storage medium, a processor, and an electronic device. Background Art
[0002] Machine translation refers to using a computer to translate a natural language into another natural language with the same semantics, which is one of the important research directions in the fields of artificial intelligence and natural language processing. Generally, the system framework of machine translation can be divided into two categories: rule-based machine translation (RBMT) and corpus-based machine translation (CBMT). Among them, CBMT can be further divided into example-based machine translation (EBMT), statistical-based machine translation (SMT), and neural machine translation (NMT) constructed using deep learning models, which has become popular in recent years.
[0003] Neural machine translation refers to a machine translation method that directly uses a neural network to perform translation modeling in an end-to-end manner. Its basic idea is to directly map the source language into the target language text using a neural network. Neural machine translation adopts an "encoder-decoder" framework: given a source language sentence, first use an encoder to map it into a continuous and dense vector, and then use a decoder to convert this vector into a target language sentence. With the development of deep learning technology, neural network machine translation models have been widely studied and have shown great advantages compared with statistical machine translation models.
[0004] (Encoder-Decoder Framework)
[0005] The basic idea of neural machine translation is to directly achieve automatic translation between natural languages through neural networks. To this end, neural machine translation usually adopts an encoder-decoder framework to implement sequence-to-sequence conversion. In specific operations, different types of neural networks can be used for the encoding and decoding of the source language and the target language. For example, the encoder can be a Convolutional Neural Network, and the decoder can be a Recurrent Neural Network. However, although recurrent neural networks can theoretically capture infinitely long historical information, they face problems such as "gradient vanishing" and "gradient explosion" during the training process and it is difficult to truly handle long-distance dependencies. To alleviate the problems of "gradient vanishing" and "gradient explosion", Long Short-Term Memory is introduced into end-to-end neural machine translation.
[0006] Figure 1 It is a schematic diagram of the "encoder-decoder" framework of neural machine translation in related technologies. Taking Figure 1 as an example, given a source language Chinese sentence "This is the secret of success", the encoder-decoder framework first generates word vector representations for each Chinese word, and then generates a vector representation of the entire Chinese sentence from left to right through a recurrent neural network. Among them, "" represents the end-of-sentence terminator. We call the recurrent neural network used at the source language end the encoder, and its role is to encode the source language sentence into a dense and continuous real number vector. Thereafter, another recurrent neural network is used at the target language end to reverse-decode the source language sentence vector into the target language English sentence "This is the secret of success". The entire decoding process generates words one by one, and when the end-of-sentence terminator "" is generated, the decoding process terminates. We call the recurrent neural network used at the target language end the decoder.
[0007] Compared with traditional statistical machine translation, neural machine translation based on the encoder-decoder framework has the advantages of directly learning features from data and being able to capture long-distance dependencies. However, the encoder-decoder framework also faces a serious problem: the dimension of the source language sentence vector representation generated by the encoder is independent of the length of the source language sentence. That is, whether it is a longer source language sentence or a shorter source language sentence, the encoder needs to map it into a vector with a fixed dimension, which poses a great challenge to achieving accurate encoding.
[0008] (Attention mechanism)
[0009] Regarding the problem of the encoder generating fixed-length vectors, end-to-end neural machine translation based on the attention mechanism has been proposed. The core idea of this mechanism is that when the decoder generates the current target language word, in fact, only a small part of the source language words are relevant, and the vast majority of the source language words are irrelevant. Therefore, instead of using a fixed-length vector representing the entire source language sentence, a relevant context vector at the source language end can be dynamically generated for each target language word.
[0010] Figure 2 It is a schematic diagram of the framework of neural network machine translation based on the attention mechanism in the related technology. As Figure 2 shown, neural machine translation based on the attention mechanism uses a completely different encoder. Its goal is no longer to generate a vector representation for the entire source language sentence, but to generate a vector representation containing global information for each source language word. Given the source language sentence X = {x1, x2,..., x n}, the bidirectional recurrent neural network encoder encodes the sentence X into a sequence of source language hidden states H = {h1, h2,..., h n}, where the forward recurrent neural network reads the sentence X sequentially to generate the forward source language hidden state sequence The backward recurrent neural network reads the sentence X in reverse order to generate the backward source language hidden state sequence The state sequences corresponding to the positions in the forward and backward hidden state sequences are concatenated to form the hidden state of the word at that position
[0011] At the decoding time t, the decoder generates the target language hidden state and the target language word at that time respectively. The target language hidden state s t at time t is determined by the target language hidden state s t-1 at time t - 1, the target language word y t-1 generated by the decoder at time t - 1, and the context vector c t at time t:
[0012] s t = g(s t-1 , y t-1 , c t )
[0013] where g is a non-linear function, such as LSTM or GRU. The context vector c t at time t is obtained by weighting the source language hidden state sequence H with the weights generated by the attention model:
[0014]
[0015] Here, the weights a of the attention modelt,j Generated from the target language implicit state s at time t-1 t and the source language implicit state sequence H:
[0016]
[0017] e t,j = f(s t , h j )
[0018] where f is a non-linear function, usually a feed-forward neural network or dot product. The weight a t,j can be understood as the correlation degree between the source language word x j and the word generated by the decoder at time t.
[0019] After obtaining the target language implicit state s t , the model estimates the probability distribution of the target language word at time t through the softmax function:
[0020] P(y t |y <t , X) = softmax(g(s t , y t-1 , c t ))
[0021] The training objective function of the neural network machine translation model is the sum of the logarithmic likelihood functions of the translation sentence pairs on the parallel corpus, expressed as:
[0022]
[0023] D represents the set of parallel sentence pairs, and the model parameters θ can be solved by optimization methods such as stochastic gradient descent (SGD), Adam or Adadelta.
[0024] In the related technology, whether using a machine translation model based on a recurrent neural network or a machine translation model based on a self-attention mechanism, when predicting the current word, only the source language information and the information of the previously generated word sequence can be utilized, and the information of the untranslated word sequence cannot be obtained. That is, the future information in the existing neural network machine translation cannot be fully utilized. Summary of the Invention
[0025] Embodiments of the present invention provide a neural network machine translation method, model and model formation method to at least solve the problem that in the process of machine translation, when predicting the current word, the information of the untranslated word sequence cannot be obtained, resulting in the future information not being fully utilized.
[0026] According to one aspect of an embodiment of the present invention, a method for forming a neural network machine translation model includes: forming an encoder, the encoder including a first multi-head attention model; forming a decoder, the decoder including a second multi-head attention model and a future information model, the future information model representing the fusion of a first attention hidden layer representation of a current predicted word and a generated word and a second attention hidden layer representation of the current predicted word and a future possible word; forming a first machine translation model through the encoder and the decoder; and performing left-to-right and right-to-left decoding training on the first machine translation model for a source language sequence to form a neural network machine translation model, wherein the first multi-head attention model and the future information model provide inputs for the second multi-head attention model.
[0027] By setting a future information model representing the fusion of a first attention hidden layer representation of a current predicted word and a generated word and a second attention hidden layer representation of the current predicted word and a future possible word in the decoder of the neural network machine translation model, and performing left-to-right and right-to-left decoding training on the machine translation model for a source language sequence, the resulting neural network machine translation model can not only utilize the historical information of the generated words before the current predicted word, but also utilize the future information of the future possible words after the current predicted word, thereby making the result of machine translation more accurate. In addition, due to the utilization of the future information of the future possible words after the current predicted word, the phenomenon of missing translation in machine translation can be effectively improved.
[0028] In a schematic implementation manner of the method for forming a neural network machine translation model, forming the encoder includes forming a first multi-head attention model, wherein forming the first multi-head attention model includes: forming a dot product attention model by using a dot product attention mechanism; setting a linear transformation model for the dot product attention model to map the input of the dot product attention model into multiple groups of vectors with a predetermined dimension through linear transformation; setting a connection model for the dot product attention model to connect the vectors obtained after being processed by the dot product attention model; and forming the first multi-head attention model through the dot product attention model, the linear transformation model, and the connection model.
[0029] By forming a multi-head attention model, the neural network machine translation model can learn relevant information in different subspaces, thereby improving the accuracy of the machine translation result.
[0030] In a schematic implementation manner of the method for forming a neural network machine translation model, forming the decoder includes forming a future information model, wherein forming the future information model includes: calculating a first attention hidden layer representation of a current predicted word and a generated word by using a dot product attention mechanism: wherein, represents the first attention hidden layer representation, represents the hidden layer state query value at the current moment, represents the historical hidden layer state key value, represents the historical hidden layer state real value, and Attention() is the mathematical function representation of the dot product attention mechanism; the second attention hidden layer representation of the current predicted word and possible future words is calculated using the dot product attention mechanism: wherein, represents the second attention hidden layer representation, represents the future hidden layer state key value, represents the future hidden layer state real value; the first attention hidden layer representation and the second attention hidden layer representation are fused to form a fused attention hidden layer representation; a dot product attention model is formed using the fused attention hidden layer representation; a linear transformation model is set for the dot product attention model to map the input of the dot product attention model into multiple groups of vectors with a predetermined dimension through linear transformation; a connection model is set for the dot product attention model to connect the vectors obtained after processing the vectors by the dot product attention model; and the future information model is formed through the dot product attention model, the linear transformation model, and the connection model.
[0031] In a schematic implementation manner of the method for forming a neural network machine translation model, fusing the first attention hidden layer representation and the second attention hidden layer representation to form a fused attention hidden layer representation includes: using a threshold mechanism to fuse the first attention hidden layer representation and the second attention hidden layer representation into a fused attention hidden layer representation: wherein, represents the fused attention hidden layer representation, r t represents the reset gate, z t represents the update gate, W g is a model parameter, σ is the sigmoid function, [;] represents the vector or matrix concatenation operation, and · is the bitwise product operation.
[0032] By fusing the first attention hidden layer representation of the current predicted word and the already generated words and the second attention hidden layer representation of the current predicted word and possible future words to form a future information model, the neural network machine translation model can obtain the historical information of the already generated words before the current predicted word and the future information of the possible future words after the current predicted word, thereby making full use of the future information for machine translation and improving the translation accuracy.
[0033] According to another aspect of the embodiments of the present invention, there is also provided a neural network machine translation model, including: an encoder, the encoder includes a first multi-head attention model; and a decoder, the decoder includes a second multi-head attention model and a future information model, the future information model represents the fusion of the first attention hidden layer representation of the current predicted word and the already generated words and the second attention hidden layer representation of the current predicted word and the future possible words, wherein the first multi-head attention model and the future information model provide inputs for the second multi-head attention model, and the neural network machine translation model has undergone decoding training for the source language sequence from left to right and from right to left.
[0034] By setting a future information model representing the fusion of the first attention hidden layer representation of the current predicted word and the already generated words and the second attention hidden layer representation of the current predicted word and the future possible words in the decoder of the neural network machine translation model, and performing decoding training on the machine translation model from left to right and from right to left for the source language sequence, the neural network machine translation model can not only utilize the historical information of the already generated words before the current predicted word, but also utilize the future information of the future possible words after the current predicted word, thereby making the result of machine translation more accurate. In addition, due to the utilization of the future information of the future possible words after the current predicted word, the phenomenon of missing translation in machine translation can be effectively improved.
[0035] In an exemplary embodiment of the neural network machine translation model, the first multi-head attention model includes: a dot product attention model configured to form using the dot product attention mechanism; a linear transformation model connected to the dot product attention model and configured to map the input of the dot product attention model into multiple groups of vectors of a predetermined dimension through linear transformation; and a connection model connected to the dot product attention model and configured to connect the vectors obtained after the vectors are processed by the dot product attention model.
[0036] Through the multi-head attention model, the neural network machine translation model can learn relevant information in different subspaces, thereby improving the accuracy of the machine translation result.
[0037] In an exemplary embodiment of the neural network machine translation model, the future information model includes: a dot product attention model; a linear transformation model connected to the dot product attention model and configured to map the input of the dot product attention model into multiple groups of vectors of a predetermined dimension through linear transformation; and a connection model connected to the dot product attention model and configured to connect the vectors obtained after the vectors are processed by the dot product attention model, wherein the dot product attention model is formed by the following steps: calculating the first attention hidden layer representation of the current predicted word and the already generated words using the dot product attention mechanism: Wherein, represents the query value of the hidden layer state at the current moment, represents the key value of the historical hidden layer state, represents the real value of the historical hidden layer state, and Attention() represents the mathematical function of the dot product attention mechanism; the second attention hidden layer representation of the current predicted word and possible future words is calculated using the dot product attention mechanism: wherein, represents the key value of the future hidden layer state, represents the real value of the future hidden layer state; the first attention hidden layer representation and the second attention hidden layer representation are fused to form a fused attention hidden layer representation; and a dot product attention model is formed using the fused attention hidden layer representation.
[0038] In an exemplary embodiment of the neural network machine translation model, fusing the first attention hidden layer representation and the second attention hidden layer representation to form a fused attention hidden layer representation includes: using a threshold mechanism to fuse the first attention hidden layer representation and the second attention hidden layer representation into a fused attention hidden layer representation: wherein, represents the fused attention hidden layer representation, r t represents the reset gate, z t represents the update gate, W g are model parameters, σ is the sigmoid function, [;] represents the vector or matrix concatenation operation, and · is the bitwise product operation.
[0039] By fusing the first attention hidden layer representation of the current predicted word and the already generated words and the second attention hidden layer representation of the current predicted word and possible future words to form a future information model, the neural network machine translation model can obtain the historical information of the already generated words before the current predicted word and the future information of the possible future words after the current predicted word, thereby making full use of the future information for machine translation and improving the translation accuracy.
[0040] According to another aspect of the present invention, there is also provided a neural network machine translation method, including: receiving a source language sequence, processing the source language sequence by a first multi-head attention model using a multi-head attention mechanism to obtain a hidden layer vector representation of the source language sequence; performing a matrix transformation on the hidden layer vector representation of the source language sequence to obtain corresponding key K and value V of the source language sequence; inputting the key K and value V into a second multi-head attention model; obtaining a fused attention hidden layer vector representation of a first attention hidden layer representation of the current predicted word and the already generated words and a second attention hidden layer representation of the current predicted word and the future possible words through a future information model, where the future information model represents the fusion of the first attention hidden layer representation of the current predicted word and the already generated words and the second attention hidden layer representation of the current predicted word and the future possible words; obtaining a query Q of the target language sequence according to the fused attention hidden layer vector representation; inputting the query Q into the second multi-head attention model; determining, by the second multi-head attention model, the probability that the target language word is the current predicted word according to the key K, value V, and query Q; and generating a text sequence with the highest probability for each word as the target language sequence of the translation result.
[0041] By fusing the first attention hidden layer representation of the current predicted word and the already generated words and the second attention hidden layer representation of the current predicted word and the future possible words, historical information of the already generated words before the current predicted word and future information of the future possible words after the current predicted word can be obtained. By utilizing the obtained historical information and future information, the result of machine translation can be made more accurate. Additionally, since the future information of the future possible words after the current predicted word is utilized, the phenomenon of missing translation in machine translation can be effectively improved.
[0042] According to another aspect of an embodiment of the present invention, there is also provided a storage medium, where the storage medium includes a stored program, and during the running of the program, a device including the storage medium is controlled to execute the above method for forming a neural network machine translation model. BRIEF DESCRIPTION OF THE DRAWINGS
[0043] The drawings described herein are used to provide a further understanding of the present invention and constitute a part of the present invention. The schematic embodiments of the present invention and their descriptions are used to explain the present invention and do not constitute an improper limitation to the present invention. In the drawings:
[0044] Figure 1 is a schematic diagram of an "encoder-decoder" framework of neural network machine translation in the related art;
[0045] Figure 2 is a schematic diagram of a framework of neural network machine translation based on an attention mechanism in the related art;
[0046] Figure 3 Shows a schematic diagram of the dot product attention mechanism;
[0047] Figure 4 Shows a schematic diagram of the multi-head attention mechanism;
[0048] Figure 5 Shows a flowchart of a method for forming a neural network machine translation model according to an embodiment of the present invention;
[0049] Figure 6 Shows a schematic diagram of forming a fused attention hidden layer representation according to an embodiment of the present invention;
[0050] Figure 7 Shows a block diagram of a neural network machine translation model according to an embodiment of the present invention;
[0051] Figure 8 Shows an internal architecture diagram of a neural network machine translation model according to an embodiment of the present invention;
[0052] Figure 9 Is an input example and an output example of a decoder shown in chronological order according to an embodiment of the present invention. Detailed implementation manners
[0053] In order to enable those skilled in the art to better understand the solution of the present invention, the technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.
[0054] It should be noted that the terms "including" and "having" and any variations thereof in the specification and claims of the present invention and the above-mentioned drawings are intended to cover non-exclusive inclusion. For example, a process, method, system, product or device including a series of steps or modules or units does not necessarily limit to those clearly listed steps or modules or units, but may include other steps or modules or units not clearly listed or inherent to these processes, methods, products or devices.
[0055] For the convenience of describing the technical solution of the present invention below, several basic concepts will be described first.
[0056] Dot product attention mechanism (Scaled Dot-Product Attention)
[0057] The Attention function maps a query and a set of key-value pairs to an output. Here, the query, key, value, and output are all vectors. The final output is a weighted sum of the values, and these weights are calculated from the query and the corresponding key.
[0058] Figure 3 Figure 4 shows a schematic diagram of the dot-product attention mechanism. As Figure 3 shown, the operation steps of the dot-product attention mechanism are as follows: First, calculate the inner product of the query (Q) and all keys (K), then divide by (dk is the dimension of the key), and use softmax to obtain the weights of the values (V). Finally, perform a weighted sum to obtain the corresponding output. Figure 3 The Mask layer in Figure 5 is to prevent the attention mechanism from paying attention to the sequence that has not been generated yet in the decoder. The specific formula is as follows:
[0059]
[0060] The dot-product attention mechanism can be used to form a dot-product attention model.
[0061] Multi-head Attention
[0062] Figure 4 Figure 6 shows a schematic diagram of the multi-head attention mechanism. As Figure 4 shown, first use linear transformation to map the query, key, and value into h groups of vectors with dimension d k , d k , d v . On these h groups of query, key, and value vectors, perform the scaled dot-product attention mechanism respectively to obtain h vectors with dimension d v . Then concatenate these vectors to obtain the final output. The specific calculation process is as follows:
[0063] MultiHead(Q, K, V) = Concat(head i ,..., head h )
[0064] where
[0065] where is a model parameter, h equals 8, d k = d v = 512 / 8 = 64.
[0066] The multi - head attention mechanism can be used to form a multi - head attention model.
[0067] According to an embodiment of the present invention, a method for forming a neural network machine translation model is provided. Figure 5 The flowchart of the method for forming a neural network machine translation model according to an embodiment of the present invention is shown. Refer to Figure 5 , according to an embodiment of the present invention, a method for forming a neural network machine translation model includes:
[0068] S502: Form an encoder, where the encoder includes a first multi - head attention model.
[0069] Specifically, forming the encoder includes forming the first multi - head attention model:
[0070] Using the dot - product attention mechanism as Figure 3 shown to form a dot - product attention model, and then setting a linear transformation model for value, key, and query for the dot - product attention model. For example, as Figure 4 shown, assuming the number of heads h of the multi - head attention mechanism is 8 and the dimension of the hidden layer representation is 512, then the input V, K, Q of the dot - product attention model can be linearly transformed through the linear transformation model to map them into 8 groups of vectors with dimensions of 64, 64, and 64. The "8 groups of vectors with dimensions of 64, 64, and 64" here is only exemplary, and the present invention is not limited thereto.
[0071] Through the dot - product attention model, the dot - product attention mechanism can be respectively executed on the above 8 groups of vectors with dimensions of 64, 64, and 64, so as to obtain 8 vectors with a dimension of 64.
[0072] Set a connection model for the dot - product attention model, and through the connection model, the 8 vectors with a dimension of 64 obtained through the dot - product attention mechanism can be connected.
[0073] The above - mentioned dot - product attention model, linear transformation model, and connection model constitute the first multi - head attention model.
[0074] As an example, another linear transformation model can be further set for the above - mentioned first multi - head attention model to linearly transform, for example, the 8 vectors with a dimension of 64 obtained through the connection process.
[0075] The first multi - head attention model can specifically pass through Figure 4The implementation of the multi-head attention mechanism shown can be a multi-head intra-attention model implemented through the multi-head attention mechanism.
[0076] As an example, forming the encoder can include forming a feed-forward neural network such that the feed-forward neural network is fully connected to the first multi-head attention model.
[0077] As an example, residual connections and layer normalization processing can be performed on the first multi-head attention model and the feed-forward neural network.
[0078] According to an embodiment of the invention, the encoder can include a plurality of identical encoding layers, each encoding layer including a first sub-layer containing the first multi-head attention model and a second sub-layer containing the feed-forward neural network. As an example, the encoder can include 6 identical encoding layers, but the present invention is not limited thereto.
[0079] S504: Form a decoder, the decoder including a second multi-head attention model and a future information model, the future information model representing the fusion of the first attention hidden layer representation of the current predicted word and the already generated words and the second attention hidden layer representation of the current predicted word and the future possible words.
[0080] Specifically, forming the decoder includes forming the future information model. Forming the future information model includes the following steps:
[0081] (1) Form a fused attention hidden layer representation.
[0082] Figure 6 The figure shows a schematic diagram of forming a fused attention hidden layer representation according to an embodiment of the present invention.
[0083] As Figure 6 shown in the left half of the figure, use the dot product attention mechanism to calculate the first attention hidden layer representation of the current predicted word and the already generated words:
[0084]
[0085] Among them, represents the first attention hidden layer representation, represents the hidden layer state query value at the current moment, represents the historical hidden layer state key value, represents the historical hidden layer state real value, Attention() is the mathematical function representation of the dot product attention mechanism, and its calculation method is as follows:
[0086]
[0087] Here, dk is the dimension of the key value K, and T represents the transpose of a vector or matrix.
[0088] As Figure 6 shown in the right half of, the second attention hidden layer representation of the current predicted word and possible future words is calculated using the dot product attention mechanism:
[0089]
[0090] where, represents the second attention hidden layer representation, represents the key value of the future hidden layer state, represents the real value of the future hidden layer state, and the calculation method of Attention is shown in formula (2).
[0091] As Figure 6 shown, the above two attention hidden layer representations are fused using a threshold mechanism. The role of this step is to fuse the historical information of the words before the current predicted word and the future information after the current predicted word. The resulting fused attention hidden layer representation is:
[0092]
[0093]
[0094] where, represents the fused attention hidden layer representation, r t represents the reset gate, z t represents the update gate, W g is a model parameter, σ is the sigmoid function, [;] represents the vector or matrix concatenation operation, and · is the operation of multiplying bit by bit. Here, formulas (1) to (5) are collectively referred to as the FutureAtt function.
[0095] It should be noted that the forward arrow here represents the sequence hidden layer state of the source language sequence decoded from left to right (left-to-right decoding), and the backward arrow represents the sequence hidden layer state of the source language sequence decoded from right to left (right-to-left decoding). For left-to-right decoding (left-to-right decoding), the word sequence generated before the current moment is the historical information, and the sequence generated from right to left is the future information.
[0096] (2) Form a dot product attention model using the mechanism of the fused attention hidden layer representation.
[0097] (3) Set a linear transformation model for the dot-product attention model obtained in step (2). The input of the dot-product attention model can be mapped into multiple groups of vectors with a predetermined dimension through linear transformation.
[0098] (4) Set a connection model for the dot-product attention model obtained in step (2). The vectors obtained after being processed by the dot-product attention model for the vectors after linear transformation can be connected through the connection model.
[0099] The future information model can be specifically implemented through the following mechanism:
[0100]
[0101] Among them,
[0102] Among them, Concat represents the concatenation operation, and W o is the model matrix parameter, and h takes the value of 16 here.
[0103] The above-mentioned dot-product attention model, linear transformation model, and connection model formed by using the fused attention hidden layer representation constitute the future information model.
[0104] Optionally, another linear transformation model can be further set for the above future information model to perform linear transformation on the vectors obtained through connection processing.
[0105] Forming the decoder further includes forming a second multi-head attention model. The structure of the second multi-head attention model is the same as that of the first multi-head attention model, and the formation process thereof will not be elaborated here. The second multi-head attention model can be a multi-head external attention (multi-head-inter-attention) model implemented through the multi-head attention mechanism. Among them, the first multi-head attention model and the future information model provide inputs for the second multi-head attention model.
[0106] Forming the decoder further includes forming a feed-forward neural network, and the feed-forward neural network is fully connected to the second multi-head attention model.
[0107] Optionally, residual connection and hierarchical normalization processing can be performed on the second multi-head attention model, the future information model, and the feed-forward neural network.
[0108] S506: Form a first machine translation model through the encoder and the decoder.
[0109] The first machine translation model includes the encoder and the decoder formed through the above steps.
[0110] S508: Perform left-to-right and right-to-left decoding training on the first machine translation model for the source language sequence to form a neural network machine translation model.
[0111] Specifically, adopt the maximum likelihood objective function and use the gradient descent method to perform parameter training on the first machine translation model. The training objective function for training the first machine translation model is the sum of the logarithmic likelihood functions of the translation sentence pairs on the parallel corpus, expressed as:
[0112]
[0113] D represents the set of parallel sentence pairs, and the model parameters θ can be solved by optimization methods such as stochastic gradient descent (SGD), Adam, or Adadelta.
[0114] In addition, when training the first machine translation model formed above in the corpus, for example, for the Chinese source language sequence "This is the secret of success", its corresponding English target language sequence has two forms, one is the forward "This is the secret of success", and the other is the reverse "success of the secret is this". When training the model, add left-to-right (l2r) and right-to-left (r2l) guiding symbols to the corpus so that the model can decode the Chinese source language sequence "This is the secret of success" in two directions, forward and reverse, i.e., "This is the secret of success" and "success of the secret is this". The first machine translation model after training has the decoding ability in both forward and reverse directions, thus forming a neural network machine translation model.
[0115] As an example, the method for forming a neural network machine translation model may further include setting an embedding layer and a position encoding layer at the input end of the encoder. Among them, the embedding layer is used to perform vector transformation on the source language sequence to be input into the encoder to convert the source language sequence into corresponding word vectors, and the position encoding layer is used to encode the positions of each word in the source language sequence one by one to obtain the position vectors of each word, and add the position vectors and word vectors. The added vector is input into the encoder.
[0116] As an example, the method for forming a neural network machine translation model may further include setting an embedding layer and a positional encoding layer on the input side of the future information model in the decoder, where the embedding layer converts the target language sequence to be input into the future information model into corresponding word vectors, and the positional encoding layer is used to encode the positions of each word in the target language sequence one by one to obtain the position vectors of each word, and the position vectors and word vectors are added, and the added vector is input into the future information model.
[0117] As an example, the method for forming a neural network machine translation model may further include setting a linear transformation layer and a normalization layer on the output side of the decoder, where the linear transformation layer is used to perform a linear transformation on the output result of the decoder, and the normalization layer is used to perform a normalization process on the linear transformation result.
[0118] According to an embodiment of the invention, the decoder may include a plurality of identical decoding layers, and each decoding layer includes a first sub-layer including a first multi-head attention model, a second sub-layer including a feed-forward neural network, and a third sub-layer including a future information model. As an example, the decoder may include 6 identical decoding layers, but the invention is not limited thereto.
[0119] According to an embodiment of the present invention, a neural network machine translation model is also provided. Figure 7 The block diagram of the neural network machine translation model according to an embodiment of the present invention is shown. Figure 8 The internal architecture diagram of the neural network machine translation model according to an embodiment of the present invention is shown, where Figure 8 The left part is the encoding layer, and the right part is the decoding layer. N represents the number of layers of the model. The number of layers may be 6 layers, and each layer includes the same structure, but the invention is not limited thereto. The number of layers of the model may be other numbers. The following will be combined with Figure 7 and Figure 8 to describe the neural network machine translation model according to an embodiment of the present invention.
[0120] As Figure 7 shown, according to an embodiment of the present invention, the neural network machine translation model 7 includes: an encoder 702 and a decoder 704.
[0121] The encoder 702 includes: a first multi-head attention model 7022, which is implemented by the multi-head attention mechanism in the left encoding layer; and a feed-forward neural network 7024 connected to the first multi-head attention model 7022, which is implemented by Figure 8 the left encoding layer Figure 8Implementation of the feedforward neural network in the left encoding layer. Among them, the first multi-head attention model can be a multi-head intra-attention model implemented through the multi-head attention mechanism.
[0122] The decoder 704 includes: a future information model 7042, which is Figure 8 Implemented by the future information model in the right decoding layer. Among them, the future information model is formed through the steps of forming the future information model in the method of forming the neural network machine translation model described above. The future information model 7042 receives the historical information of the generated words before the current predicted word and the future information of the possible words after the current predicted word, and processes the historical information and the future information to form a processing result; a second multi-head attention model 7044 connected to the future information model 7042 and the feedforward neural network 7024 in the encoder 702, which is Figure 8 Implemented by the multi-head attention mechanism in the right decoding layer. The second multi-head attention model 7044 receives inputs from the feedforward neural network 7024 and the future information model 7042; and a feedforward neural network 7046 connected to the second multi-head attention model 7044, which is Figure 8 Implemented by the feedforward neural network in the right decoding layer. Among them, the second multi-head attention model can be a multi-head inter-attention model implemented through the multi-head attention mechanism.
[0123] Among them, the first multi-head attention model 7022 and the second multi-head attention model 7044 have the same structure. The first multi-head attention model 7022, the future information model 7042, and the second multi-head attention model 7044 in the neural network machine translation model according to the embodiment of the present invention can be Figure 5 Formed by the method of forming the neural network machine translation model according to the embodiment of the present invention as shown, and their specific structures will not be elaborated here.
[0124] As Figure 7 Shown, the neural network machine translation model 7 according to the embodiment of the present invention may further include an embedding model 706 and a position encoding model 708 provided at the input end of the encoder, which are respectively Figure 8 Implemented by the embedding layer and the position encoding layer at the input end of the left encoding layer. The functions of the embedding layer and the position encoding layer have been introduced above and will not be elaborated here.
[0125] The neural network machine translation model 7 according to the embodiment of the present invention may further include an embedding model 710 and a position encoding model 712 provided at the input end of the decoder, which are respectively Figure 8The embedding layer and the positional encoding layer at the input end of the right decoding layer are implemented. The functions of the embedding layer and the positional encoding layer have been introduced above and will not be elaborated here.
[0126] According to an embodiment of the present invention, the neural network machine translation model 7 may further include a linear transformation model 714 and a normalization model 716 provided at the output side of the decoder, which are respectively implemented by Figure 8 the linear transformation layer and the normalization layer at the output of the right decoding layer. The functions of the linear transformation layer and the normalization layer have been introduced above and will not be elaborated here.
[0127] According to an embodiment of the invention, the encoder may include a plurality of identical encoding layers, and each encoding layer includes a first sub-layer including a first multi-head attention model 7022 and a second sub-layer including a feed-forward neural network 7024. As an example, the encoder may include 6 identical encoding layers, but the present invention is not limited thereto.
[0128] According to an embodiment of the invention, the decoder may include a plurality of identical decoding layers, and each decoding layer includes a first sub-layer including a second multi-head attention model 7044, a second sub-layer including a feed-forward neural network 7046, and a third sub-layer including a future information model 7042. As an example, the decoder may include 6 identical decoding layers, but the present invention is not limited thereto.
[0129] According to an embodiment of the present invention, a neural network machine translation method is provided. For ease of understanding, the following is combined with Figure 8 to specifically describe the neural network machine translation method according to an embodiment of the invention.
[0130] As Figure 8 shown, at the input side of the encoder, an input of a source language sequence is received, and each word in the source language sequence is formed into a corresponding word vector through matrix transformation by using the embedding layer ( Figure 8 input embedding in it), and the positions of the words in the source language sequence are encoded by using the positional encoding layer to form corresponding position vectors. The obtained position vectors are added to the word vectors so that the word vectors have position information, and a first hidden layer vector representation of the source language sequence is obtained from the word vectors with position information.
[0131] The first hidden layer vector representation of the obtained source language sequence is subjected to matrix transformation to obtain the corresponding original query (Q), key (K), and value (V) of the source language sequence. The original Q, K, and V are input into the first multi-head attention model, and after the first multi-head attention model processes Q, K, and V by using the multi-head attention mechanism, a second hidden layer vector representation of the source language sequence is obtained.
[0132] Perform residual connection and layer normalization on the second hidden layer vector representation. For example, add the first hidden layer vector representation and the second hidden layer vector and perform normalization to optimize the model.
[0133] Use a feed-forward neural network to perform a non-linear transformation on the second hidden layer vector representation that has undergone residual connection and layer normalization to obtain a third hidden layer vector representation.
[0134] Perform residual connection and layer normalization on the third hidden layer vector representation. For example, add the second hidden layer vector representation that has undergone residual connection and layer normalization and the third hidden layer vector representation and perform normalization to further optimize the model.
[0135] Perform a matrix transformation on the third hidden layer vector that has undergone residual connection and layer normalization to obtain the processed K and V of the source language sequence after the above processing.
[0136] Input the processed K and V into the second multi-head attention model.
[0137] Receive, at the input side of the decoder, an input of a target language sequence that has been shifted with respect to the target language sequence output from the output side of the decoder. Here, the target language sequence includes a forward target language sequence and an inverted target language sequence corresponding to the source language sequence.
[0138] Figure 9 Examples of the input and output of the decoder shown in chronological order according to an embodiment of the present invention, where (a) represents an example of the input of the decoder and (b) represents an example of the output of the decoder. In Figure 9 Among them, <pad>It is a placeholder, and its vector representation is all zeros. <l2r>And <r2l>are respectively used to guide the decoding directions from left to right (left-to-right decoding) and from right to left (right-to-right decoding), <eos>Indicates the end of a sentence.
[0139] For example, taking the source language sequence "I love China" as an example, the target language sequences output on the decoder side are "I love China." and ".China love I", as shown in (b) of Figure 9 .
[0140] Taking the third moment as an example, the words input on the input side of the decoder are "I" and ".", and the words output on the output side of the decoder are "love" and "China". That is, on the decoder side, when predicting the current word "love", the target language words output in the forward and backward directions at the previous moment of this decoded word in the source language sequence are utilized.
[0141] The target language sequence received on the input side of the decoder is transformed into corresponding word vectors through matrix transformation using an embedding layer ( Figure 8 which is output embedding in this case). The position encoder is used to encode the positions of each word in the received target language sequence to form corresponding position vectors, and the obtained position vectors are added to the word vectors so that the word vectors have position information, and the first hidden layer representation of the target language sequence is obtained from this word vector with position information.
[0142] The first hidden layer representation of the target language sequence obtained is subjected to matrix transformation to obtain the corresponding original query (Q), key (K), and value (V) of the target language sequence, and the original Q, K, and V are input into the future information model. The future information model has been introduced in the previous text and will not be elaborated here.
[0143] The future information model processes the input Q, K, and V to obtain the second hidden layer vector representation of the target language sequence.
[0144] Residual connection and layer normalization processing are performed on the second hidden layer vector representation of the target language sequence. For example, the first hidden layer vector representation and the second hidden layer vector representation of the target language sequence are added and normalized to optimize the model.
[0145] Matrix transformation is performed on the second hidden layer vector of the target language sequence that has undergone residual connection and layer normalization processing to obtain the processed Q of the target language sequence after the above processing.
[0146] The processed Q above is input into the second multi-head attention model.
[0147] The second multi-head attention model uses the multi-head attention mechanism to process the K and V of the source language sequence input from the decoder and the Q of the target language sequence input from the future information model to obtain the fourth hidden layer vector representation.
[0148] Perform residual connection and layer normalization on the fourth hidden layer vector representation. For example, add the second hidden layer vector representation that has undergone residual connection and layer normalization from the future information model to the fourth hidden layer vector representation and perform normalization to further optimize the model.
[0149] Use a feed-forward neural network to perform a non-linear transformation on the fourth hidden layer vector representation that has undergone residual connection and layer normalization to obtain the fifth hidden layer vector representation.
[0150] Perform residual connection and layer normalization on the fifth hidden layer vector representation. For example, add the fourth hidden layer vector representation and the fifth hidden layer vector representation that have undergone residual connection and layer normalization and perform normalization to further optimize the model.
[0151] Linearly process the fifth hidden layer vector representation that has undergone residual connection and layer normalization, and use the Softmax function to calculate the predicted probability of the current word.
[0152] Generate the text sequence with the highest probability for each word as the target language sequence of the translation result.
[0153] In this article, the role of the feed-forward neural network is to perform a non-linear transformation on the hidden layer vector representation of the model, and its calculation method is as follows:
[0154] FFN(x) = max(0, xW1 + b1)W2 + b2
[0155] Here, W1, W2, b1, and b2 are model parameters.
[0156] In the neural network machine translation method according to the embodiments of the present invention, the multi-head attention mechanism is used in three ways: (i) the encoder-decoder attention layer, that is Figure 7 the second multi-head attention model in, where the query Q input to the second multi-head attention model comes from the future information model formed by using the fused multi-head attention mechanism in the decoder, and the key K and value V input to the second multi-head attention model come from the output of the encoder; (2) the encoder layer separately uses the first multi-head attention model, and the query Q, key K, and value V input to the first multi-head attention model all come from the input of the previous layer of the encoder; (3) the decoder layer uses the future information model, and the query Q, key K, and value V input to the future information model all come from the input of the previous layer of the decoder.
[0157] Experimental results
[0158] The inventors conducted experiments on the neural network machine translation model according to the embodiments of the present invention. In the experiments, 2 million aligned sentence pairs were extracted from the Chinese-English training data released by the Linguistic Data Consortium as the Chinese-English training corpus, and all test sets MT03 - MT06 from 2003 to 2006 in the NIST MT Evaluation were used as the development set and the test set. Among them, MT03 was used as the development set. In the comparative experiment, case-insensitive BLEU-4 was used as the evaluation metric.
[0159] Table 1 shows the performance of the present invention, the standard deep neural machine translation system, and the statistical machine translation system on 4 sets of test data (MT03, MT04, MT05, MT06), where "future" represents incorporating future information.
[0160]
[0161] Table 1
[0162] It can be seen that after incorporating future information, the present invention has an improvement of 1.32 BLEU values compared to the standard deep neural machine translation system in the evaluation metric (BLEU) automatically given by the machine. In addition, a manual evaluation was conducted on the omission situation of the test set. The results show that after incorporating future information, the omission error rate of the model decreased by 30.6%, greatly improving the serious problem of model omission and enhancing the translation quality of the model.
[0163] According to an embodiment of the present invention, there is also provided a storage medium. The storage medium includes a stored program, wherein when the program runs, it controls a device including the storage medium to execute the method of forming the neural network machine translation model or execute the neural network machine translation method.
[0164] According to an embodiment of the present invention, there is also provided a processor. The processor is used to run a program, wherein when the program runs, it executes the method of forming the neural network machine translation model or executes the neural network machine translation method.
[0165] According to an embodiment of the present invention, there is also provided an electronic device, including: one or more processors, a memory, a display device, and one or more programs, wherein the one or more programs are stored in the memory and are configured to be executed by the one or more processors so that the electronic device executes the method of forming the neural network machine translation model or executes the neural network machine translation method.
[0166] In addition, the present disclosure includes embodiments according to the following items:
[0167] Item 1. A method for forming a neural network machine translation model, characterized by comprising:
[0168] Forming an encoder, the encoder including a first multi-head attention model;
[0169] Forming a decoder, the decoder including a second multi-head attention model and a future information model, the future information model representing the fusion of the first attention hidden layer representation of the current predicted word and the already generated words and the second attention hidden layer representation of the current predicted word and the future possible words;
[0170] Forming a first machine translation model through the encoder and the decoder; and
[0171] Performing decoding training on the first machine translation model from left to right and from right to left for the source language sequence to form the neural network machine translation model,
[0172] wherein, the first multi-head attention model and the future information model provide inputs for the second multi-head attention model.
[0173] Item 2. The method according to Item 1, characterized in that forming the encoder includes forming the first multi-head attention model, wherein forming the first multi-head attention model includes:
[0174] Using the dot product attention mechanism to form a dot product attention model;
[0175] Setting a linear transformation model for the dot product attention model to map the input of the dot product attention model into multiple groups of vectors of a predetermined dimension through linear transformation;
[0176] Setting a connection model for the dot product attention model to connect the vectors obtained after the vectors are processed by the dot product attention model; and
[0177] Forming the first multi-head attention model through the dot product attention model, the linear transformation model and the connection model.
[0178] Item 3. The method according to Item 1, characterized in that forming the decoder includes forming the future information model, wherein forming the future information model includes:
[0179] Using the dot product attention mechanism to calculate the first attention hidden layer representation of the current predicted word and the already generated words: wherein, represents the first attention hidden layer representation, represents the hidden layer state query value at the current moment, represents the historical hidden layer state key value, Represents the real value of the hidden state of history, and Attention() is represented by the mathematical function of the dot product attention mechanism;
[0180] Using the dot product attention mechanism to calculate the second attention hidden layer representation of the current predicted word and future possible words: Wherein, Represents the second attention hidden layer representation, Represents the key value of the hidden state in the future, Represents the real value of the hidden state in the future;
[0181] Fusing the first attention hidden layer representation and the second attention hidden layer representation to form a fused attention hidden layer representation;
[0182] Using the fused attention hidden layer representation to form a dot product attention model;
[0183] Setting a linear transformation model for the dot product attention model to map the input of the dot product attention model into multiple groups of vectors with a predetermined dimension through linear transformation;
[0184] Setting a connection model for the dot product attention model to connect the vectors obtained after the vectors are processed by the dot product attention model; and
[0185] Forming the future information model through the dot product attention model, the linear transformation model and the connection model.
[0186] Item 4. The method according to item 3, wherein fusing the first attention hidden layer representation and the second attention hidden layer representation to form a fused attention hidden layer representation includes:
[0187] Using a threshold mechanism to fuse the first attention hidden layer representation and the second attention hidden layer representation into a fused attention hidden layer representation:
[0188]
[0189]
[0190] Wherein, Represents the fused attention hidden layer representation, r t Represents the reset gate, z t Represents the update gate, W g Is a model parameter, σ is the sigmoid function, [;] represents vector or matrix concatenation operation, · is the bitwise product operation.
[0191] Item 5. The method according to any one of Items 1 to 4, characterized in that the first multi-head attention model and the second multi-head attention model have the same structure.
[0192] Item 6. The method according to any one of Items 1 to 4, characterized in that the encoder comprises a plurality of encoding layers each containing the first multi-head attention model, and the decoder comprises a plurality of decoding layers each containing the second multi-head attention model and the future information model.
[0193] Item 7. The method according to Item 6, characterized in that each layer of the encoder and the decoder further comprises a feed-forward neural network.
[0194] Item 8. A neural network machine translation model, characterized by comprising:
[0195] An encoder comprising a first multi-head attention model; and
[0196] A decoder comprising a second multi-head attention model and a future information model, the future information model representing the fusion of a first attention hidden layer representation of the current predicted word and the already generated words and a second attention hidden layer representation of the current predicted word and the future possible words,
[0197] wherein the first multi-head attention model and the future information model provide inputs to the second multi-head attention model, and the neural network machine translation model has been trained by decoding the source language sequence from left to right and from right to left.
[0198] Item 9. The neural network machine translation model according to Item 8, characterized in that the first multi-head attention model comprises:
[0199] A dot product attention model configured to form using a dot product attention mechanism;
[0200] A linear transformation model connected to the dot product attention model and configured to map the input of the dot product attention model into multiple groups of vectors of a predetermined dimension through linear transformation; and
[0201] A connection model connected to the dot product attention model and configured to connect the vectors obtained after the vectors are processed by the dot product attention model.
[0202] Item 10. The neural network machine translation model according to Item 8, characterized in that the future information model comprises:
[0203] A dot product attention model;
[0204] A linear transformation model, connected to the dot-product attention model and configured to map the input of the dot-product attention model into multiple groups of vectors of a predetermined dimension through linear transformation; and
[0205] A connection model, connected to the dot-product attention model and configured to concatenate the vectors obtained after processing the vectors by the dot-product attention model,
[0206] wherein, the dot-product attention model is formed through the following steps:
[0207] Using the dot-product attention mechanism to calculate the first attention hidden layer representation of the current predicted word and the already generated words: wherein, represents the hidden layer state query value at the current moment, represents the hidden layer state key value of the history, represents the hidden layer state real value of the history, and Attention() is the mathematical function representation of the dot-product attention mechanism;
[0208] Using the dot-product attention mechanism to calculate the second attention hidden layer representation of the current predicted word and the possible future words: wherein, represents the hidden layer state key value of the future, represents the hidden layer state real value of the future;
[0209] Fusing the first attention hidden layer representation and the second attention hidden layer representation to form a fused attention hidden layer representation; and
[0210] Using the fused attention hidden layer representation to form the dot-product attention model.
[0211] Item 11. The neural network machine translation model according to item 10, wherein fusing the first attention hidden layer representation and the second attention hidden layer representation to form a fused attention hidden layer representation includes:
[0212] Using a threshold mechanism to fuse the first attention hidden layer representation and the second attention hidden layer representation into a fused attention hidden layer representation:
[0213]
[0214]
[0215] wherein, represents the fused attention hidden layer representation, r t represents the reset gate, z t represents the update gate, W g is a model parameter, σ is the sigmoid function, [;] represents the vector or matrix concatenation operation, and · represents the bitwise product operation.
[0216] Item 12. The neural network machine translation model according to any one of Items 8 to 11, characterized in that the first multi-head attention model and the second multi-head attention model have the same structure.
[0217] Item 13. The neural network machine translation model according to any one of Items 8 to 11, characterized in that the encoder includes multiple layers each containing the first multi-head attention model, and the decoder includes multiple layers each containing the second multi-head attention model and the future information model.
[0218] Item 14. The neural network machine translation model according to Item 13, characterized in that each layer of the encoder and the decoder further includes a feed-forward neural network.
[0219] Item 15. A neural network machine translation method, characterized in that it includes:
[0220] Receiving a source language sequence, and using a first multi-head attention model to process the source language sequence by means of the multi-head attention mechanism to obtain a hidden layer vector representation of the source language sequence;
[0221] Performing a matrix transformation on the hidden layer vector representation of the source language sequence to obtain the corresponding key K and value V of the source language sequence;
[0222] Inputting the key K and the value V into a second multi-head attention model;
[0223] Obtaining a fused attention hidden layer vector representation of the first attention hidden layer representation of the current predicted word and the already generated words and the second attention hidden layer representation of the current predicted word and the possible future words through a future information model, where the future information model represents the fusion of the first attention hidden layer representation of the current predicted word and the already generated words and the second attention hidden layer representation of the current predicted word and the possible future words;
[0224] Obtaining a query Q of the target language sequence according to the fused attention hidden layer vector representation;
[0225] Inputting the query Q into the second multi-head attention model;
[0226] Determining, by the second multi-head attention model, the probability that the target language word is the current predicted word according to the key K, the value V, and the query Q; and
[0227] Generating a text sequence with the highest probability for each word as the target language sequence of the translation result.
[0228] Item 16. A storage medium, characterized in that the storage medium includes a stored program, wherein when the program runs, it controls a device including the storage medium to execute the method according to any one of Items 1 to 7 and 15.
[0229] Item 18. A processor, characterized in that the processor is used to run a program, wherein when the program runs, it executes the method according to any one of Items 1 to 7.
[0230] Item 19. An electronic device, characterized in that it includes: one or more processors, a memory, a display device, and one or more programs, wherein the one or more programs are stored in the memory and are configured to be executed by the one or more processors, so that the electronic device executes the method according to any one of Items 1 to 7.
[0231] In the above embodiments of the present invention, the descriptions of the respective embodiments have their own emphases. For parts not detailed in a certain embodiment, reference may be made to the relevant descriptions of other embodiments.
[0232] In the several embodiments provided by the present invention, it should be understood that the disclosed technical content can be implemented in other ways. Among them, the device embodiments described above are merely illustrative. For example, the division of the units or modules is only a logical function division. In actual implementation, there may be other division methods. For example, multiple units or modules or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed couplings or direct couplings or communication connections to each other can be through some interfaces, and the indirect couplings or communication connections of the modules or units can be in an electrical or other form.
[0233] The units or modules described as separate components may or may not be physically separated. The components shown as units or modules may or may not be physical units or modules, that is, they can be located in one place, or they can be distributed to multiple network units or modules. Some or all of the units or modules can be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0234] In addition, the functional units or modules in the respective embodiments of the present invention can be integrated into one processing unit or module, or each unit or module can exist physically alone, or two or more units or modules can be integrated into one unit or module. The above integrated units or modules can be implemented in the form of hardware or in the form of software functional units or modules.
[0235] When the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or all or part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in various embodiments of the present invention. The aforementioned storage medium includes: various media that can store program codes, such as USB flash drives, read-only memories (ROMs), random access memories (RAMs), mobile hard disks, magnetic disks, or optical discs.
[0236] The above are only the preferred embodiments of the present invention. It should be noted that for those of ordinary skill in the art, without departing from the principle of the present invention, several improvements and refinements can be made, and these improvements and refinements should also be regarded as the protection scope of the present invention.< / eos> < / pad>
Claims
1. A method for forming a neural network machine translation model, characterized in that, Including: Forming an encoder, the encoder including a first multi-head attention model; Forming a decoder, the decoder including a second multi-head attention model and a future information model, the future information model representing the fusion of a first attention hidden layer representation of a current predicted word and a generated word and a second attention hidden layer representation of the current predicted word and a future possible word; Forming a first machine translation model through the encoder and the decoder; And Performing decoding training on the first machine translation model from left to right and from right to left for a source language sequence to form the neural network machine translation model, wherein the first multi-head attention model and the future information model provide inputs for the second multi-head attention model, wherein forming the decoder includes forming the future information model, and forming the future information model includes: Calculate the first attention hidden layer representation of the current predicted word and the already generated words using the dot product attention mechanism: Among them, represents the first attention hidden layer representation, represents the hidden layer state query value at the current moment, represents the hidden layer state key value of history, represents the hidden layer state real value of history, and Attention() is the mathematical function representation of the dot product attention mechanism; Calculate the second attention hidden layer representation of the current predicted word and possible future words using the dot product attention mechanism: Among them, represents the second attention hidden layer representation, represents the hidden layer state key value in the future, represents the hidden layer state real value in the future; Fusing the first attention hidden layer representation and the second attention hidden layer representation to form a fused attention hidden layer representation; Using the fused attention hidden layer representation to form a dot product attention model; Setting a linear transformation model for the dot product attention model to map the input of the dot product attention model into multiple groups of vectors of a predetermined dimension through linear transformation; Setting a connection model for the dot product attention model to connect the vectors obtained after processing the vectors by the dot product attention model; and Forming the future information model through the dot product attention model, the linear transformation model, and the connection model.
2. The method according to claim 1, characterized in that, Forming the encoder includes forming the first multi-head attention model, wherein forming the first multi-head attention model includes: Using a dot product attention mechanism to form a dot product attention model; Setting a linear transformation model for the dot product attention model to map the input of the dot product attention model into multiple groups of vectors of a predetermined dimension through linear transformation; Setting a connection model for the dot product attention model to connect the vectors obtained after processing the vectors by the dot product attention model; and Forming the first multi-head attention model through the dot product attention model, the linear transformation model, and the connection model.
3. The method according to claim 1, characterized in that Fusing the first attention hidden layer representation and the second attention hidden layer representation to form a fused attention hidden layer representation includes: Using a threshold mechanism to fuse the first attention hidden layer representation and the second attention hidden layer representation into a fused attention hidden layer representation: Among them, represents the fused attention hidden layer representation, r t represents the reset gate, z t represents the update gate, W g is the model parameter, σ is the sigmoid function, [;] represents the vector or matrix concatenation operation, and · is the bitwise product operation.
4. A neural network machine translation model, characterized in that, Including: An encoder, the encoder including a first multi-head attention model; And A decoder, the decoder including a second multi-head attention model and a future information model, the future information model representing the fusion of a first attention hidden layer representation of a current predicted word and a generated word and a second attention hidden layer representation of the current predicted word and a future possible word, wherein the first multi-head attention model and the future information model provide inputs for the second multi-head attention model, and the neural network machine translation model has undergone decoding training from left to right and from right to left for a source language sequence, wherein the future information model includes: A dot product attention model; A linear transformation model, connected to the dot-product attention model and configured to map the input of the dot-product attention model into multiple groups of vectors of a predetermined dimension through linear transformation; and A connection model, connected to the dot-product attention model and configured to concatenate the vectors obtained after processing the vectors by the dot-product attention model, wherein, the dot-product attention model is formed through the following steps: Calculate the first attention hidden layer representation of the current predicted word and the already generated words using the dot product attention mechanism: Among them, represents the hidden layer state query value at the current moment, represents the hidden layer state key value (key) of history, represents the hidden layer state real value of history, and Attention() is the mathematical function representation of the dot product attention mechanism; Calculate the second attention hidden layer representation of the current predicted word and possible future words using the dot product attention mechanism: Among them, represents the key value of the future hidden layer state, represents the real value of the future hidden layer state; Fusing the first attention hidden layer representation and the second attention hidden layer representation to form a fused attention hidden layer representation; and Using the fused attention hidden layer representation to form the dot-product attention model.
5. The neural network machine translation model according to claim 4, wherein The first multi-head attention model includes: A dot-product attention model, configured to be formed using the dot-product attention mechanism; A linear transformation model, connected to the dot-product attention model and configured to map the input of the dot-product attention model into multiple groups of vectors of a predetermined dimension through linear transformation; and A connection model, connected to the dot-product attention model and configured to concatenate the vectors obtained after processing the vectors by the dot-product attention model.
6. The neural network machine translation model according to claim 4, wherein Fusing the first attention hidden layer representation and the second attention hidden layer representation to form a fused attention hidden layer representation includes: Using a threshold mechanism to fuse the first attention hidden layer representation and the second attention hidden layer representation into a fused attention hidden layer representation: Among them, represents the fused attention hidden layer representation, r t represents the reset gate, z t represents the update gate, W g is the model parameter, σ is the sigmoid function, [;] represents the vector or matrix concatenation operation, and · is the bitwise product operation.
7. A neural network machine translation method, characterized in that, including: Receiving a source language sequence, and processing the source language sequence by a first multi-head attention model using the multi-head attention mechanism to obtain a hidden layer vector representation of the source language sequence; Performing a matrix transformation on the hidden layer vector representation of the source language sequence to obtain corresponding key K and value V of the source language sequence; Inputting the key K and the value V into a second multi-head attention model; Obtaining a fused attention hidden layer vector representation of the first attention hidden layer representation of the current predicted word and the already generated words and the second attention hidden layer representation of the current predicted word and the future possible words through a future information model, where the future information model represents the fusion of the first attention hidden layer representation of the current predicted word and the already generated words and the second attention hidden layer representation of the current predicted word and the future possible words; Obtaining a query Q of the target language sequence according to the fused attention hidden layer vector representation; Inputting the query Q into the second multi-head attention model; Determining, by the second multi-head attention model, the probability that the target language word is the current predicted word according to the key K, the value V, and the query Q; and Generating a text sequence with the highest probability for each word as the target language sequence of the translation result.
8. A storage medium, characterized in that, The storage medium includes a stored program, wherein, when the program runs, it controls a device including the storage medium to execute the method according to any one of claims 1 to 3.
Citation Information
Patent Citations
Neural machine translation method and device based on word vector connection technology
CN107729329A
Text sentiment classification algorithm based on convolutional neural network and attention mechanism
CN108664632A