A Neural Machine Translation Method Based on Feature Pyramid

By introducing feature pyramid structure and multi-scale encoding and decoding attention in the Transformer model, the problems of encoder layer redundancy and dimension inconsistency are solved, and the translation quality of neural machine translation is improved.

CN114528854BActive Publication Date: 2025-07-08XIAONIU FANYI
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202210073567.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-01-21
Publication Date
2025-07-08
Estimated Expiration
2042-01-21

AI Technical Summary

Technical Problem

The existing Transformer model has the problem of encoder layer redundant calculation and decoder focusing only on the output of the last layer of the encoder in neural machine translation, which makes the model unable to effectively utilize fine-grained source language context information.

Method used

Using a feature pyramid-based method, the hidden layer state is scaled into a pyramid type through a fully connected layer in the encoder, divided into different sub-blocks, and multi-scale multi-head segmentation and weighted average encoding and decoding attention calculation are performed in the decoder, focusing on the encoding information of different layers.

Benefits of technology

The redundant calculation of the encoder layer is reduced, the model's ability to capture semantic information of different granularity is improved, and the translation quality is improved without increasing model parameters.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114528854B_ABST
    Figure CN114528854B_ABST
Patent Text Reader

Abstract

The present invention discloses a neural machine translation method based on a feature pyramid. The steps are as follows: The preprocessed source language input is sent to the encoder end of the translation model to be encoded into context vectors of different dimensions; the hidden layer dimensions of the encoder are scaled in a pyramid shape during the feed-forward process and divided into different sub-blocks; during the calculation of the encoder-decoder attention weights, the encoded key vectors and decoder query vectors of different dimensions are subjected to multi-head segmentation at different scales; the encoder-decoder attention calculation results of different dimensions are weighted and averaged to obtain the final encoder-decoder attention output vector. The decoder decodes the source language context vector into the target language translation through the stacked decoder layers, and the gradient is updated through the cross-entropy loss function to optimize the weights of the translation model. The present invention reduces the redundant calculation in the encoder-decoder attention calculation, solves the problem of inconsistent dimensions in the encoder-decoder attention calculation, effectively integrates source language features of different scales, and improves the translation quality of the model.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to a machine language translation technology, specifically a neural machine translation method based on a feature pyramid. Background Art

[0002] Machine Translation (MT for short) is an experimental discipline that uses computers to translate between natural languages. Using machine translation technology, a source language can be automatically converted into a target language. As a key technology to eliminate the barriers of cross - language communication, machine translation has always been an important part of natural language processing research. Compared with manual translation, machine translation is more efficient and less costly, which is of great significance for promoting cultural dissemination and communication. Machine translation technology can be generally classified into two methods: the rational - based method and the empirical - based method. Since it was proposed in the 1940s, machine translation has experienced nearly 70 years of development, and its development history can be roughly divided into three stages: rule - based machine translation, statistical - based machine translation, and neural network - based machine translation.

[0003] The rule - based machine translation technology uses the method of manually constructing rules to perform corresponding conversions on the source language input to obtain the target translation result. The disadvantage of this method is that it requires a large amount of manual effort to construct rules, the rule coverage is limited and conflicts may occur, resulting in poor system scalability and robustness. Subsequently, researchers adopted statistical - based machine translation technology, using statistical methods for modeling and completely abandoning the dependence on manual rules. Statistical machine translation needs to perform statistical analysis on a large amount of bilingual parallel corpora to construct a statistical translation model to complete the translation. In recent years, neural networks have received extensive attention in the field of machine translation. Neural machine translation adopts an end - to - end encoder - decoder framework. The encoder encodes the source language input into a dense source language context vector, and the decoder is responsible for autoregressive decoding with reference to the source language context vector to generate the final translation result.

[0004] As the most mainstream neural machine translation model currently, the Transformer model still adopts an encoder - decoder structure, using the attention mechanism to encode and decode the source language context information and the position information of words. During the attention calculation process, a multi - head splitting method is adopted so that different heads can focus on information in different semantic spaces. In addition to self - attention, there is also encoder - decoder attention in the attention mechanism. The difference between them is that the query vector, key vector, and value vector of self - attention are all intermediate vectors of the same layer, while the query vector of encoder - decoder attention is the intermediate vector of the decoding end, and the key vector and value vector are the source language encoding vectors output by the encoding end. In the Transformer, the size of the hidden layer dimension is fixed (usually 512).

[0005] The encoder and decoder of the Transformer are constructed by stacking layers. Due to the stacking of layers with the same structure, there is a certain degree of redundancy in the outputs of different layers, resulting in low encoding efficiency of the model. At the same time, the decoder only focuses on the output of the last layer of the encoder, causing the model to be unable to utilize the source language context information from a more fine-grained aspect.

[0006] The feature pyramid model was first proposed in the field of images. Tasks in the field of images generally use convolutional neural networks, and due to the existence of pooling layers, convolutional neural networks naturally present a pyramid shape. In object detection tasks, since the scale sizes of different objects are different, and the features extracted by different layers in the convolutional network have different coarseness and fineness, it is necessary to reversely transfer the features of the last layer and fuse them with the underlying information as the feature information finally required for the detection task. Summary of the Invention

[0007] The present invention proposes a neural machine translation method based on a feature pyramid. When calculating the encoder-decoder attention at the decoder end in the Transformer model, it not only focuses on the final source language encoding information at the encoder end but also focuses on the hidden layer information of different layers at the encoder end.

[0008] To solve the above technical problems, the technical solution adopted by the present invention is:

[0009] The present invention provides a neural machine translation method based on a feature pyramid, including the following steps:

[0010] 1) Send the preprocessed source language input into the encoder end of the translation model, and encode it into context vectors of different dimensions through the word embedding layer and the stacked encoder layers;

[0011] 2) The hidden layer dimensions of the encoder are scaled in a pyramid shape during the feed-forward process. According to the dimension sizes, the encoder layers are divided into different sub-blocks, and the output vectors of different sub-blocks are saved in the source language encoding vector sequence;

[0012] 3) Send the source language encoding vector sequence of the encoder into the encoder-decoder attention module at the decoder end of the translation model. During the calculation of the encoder-decoder attention weights, perform multi-head splitting of different scales on the encoding key vectors and decoding query vectors of different dimensions, so that the encoded sub-key vectors and decoding sub-query vectors have the same dimension;

[0013] 4) Weightedly average the calculation results of the encoder-decoder attention of different dimensions to obtain the final encoder-decoder attention output vector. The decoder decodes the source language context vector into the target language translation through the stacked decoder layers and updates the gradient through the cross-entropy loss function to optimize the weights of the translation model.

[0014] In step 1), the training data is preprocessed, and the source language input is fed into the encoder end of the translation model to encode the source language information into context vectors of different dimensions. The encoder consists of stacked encoder layers, and each encoder layer includes a self-attention sub-layer and a fully-connected sub-layer. Among them, the calculation method of the self-attention sub-layer is as follows:

[0015]

[0016]

[0017] where Self_Att represents the self-attention sub-layer, is the input hidden layer vector, W q , W k , W v are the parameters of the self-attention sub-layer, softmax(·) is the attention weight calculation function, L s is the length of the source sentence, d is the corresponding hidden layer vector dimension, a is the self-attention weight,

[0018] In step 2), the hidden layer dimension of the encoder is scaled in a pyramid shape during the feed-forward process, and the scaling process occurs in each fully-connected sub-layer:

[0019] h o = W2ReLU(h i W1 + b1) + b2

[0020] where h i is the input vector of the fully-connected sub-layer, that is, the output vector of the previous self-attention sub-layer, H o is the output vector of the fully-connected layer, W1 ∈ R d×4d , W2 ∈ R 4d×d / 2 , b1 ∈ R 4d , b2 ∈ R d / 2 are the parameters of the fully-connected sub-layer, ReLU is the activation function, and finally the fully-connected sub-layer scales the input from the corresponding hidden layer vector dimension d to d / 2;

[0021] According to the dimension size, the encoder is divided into different sub-blocks, and the output vectors of different sub-blocks are saved in the source language encoding vector sequence as a set of source language encoding information of different dimensions.

[0022] In step 3), the source language encoding vector sequence of the encoder is fed into the decoder end of the model. The decoder introduces the source language context information through the multi-head encoder-decoder attention module, that is, the source language encoding vector sequence of the encoder; the multi-scale segmentation method is used during multi-head segmentation; for the decoding query vector with dimension D q and the source language key vector with dimension D k , the calculation steps of the attention score are as follows:

[0023] 301) Cut the encoded key vector into source language sub-key vectors with head_key dimensions of head_dim, where head_key is the number of key heads, head_dim is the dimension of the head, and head_key = D k / head_dim;

[0024] 302) Cut the decoded query vector into decoded sub-query vectors with head_query dimensions of head_dim, where head_query is the number of query heads, head_dim is the dimension of the head, and head_query = D q / head_dim, where D q = n × D k At this time, head_query = n × head_key, and n is the proportionality coefficient of D q and D k ;

[0025] 303) When n > 1, n decoded sub-query vectors correspond to 1 encoded sub-key vector; when n = 1, 1 decoded sub-query vector corresponds to 1 encoded sub-key vector; when n < 1, 1 decoded sub-query vector corresponds to n encoded sub-key vectors; the calculation method of the encoder-decoder attention is as follows:

[0026]

[0027] where is the input hidden layer vector, which serves as the decoded query vector here, and L t is the length of the target sentence, is the encoded key vector, which serves as the decoded key vector here, and L s is the length of the source sentence.

[0028] In step 4), the calculation results of the encoder-decoder attention with different dimensions are weighted and averaged, and the weights are learnable. The final calculation method of the encoder-decoder attention result is as follows:

[0029]

[0030] where L is the number of layers of the encoder, corresponding to the outputs of L encoder layers with different dimensions, and v i represents the calculation result of the encoder-decoder attention for the output of the i-th encoder layer, and γ i is the weight parameter learnable by the model;

[0031] The decoder decodes the source language context vector into the target language translation through stacked decoder layers and updates the gradient through the cross-entropy loss function to optimize the weights of the model.

[0032] The present invention has the following beneficial effects and advantages:

[0033] 1. The present invention uses the form of a feature pyramid to reduce the redundancy of features at different layers, enabling the encoder to capture semantic information with different levels of granularity.

[0034] 2. When calculating the encoder-decoder attention, the present invention performs multi-head cutting at different scales according to the dimension size, solving the problem of inconsistent dimensions between the query vector and the key vector, so that the decoder can focus on source language context vectors of different dimensions.

[0035] 3. Through multi-scale encoder-decoder attention calculation, the present invention can effectively improve the translation quality of the model with almost no increase in model parameters. BRIEF DESCRIPTION OF THE DRAWINGS

[0036] Figure 1 It is a diagram of the original Transformer model;

[0037] Figure 2 It is a diagram of the improved Transformer model based on the feature pyramid in the present invention;

[0038] Figure 3A It is a diagram (I) of the multi-scale encoder-decoder attention calculation method in the present invention;

[0039] Figure 3B It is a diagram (II) of the multi-scale encoder-decoder attention calculation method in the present invention. DETAILED DESCRIPTION OF THE INVENTION

[0040] The present invention will be further described below in conjunction with the accompanying drawings of the specification.

[0041] The present invention proposes a Transformer model based on a feature pyramid, which simultaneously focuses on the hidden layer information of different encoding layers at the encoding end when calculating the encoder-decoder attention. Since the way of stacking the same structural layers has redundant calculations, the hidden layer state of the encoder is scaled in a pyramid shape through the fully connected layer of the encoding layer, and the higher the hidden layer state, the denser the semantic information it contains. At the same time, multi-scale multi-head segmentation is used to avoid the problem of inconsistent dimensions between the query vector and the key vector in the encoder-decoder attention calculation.

[0042] A neural machine translation method based on a feature pyramid according to the present invention includes the following steps:

[0043] 1) Send the preprocessed source language input into the encoder end of the translation model, and encode it into context vectors of different dimensions through the word embedding layer and the stacked encoder layers;

[0044] 2) The hidden layer dimensions of the encoder are scaled in a pyramid shape during the feed-forward process. According to the dimension sizes, the encoder layers are divided into different sub-blocks, and the output vectors of different sub-blocks are saved in the source language encoding vector sequence;

[0045] 3) The source language encoding vector sequence of the encoder is fed into the encoder-decoder attention module at the decoder end of the translation model. During the calculation of the encoder-decoder attention weights, multi-head splitting of different scales is performed on the encoding key vectors and decoding query vectors of different dimensions, so that the encoded sub-key vectors and decoded sub-query vectors have the same dimension;

[0046] 4) The calculation results of the encoder-decoder attention of different dimensions are weighted and averaged to obtain the final encoder-decoder attention output vector. The decoder decodes the source language context vector into the target language translation through stacked decoder layers, and performs gradient update through the cross-entropy loss function to optimize the weights of the translation model.

[0047] In step 1), first, the training data is preprocessed. For example, for the Chinese sentence "The cat likes to eat fish.", after word segmentation, it becomes "The cat likes to eat fish.", and then the source language input is fed into the encoder end of the translation model to encode the source language information into context vectors of different dimensions. As Figure 1 shown, the encoder is composed of stacked encoder layers, and each encoder layer includes a self-attention sub-layer and a fully-connected sub-layer. Among them, the calculation method of the self-attention module is as follows:

[0048]

[0049]

[0050] Among them, Self_Att represents the self-attention sub-layer, is the input hidden layer vector, W q , W k , W v are the parameters of the self-attention sub-layer, softmax(·) is the attention weight calculation function, L s is the length of the source sentence, d is the dimension of the corresponding hidden layer vector, a is the self-attention weight,

[0051] In the original Transformer model, the hidden layer dimensions of different encoder layers are the same (usually 512 dimensions), so the network structures of different layers are the same, resulting in redundant calculations between adjacent layers.

[0052] In step 2), in order to reduce the redundant calculations of the encoder layers and at the same time capture the source language semantic information of different dimensions, the hidden layer dimensions of the encoder are scaled in a pyramid shape during the feed-forward process, and the scaling process occurs in each fully-connected sub-layer:

[0053] h o = W2ReLU(h i W1 + b1) + b2

[0054] where h i is the input vector of the fully - connected sub - layer, i.e., the output vector of the previous self - attention sub - layer, H o is the output vector of the fully - connected layer, W1 ∈ R d×4d , W2 ∈ R 4d×d / 2 , b1 ∈ R 4d , b2 ∈ R d / 2 are the parameters of the fully - connected sub - layer, ReLU is the activation function, and finally the fully - connected sub - layer scales the input from the corresponding hidden - layer vector dimension d to d / 2;

[0055] Meanwhile, different from the original Transformer model which takes the output of the last layer of the encoder as the only source - language encoding vector, the present invention believes that the outputs of different sub - layers of the encoder contain different semantic information and there is a complementary relationship between them. Therefore, the encoder is divided into different sub - blocks according to the dimension size, and the output vectors of different sub - blocks are saved in the source - language encoding vector sequence as the encoding information set of different dimensions of the source language. As Figure 2 shown, when calculating the encoder - decoder attention, the decoder can simultaneously focus on the encoding information of different layers and different dimensions of the encoding layer.

[0056] In step 3), the source - language encoding vector sequence of the encoder is sent to the decoder end of the model. The decoder introduces the source - language context information, i.e., the source - language encoding vector sequence of the encoder, through the multi - head encoder - decoder attention module. Since the calculation of the attention weights requires the decoding query vector and the source - language key vector to have the same dimension, a multi - scale segmentation method is used during the multi - head segmentation. For the decoding query vector with dimension D q and the source - language key vector with dimension D k , the calculation steps of the attention score are as follows:

[0057] 301) Cut the encoded key vector into head_key source - language sub - key vectors with dimension head_dim, where head_key is the number of key heads and head_dim is the dimension of the head, head_key = D k / head_dim;

[0058] 302) Cut the decoding query vector into head_query decoding sub - query vectors with dimension head_dim, where head_query is the number of query heads and head_dim is the dimension of the head, head_query = D q / head_dim, where D q = n × D k, at this time, head_query = n × head_key, where n is the proportionality coefficient of D q and D k 's proportionality coefficient;

[0059] 303) When n > 1, as Figure 3A shown, n decoded sub-query vectors correspond to 1 encoded sub-key vector; when n = 1, similar to the original Transformer model, 1 decoded sub-query vector corresponds to 1 encoded sub-key vector; when n < 1, as Figure 3B shown, 1 decoded sub-query vector corresponds to n encoded sub-key vectors, and the final attention calculation result is the average of the corresponding multiple outputs. The calculation method of encoder-decoder attention is slightly different from that of self-attention, and its calculation method is as follows:

[0060]

[0061] where is the input hidden layer vector, which serves as the decoding query vector here, L t is the length of the target sentence, is the encoded key vector, which serves as the decoding key vector here, L s is the length of the source sentence.

[0062] In step 4), the encoder-decoder attention calculation results of different dimensions are weighted and averaged, and the weights are learnable. The calculation method of the final encoder-decoder attention result is as follows:

[0063]

[0064] where L is the number of encoder layers, corresponding to the outputs of L encoder layers of different dimensions, v i represents the encoder-decoder attention calculation result of the output of the i-th encoder layer, γ i is the learnable weight parameter of the model.

[0065] In the training stage, the model updates the gradient through the cross-entropy loss function to optimize the weights of the model.

[0066] In the decoding stage, the trained model can generate the corresponding target language translation "Cats like to eat fish." in an autoregressive manner based on the source language input "Cats like to eat fish."

[0067] To verify the effectiveness of the method, the present invention applies the neural machine translation model based on the feature pyramid to two tasks of the machine translation standard datasets WMT 14 English-German and IWSLT14 German-English. Among them, the WMT 14 English-German dataset has approximately 4.5 million training data, and the IWSLT14 German-English dataset has approximately 0.15 million training data. It can be found from Table 1 that compared with the original Transformer model, the BLEU value of the model translation of the method proposed by the present invention has been significantly improved, proving that the method proposed by the present invention can effectively improve the translation performance of the model.

[0068]

[0069] Table 1 Comparison of experimental results

[0070] The present invention adopts the feature pyramid method, which reduces the redundant calculation in the encoder-decoder attention calculation, and at the same time adopts the multi-scale segmentation method to solve the problem of inconsistent dimensions in the encoder-decoder attention calculation. The final experimental results prove that the feature pyramid neural machine translation model proposed by the present invention can effectively fuse source language features of different scales, thereby improving the translation quality of the model.

Claims

1. A neural machine translation method based on a feature pyramid, characterized in that Including the following steps: 1) Feed the preprocessed source language input into the encoder end of the translation model, and encode it into context vectors of different dimensions through the word embedding layer and the stacked encoder layers; 2) The hidden layer dimension of the encoder scales in a pyramid shape during the feed-forward process. According to the dimension size, the encoder layers are divided into different sub-blocks, and the output vectors of different sub-blocks are saved in the source language encoding vector sequence; 3) Feed the source language encoding vector sequence of the encoder into the encoder-decoder attention module at the decoder end of the translation model. During the calculation of the encoder-decoder attention weights, perform multi-head segmentation of different scales on the encoding key vectors and decoding query vectors of different dimensions, so that the encoded sub-key vectors and decoding sub-query vectors have the same dimension; 4) Weightedly average the calculation results of the encoder-decoder attention of different dimensions to obtain the final encoder-decoder attention output vector. The decoder decodes the source language context vector into the target language translation through the stacked decoder layers, and updates the gradient through the cross-entropy loss function to optimize the weights of the translation model.

2. The neural machine translation method based on a feature pyramid according to claim 1, wherein: In step 1), preprocess the training data, feed the source language input into the encoder end of the translation model, and encode the source language information into context vectors of different dimensions; the encoder is composed of stacked encoder layers, and each encoder layer includes a self-attention sub-layer and a fully connected sub-layer, where the calculation method of the self-attention sub-layer is as follows: Among them, Self_Att represents the self-attention sub-layer, is the input hidden layer vector, W q , W k , W v are the parameters of the self-attention sub-layer, softmax(·) is the attention weight calculation function, L s is the length of the source sentence, d is the dimension of the corresponding hidden layer vector, and a is the self-attention weight, 3. The neural machine translation method based on a feature pyramid according to claim 1, wherein: In step 2), the hidden layer dimension of the encoder scales in a pyramid shape during the feed-forward process, and the scaling process occurs in each fully connected sub-layer: h o = W2ReLU(h i W1 + b1) + b2 where h i is the input vector of the fully-connected sub-layer, i.e., the output vector of the previous self-attention sub-layer, and H o is the output vector of the fully-connected layer, W1 ∈ R d×4d and W2 ∈ R 4d×d / 2 and b1 ∈ R 4d and b2 ∈ R d / 2 are the parameters of the fully-connected sub-layer, ReLU is the activation function, and finally the fully-connected sub-layer scales the input from the corresponding hidden layer vector dimension d to d / 2; Divide the encoder into different sub-blocks according to the dimension size, and save the output vectors of different sub-blocks in the source language encoding vector sequence as the encoding information set of different dimensions of the source language.

4. The neural machine translation method based on a feature pyramid according to claim 1, wherein: In step 3), the source language encoding vector sequence of the encoder is fed into the decoder end of the model. The decoder introduces the source language context information through the multi-head encoder-decoder attention module, that is, the source language encoding vector sequence of the encoder; a multi-scale segmentation method is used during multi-head segmentation; for the decoding query vector with dimension D q and the source language key vector with dimension D k , the calculation steps of the attention scores are as follows: 301) Cut the encoded key vector into source language sub-key vectors with head_key dimensions of head_dim, where head_key is the number of key heads, head_dim is the dimension of the head, and head_key = D k / head_dim; (302) Cut the decoded query vector into head_query decoded sub-query vectors with a dimension of head_dim, where head_query is the number of query heads, head_dim is the dimension of the head, and head_query = D q / head_dim, where D q = n × D k , at this time head_query = n × head_key, n is the proportionality coefficient of D q and D k ; 303) When n > 1, n decoding sub-query vectors correspond to 1 encoding sub-key vector; when n = 1, 1 decoding sub-query vector corresponds to 1 encoding sub-key vector; when n < 1, 1 decoding sub-query vector corresponds to n encoding sub-key vectors; the calculation method of the encoder-decoder attention is as follows: Among them is the input hidden layer vector, which serves as the decoding query vector here, L t is the length of the target sentence, is the encoding key vector, which serves as the decoding key vector here, L s is the length of the source sentence.

5. The neural machine translation method based on a feature pyramid according to claim 1, wherein: In step 4), weightedly average the calculation results of the encoder-decoder attention of different dimensions, and its weights are learnable. The calculation method of the final encoder-decoder attention result is as follows: where L is the number of encoder layers, corresponding to the outputs of L encoder layers with different dimensions, v i represents the result of the encoding-decoding attention calculation for the output of the i-th encoder layer, γ i is a weight parameter that can be learned by the model; The decoder decodes the source language context vector into the target language translation through the stacked decoder layers, and updates the gradient through the cross-entropy loss function to optimize the weights of the model.

Citation Information

Patent Citations

  • Deep neural machine translation system based on random residual algorithm

    CN111353315A

  • Chinese-Mongolian translation method combining Bert language model and fine-grained compression

    CN112395891A