A Mongolian-Chinese machine translation method based on deep residual shrinkage network and seq2seq

By introducing deep residual shrinking network and seq2seq model in Mongolian and Chinese machine translation, and combining soft thresholding and GAN adversarial attacks, the problems of scarcity and data sparseness of Mongolian and Chinese corpus are solved, achieving high accuracy and efficient Mongolian and Chinese translation.

CN115577720BInactive Publication Date: 2025-05-06INNER MONGOLIA UNIV OF TECH
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202211137746.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-09-19
Publication Date
2025-05-06
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

Due to the scarcity of Mongolian Chinese corpus and sparse data, existing Mongolian Chinese machine translation methods are difficult to achieve high-accurate translation.

Method used

The Mongolian-Chinese machine translation method based on deep residual shrinking network (DRSN) and seq2seq is adopted. By introducing soft thresholding as the shrinking layer, features that are independent of Mongolian text are eliminated, and the GAN-based adversarial attack method is used to improve the robustness of the model.

Benefits of technology

It effectively solves the problem of underfitting due to sparse data, improves the accuracy and efficiency of Mongolian and Chinese translation, and realizes multimodal complementary translation of images and text.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115577720B_ABST
    Figure CN115577720B_ABST
Patent Text Reader

Abstract

A Mongolian-Chinese machine translation method based on a deep residual shrinkage network and seq2seq, establishes a Mongolian dictionary and a Chinese dictionary, uses YOLO-V7 to detect targets on images, obtains a feature map matrix containing text information, uses CBOW to generate word vectors from Mongolian text, and weights the feature map matrix; sends the weighted feature map matrix to a deep residual shrinkage network to remove features in the image that are not related to the Mongolian text; uses an LSTM encoder to encode the word vector, fuses it with the feature output of the deep residual shrinkage network, inputs an LSTM decoder based on an attention mechanism, and uses the attention mechanism to continuously update the state until the output result; uses a GAN-based SN-GAN attack method to attack the trained model until the generated adversarial sample is real and natural. The present invention solves the underfitting problem caused by data sparsity and improves the quality of Mongolian-Chinese translation.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of machine translation, and in particular relates to a Mongolian-Chinese machine translation method based on a deep residual shrinkage network and seq2seq. Background Art

[0002] Since the introduction of machine translation, it has gone through three important stages. The first is the rule-based machine translation method. Driven by rules, this method completes the analysis of the sentence to be translated and generates its corresponding target language translation. However, since this method requires manual writing of translation rules, the cost is extremely high. Later, a corpus-based machine translation method was proposed. This method is based on a large-scale bilingual corpus and requires training. Due to its inherent properties, the accuracy of its translation is not very high. With the development of deep learning, the next step is the deep learning-based machine translation method. This method is currently very popular, and with the development of deep learning, it has made very significant progress.

[0003] In the field of machine translation, the seq2seq model proposed by Sutskever et al. and the seq2seq model based on the attention mechanism proposed by Bahdanau et al. and Luong et al. are both relatively mainstream and have achieved good results. However, for Mongolian-Chinese machine translation, due to the scarcity of corpora, bilingual corpora are even more difficult to find, and the scale is small, there is a serious data sparsity problem, and there are some wrong words in the corpus, which seriously affects the translation accuracy. Summary of the invention

[0004] In order to overcome the shortcomings of the above-mentioned prior art, the purpose of the present invention is to provide a Mongolian-Chinese machine translation method based on Deep Residual Shrinkage Networks (DRSN) and seq2seq. DRSN is used to introduce "soft thresholding" as a "shrinkage layer" into the residual module when facing feature information irrelevant to the current task, and DRSN proposes a method for adaptively setting the threshold. At the same time, a GAN-based adversarial attack is used to continuously attack the training model until the output result is changed, so as to improve the robustness of the model, solve the underfitting problem caused by data sparsity, and improve the quality of Mongolian-Chinese translation.

[0005] In order to achieve the above object, the technical solution adopted by the present invention is:

[0006] A Mongolian-Chinese machine translation method based on deep residual shrinkage network and seq2seq, comprising:

[0007] Step 1: Establish a Mongolian dictionary and a Chinese dictionary, the Mongolian dictionary contains a number of Mongolian characters, and the Chinese dictionary contains a number of Chinese characters; for Mongolian text and images, use YOLO-V7 to perform target detection on the image to obtain a feature map matrix containing text information, and then use the CBOW model to generate word vectors for the Mongolian text, and use the word vectors to weight the feature map matrix; wherein the text is a partial description of the image content;

[0008] Step 2: Send the weighted feature map matrix to the deep residual shrinkage network, adjust the threshold in the deep residual shrinkage network, and use the soft thresholding function to set the threshold of the part that is not related to the Mongolian text to 0, so as to eliminate the features in the image that are not related to the Mongolian text, so that the output features only contain features related to the Mongolian text;

[0009] Step 3: Use the LSTM encoder to encode the word vector, then fuse its encoded output with the feature output of the deep residual contraction network, input the fused features into the LSTM decoder based on the attention mechanism, and use the attention mechanism to continuously update the state until the output result is obtained;

[0010] Step 4: Use the GAN-based SN-GAN attack method to attack the trained model until the generated adversarial samples are real and natural, and the generator can stably generate corresponding adversarial samples for any input, and then obtain the final translation model. Use the final translation model to perform Mongolian-Chinese translation.

[0011] In one embodiment, the YOLO-V7 is mainly composed of a Backbone structure, an SPP-CSP and a Head structure;

[0012] The Backbone structure includes a Stem layer, an ELAN layer and a DS layer. The Stem layer is composed of three stacked Conv convolution layers, each Conv layer uses standard convolution + BN + SiLU, and the Stem layer outputs a feature map with two times downsampling;

[0013] In the ELAN layer, a convolution with a stride of 2 is first used to obtain a four-fold sampling image, and then a series of convolutions are used to process the four-fold downsampled feature map. In a branch operation of the ELAN layer, except for the two 1*1 convolution channels at the beginning, the number of remaining convolution kernels is 1, which can ensure that the input channel and the output channel are consistent.

[0014] The DS layer further performs a downsampling operation on the four-fold downsampled feature map, which includes a left branch and a right branch, wherein the left branch uses maxpooling to achieve spatial downsampling, followed by a 1*1 convolution compression channel; the right branch first uses a 1*1 convolution compression channel, and then uses a 3*3 convolution with a step size of 2 to complete the downsampling, and finally merges the results of the two branches to output a feature map with a channel number equal to the input channel number but a spatial resolution reduced by twice, and then repeats the stacking of the DS layer and the ELAN layer twice;

[0015] The Backbone structure outputs a 32-fold downsampled feature map with 1024 channels, which is first processed by SPP-CSP. After the processing, the number of channels of the 32-fold downsampled feature map is reduced from 1024 to 512, and then enters the Head structure;

[0016] The Head structure includes an ELAN layer and a DS layer. The ELAN layer and the DS layer in the Head structure are the same as those in the Backbone, but the parameters are different. Finally, the ELAN layer in the Head structure outputs a feature map matrix containing all the information of the image.

[0017] In one embodiment, the CBOW model includes an input layer, a mapping layer, and an output layer, and the output layer is a Huffman tree. The training process of the CBOW model is as follows:

[0018] First, use the Mongolian dictionary created in step 1 to encode the Mongolian text into a one-hot vector. The one-hot vector of the context word is used as input. The one-hot vector of each word in the input layer is multiplied by the weight matrix W. VN Get the corresponding vector, where the weight matrix W VN The initial value is the unit matrix, and its parameters are obtained by training the CBOW network; V is the number of words in the Mongolian dictionary, that is, the dimension of the one-hot vector is V, and N is the number of neurons in the hidden layer;

[0019] Then the obtained vectors are added and averaged as the input of the output layer. In the output layer, the averaged vector is multiplied by the weight matrix in the output layer to obtain the output vector. The softmax function is applied to the output vector to obtain the probability distribution of each word.

[0020] Finally, the loss metric function is used to calculate the loss value between the probability distribution and the expected output one-hot vector. The loss value is used for back propagation to update the network parameters. The matrix W is obtained after multiple iterations. VN That is the final word vector matrix.

[0021] In one embodiment, Mongolian text word vectors are used to weight the feature graph. When the dimensions of two vectors are different, a zero-filling operation is performed. The calculation formula is as follows:

[0022]

[0023] Where D is the weighted feature map vector, and im is the feature map vector output by YOLO-V7.

[0024] In one embodiment, in step 2, the deep residual shrinkage network includes several improved residual modules; one improved residual module includes two batch normalizations, two rectified linear unit activation functions, two convolutional layers and an identity mapping.

[0025] In one embodiment, the threshold in the deep residual shrinkage network is a vector, and each channel of the weighted feature map vector corresponds to a threshold; the residual module is used to calculate the absolute value of all features in the input feature map, and then a feature is obtained after global mean pooling and averaging. At the same time, the feature map after global mean pooling is input into a fully connected network, which uses the Sigmoid function as the last layer, normalizes the output to between 0 and 1, and obtains a coefficient a. The feature obtained after global mean pooling and averaging is multiplied by a, that is, the final threshold.

[0026] In one embodiment, the step 2 adjusts the convolution layer in the deep residual shrinkage network to replace it with a dilated convolution to obtain a larger receptive field without affecting the pixels.

[0027] In one embodiment, in step 3, the LSTM encoder first represents a Mongolian text sequence as a fixed-length context vector c, which is calculated as follows:

[0028] c=f(h1,h2,…,h N )

[0029] where h1,h2,…,h N is the hidden state of the neural unit, N is the length of the input Mongolian text;

[0030] The last state h output by the LSTM decoder N It is fused with the feature map vector img output by DRSN that only contains text features to obtain the fusion result h′ N , the formula is as follows:

[0031]

[0032] Then update h N state, let h N =h′ N , that is, using h′N The value of h N ;

[0033] The attention mechanism is when the encoder outputs the last state h N After that, let the first input s0 of the decoder, i.e. the first state of the encoder, be equal to the last state h of the encoder. N , the attention mechanism and the decoder start working at the same time, calculating s0 and h i The correlation formula is as follows:

[0034] α1=align(h i ,s0)

[0035] Where α1 is h i The weight between α1 and s0, α1 has N values, align is the correlation calculation function, the formula is as follows:

[0036]

[0037] The Softmax function formula is as follows:

[0038]

[0039] At this time, α1=[α′1,…,α′ N ]

[0040] where w k With w Q is the parameter matrix, learned from the training data, h i is the i-th state in the LSTM encoder, is the matrix (w k ×h i ) and the matrix (w Q ×s0), α′ i is the weight α i The i-th value in ;

[0041] The weighted average is calculated for the obtained weights, and the calculation formula is as follows:

[0042] c0=α′ i h1+…+α′ N h N

[0043] Where c0 is the weighted average of the weights of the first state s0 of the LSTM decoder. The new state vector s1 is calculated using the hyperbolic tangent function. The calculation formula is as follows:

[0044]

[0045] Where x′1 is the first input of the decoder, that is, according to hN The first predicted value, A′ and b, are learned from the model.

[0046] α2=align(h i ,s1)

[0047] Calculate the N weights (α′1,...,α′) in α2 N ), and then use the formula

[0048] c1=α′ i h1+…+β′ N h N

[0049] After calculating c1, the decoder receives the second input x′2 and then calculates s2. The above steps are repeated until the decoder receives the termination symbol and stops.

[0050] Compared with the prior art, the beneficial effects of the present invention are as follows: the multimodal complementary translation of images and texts can effectively solve the problem of low translation accuracy when the semantic expression of texts is incomplete. The DRSN network removes "noise" signals, which can avoid interference from irrelevant signals and improve translation efficiency. Due to the scarcity of Mongolian corpus, the generator in GAN can generate a large number of pseudo data sets to achieve the purpose of expanding the corpus and avoid underfitting problems caused by data sparsity. BRIEF DESCRIPTION OF THE DRAWINGS

[0051] Figure 1 It is a flow chart of the present invention.

[0052] Figure 2 This is the Stem layer structure diagram in Backbone.

[0053] Figure 3 This is the ELAN layer structure diagram in Backbone.

[0054] Figure 4 This is the DS layer structure diagram in Backbone.

[0055] Figure 5 This is the ELAN layer structure diagram in the Head.

[0056] Figure 6 This is the DownSample structure diagram in Head.

[0057] Figure 7 It is the improved residual module.

[0058] Figure 8 is the soft thresholding and its derivative.

[0059] Fig. 9 This is the attack principle based on GAN.

[0060] Fig.10 For input sample graph DETAILED DESCRIPTION

[0061] The embodiments of the present invention are described in detail below with reference to the accompanying drawings and examples.

[0062] refer to Figure 1 The present invention provides a Mongolian-Chinese machine translation method based on DRSN and seq2seq, adopts YOLO-V7 for target detection, and uses the word vector generated by the CBOW model to weight the feature map matrix output by YOLO-V7, then uses the DRSN network to remove the "noise" signal in the feature map matrix, and fuses the state of the LSTM decoder with the output of the DRSN, and then uses GAN for model training.

[0063] Specifically, the present invention comprises the following steps:

[0064] Step 1: First, create two dictionaries, a Mongolian dictionary and a Chinese dictionary. The Mongolian dictionary contains several Mongolian characters, and the Chinese dictionary contains several Chinese characters. Use YOLO-V7 to detect the image and obtain a feature map matrix containing text information. Then use the CBOW model to generate word vectors for Mongolian text, and use the Mongolian text word vectors to weight the feature map matrix generated by YOLO-V7.

[0065] Wherein, the Mongolian dictionary comprises 5 vowels, 24 consonants, 29 characters in total, and a start symbol and a terminator are added in both the Mongolian dictionary and the Chinese dictionary. The purpose of adding the start symbol and the terminator in the Mongolian dictionary is to tell the computer when to start and when to end. In the present invention, the Mongolian text is a partial description of the image content.

[0066] In this step, YOLO-V7 is used to detect the target of the input image and obtain a feature map matrix containing all the information of the image.

[0067] For example, YOLO-V7 is mainly composed of Backbone structure, SPP-CSP and Head structure, and the Backbone structure is mainly composed of Stem layer, ELAN layer and DS layer.

[0068] The first is the initial Stem layer, such as Figure 2 As shown in the figure, it consists of three layers of stacked Conv convolutions, each of which uses standard convolution + BN + SiLU. The Stem layer outputs a feature map with two times downsampling.

[0069] Next is the ELAN layer, such as Figure 3As shown in the figure, in the ELAN layer, a convolution with a stride of 2 is first used to obtain a four-fold sampling image, and then a series of convolutions are used to process the four-fold downsampled feature map. In a branch operation of the ELAN layer, except for the two 1*1 convolution channels at the beginning, the number of remaining convolution kernels is 1, which can ensure that the input channel and the output channel are consistent.

[0070] Finally, there is the DS layer, such as Figure 4 As shown in the figure, the DS layer is used to downsample the four-fold downsampled feature map, which includes the left branch and the right branch. The left branch uses maxpooling (MP) to achieve spatial downsampling, followed by a 1*1 convolution compression channel: the right first uses a 1*1 convolution compression channel, and then uses a 3*3 convolution with a step size of 2 to complete the downsampling. Finally, the results of the two branches are merged to output a feature map with the same number of channels as the input channels, but the spatial resolution is reduced by twice, and then the DS layer and the ELAN layer are stacked twice. Finally, Backbone outputs a 32-fold downsampled feature map with a channel number of 1024.

[0071] The 32-fold downsampled feature map is processed by SPP-CSP. After processing, the number of channels is reduced from 1024 to 512, and then enters the Head structure.

[0072] The Head structure includes the ELAN layer and the DS layer, such as Figure 5 and Figure 6 As shown in the figure, the ELAN layer and DS layer in the Head structure have the same structure as those in the Backbone, but the parameters (convolution step) are different. For example, in the Backbone structure, the step size of the first 3*3 convolution of the ELAN layer is 1, while in the Head structure, the step size of the first 3*3 convolution of the ELAN layer is 2. Finally, the ELAN layer in the Head structure will output a feature map matrix containing all the information of the image.

[0073] For example, SPP-CSP consists of two parts: SPP and CSP. SPP is spatial pyramid pooling, which can increase the receptive field and adapt the algorithm to images of different resolutions (i.e. large objects, small objects); dynamically changing the feature map can adjust the size of the feature map to be consistent, which is convenient for stacking together later. CSP divides the feature into two parts, only one of which is processed conventionally, and finally the two parts are merged together, reducing the amount of calculation by half. This method greatly improves the speed, and the accuracy is slightly improved instead of reduced.

[0074] Word2Vec is a shallow neural network, which mainly uses a neural network to map sparse word vectors to dense word vectors. The obtained word vectors contain context information and semantic information. Word2Vec contains two models, Skip-gram and CBOW. The present invention adopts the CBOW model, which can calculate the probability distribution of the central word according to the context vector.

[0075] The CBOW model has no hidden layers, and only includes an input layer, a mapping layer, and an output layer. The output layer is a Huffman tree. The word to be predicted is found along the Huffman tree, and the value of the multiplication of the probabilities on the path from the root node to the leaf node of the predicted word is maximized. The Huffman tree is built based on the words and their frequencies in the Mongolian corpus. Assuming that the context words generated by the Mongolian central word are independent of each other, the training process of the CBOW model is as follows:

[0076] First, use the Mongolian dictionary created in step 1 to encode the Mongolian text into a one-hot vector. The one-hot vector of the context word is used as input. The one-hot vector of each word in the input layer is multiplied by the weight matrix W. VN Get the corresponding vector, W VN Initially, it is a unit matrix, and its parameters are obtained through training of the CBOW network; V is the number of words in the Mongolian dictionary, that is, the dimension of the one-hot vector is V; N is the number of hidden layer neurons, that is, the dimension of the word vector we hope to get in the end is N.

[0077] Then the obtained vectors are added and averaged as the input of the output layer. In the output layer, the averaged vector is multiplied by the weight matrix in the output layer to obtain the output vector. The softmax function is applied to the output vector to obtain the probability distribution of each word. Here, the softmax function formula is as follows:

[0078]

[0079] Among them, H j is the jth value in the one-hot vector of the context word; H i is the jth value in the one-hot vector of the context word.

[0080] Finally, the loss metric function is used to calculate the loss value between the output probability distribution and the expected output one-hot vector. This loss value is used for back propagation to update the network parameters. The matrix W is obtained after multiple iterations. VN That is the final word vector matrix.

[0081] At this point, the feature map matrix output by YOLO-V7 has been obtained. The Mongolian text word vector is used to weight the feature map matrix. When the two vector dimensions are different, a 0-filling operation is performed. The calculation formula is as follows:

[0082]

[0083] Where D is the weighted feature map matrix, and im is the feature map matrix output by YOLO-V7.

[0084] Step 2: Send the weighted feature map matrix to DRSN, adjust the threshold in DRSN, and use the soft thresholding function to set the threshold of the part that is not related to the Mongolian text to 0 to eliminate the features in the image that are not related to the text, so that the output feature map matrix only contains features related to the Mongolian text.

[0085] The feature map matrix contains some features that are irrelevant to Mongolian text, which will cause great interference to Mongolian translation. After the weighted feature map matrix is ​​sent to DRSN, the "soft thresholding" unique to the improved residual module of DRSN will adjust the threshold according to the weight of the feature map matrix, and set the threshold of the low-weight part to 0, so as to eliminate the part of the feature map matrix that is irrelevant to the Mongolian text content. In addition, the convolution layer in the deep residual shrinkage network can be adjusted to replace it with a hollow convolution to obtain a larger receptive field without affecting the pixels.

[0086] A deep residual network consists of a convolutional layer (Conv), several residual blocks (RBUs), a batch normalization (BN), a ReLU activation function, a global average pooling layer (GAP), and a fully connected output layer (FC) from front to back; the main part of the deep residual network is composed of many residual modules. When the deep residual network is trained based on back-propagation, its loss can not only be back-propagated layer by layer through the convolutional layer, but also can be more conveniently back-propagated through the identity mapping, so that it is easier to train and obtain a better model.

[0087] Specifically, in this step, an improved residual module is used, thereby obtaining a deep residual shrinkage network. The improved residual module is a basic component of DRSN, such as Figure 7 As shown in the figure, a residual module contains two batch normalizations, two rectified linear unit activation functions, two convolutional layers, and identity mapping; identity mapping is the core contribution of deep residual networks, which greatly reduces the difficulty of deep neural network training. When a feature map with a channel number of C, a width of W, and a height of 1 is input, the number of channels and width of the output image can be controlled by adjusting the number of convolution kernels and the moving step size. The small subnetwork on the right is used to obtain the threshold.

[0088] In the deep residual shrinkage network of the present invention, its threshold is not a value, but a vector, that is, each channel of the weighted feature map vector corresponds to a threshold; DRSN proposes a method for obtaining an adaptive threshold, and uses a soft thresholding function to automatically adjust the threshold according to the training situation. The method is as follows: the feature map matrix output by YOLO-V7 is used as the input of DRSN, and the residual module is used to calculate the absolute value of all features in the input feature map, and then a feature is obtained after global mean pooling and averaging. At the same time, the feature map after global mean pooling is input into a fully connected network. The fully connected network uses the Sigmoid function as the last layer, normalizes the output to between 0 and 1, and obtains a coefficient a. Multiply a by the feature obtained after global mean pooling and averaging, that is, the final threshold. For example, the Sigmoid function formula here is as follows:

[0089]

[0090] Where x is the eigenvalue in the feature map matrix.

[0091] The soft thresholding function and its derivative of the present invention are as follows: Figure 8 As shown:

[0092] The soft threshold function is a function that shrinks the input data toward zero. It is often used in signal noise reduction algorithms. Its formula is as follows:

[0093]

[0094] x represents the input feature, that is, the eigenvalue in the feature map matrix, y represents the output feature, that is, the eigenvalue after being changed by the soft thresholding function, and τ represents the threshold. Here, the threshold needs to be a positive number and cannot be too large. If the threshold is larger than the absolute value of all input features, then the output feature y can only be zero, so soft thresholding is meaningless. At the same time, the derivative formula of the soft thresholding function is as follows:

[0095]

[0096] This property is the same as the ReLU activation function, so the soft thresholding function is also helpful in preventing "gradient disappearance" and "gradient explosion".

[0097] Step 3: Use the LSTM encoder to encode the Mongolian text, and then fuse its encoded output with the feature output of the deep residual shrinkage network. Since the feature map matrix output by DRSN has deleted the features irrelevant to the Mongolian text, the two features are extremely similar at this time. Fully fusing these two features can maximize the accuracy of the model. Input the fused feature vector into the LSTM decoder based on the attention mechanism, and use the attention mechanism to continuously update the state until the terminator is received and the result is output.

[0098] Specifically, in the LSTM encoder, a Mongolian text sequence is first represented as a fixed-length context vector c, which is calculated as follows:

[0099] c=f(h1,h2,…,h N )

[0100] where h1,h2,…,h N is the hidden state of the neural unit, N is the length of the input Mongolian text;

[0101] At this point, the LSTM decoder will output the last state h N At the same time, DRSN also outputs a feature map vector containing only text features, denoted as img. The two vectors h N Fuse with img to get the fusion result h′ N , update h N The state calculation formula is as follows:

[0102]

[0103] Where img is the output vector of DRSN, then h′ N The state is the result of the fusion of the feature vector output by DRSN and the last state output by LSTM decoder, and then update h N state, let h N =h′ N , that is, using h′ N The value of h N .

[0104] In this step, the attention mechanism is when the encoder outputs the last state h N After that, let the first input s0 of the decoder, that is, the first state of the encoder, be equal to the last state h of the encoder N , the attention mechanism and the decoder start working at the same time, calculating the states s0 and h i The correlation formula is as follows:

[0105] α1=align(h i ,s0)

[0106] Where α1 is h i The weight between s0 and s1, because there are N states in the LSTM encoder, the weight α1 has N values, h i ∈[h1,h N ]. align is the correlation calculation function, the formula is as follows:

[0107]

[0108] The Softmax function formula in this step is as follows:

[0109]

[0110] At this time, α1=[α′1,…,α′ N ]

[0111] where w k With w Q is the parameter matrix, learned from the training data, h i is the i-th state in the LSTM encoder, s0 is the first state of the LSTM decoder, is the matrix (w k ×h i ) and the matrix (w Q ×s0), α′ i is the weight α i The i-th value in .

[0112] Then calculate the weighted average of the weights obtained above, and the calculation formula is as follows:

[0113] c0=α′ i h1+…+α′ N h N

[0114] Where c0 is the weighted average of the weights of the first state s0 of the LSTM decoder. At this point, the new state vector s1 can be calculated using the hyperbolic tangent function, and the calculation formula is as follows:

[0115]

[0116] Where x′1 is the first input of the decoder, that is, the LSTM decoder is based on h N The first predicted value (i.e., the word vector of the first Chinese character), A′ and b are learned from the model, and then let

[0117] α2=align(h i ,s1)

[0118] Calculate the N weights (α′1,...,α′) in α2N ), and then use the formula

[0119] c1=α′ i h1+…+α′ N h N

[0120] Calculate c1, merge the first character calculated above with the word vector of the next character predicted by the LSTM decoder based on the new state s1 and c1 as the second input x′2 received by the decoder. The decoder receives the second input x′2 and then calculates s2. Repeat the above steps until the decoder receives the terminator and stops.

[0121] Step 4: Use the GAN-based SN-GAN attack method to attack the trained model until the generated adversarial samples are real and natural, and the generator can stably generate corresponding adversarial samples for any input, and then obtain the final translation model. Use the final translation model to perform Mongolian-Chinese translation.

[0122] The GAN-based attack method consists of three parts: the generator, the judge, and the model. This attack method does not require the knowledge of the specific information of the entire model, but can generate relatively similar adversarial samples. Using these samples with small differences to attack the trained model can greatly improve the robustness of the model. When the adversarial samples generated by GAN are very similar to the original data samples and are real and natural, it means that the model has already accurately recognized the data, and even if the input data contains a small amount of "noise", the model can still recognize it well.

[0123] Specifically, the attack principle based on GAN is as follows: Fig. 9 As shown:

[0124] The generator is a generative network that tries its best to generate fake data that is close to reality from random noise. The entire network structure has no pooling layer. It inputs an n-dimensional noise. The neural network obtains the feature information of the input image step by step based on the input vector information, such as lines, style, etc., and then continuously optimizes the details of the image based on the depth of the network.

[0125] The discriminator is a discriminator network. The stability of the discriminator greatly affects the performance of GAN. The basic structure includes convolutional layer, fully connected layer and densely connected layer. The discriminator will try its best to distinguish real pictures from fake pictures. Adding spectral normalization to the discriminator and applying Lipschitz constraints to the convolution kernel weights can make parameter changes more stable to the greatest extent and prevent gradient explosion.

[0126] The training process of GAN is as follows:

[0127] 1. First, send an image into the discriminator, mark the sample as true, and train the discriminator.

[0128] 2. The generator generates a fake image and sends it into the discriminator, marks the sample as false, and trains the discriminator.

[0129] 3. The generator generates a fake image and sends it into the discriminator, and trains the generator according to the judgment result.

[0130] 4. Continuously repeat the above steps to complete the training of the generator and the discriminator.

[0131] The process of the present invention can also be described as follows:

[0132] 1). Use YOLO-V7 to perform object detection on the image and generate a feature map matrix;

[0133] 2). Use the CBOW algorithm to generate Mongolian text word vectors;

[0134] 3). Use the Mongolian text word vectors to weight the feature map matrix generated by YOLO-V7;

[0135] 4). Use DRSN to remove irrelevant features from the weighted image and perform feature extraction;

[0136] 5). Use the LSTM encoder to encode the Mongolian text;

[0137] 6). Use the fusion layer to perform feature fusion on the encoded Mongolian text and the features extracted by DRSN, denoted as (h, s);

[0138] 7). Then the encoder and the attention mechanism start to work, calculate the correlation align(hi, s) between each s and h, and the result is denoted as αi. αi is a real number between 0 and 1, and the sum of all αs is 1;

[0139] 8). Update each state s in turn according to α until the decoder receives the termination symbol;

[0140] 9). Use the GAN-based attack method to attack the established model to improve the robustness.

[0141] In a specific embodiment of the present invention, assume that the established Chinese dictionary is {"@", "sheep", "is", "eating", "grass", "#"}, and the Mongolian dictionary is (The Chinese translation is "The sheep is eating grass."). At this time, the length of the Mongolian text is 4.

[0142] First, input the image as Fig.10 shown, and input the text (The Chinese translation is "The sheep is eating grass.") After the image is input into YOLO-V7, YOLO-V7 will perform object detection on the image. At this time, the output feature matrix includes not only the sheep and the grass, but also the irrelevant sky, the tree behind, etc. Suppose the obtained feature map matrix im is:

[0143]

[0144] Then the Mongolian text (The Chinese translation is "The sheep is eating grass.") is used to generate a one-hot vector, and the generated vector is:

[0145]

[0146] Among them, taking "is" as the central word, then "sheep", "eat", and "grass" are context words.

[0147] At this time, the CBOW model is used to encode the one-hot vector into a word vector. Suppose the weight matrix W VN obtained by model training is:

[0148]

[0149] The weight matrix W VN is multiplied by the one-hot vectors of the context words respectively, and the obtained word vectors are as follows:

[0150]

[0151] However, the obtained vectors are weighted and averaged to obtain

[0152]

[0153] Suppose the weight matrix U of the output layer is:

[0154]

[0155] The weighted and averaged vector is multiplied by the weight matrix to obtain:

[0156]

[0157] The softmax function is applied to the above vector to obtain the probability distribution of each word, as follows:

[0158]

[0159] Then the CBOW model performs backpropagation to obtain the final word vector W VN . Suppose the final W VN is as follows:

[0160]

[0161] The word vector W VN The weighted average is taken with the feature map vector im output by YOLO-V7. Since im is a 3*3 matrix, a 0-filling operation is performed to obtain:

[0162]

[0163] The matrix obtained above is used as the input of DRSN. DRSN will automatically adjust the threshold according to the training situation, so as to obtain a feature map matrix containing only the sheep and grass in the text information, that is, DRSN will remove the features of the sky, trees, etc. in the image. Assume that the feature map matrix img obtained is:

[0164]

[0165] The word vector obtained from the CBOW model is used as the input of the LSTM encoder, and the N states generated by the LSTM encoder are equal to 4 because the Mongolian text length is 4, that is, the LSTM encoder will generate 4 states, but will output the last state h4. Assume that h1 is:

[0166]

[0167] h2 is:

[0168]

[0169] h3 is:

[0170]

[0171] h4 is:

[0172] ,,,

[0173] Take the weighted average of the feature map matrix of h4 and DRSN output, and fill in 0 when the dimensions are different, and we get:

[0174]

[0175] Update the status of h4 to:

[0176]

[0177] Then the LSTM decoder receives the start symbol “@”, and the LSTM decoder and attention mechanism start working to calculate h i The weight α between s0 and s1, let the first state s0 of the decoder equal to h4, and set w k and w Q are all characteristic matrices, we can get Both are 5. For Using the Softmax function, we get:

[0178]

[0179] Similarly, we can get

[0180] Then calculate the required weight α′ i The weighted average c0 of, substitute α′ i Substitute it in, and we can get

[0181]

[0182] At this time, the LSTM decoder predicts that the first character is "sheep" according to s0 and c0. Use the word vector of the character "sheep" as the input x′1 of the decoder. Concatenate x′1 with s0 and c0, and use the hyperbolic tangent function for the concatenated matrix, as follows:

[0183]

[0184] Among them, both A′ and b are parameters to be learned in the model. Assume that the result of compressing the vector s1 into a 4*4 matrix is:

[0185]

[0186] Then calculate h according to the above method i The weight between and s1, and obtain the new weighted average c1 of the weight. At this time, the LSTM decoder predicts that the second character is "in" according to s1 and c1. Combine the first character "sheep" and "in", and use the word vector of "sheep in" as the second input x′2 of the LSTM decoder. Then concatenate x′2, s1, and c1 and use the hyperbolic tangent function to get s2. Similarly, obtain s3 and s4 until the LSTM decoder receives the termination symbol "#". At this time, it stops and outputs the last state s4.

[0187] The result of s4 is a 16-dimensional vector. Divide this vector into a 4*4 matrix every 4 numerical values, which contains the probability of each word. For example, the value of s4 is as follows:

[0188]

[0189] It can be seen that the first word with the highest probability is the first in the Chinese dictionary, the second word with the highest probability is the second, and so on. The final result is "The sheep is eating grass".

Claims

1. A Mongolian-Chinese machine translation method based on deep residual contraction network and seq2seq, characterized in that: include: Step 1: Establish a Mongolian dictionary and a Chinese dictionary, the Mongolian dictionary contains a number of Mongolian characters, and the Chinese dictionary contains a number of Chinese characters; for Mongolian text and images, use YOLO-V7 to perform target detection on the image to obtain a feature map matrix containing text information, and then use the CBOW model to generate word vectors for the Mongolian text, and use the word vectors to weight the feature map matrix; wherein the text is a partial description of the image content; Step 2: Send the weighted feature map matrix to the deep residual shrinkage network, adjust the threshold in the deep residual shrinkage network, and use the soft thresholding function to set the threshold of the part that is not related to the Mongolian text to 0, so as to eliminate the features in the image that are not related to the Mongolian text, so that the output features only contain features related to the Mongolian text; Step 3: Use the LSTM encoder to encode the word vector, then fuse its encoded output with the feature output of the deep residual contraction network, input the fused features into the LSTM decoder based on the attention mechanism, and use the attention mechanism to continuously update the state until the output result is obtained; Step 4: Use the GAN-based SN-GAN attack method to attack the trained model until the generated adversarial samples are real and natural, and the generator can stably generate corresponding adversarial samples for any input, and then obtain the final translation model. Use the final translation model to perform Mongolian-Chinese translation.

2. The Mongolian-Chinese machine translation method based on deep residual shrinkage network and seq2seq according to claim 1, characterized in that: The YOLO-V7 is mainly composed of Backbone structure, SPP-CSP and Head structure; The Backbone structure includes a Stem layer, an ELAN layer and a DS layer. The Stem layer is composed of three stacked Conv convolution layers, each Conv layer uses standard convolution + BN + SiLU, and the Stem layer outputs a feature map with two times downsampling; In the ELAN layer, a convolution with a stride of 2 is first used to obtain a four-fold sampling image, and then a series of convolutions are used to process the four-fold downsampled feature map. In a branch operation of the ELAN layer, except for the two 1*1 convolution channels at the beginning, the number of remaining convolution kernels is 1, which can ensure that the input channel and the output channel are consistent; The DS layer further performs a downsampling operation on the four-fold downsampled feature map, which includes a left branch and a right branch, wherein the left branch uses maxpooling to achieve spatial downsampling, followed by a 1*1 convolution compression channel; the right branch first uses a 1*1 convolution compression channel, and then uses a 3*3 convolution with a step size of 2 to complete the downsampling, and finally merges the results of the two branches to output a feature map with a channel number equal to the input channel number but a spatial resolution reduced by twice, and then repeats the stacking of the DS layer and the ELAN layer twice; The Backbone structure outputs a 32-fold downsampled feature map with 1024 channels, which is first processed by SPP-CSP. After the processing, the number of channels of the 32-fold downsampled feature map is reduced from 1024 to 512, and then enters the Head structure; The Head structure includes an ELAN layer and a DS layer. The ELAN layer and the DS layer in the Head structure are the same as those in the Backbone, but the parameters are different. Finally, the ELAN layer in the Head structure outputs a feature map matrix containing all the information of the image.

3. The Mongolian-Chinese machine translation method based on deep residual shrinkage network and seq2seq according to claim 1, characterized in that: The CBOW model includes an input layer, a mapping layer, and an output layer, and the output layer is a Huffman tree. The training process of the CBOW model is as follows: First, use the Mongolian dictionary created in step 1 to encode the Mongolian text into a one-hot vector. The one-hot vector of the context word is used as input. The one-hot vector of each word in the input layer is multiplied by the weight matrix W. VN Get the corresponding vector, where the weight matrix W VN The initial value is the unit matrix, and its parameters are obtained by training the CBOW network; V is the number of words in the Mongolian dictionary, that is, the dimension of the one-hot vector is V, and N is the number of neurons in the hidden layer; Then the obtained vectors are added and averaged as the input of the output layer. In the output layer, the averaged vector is multiplied by the weight matrix in the output layer to obtain the output vector. The softmax function is applied to the output vector to obtain the probability distribution of each word. Finally, the loss metric function is used to calculate the loss value between the probability distribution and the expected output one-hot vector. The loss value is used for back propagation to update the network parameters. The matrix W is obtained after multiple iterations. VN That is the final word vector matrix.

4. The Mongolian-Chinese machine translation method based on deep residual shrinkage network and seq2seq according to claim 3 is characterized in that: The Mongolian text word vector is used to weight the feature map. When the dimensions of the two vectors are different, a zero-filling operation is performed. The calculation formula is as follows: Where D is the weighted feature map vector, and im is the feature map matrix output by YOLO-V7.

5. The Mongolian-Chinese machine translation method based on deep residual shrinkage network and seq2seq according to claim 1, characterized in that: In step 2, the deep residual shrinkage network includes several improved residual modules; one improved residual module includes two batch normalizations, two rectified linear unit activation functions, two convolutional layers and an identity mapping.

6. The Mongolian-Chinese machine translation method based on deep residual shrinkage network and seq2seq according to claim 1 or 5, characterized in that: The threshold in the deep residual shrinkage network is a vector, and each channel of the weighted feature map vector corresponds to a threshold; the residual module is used to calculate the absolute value of all features in the input feature map, and then a feature is obtained after global mean pooling and averaging. At the same time, the feature map after global mean pooling is input into a fully connected network, which uses the Sigmoid function as the last layer, normalizes the output to between 0 and 1, and obtains a coefficient a. The feature obtained after global mean pooling and averaging is multiplied by a, that is, the final threshold.

7. The Mongolian-Chinese machine translation method based on deep residual shrinkage network and seq2seq according to claim 6, characterized in that: The Sigmoid function formula is as follows: Where x is the eigenvalue in the feature map matrix; The formula of the soft thresholding function is as follows: x represents the input feature, y represents the output feature, and τ represents the threshold, which is a positive number.

8. The Mongolian-Chinese machine translation method based on deep residual shrinkage network and seq2seq according to claim 1, characterized in that: In step 2, the convolution layer in the deep residual shrinkage network is adjusted to be replaced by a dilated convolution to obtain a larger receptive field without affecting the pixels.

9. The Mongolian-Chinese machine translation method based on deep residual shrinkage network and seq2seq according to claim 1, characterized in that: In step 3, the LSTM encoder first represents a Mongolian text sequence as a fixed-length context vector c, which is calculated as follows: c=f(h1,h2,…,h N ) where h1,h2,…,h N is the hidden state of the neural unit, N is the length of the input Mongolian text; The last state h output by the LSTM decoder N It is fused with the feature map vector img output by the deep residual shrinkage network that only contains text features to obtain the fusion result h′ N , the formula is as follows: Then update h N state, let h N =h′ N , that is, using h′ N The value of h N ; The attention mechanism is when the encoder outputs the last state h N After that, let the first input s0 of the decoder, i.e. the first state of the encoder, be equal to the last state h of the encoder. N , the attention mechanism and the decoder start working at the same time, calculating s0 and h i The correlation formula is as follows: α1=align(h i ,s0) Where α1 is h i The weight between α1 and s0, α1 has N values, align is the correlation calculation function, the formula is as follows: The Softmax function formula is as follows: At this time α1=[α1′,…,α′ N ] where w k With w Q is the parameter matrix, learned from the training data, h i is the i-th state in the LSTM encoder, is the matrix (w k ×h i ) and the matrix (w Q ×s0), α i ′ is the weight α i The i-th value in ; The weighted average is calculated for the obtained weights, and the calculation formula is as follows: c0=a i ′h1+…+α′ N h N Where c0 is the weighted average of the weights of the first state s0 of the LSTM decoder. The new state vector s1 is calculated using the hyperbolic tangent function. The calculation formula is as follows: Where x1′ is the first input of the decoder, that is, according to h N The first predicted value, A′ and b, are learned from the model. α2=align(h i ,s1) Calculate the N weights (α1′,...,α′) in α2 N ), and then use the formula c1=a i ′h1+…+α′ N h N After calculating c1, the decoder receives the second input x2′, and then calculates s2. The above steps are repeated until the decoder receives the termination symbol and stops.

Citation Information

Patent Citations

  • Variation text generation method and device, translation model training method and device and text classification method and device

    CN113468856A