A Chinese-Vietnamese Neural Machine Translation Method Integrating Dual Representations of BERT and Word Embeddings
By fusing the dual representation of BERT and word embedding in Hanyue neural machine translation and using attention mechanism for adaptive dynamic fusion, the problem of unsatisfactory performance of Hanyue neural machine translation is solved, and more efficient language information integration and performance improvement is achieved.
Patent Information
- Application Number
- CN202111042653.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-09-07
- Publication Date
- 2025-05-30
- Estimated Expiration
- 2041-09-07
AI Technical Summary
Due to the small size of bilingual parallel corpus, the neural machine translation effect is not ideal. The existing methods rely on parameter initialization of pre-trained machine translation models, and the feature fusion method is simple.
Using the method of fusion BERT and word embedding dual representation, BERT pre-trained language model representation and word embedding representation are performed on the source language sequence, and then adaptive dynamic fusion of dual representation is achieved through the attention mechanism to enhance the representation learning ability of the source language.
The language information in the BERT pre-trained language model is effectively integrated into the neural machine translation model, which improves the performance of the Hanyue neural machine translation model, solves the problem of poor performance of low-resource neural machine translation, and reduces the complexity of the model.
Smart Images

Figure CN113901843B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to a Chinese-Vietnamese neural machine translation method that fuses the dual representations of BERT and word embeddings, and belongs to the technical field of natural language processing. Background Art
[0002] The demand for Chinese-Vietnamese machine translation is increasing continuously. Neural machine translation is the current mainstream machine translation method. However, in low-resource machine translation tasks such as Chinese-Vietnamese, due to the small scale of bilingual parallel corpora, the effect of neural machine translation is not ideal. Considering that monolingual corpora are rich, a large amount of monolingual corpora can be used for self-supervised learning to obtain a pre-trained language model containing rich language information. Incorporating this pre-trained language model into the neural machine translation system is of great significance for low-resource machine translation. Therefore, a Chinese-Vietnamese neural machine translation method that fuses the dual representations of BERT and word embeddings is proposed.
[0003] Currently, the BERT pre-trained language model has achieved excellent results in NLP tasks such as syntactic analysis and text classification, proving that the language model contains rich language information. These language information are contained in the representation vectors obtained after encoding and cannot be directly observed. Therefore, Jinhua Zhu et al. proposed the BERT-fused algorithm to randomly incorporate the hidden states output by the BERT pre-trained language model into the encoder and decoder structures of the Transformer model. The hidden state vectors output by the BERT pre-trained language model and the hidden state vectors output by the word embedding layer are fused by random probability weighting, so as to generate hidden states containing the language information in the pre-trained language model and the language information in the word embedding layer, and realize the use of the language information contained in the BERT pre-trained language model for neural machine translation. This method has achieved a large improvement compared with the Transformer model in the translation tasks of multiple public datasets, proving the feasibility of incorporating the BERT pre-trained language model as an external knowledge base into the neural machine translation model. However, the method of Jinhua Zhu et al. depends on initializing the parameters of the pre-trained machine translation model, and the knowledge of the pre-trained language model needs to be introduced for each layer. Moreover, their feature fusion method is simple splicing. The cross-attention mechanism makes the pre-trained language model information affected by the word embedding information, and finally the random weight addition method is used for feature fusion.
[0004] Therefore, the present invention conducts research work on how to effectively incorporate the language information in the BERT pre-trained language model in low-resource neural machine translation. Summary of the Invention
[0005] Aiming at the problem that the translation performance of Chinese-Vietnamese neural machine translation is restricted by insufficient bilingual parallel sentence pair data, this invention proposes a Chinese-Vietnamese neural machine translation method that combines the dual representations of BERT and word embeddings. This method respectively performs BERT pre-trained language model representation and word embedding representation on the source language sequence, and then uses the attention mechanism to achieve the adaptive dynamic fusion of the dual representations, enhancing the representation learning ability of the source language. Multiple groups of experiments have been carried out on Chinese-Vietnamese and English-Vietnamese translation tasks. The results show that the use of the adaptive dynamic fusion of BERT pre-trained model representation and word embedding representation can effectively integrate the language information in the BERT pre-trained language model into the neural machine translation model, effectively improving the performance of the Chinese-Vietnamese neural machine translation model.
[0006] The technical solution of this invention is as follows: Based on the Chinese-Vietnamese neural machine translation method that combines the dual representations of BERT and word embeddings, the specific steps of the Chinese-Vietnamese neural machine translation method that combines the dual representations of BERT and word embeddings are as follows:
[0007] Step1. Collect Chinese-Vietnamese parallel corpora for training the parallel sentence pair extraction model;
[0008] Step2. Collect the pre-trained Chinese BERT pre-trained language model parameters and dictionaries;
[0009] Step3. Respectively perform BERT pre-trained language model pre-training representation and word embedding representation on the source language sequence;
[0010] Step4. Use the cross-attention mechanism to make the source language sequence representation pre-trained by the BERT pre-trained language model be constrained by the word embedding representation, and splice and fuse the source language sequence representation and the word embedding representation after being trained by the BERT pre-trained language model to obtain a fused representation as the input of the encoder;
[0011] Step5. Use the encoder to enable the two different-source representations in the fused representation to achieve deep dynamic interactive fusion;
[0012] Step6. Use the dual representations of the BERT pre-trained language model and word embeddings to train the neural machine translation model.
[0013] As a further solution of this invention, in Step1, crawler technology is used to collect Chinese-Vietnamese bilingual parallel sentence pairs on the Internet, and the collected data is cleaned and Tokenize processed to construct a dataset of Chinese-Vietnamese bilingual parallel sentence pairs, which is used as experimental training, testing, and verification data.
[0014] As a further solution of the present invention, in Step2, collect the Chinese BERT pre-trained language model parameters and dictionaries released by Google, and instantiate the model parameters and dictionaries into the BERT pre-trained language model under the Pytorch framework.
[0015] As a further solution of the present invention, the specific steps of Step3 are as follows:
[0016] Step3.1: Segment the Chinese-Vietnamese monolingual corpus according to the BERT pre-trained language model dictionary and the training corpus dictionary; obtain two ID sequences of the input sequence;
[0017] Step3.2: Input the two text IDs obtained after segmentation into word embedding and the BERT pre-trained language model for representation.
[0018] As a further solution of the present invention, the specific steps of Step4 are as follows:
[0019] Step4.1: Use the cross-attention mechanism to calculate with the BERT pre-trained language model representation and the word embedding representation. Use the word embedding representation as the query condition, calculate the attention weights through the BERT pre-trained language model representation, and then calculate with the weights and the BERT pre-trained language model representation to make the BERT pre-trained language model representation be constrained by the word embedding representation;
[0020] Step4.2: Perform self-attention mechanism calculation on the word embedding representation to strengthen the internal connection of this representation;
[0021] Step4.3: Concatenate the representations obtained in Step4.1 and Step4.2 to obtain a fused representation;
[0022] As a further solution of the present invention, in Step5, the encoder designs a self-attention mechanism to enable deep dynamic interaction and fusion of the two differently sourced representations in the fused representation.
[0023] As a further solution of the present invention, in Step6, the representation obtained after the self-attention mechanism in Step5 participates in the training of the Transformer model to achieve the fusion of the BERT pre-trained language model and the word embedding part trained by the Transformer language model.
[0024] The present invention proposes a Chinese-Vietnamese neural machine translation method that fuses dual representations of BERT and word embeddings. Compared with the method proposed by Jinhua Zhu et al., the method proposed in the present invention only uses the pre-trained language model once, and the model structure is simpler. It solves the problem that the method of Jinhua Zhu et al. depends on initializing the parameters of the pre-trained machine translation model. The present invention does not require pre-training the machine translation model. In terms of information fusion, it uses an adaptive fusion method to replace the random weighted fusion method, achieving greater performance improvement in the Chinese-Vietnamese neural machine translation task. Moreover, their feature fusion method is simple splicing. Although the method of the present invention uses the cross-attention mechanism proposed by Jinhua Zhu et al. to constrain the information of the pre-trained language model by the word embedding information, when Jinhua Zhu et al. finally perform feature fusion, they use the method of adding random weights. In contrast, in the present invention, after splicing the two feature vectors, the self-attention mechanism is used to perform internal information interaction and fusion on the fused vector. Compared with previous work, the present invention not only reduces the model complexity but also improves the performance.
[0025] The beneficial effects of the present invention are as follows:
[0026] 1. The present invention uses a Chinese-Vietnamese neural machine translation method that fuses dual representations of BERT and word embeddings, and its effect is significantly better than that of the Transformer-based model, improving the performance of the overall machine translation model.
[0027] 2. The present invention adopts multiple groups of attention mechanisms to achieve the fusion of two different sources of representations. Experiments prove that this fusion method has a greater improvement in the BLEU index compared with the fusion method proposed by the BERT-fused algorithm;
[0028] 3. The present invention not only reduces the model complexity but also improves the performance;
[0029] 4. The method of the present invention respectively performs BERT pre-trained language model representation and word embedding representation on the source language sequence, and then uses the attention mechanism to achieve the adaptive dynamic fusion of the dual representations, enhancing the representation learning ability of the source language. Multiple groups of experiments have been carried out on Chinese-Vietnamese and English-Vietnamese translation tasks. The results show that using the adaptive dynamic fusion of BERT pre-trained model representation and word embedding representation can effectively integrate the language information in the BERT pre-trained language model into the neural machine translation model, effectively improving the performance of the Chinese-Vietnamese neural machine translation model, and solving the problem that the performance of Chinese-Vietnamese neural machine translation is not ideal due to Vietnamese being a low-resource language. Description of the Drawings
[0030] Figure 1Flowchart of the Chinese-Vietnamese neural machine translation method that integrates the dual representations of BERT and word embeddings proposed by the present invention. Detailed implementation manners
[0031] Example 1: As Figure 1 shown, the Chinese-Vietnamese neural machine translation method that integrates the dual representations of BERT and word embeddings
[0032] The specific steps of the Chinese-Vietnamese neural machine translation method based on the integration of the dual representations of BERT and word embeddings are as follows:
[0033] Step1. Collect Chinese-Vietnamese parallel corpora for training the parallel sentence pair extraction model;
[0034] Step2. Collect pre-trained Chinese BERT pre-trained language model parameters and dictionaries;
[0035] Step3. Perform pre-training representations of the BERT pre-trained language model and word embedding representations on the source language sequence respectively;
[0036] Step4. Use the cross-attention mechanism to make the source language sequence representation pre-trained by the BERT pre-trained language model be constrained by the word embedding representation, and splice and fuse the source language sequence representation and the word embedding representation after being trained by the BERT pre-trained language model to obtain a fused representation as the input of the encoder;
[0037] Step5. Use the encoder to enable deep dynamic interaction and fusion of the two representations from different sources in the fused representation;
[0038] Step6. Use the dual representations of the BERT pre-trained language model and word embeddings to train the neural machine translation model.
[0039] As a further solution of the present invention, in Step1, crawler technology is used to collect Chinese-Vietnamese bilingual parallel sentence pairs on the Internet, and the collected data is cleaned and Tokenize processed to construct a dataset of Chinese-Vietnamese bilingual parallel sentence pairs, which is used as experimental training, testing, and verification data.
[0040] As a further solution of the present invention, in Step2, the Chinese BERT pre-trained language model parameters and dictionaries released by Google are collected, and the model parameters and dictionaries are instantiated as the BERT pre-trained language model in the Pytorch framework.
[0041] As a further solution of the present invention, the specific steps of Step3 are:
[0042] Step3.1. Tokenize the Chinese-Vietnamese monolingual corpus according to the BERT pre-trained language model dictionary and the training corpus dictionary; obtain two ID sequences of the input sequence.
[0043] Step3.2. Input the text IDs obtained after tokenization into word embedding and the BERT pre-trained language model for representation respectively.
[0044] As a further solution of the present invention, the specific steps of Step4 are as follows:
[0045] Step4.1. Use the cross-attention mechanism to calculate with the BERT pre-trained language model representation and the word embedding representation. Use the word embedding representation as the query condition, calculate the attention weights through the BERT pre-trained language model representation, and then calculate with the weights and the BERT pre-trained language model representation to make the BERT pre-trained language model representation be constrained by the word embedding representation.
[0046] Step4.2. Calculate the self-attention mechanism for the word embedding representation to strengthen the internal connection of this representation.
[0047] Step4.3. Concatenate the representations obtained in Step4.1 and Step4.2 to obtain a fused representation.
[0048] As a further solution of the present invention, in Step5, the encoder designs a self-attention mechanism to enable deep dynamic interaction and fusion of the two representations from different sources in the fused representation.
[0049] As a further solution of the present invention, in Step6, the representation obtained after the self-attention mechanism in Step5 participates in the training of the Transformer model to realize the fusion of the BERT pre-trained language model and the word embedding part trained by the Transformer language model.
[0050] To verify the effectiveness of the Chinese-Vietnamese neural machine translation that fuses the dual representations of BERT and word embedding in the above embodiments, the following 5 comparative experiments on the translation performance of Chinese-Vietnamese neural machine translation methods are carried out:
[0051] ⑴ RNNSearch: A neural machine translation method based on the recurrent neural network structure.
[0052] ⑵ CNN: A neural machine translation method based on the convolutional neural network structure.
[0053] ⑶ Transformer: A neural machine translation method based on the Transformer network structure.
[0054] ⑷BERT-fused: A neural machine translation method that incorporates the BERT pre-trained language model into the Transformer encoder and decoder.
[0055] ⑸Ours: A neural machine translation method that fuses the dual representations of BERT and word embeddings.
[0056] The above methods use the same training set, test set, and validation set in the experiment. The BERT-fused and ours methods use the same pre-trained language model. The experimental results are shown in Table 1.
[0057] Comparison experiment results of neural machine translation in Table 1
[0058]
[0059] As can be seen from the experimental results in Table 1, after the present invention fuses the pre-training of the source language sequence with the BERT pre-trained language model and the dual representation of word embeddings, compared with the Transformer model, a performance improvement of 1.99 BLEU values is obtained on Chinese-Vietnamese data. This shows that using the BERT pre-trained language model can supplement the language information capture ability of the neural machine translation model in low-resource scenarios, achieving the purpose of improving the performance of the Chinese-Vietnamese neural machine translation model. The present invention has a 1.26 BLEU value improvement compared with the BERT-fused method on the Chinese-Vietnamese dataset, indicating that the present invention can make more effective use of the language information in the BERT pre-trained language model compared with the BERT-fused method in the low-resource Chinese-Vietnamese neural machine translation task.
[0060] To verify the effectiveness of the present invention in neural machine translation with different amounts of low-resource data, a comparative experiment on the improvement range of BLEU values of the Ours method relative to the Transformer method under 3 groups of different amounts of data was designed:
[0061] ⑴ Using 127.4k Chinese-Vietnamese data as the training data, compare the change range of BLEU values between the two methods.
[0062] ⑵ Randomly select 100k Chinese-Vietnamese data as the training data, compare the change range of BLEU values between the two methods.
[0063] ⑶ Randomly select 70k Chinese-Vietnamese data as the training data, compare the change range of BLEU values between the two methods.
[0064] The same validation set, test set, model hyperparameters, and the same Chinese BERT pre-trained language model are used in the three groups of experiments. The experimental results are shown in Table 2.
[0065] Comparison experiment results of different amounts of Chinese-Vietnamese data in Table 2
[0066]
[0067] As can be seen from the experimental results in Table 2, in the experiments with data of 70k, 100k, and 127.4k, the improvement amplitudes of the BLEU value of the present invention relative to Transformer are 4.34, 2.12, and 1.99 respectively, showing a gradually decreasing trend. This change trend indicates that the improvement of the present invention relative to the Transformer model in terms of the BLEU value decreases continuously with the increase of the training data. It proves that the BERT pre-trained language model has a greater supplementary effect on the neural machine translation model when the training data is less, and can achieve better performance in low-resource neural machine translation tasks with only tens of thousands of data volumes.
[0068] To explore the influence of introducing the pre-trained language model into the translation model by using the representation fusion method proposed in the present invention in the encoder, the following three groups of ablation experiments were designed:
[0069] ⑴ Only fuse the dual representations of the BERT pre-trained language model and word embeddings as the input of the first layer of the encoder.
[0070] ⑵ Incorporate the BERT pre-trained language model into the inputs of the first three layers of the encoder.
[0071] ⑶ Incorporate the BERT pre-trained language model into the inputs of all layers of the encoder.
[0072] In the three groups of experiments, the same 127.4k Chinese-Vietnamese data was used as the training set, and the validation set, test set, model hyperparameters, and Chinese BERT pre-trained language model were the same. The experimental results are shown in Table 3.
[0073] Table 3 Results of ablation experiments on multi-layer incorporation of pre-trained language models
[0074]
[0075] As can be seen from the experimental results, using the result of fusing BERT and word embeddings in the present invention as the input of the first layer of the encoder can achieve the best performance. Incorporating the BERT pre-trained language model into the inputs of the first three layers and all layers of the encoder does not significantly improve the performance of the neural machine translation model. The BERT pre-trained language model has a good supplementary ability for the neural machine translation model, indicating that the proposed representation fusion method in the present invention can fully utilize the language knowledge of the pre-trained language model in the shallow network to achieve the purpose of improving the performance of the neural machine translation model.
[0076] To explore the influence of incorporating the pre-trained language model information in the decoding stage using the present invention on the performance of the translation model, we designed the following ablation experiments:
[0077] ⑴ The BERT pre-trained language model is only fused with the encoder output hidden state vector as the decoder input.
[0078] ⑵ The BERT pre-trained language model is only fused with the word embedding as the encoder input.
[0079] ⑶ The BERT pre-trained language model is fused with the word embedding as the encoder input. After the encoding stage, the BERT pre-trained language model is fused with the hidden state vector output by the encoder as the decoder input.
[0080] The same 127.4k Chinese-Vietnamese data is used as the training set in the three groups of experiments. The validation set, test set, model hyperparameters, and Chinese BERT pre-trained language model used are the same. The experimental results are shown in Table 4.
[0081] Table 4 Ablation experiment results of integrating the pre-trained language model in the decoding stage
[0082]
[0083] It can be seen from the experimental results that using the present invention to integrate the BERT pre-trained language model in the decoding stage has a negative impact on the performance of the neural machine translation model. Only integrating the BERT pre-trained language model in the decoding stage results in the performance of the neural machine translation being lower than that of the baseline model Transformer. The performance of integrating the BERT pre-trained language model in both the encoding stage and the decoding stage is also lower than that of only integrating the BERT pre-trained language model in the encoding stage. It is proved that using the proposed representation fusion method in the decoding stage to integrate the BERT pre-trained language model does not improve the performance of the neural machine translation model.
[0084] To verify the effectiveness of the present invention in other language translation tasks, experiments were also conducted on the IWSLT15 English-Vietnamese translation dataset. The data scale of this dataset is shown in Table 5.
[0085] Table 5 English-Vietnamese dataset
[0086]
[0087] Comparative experiments of RNNSearch, CNN, Transformer, BERT-fused method, and Ours method were conducted on this dataset. The experimental results are shown in Table 6.
[0088] Table 6 Comparative experiment results of English-Vietnamese neural machine translation
[0089]
[0090] As can be seen from the experimental results in Table 6, the Chinese-Vietnamese neural machine translation method that combines BERT and word embedding dual representations proposed in the present invention has a performance improvement of 1.56 BLEU values compared with the Transformer model on English-Vietnamese data, and a 0.41 BLEU value improvement compared with the BERT-fused method. This shows that this method is not only applicable to Chinese-Vietnamese neural machine translation, but also using the pre-trained language model and word embedding layer of the source language for dual representation in other low-resource neural machine translation tasks can also improve the performance of the neural machine translation model.
[0091] Example 2: As Figure 1 shown, the Chinese-Vietnamese neural machine translation method that combines BERT and word embedding dual representations is specifically as follows:
[0092] Step1. First, use web crawler technology to collect a large number of Chinese-Vietnamese parallel sentence pairs on the Internet, clean and Tokenize the collected data, thus constructing a dataset of Chinese-Vietnamese bilingual parallel sentence pairs, and use this dataset as the experimental training, testing, and validation data;
[0093] Step2. Perform word embedding on the processed dataset. Without additional design in this part, the input text is segmented according to the word embedding dictionary and then input into the word embedding module to obtain the word embedding representation E of the input text embedding .
[0094] Step3. After segmenting the input text according to the BERT pre-trained language model dictionary, obtain the input sequence x = (x 1 ,..., x n ). After inputting the input sequence into the BERT pre-trained language model, a hidden state vector will be output at each layer of the model. In this method, the hidden state vector h output by the last layer 6 is used as the output E of this part bert-out .
[0095] Step4. Use E bert-out and the word embedding representation E embedding to perform cross-attention mechanism calculation. Take the output E of the word embedding part embedding as Query, E bert-out as Key to calculate the attention weight, multiply E bert-out as Value by the attention weight, so that the source language sequence representation pre-trained by the BERT pre-trained language model is constrained by the word embedding representation. The calculation process is shown in equations (1)(2)(3)(4). After applying the cross-attention mechanism, after E bert-out is constrained by E embedding , a new representation E' bert-out is obtained.
[0096] Query = E embedding (1)
[0097] Value = Key = E bert-out (2)
[0098]
[0099] E' bert-out = Attention(Query, Key, Value) (4)
[0100] Step 5. Perform self-attention mechanism calculation on E embedding for representation enhancement. The calculation process is shown in Equations (5) and (6), and E' is obtained embedding .
[0101] Query = Value = Key = E embedding (5)
[0102] E' embedding = Attention(Query, Key, Value) (6)
[0103] Step 6. Concatenate E' bert-out and E' embedding . After dimensionality transformation through linear transformation, a new text sequence hidden state vector E bert-embedding is obtained. The calculation process is shown in Equations (7) and (8).
[0104] E contact = contact(E' bert-out , E' embedding ) (7)
[0105] E bert-embedding = Linear(E contact ) (8)
[0106] Step 7. The BERT pre-trained language model representation and word embedding representation fusion module obtains a representation vector E bert-out containing the information of E' embedding and E' bert-embedding . No connection is established between the two parts of information. When E bert-embedding enters the first layer of the encoder, a self-attention mechanism calculation is performed once, enabling the two originally independent parts to establish a connection and obtaining E' bert-embedding . The calculation process is shown in Equations (9) and (10).
[0107] Query = Value = Key = E bert-embedding (9)
[0108] E' bert-embedding = Attention(Query, Key, Value) (10)
[0109] Step8. After being calculated by the self-attention mechanism, E' is obtained bert-embedding , realizing E bert-out and E embedding for dynamic fusion. E' bert-embedding passes through a feed-forward neural network to obtain the output H of the first layer of the encoder 1 , and finally obtains the final output of the encoder after passing through multiple encoding layers. The calculation process is shown in Equations (11), (12), and (13).
[0110] H 1 = FNN(E' bert-embedding ) (11)
[0111] h t = Attention(H t-1 , H t-1 , H t-1 ), t > 1 (12)
[0112] H t = FNN(h t ), t > 1 (13)
[0113] Step9. To verify the performance of the audit machine translation, the BLEU value is used as an evaluation index. The calculation method of BLEU is shown in Equation (14).
[0114]
[0115] The specific implementation manners of the present invention have been described in detail above with reference to the accompanying drawings. However, the present invention is not limited to the above implementation manners. Various changes can be made without departing from the spirit of the present invention within the knowledge scope of those of ordinary skill in the art.
Claims
1. A Chinese-Vietnamese neural machine translation method that fuses the dual representations of BERT and word embeddings, characterized in that: The method includes: Step1. Collect Chinese-Vietnamese parallel corpora for training the parallel sentence pair extraction model; Step2. Collect the pre-trained Chinese BERT pre-trained language model parameters and dictionaries; Step3. Perform pre-training representations of the BERT pre-trained language model and word embedding representations on the source language sequence respectively; Step4. Use the cross-attention mechanism to make the source language sequence representation pre-trained by the BERT pre-trained language model be constrained by the word embedding representation, and splice and fuse the source language sequence representation and the word embedding representation after being trained by the BERT pre-trained language model to obtain a fused representation as the input of the encoder; Step5. Use the encoder to enable deep dynamic interaction and fusion of the two different-source representations in the fused representation; Step6. Use the dual representations of the BERT pre-trained language model and word embeddings to train the neural machine translation model.
2. The Chinese-Vietnamese neural machine translation method that fuses the dual representations of BERT and word embeddings according to claim 1, characterized in that: In the Step1, crawler technology is used to collect Chinese-Vietnamese bilingual parallel sentence pairs on the Internet, and the collected data is cleaned and Tokenize processed to construct a dataset of Chinese-Vietnamese bilingual parallel sentence pairs, and this dataset is used as experimental training, testing, and validation data.
3. The Chinese-Vietnamese neural machine translation method that fuses the dual representations of BERT and word embeddings according to claim 1, characterized in that: In the Step2, the Chinese BERT pre-trained language model parameters and dictionaries released by google are collected, and the model parameters and dictionaries are instantiated as the BERT pre-trained language model under the Pytorch framework.
4. The Chinese-Vietnamese neural machine translation method that fuses the dual representations of BERT and word embeddings according to claim 1, characterized in that: The specific steps of the Step3 are: Step3.
1. Segment the Chinese-Vietnamese monolingual corpus according to the BERT pre-trained language model dictionary and the training corpus dictionary; Step3.
2. Input the text IDs obtained after the two segmentations into word embeddings and the BERT pre-trained language model respectively for representation.
5. The Chinese-Vietnamese neural machine translation method that fuses the dual representations of BERT and word embeddings according to claim 1, characterized in that: The specific steps of the Step4 are: Step4.
1. Use the BERT pre-trained language model representation and the word embedding representation to perform cross-attention mechanism calculation, use the word embedding representation as the query condition, calculate the attention weights through the BERT pre-trained language model representation, and then calculate with this weight and the BERT pre-trained language model representation to make the BERT pre-trained language model representation be constrained by the word embedding representation; Step4.
2. Perform self-attention mechanism calculation on the word embedding representation to strengthen the internal connection of this representation; Step4.
3. Splice the representations obtained in Step4.1 and Step4.2 to obtain a fused representation.
6. The Chinese-Vietnamese neural machine translation method integrating dual representations of BERT and word embeddings according to claim 1, characterized in that: In the said Step 5, the encoder designs a self-attention mechanism to enable deep dynamic interaction and fusion between the two representations from different sources in the fused representation.
7. The Chinese-Vietnamese neural machine translation method integrating dual representations of BERT and word embeddings according to claim 6, characterized in that: In the said Step 6, the representation obtained after the self-attention mechanism in Step 5 participates in the training of the Transformer model, realizing the fusion of the BERT pre-trained language model and the word embedding part trained by the Transformer language model.
Citation Information
Patent Citations
Mongolian-Chinese neural machine translation method based on combination of distillation BERT and improved Transformer
CN112347796A
Chinese entity identification method based on BERT and Word2Vec vector fusion
CN112632997A