A Design Method for Machine Translation Based on Deep Learning

By preprocessing, pre-training and fine-tuning the multilingual machine translation model, the problem of poor translation effect in small languages ​​is solved, and efficient translation effect is achieved with less corpus data, which is suitable for the rapid iteration and development of multilingual machine translation.

CN116187350BActive Publication Date: 2025-05-30TOEC TECHNOLOGLY CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202211454321.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-11-21
Publication Date
2025-05-30
Estimated Expiration
2042-11-21

AI Technical Summary

Technical Problem

The existing multilingual machine translation model lacks high-quality corpus data, especially for small languages, has poor translation effect.

Method used

The machine translation design method based on deep learning is adopted, and the input corpus data is pre-processed, and the corpus data of multiple languages ​​is pre-trained, and the corpus of the source and target languages ​​is fine-tuned based on pre-training to reduce the dependence on corpus data.

Benefits of technology

It significantly improves the effect of multilingual machine translation, especially the translation effect of small languages, can obtain better translation results with less corpus data, and the model can also be iterated and developed quickly.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116187350B_ABST
    Figure CN116187350B_ABST
Patent Text Reader

Abstract

A machine translation design method based on deep learning, which includes data preprocessing, data pre-training, and data fine-tuning; first, preprocess the input corpus data, and use corpus data in multiple languages for data pre-training. On the basis of pre-training, use the corpus of the source language and the target language to fine-tune the data. The present invention breaks the limitations of languages and language families, has less dependence on corpus data, and can obtain better translation results with less corpus data, which can significantly improve the effect of multilingual machine translation, especially greatly improve the translation effect of minority languages.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of computer natural language processing, and more specifically, to a machine translation design method based on deep learning. Background Art

[0002] Artificial intelligence translation supplements the shortage of human translation, improves the translation efficiency, and is convenient for individual users to operate. It has great dominance and superiority in low to medium-level translation tasks. However, in practical applications, artificial intelligence translation requires a long time for training, and the effect depends greatly on the quantity and quality of corpus data. In the case of lack of high-quality corpus data, it is easy to obtain unsatisfactory translation results.

[0003] The currently commonly used multilingual machine translation model is the mBART model proposed in 2020. This model has a good effect when the training dataset is large, but for small languages with fewer datasets, the translation effect is poor. For example, when using an English-German dataset of 4.5M, its BLEU evaluation index is 30.5, while in the case of the small language Kazakh, when using an English-Kazakh dataset of 128k, its BLEU evaluation index is only 2.5, and the translation effect is poor. Summary of the Invention

[0004] In view of the problem that the existing technology has a large dependence on the quantity and quality of corpus data, especially for small languages with fewer datasets and poor translation effects, the present invention provides a machine translation design method based on deep learning. This method preprocesses the input corpus data and uses corpus data in multiple languages for pre-training. On the basis of pre-training, the corpus of the source language and the target language is used to fine-tune the data. The present invention has less dependence on corpus data, and a good translation effect can be obtained by making minor adjustments on the basis of pre-training, which can significantly improve the effect of multilingual machine translation, especially greatly improve the translation effect of small languages.

[0005] The technical solution adopted by the present invention is: a machine translation design method based on deep learning, using a Django service carrier as the implementation platform. This method includes data preprocessing, data pre-training, and data fine-tuning;

[0006] Step 1, the data preprocessing is to select different tokenizers according to different languages to tokenize the input corpus data, and use the bpe tokenization method to split the tokenized corpus data into tokens. Language markers are added to the input corpus to distinguish multiple languages in the mixed corpus, and RAS random replacement is used to replace some words with synonyms in other languages;

[0007] Step 2, the data pre-training is to perform pre-training on the corpus data of multiple languages that have undergone data pre-processing. The corpus data is placed into the Transformer network structure and trained in an encoding-decoding manner to obtain a pre-trained model;

[0008] Step 3, the data fine-tuning is to fine-tune the effect of the pre-trained model. The source language and target language data that have undergone data pre-processing are used for training to obtain a fine-tuned model, and the model undertakes the translation tasks for specified language pairs and specified fields;

[0009] Through the above steps, the construction of the deep learning model and the rapid iteration of the translation model for different language pairs are realized.

[0010] The Transformer network structure described in Step 2 is twelve Transformer_Big structures, consisting of an encoder and a decoder stacked by twelve Blocks.

[0011] The Transformer network structure described in Step 2 introduces the Self-Attention mechanism to establish the connection between Blocks, and adopts a design pattern of calculating with a one-time input sequence, and calculates multiple inputs in batches.

[0012] The fine-tuning of the model effect described in Step 3 is to input the source language and target language data, and fine-tune the model effect on the basis of pre-training. That is, first, a basic model is pre-trained using a large amount of computing power and a large amount of multilingual parallel corpus, and then, starting from the basic model, it is fine-tuned with a small amount of monolingual or bilingual corpus to evolve a machine translation model adapted to a specific language pair.

[0013] The technical effects produced by the present invention are as follows: The present invention is a technical method for solving the mutual translation between multiple languages based on deep learning. By transferring knowledge from resource-rich pre-training tasks to resource-low / zero-resource downstream tasks, great success has been achieved in language understanding and text generation. Using a large amount of bilingual parallel corpus of multiple languages to jointly train a unified model, and then fine-tuning based on this, the role of the pre-trained model is given full play. This method not only has strong versatility, but also breaks the limitations of languages and language families, and better translation results can be obtained with less corpus data; moreover, the pre-training-fine-tuning mode enables the model to be quickly iterated and developed, bringing convenience to the expansion of language families and the extension of professional fields.

[0014] This method has less dependence on corpus data. With minor adjustments based on pre-training, better translation results can be obtained, which can significantly improve the effect of multilingual machine translation, especially greatly improving the translation effect of minority languages. For example, when using an English-German dataset of 4.5M in size, the BLEU evaluation index of the mBART model is 30.5, while that of the present invention is 35.2; when using a Kazakh-English dataset of 128k in size, the BLEU evaluation index of the mBART model is 7.4, while that of the present invention is 12.3; when using an English-Kazakh dataset of 128k in size, the BLEU evaluation index of the mBART model is 2.5, while that of the present invention is 8.2. It can be seen that after using the method of the present invention, the translation effects of the three corpora of the present invention are all better than those of the mBART model, especially for minority languages. For example, the translation effect of English-Kazakh of the present invention is more than 3.2 times higher than that of the mBART model.

[0015] This method can significantly improve the effect of multilingual machine translation and can be applied in various scenarios such as online deployment, offline deployment, and hardware integration to improve the intelligence level of language understanding. Brief Description of the Drawings

[0016] Figure 1 is a flowchart of a machine translation design method based on deep learning according to the present invention;

[0017] Figure 2 is a Transformer network encoding-decoding structure diagram of a machine translation design method based on deep learning according to the present invention;

[0018] Figure 3 is a network structure diagram of a machine translation design method based on deep learning according to the present invention. Detailed Embodiments

[0019] The following will further describe in detail the specific embodiments of the present invention in conjunction with the drawings and embodiments.

[0020] Figure 1 is a flowchart of a machine translation design method based on deep learning according to the present invention, including preprocessing, pre-training, and fine-tuning; among them, preprocessing includes start, input data, tokenization using a tokenizer and bpe tokenization method, adding language markers and performing RAS random replacement process; pre-training includes corpus data in multiple languages, pre-training the corpus data and pre-training model process; fine-tuning includes corpus data in the source language and target language, fine-tuning the pre-training model and fine-tuning model process.

[0021] Refer to Figure 1 the flowchart of a machine translation design method based on deep learning according to the present invention, and the specific implementation steps of the present invention are as follows.

[0022] Step 1, preprocess the input data.

[0023] After starting to input the data, use different tokenizers to tokenize the collected corpus data according to the language. Among them, use the jieba tokenizer for Chinese, the mecab tokenizer for Japanese, and the mosestokenizer for other languages such as English and German. Then, use the BPE tokenization method to further split individual words in the data into tokens. Add language markers to the input corpus to distinguish multiple languages in the mixed corpus. Finally, use RAS random replacement on the input corpus to replace some words with synonyms in other languages.

[0024] The following takes English-German as an example to preprocess the data of the source language English.

[0025] Table 1 is the data preprocessing step table. As shown in Table 1, the original corpus (English) is "I like singing andswimming". After being tokenized by the mosestokenizer, it becomes "I \t like \t singing \t and \tswimming". After BPE tokenization, it becomes "I \t like \t sing \t ing \t and \t swim \tming". Finally, after RAS random replacement (replaced with German), it becomes "I \t like \t singen \t ing \t and\t schwimmen \t ming". After these preprocessing steps, it becomes a corpus format that can be used for model training.

[0026] Table 1 Data Preprocessing Step Table

[0027] Original corpus (English) I like singing and swimming. mosestokenizer tokenization I \t like \t singing \t and \t swimming. BPE tokenization I \t like \t sing \t ing \t and \t swim \t ming. RAS random replacement (replaced with German) I \t like \t singen \t ing \t and \t schwimmen \t ming.

[0028] Step 2, data pre-training.

[0029] Perform corpus data pre-training on the corpus data of multiple languages after data preprocessing. At the same time, construct a Transformer network structure, place the preprocessed corpus data of multiple languages into the Transformer network, and train it in an encoding-decoding manner to obtain a pre-trained model.

[0030] Figure 2It is the structure diagram of the encoder-decoder of the Transformer network, which is a machine translation design method based on deep learning in the present invention. The Transformer network structure consists of an encoder and a decoder. The role of the encoder is to map natural language into feature vectors. The vector representation should contain all semantic information as much as possible, and its position features should conform to the grammar rules unique to the language. The dimension of the feature vector represents the complexity of the model and also represents the representational ability of the model. The role of the decoder is to restore the feature vector after encoding the source language into the natural language representation of the target language. The decoder has the same structure as the encoder, but its parameter distribution is different from that of the encoding end. Its essence is to learn the expression pattern of the target language and concretize the abstract semantic features into the target language.

[0031] The input corpus Source "I read a book on Sunday" is mapped into a feature vector by the encoder, and the feature vector is then mapped by the decoder into the corresponding translation Target "I read a book on Sunday".

[0032] The present invention adopts the Transformer_Big structure, and its basic unit is called a Transformer Block. The model is stacked with twelve Blocks for the encoding end and the decoding end. Its design structure is determined by comprehensively considering the complexity of the task and the representational ability of the model.

[0033] Figure 3 It is the network structure diagram of a machine translation design method based on deep learning in the present invention, including the processes of source language data, target language data, encoder, decoder, translation result, and probability. After the source language data is input into the encoder and the target language data is input into the decoder, the data passes through the Transformer network model together to obtain the source language translation result and the probability corresponding to the result. The optimal translation result is selected according to the corresponding probability, and the network model is continuously iterated according to the translation effect.

[0034] The Transformer network structure introduces the Self-Attention mechanism and adopts a design pattern of calculating with a one-time input sequence, which improves the parallelizability of the model. Therefore, multiple inputs can be calculated in batches, and its absolute position encoding mechanism is also superior to the sequential input method of the RNN structure.

[0035] Step 3, fine-tune the model effect.

[0036] Input the corpus data of the source language and the target language, and fine-tune the model effect on the basis of the pre-trained model to obtain a fine-tuned model.

[0037] First, a large amount of computing power and a large amount of multilingual parallel corpora are used to pre-train a basic model. Then, starting from the basic model, it is fine-tuned with a small amount of monolingual or bilingual corpora to evolve a machine translation model adapted to a specific language pair. Pre-training and fine-tuning have achieved great success in language understanding by transferring knowledge from resource-rich pre-training tasks to resource-scarce / zero-resource downstream tasks. This method of pre-training plus fine-tuning can help with rapid development and iteration, facilitate users to add and expand languages by themselves, and can achieve good translation results through relatively easy training and a small amount of parallel corpora. Fine-tuning is of great help in expanding translation in professional fields.

[0038] Through the above steps, the construction of the learning model of the machine translation design method based on deep learning according to the present invention and the rapid iteration of the translation models for different languages can be achieved.

Claims

1. A machine translation design method based on deep learning, which uses the Django service carrier as the implementation platform, characterized in that: This method includes data preprocessing, data pre-training, and data fine-tuning; Step 1, for the data preprocessing, different tokenizers are selected according to different languages to tokenize the input corpus data, and the tokenized corpus data is split into tokens using the bpe tokenization method. Language markers are added to the input corpus to distinguish multiple languages in the mixed corpus, and RAS random replacement is used to replace some tokens with synonyms in other languages; Step 2, for the data pre-training, the corpus data of multiple languages after data preprocessing is pre-trained. The corpus data is placed into the Transformer network structure and trained in an encoder-decoder manner to obtain a pre-trained model; Step 3, for the data fine-tuning, the effect of the pre-trained model is fine-tuned. The source language and target language data after data preprocessing are used for training to obtain a fine-tuned model, and the model undertakes the translation task for a specified language pair in a specified domain; Through the above steps, the construction of a deep learning model and the rapid iteration of translation models for different language pairs are realized.

2. A machine translation design method based on deep learning according to claim 1, characterized in that: The Transformer network structure described in Step 2 is twelve Transformer_Big structures, consisting of an encoder and a decoder stacked by twelve Blocks.

3. A machine translation design method based on deep learning according to claim 1, characterized in that: The Transformer network structure described in Step 2 introduces the Self-Attention mechanism to establish the connection between Blocks, and adopts a design pattern of calculating with a one-time input sequence, and calculates multiple inputs in batches.

4. A machine translation design method based on deep learning according to claim 1, characterized in that: For the fine-tuning of the model effect described in Step 3, the source language and target language data are input, and the model effect is fine-tuned on the basis of pre-training. That is, first, a basic model is pre-trained using a large amount of computing power and a large amount of multilingual parallel corpus, and then starting from the basic model, it is fine-tuned with a small amount of monolingual or bilingual corpus, so as to evolve into a machine translation model adapted to a specific language.

Citation Information

Patent Citations

  • Neural machine translation method based on Transform model optimization

    CN114722843A

  • Machine translation method, target translation model training method, and related program and device

    CN115130479A