An Automatic Post-Editing Method for Machine Translation Based on Large Model Data Augmentation

By using domain screening, forward translation and large language models to generate pseudo-data in the automatic post-editing task, and training in combination with mBART model, data scarcity problem is solved and the quality of machine-translated translations is significantly improved.

CN117556833BActive Publication Date: 2025-06-17HARBIN INST OF TECH
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202311332992.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-10-16
Publication Date
2025-06-17
Estimated Expiration
2043-10-16

AI Technical Summary

Technical Problem

Automatic post-editing tasks face data scarcity, resulting in poor model performance.

Method used

Pseudo-data is generated through domain screening and forward translation, and a large language model is used to generate auxiliary machine-translated translations, and the data is passed into the cross-language pre-trained model mBART for training for data augmentation.

Benefits of technology

It effectively improves the quality of machine translation, solves the problem of data scarcity, and shows better results in the automatic post-editing task on multilingual pairs.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN117556833B_ABST
    Figure CN117556833B_ABST
Patent Text Reader

Abstract

The present invention is a method for automatic post-editing of machine translation based on large model data augmentation. The present invention relates to the technical fields of automatic post-editing of machine translation and data augmentation. The present invention generates a large amount of pseudo-data that can be used for training through domain screening and forward translation, and generates additional auxiliary machine translation translations with the help of a large language model to solve the problem of data scarcity faced by the automatic post-editing task. Then, all the data obtained after data augmentation is input into the cross-lingual pre-trained model mBART for training, effectively improving the quality of machine translation translations. The method proposed by the present invention reasonably utilizes the language ability of the large language model, can simply and efficiently solve the problem of data scarcity faced by the automatic post-editing task, and at the same time, this method can be directly applied to the automatic post-editing task on multiple language pairs without having to train multiple machine translation models for data augmentation on different language pairs.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of machine translation automatic post - editing and data augmentation, and is a machine translation automatic post - editing method based on large - model data augmentation. Background Art

[0002] Automatic post - editing is an important technology for processing the output translations of machine translation systems to correct various defects in the translations. From an application perspective, this technology plays a very important role. It can provide machine - translated texts with improved quality for professional translators, reduce the (manual) post - editing workload, and also adjust the output of general machine translation systems to meet the format or specialized vocabulary requirements of a specific application field.

[0003] For current mainstream automatic post - editing methods, their implementation processes are very similar to machine translation. They usually are based on the Transformer architecture, with an encoder and a decoder structure. Taking the original text and the corresponding machine - translated text as inputs and the post - edited translation as the output, during the training stage, they learn to distinguish various errors in machine - translated texts through learning a large amount of manually annotated data, so as to correct various errors in newly emerging machine - translated texts during the testing stage.

[0004] As a supervised task, in addition to the original text and machine - translated texts, automatic post - editing also requires post - edited translations obtained by translation experts editing based on the machine - translated texts with reference to the original text. However, the process of obtaining post - edited translations through translation experts' editing consumes a large amount of manpower and time, and it is difficult to obtain a large amount of post - edited data. Therefore, the scale of datasets related to the automatic post - editing task is often small, facing a relatively serious data scarcity problem. The performance of models trained solely based on these data is often not ideal. Thus, in order to further improve the performance of various automatic post - editing methods, researchers from different teams have proposed a variety of feasible and effective automatic post - editing data augmentation methods.

[0005] As one of the most important recent developments in the field of natural language processing, pre - trained language models can automatically capture the linguistic rules of texts through pre - training and provide basic support for downstream tasks. Large - scale pre - trained language models represented by ChatGPT (hereinafter referred to as large language models) have shown very excellent language understanding, generation, and knowledge reasoning abilities. In order to utilize the powerful language capabilities of large language models, solve the data scarcity problem faced by the automatic post - editing task, and improve the performance of the automatic post - editing task, the present invention proposes a machine translation automatic post - editing method based on large - language - model data augmentation. Summary of the Invention

[0006] The present invention generates a large amount of pseudo-data that can be used for training through domain screening and forward translation, generates additional auxiliary machine translation translations with the help of a large language model, solves the problem of data scarcity faced by the automatic post-editing task, and then inputs all the data obtained after data augmentation into the cross-lingual pre-trained model mBART for training, effectively improving the quality of machine translation translations. Therefore, the present invention provides an automatic post-editing method for machine translation based on large model data augmentation.

[0007] It should be noted that in this article, relational terms such as first and second are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Moreover, the term "comprising", "including" or any other variant thereof is intended to cover non-exclusive inclusion, so that a process, method, article or device comprising a series of elements not only includes those elements, but also includes other elements not expressly listed, or also includes elements inherent to such process, method, article or device.

[0008] The present invention provides an automatic post-editing method for machine translation based on large model data augmentation, and the present invention provides the following technical solutions:

[0009] An automatic post-editing method for machine translation based on large model data augmentation, the method comprising the following steps:

[0010] Step 1: Collect bilingual parallel corpora for pseudo-data generation;

[0011] Step 2: Train a classifier for discriminating the domain to which the bilingual parallel corpus belongs according to the bilingual parallel corpus and the real training set data;

[0012] Step 3: Screen out the part of the bilingual parallel corpus that is closest to the domain of the training set data as the basis for pseudo-data generation;

[0013] Step 4: Use the method of forward translation to construct the triples required for the automatic post-editing task from the bilingual parallel corpus screened by the domain;

[0014] Step 5: Input the original texts in the generated pseudo-data and the real training set data into the large language model respectively to generate additional auxiliary machine translation translations, and merge them with the original triples into quadruples;

[0015] Step 6: Upsample the real training set data and merge it with the pseudo-data to form the final training set data;

[0016] Step 7: Perform preprocessing such as word segmentation and sub-word segmentation on the training set data. Concatenate the original text, the machine translation, and the assisted machine translation as the encoder-side input, and the post-edited translation as the decoder-side input to train the mBART model and complete the automatic post-editing task.

[0017] Preferably, step 1 is specifically as follows:

[0018] Collect publicly available Chinese-English parallel corpora, including CCMT, UN, ParaCrawl, Wiki-Matrix, WikiTitles, and NewsCommentary. Through data cleaning operations such as removing traditional Chinese characters, duplicates, and blank lines, obtain 20 million bilingual parallel corpora.

[0019] Preferably, step 2 is specifically as follows:

[0020] Randomly sample 50,000 from the collected bilingual parallel corpora as negative examples, and use 5,000 training sets of the CCMT2023 automatic post-editing task as positive examples to train a classifier based on mBERT. Learn the domain information of the real training set data through the mBERT classifier to determine whether the domain of a piece of data is the same as that of the training set data.

[0021] Preferably, step 3 is specifically as follows:

[0022] Score the domain of all 20 million bilingual parallel corpus data on the interval [0, 1] according to the trained mBERT classifier. 0 points indicate that it is completely irrelevant to the domain of the training set data, and 1 point indicates that it is completely the same as the domain of the training set data. After sorting by score from high to low, select the top 200,000 that are most similar to the training set positive examples as the basis for generating pseudo-data later.

[0023] Preferably, step 4 is specifically as follows:

[0024] Input all the Chinese in the selected 200,000 bilingual parallel corpora into the trained Transformer-based neural machine translation model. Use the English translation output by the model as the machine translation, the original Chinese in the original bilingual parallel corpus as the original text, and the English as the post-edited translation to finally obtain 200,000 high-quality machine translation automatic post-editing pseudo-data that belong to a similar domain as the real training set data.

[0025] Preferably, step 5 is specifically as follows:

[0026] Combine the 200,000 pieces of pseudo-data obtained in step 4 with 5,000 pieces of real training set data, and send the original texts in these data to ChatGPT. Use "Translate "{src}" into English:" as the prompt to obtain the translations of each original text by ChatGPT. Take these obtained translations as auxiliary machine translation translations, and merge them with the original triples to obtain quadruples for subsequent model training.

[0027] Preferably, step 6 is specifically as follows:

[0028] Upsample the 5,000 training sets of the CCMT2023 automatic post-editing task by 10 times, and combine them with the obtained 200,000 pieces of pseudo-data as the training set.

[0029] Preferably, step 7 is specifically as follows:

[0030] Use sentencepiece to tokenize and subword segment all the data, and then concatenate the processed original text, machine translation translation, and auxiliary machine translation translation together as the input to the encoder side of the mBART architecture, and use the processed post-editing translation as the input to the decoder side of the mBART architecture;

[0031] Before starting the formal training, it is necessary to first call fairseq for data preprocessing, build a dictionary and binarize the training data, and then use fairseq to call the parameters of the pre-trained mbart-cc25 to initialize the parameters of the mBART model. Then, perform the training of the automatic post-editing task on the above training set to complete the automatic post-editing.

[0032] A computer-readable storage medium stores a computer program thereon, and the program is executed by a processor to implement a machine translation automatic post-editing method based on large model data augmentation.

[0033] A computer device includes a memory and a processor. The memory stores a computer program, and when the processor executes the computer program, it implements a machine translation automatic post-editing method based on large model data augmentation.

[0034] The present invention has the following beneficial effects:

[0035] Compared with the prior art, the present invention:

[0036] The present invention screens a large amount of bilingual parallel corpora by training a classifier capable of discriminating data domains, and performs data augmentation on the basis of the screening by combining forward translation, upsampling, and auxiliary translation based on large language models. Finally, automatic post-editing is achieved based on the mBART model. Compared with the data augmentation methods in traditional automatic post-editing methods, the method proposed by the present invention reasonably utilizes the language capabilities of large language models, can simply and efficiently solve the data scarcity problem faced by automatic post-editing tasks, and at the same time, this method can be directly applied to automatic post-editing tasks on multilingual pairs without having to train multiple machine translation models for data augmentation on different language pairs. In addition, this method can also be directly applied to low-resource language pairs and achieve better results than traditional data augmentation methods. Description of the Drawings

[0037] In order to more clearly illustrate the specific embodiments of the present invention or the technical solutions in the prior art, the following will briefly introduce the drawings required for the description of the specific embodiments or the prior art. Obviously, the drawings in the following description are some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.

[0038] Figure 1 Flowchart of the automatic post-editing method for machine translation based on large model data augmentation;

[0039] Figure 2 Automatic post-editing framework diagram based on mBART. Detailed Embodiments

[0040] The following will clearly and completely describe the technical solutions of the present invention in conjunction with the drawings. Obviously, the described embodiments are some embodiments of the present invention, rather than all embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts fall within the scope of protection of the present invention.

[0041] In the description of the present invention, it should be noted that the orientation or positional relationship indicated by the terms "center", "upper", "lower", "left", "right", "vertical", "horizontal", "inner", "outer", etc. is based on the orientation or positional relationship shown in the drawings, and is only for the convenience of describing the present invention and simplifying the description, rather than indicating or implying that the device or element referred to must have a specific orientation, be constructed and operated in a specific orientation, and therefore should not be construed as a limitation of the present invention. In addition, the terms "first", "second", and "third" are only used for descriptive purposes and cannot be construed as indicating or implying relative importance.

[0042] In the description of the present invention, it should be noted that unless otherwise clearly specified and defined, the terms "installation", "connection", and "coupling" should be understood in a broad sense. For example, it can be a fixed connection, a detachable connection, or an integral connection; it can be a mechanical connection or an electrical connection; it can be a direct connection or an indirect connection through an intermediate medium, and it can be the communication inside two components. For those of ordinary skill in the art, the specific meanings of the above terms in the present invention can be understood according to specific situations.

[0043] In addition, the technical features involved in different embodiments of the present invention described below can be combined with each other as long as they do not conflict with each other.

[0044] The present invention will be described in detail below in conjunction with specific embodiments. Specific Embodiment 1:

[0046] According to Figures 1 to 2 as shown, the specific optimized technical solution adopted by the present invention to solve the above technical problems is: the present invention relates to a method for automatic post-editing of machine translation based on large model data augmentation.

[0047] A method for automatic post-editing of machine translation based on large model data augmentation, the method comprising the following steps:

[0048] Step 1: Collect bilingual parallel corpora for pseudo-data generation;

[0049] Step 2: Train a classifier for discriminating the domain to which the bilingual parallel corpus belongs according to the bilingual parallel corpus and the real training set data;

[0050] Step 3: Screen out the part of the bilingual parallel corpus that is closest to the domain of the training set data as the basis for pseudo-data generation;

[0051] Step 4: Use the method of forward translation to construct the triples required for the automatic post-editing task from the bilingual parallel corpus screened by the domain;

[0052] Step 5: Input the original texts in the generated pseudo-data and the real training set data into the large language model ChatGPT respectively, use Translate "{src}" into English: as the prompt, obtain the translations of ChatGPT for each original text, and use these obtained translations as additional auxiliary machine translation translations, and merge them with the original triples to obtain quadruples;

[0053] Step 6: Upsample the real training set data and merge it with the pseudo-data to form the final training set data;

[0054] Step 7: Perform preprocessing such as word segmentation and subword segmentation on the training set data. Concatenate the original text, machine translation, and auxiliary machine translation as the encoder-side input, and the post-edited translation as the decoder-side input to train the mBART model and complete the automatic post-editing task. Specific Embodiment 2:

[0056] The difference between Embodiment 2 and Embodiment 1 of this application is only that:

[0057] Taking the Chinese-English machine translation automatic post-editing task of CCMT2023 as an example, the following steps are described:

[0058] Step 1: Collect bilingual parallel corpora for subsequent pseudo-data generation.

[0059] In this step, a large number of publicly available Chinese-English parallel corpora are collected, including CCMT, UN, ParaCrawl, Wiki-Matrix, WikiTitles, and NewsCommentary, etc. After data cleaning operations such as removing traditional Chinese characters, duplicates, and blank lines, approximately 20 million bilingual parallel corpora are obtained.

[0060] Step 2: Train a classifier for discriminating the domain to which the bilingual parallel corpus belongs based on the bilingual parallel corpus and real training set data.

[0061] In this step, 50,000 samples are randomly selected from the above-mentioned collected bilingual parallel corpora as negative examples, and 5,000 training sets of the CCMT2023 automatic post-editing task are used as positive examples to train a classifier based on mBERT. This classifier can learn the domain information of the real training set data and discriminate whether the domain to which a piece of data belongs is the same as the domain to which the training set data belongs.

[0062] Step 3: Screen out the part of the bilingual parallel corpus that is closest to the domain of the training set data as the basis for pseudo-data generation.

[0063] In this step, the classifier obtained from training is used to score the domain to which all 20 million pieces of data belong in the interval [0, 1]. A score of 0 indicates that it is completely irrelevant to the domain of the training set data, and a score of 1 indicates that it is completely the same as the domain of the training set data. After sorting in descending order of scores, the first 200,000 pieces that are most similar to the training set positive examples are selected as the basis for subsequent pseudo-data generation.

[0064] Step 4: Use the forward translation method to construct the triples required for the automatic post-editing task from the bilingual parallel corpus screened by the domain.

[0065] In this step, all the Chinese texts in the selected 200,000 bilingual parallel corpora are input into a pre-trained neural machine translation model based on Transformer. The English translations output by the model are used as machine translation translations, the Chinese texts in the original bilingual parallel corpora are used as the source texts, and the English texts are used as the post-edited translations. Finally, 200,000 pieces of machine translation automatic post-editing pseudo-data with high quality and belonging to a similar field as the real training set data are obtained.

[0066] Step 5: Input the source texts in the generated pseudo-data and the real training set data into the large language model ChatGPT respectively, using "Translate "{src}" into English:" as the prompt to obtain the translations of ChatGPT for each source text. These obtained translations are used as additional auxiliary machine translation translations, and are combined with the original triple to obtain a quadruple.

[0067] In this step, the 200,000 pieces of pseudo-data constructed in Step 4 and 5,000 pieces of real training set data are combined together, and the source texts in these data are all input into ChatGPT, using "Translate "{src}" into English:" as the prompt to obtain the translations of ChatGPT for each source text. These obtained translations are used as auxiliary machine translation translations, and are combined with the original triple to obtain a quadruple for subsequent model training.

[0068] Step 6: Upsample the real training set data and combine it with the pseudo-data to form the final training set data.

[0069] In this step, the 5,000 training sets of the CCMT2023 automatic post-editing task are upsampled by 10 times and combined with the 200,000 pieces of pseudo-data obtained from the above operations as the training set.

[0070] Step 7: Perform preprocessing such as tokenization and subword segmentation on the training set data. Concatenate the source text, machine translation translation, and auxiliary machine translation translation as the encoder-side input, and the post-edited translation as the decoder-side input to train the mBART model.

[0071] In this step, sentencepiece is used to tokenize and segment all the data into subwords. Then, the processed source text, machine translation translation, and auxiliary machine translation translation are concatenated together as the encoder-side input of the mBART architecture, and the processed post-edited translation is used as the decoder-side input of the mBART architecture. Before starting the formal training, it is necessary to first call fairseq for data preprocessing, build a dictionary and binarize the training data, and then use fairseq to call the parameters of the pre-trained mbart-cc25 to initialize the parameters of the mBART model. Then, the training of the automatic post-editing task is carried out on the above training set.

[0072] To verify the performance improvement of the machine translation automatic post - editing method based on large - language model data augmentation proposed by the present invention, the present invention conducted experiments on the Chinese - English automatic post - editing task of CCMT2023. The scale of the dataset provided by the task is shown in Table 1.

[0073] Table 1. Scale of the Chinese - English Automatic Post - editing Task Dataset of CCMT2023

[0074]

[0075] In the experiment, the present invention first obtained 200,000 pieces of data from Chinese - English parallel corpora such as CCMT, UN, ParaCrawl, Wiki - Matrix, WikiTitles, and NewsCommentary through domain screening, and then generated the final training set through forward translation, up - sampling, and ChatGPT - based assisted translation. The model used during training was mBART.cc25, the optimizer was AdamW, the learning rate was set to 3e - 5, the batchsize was set to 1024 tokens, the warmup steps were set to 2500, and it was trained for 20 rounds. The machine translation automatic post - editing method based on large - language model data augmentation was compared with the reference baseline of this task. The final experimental results are shown in Table 2.

[0076] Table 2. Experimental Results of the Chinese - English Automatic Post - editing Task of CCMT2023

[0077]

[0078] The experiment proves that the machine translation automatic post - editing method based on large - language model data augmentation proposed by the present invention has better performance compared with the reference baseline, with an improvement of - 4.1 in the TER metric and + 5.83 in the BLEU metric.

[0079] The present invention screens a large number of bilingual parallel corpora by training a classifier that can distinguish data domains, and on this basis, combines forward translation, up - sampling, and large - language - model - based assisted translation for data augmentation, and finally realizes automatic post - editing based on the mBART model. Compared with the data augmentation methods in traditional automatic post - editing methods, the method proposed by the present invention rationally utilizes the language ability of large - language models, can simply and efficiently solve the data scarcity problem faced by automatic post - editing tasks. At the same time, this method can be directly applied to automatic post - editing tasks on multiple language pairs without having to train multiple machine translation models for data augmentation on different language pairs. In addition, this method can also be directly applied to low - resource language pairs and achieve better results compared with traditional data augmentation methods. Specific Embodiment Three:

[0081] The difference between the third embodiment and the second embodiment of this application is only that:

[0082] The present invention provides a computer-readable storage medium, on which a computer program is stored, and the program is executed by a processor to implement a machine translation automatic post-editing method based on large model data enhancement. Specific Embodiment 4:

[0084] The difference between the fourth embodiment and the third embodiment of this application is only that:

[0085] The present invention provides a computer device, including a memory and a processor, the memory stores a computer program, and when the processor executes the computer program, a machine translation automatic post-editing method based on large model data enhancement is implemented.

[0086] In the description of this specification, the descriptions referring to terms such as "one embodiment", "some embodiments", "examples", "specific examples", or "some examples", etc. mean that the specific features, structures, materials, or characteristics described in connection with the embodiment or example are included in at least one embodiment or example of the present invention. In this specification, the schematic representations of the above terms do not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials, or characteristics described can be combined in any one or N embodiments or examples in a suitable manner. In addition, without contradiction, those skilled in the art can combine and combine the different embodiments or examples described in this specification and the features of different embodiments or examples. Furthermore, the terms "first" and "second" are only used for descriptive purposes and cannot be understood as indicating or implying relative importance or implicitly specifying the quantity of the indicated technical features. Thus, the features defined with "first" and "second" can explicitly or implicitly include at least one of such features. In the description of the present invention, the meaning of "N" is at least two, such as two, three, etc., unless otherwise specifically and clearly defined. Any process or method description represented in a flowchart or described in other ways herein can be understood as representing a module, segment, or part of code including one or more N executable instructions for implementing a customized logical function or process, and the scope of the preferred embodiments of the present invention includes additional implementations, where the functions can be executed in a manner that is not in the order shown or discussed, including in a substantially simultaneous manner according to the functions involved or in a reverse order, which should be understood by those skilled in the art to which the embodiments of the present invention pertain. The logic and / or steps represented in a flowchart or described in other ways herein, for example, can be considered as a sequenced list of executable instructions for implementing a logical function, and can be specifically implemented in any computer-readable medium for use by an instruction execution system, apparatus, or device (such as a computer-based system, a system including a processor, or other systems that can fetch and execute instructions from the instruction execution system, apparatus, or device), or in combination with such instruction execution systems, apparatus, or devices. For the purposes of this specification, a "computer-readable medium" can be any device that can contain, store, communicate, propagate, or transmit a program for use by or in combination with an instruction execution system, apparatus, or device. More specific examples (non-exhaustive list) of computer-readable media include the following: an electrical connection portion (electronic device) having one or N wirings, a portable computer diskette (magnetic device), a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber device, and a portable compact disc read-only memory (CDROM).In addition, the computer-readable medium can even be paper or other suitable media on which the program can be printed, because the program can be obtained electronically, for example, by optically scanning the paper or other media, followed by editing, interpretation, or other appropriate processing if necessary, and then storing it in a computer memory. It should be understood that the various parts of the present invention can be implemented using hardware, software, firmware, or a combination thereof. In the above-described embodiments, the N steps or methods can be implemented using software or firmware stored in a memory and executed by a suitable instruction execution system. For example, if implemented using hardware, as in another embodiment, any one or a combination of the following techniques well known in the art can be used: discrete logic circuits having logic gate circuits for implementing logical functions on data signals, application specific integrated circuits having appropriate combinational logic gate circuits, programmable gate arrays (PGAs), field programmable gate arrays (FPGAs), and the like.

[0087] The above is only a preferred embodiment of an automatic post-editing method for machine translation based on large model data augmentation. The protection scope of an automatic post-editing method for machine translation based on large model data augmentation is not limited to the above embodiments. Any technical solutions falling within this concept belong to the protection scope of the present invention. It should be noted that for those skilled in the art, several improvements and changes made without departing from the principle of the present invention should also be regarded as within the protection scope of the present invention.

Claims

1. An automatic post-editing method for machine translation based on large model data augmentation, characterized in that: The method includes the following steps: Step 1: Collect bilingual parallel corpora for pseudo-data generation; Step 2: Train a classifier for discriminating the field to which the bilingual parallel corpus belongs based on the bilingual parallel corpus and real training set data; Step 3: Screen out the part of the bilingual parallel corpus that is closest to the field of the real training set data as the basis for pseudo-data generation; Step 4: Use the method of forward translation to construct the triples required for the automatic post-editing task from the bilingual parallel corpus screened by the field; including: input all the Chinese in the screened bilingual parallel corpus into the trained neural machine translation model based on Transformer, use the English translation output by the model as the machine translation, the original Chinese in the original bilingual parallel corpus as the source text, and the English as the post-edited translation, finally obtaining high-quality machine translation automatic post-editing pseudo-data that belongs to a similar field as the real training set data; the triples are the Chinese and English in the bilingual parallel corpus, and the machine translation; Step 5: Combine the generated machine translation automatic post-editing pseudo-data and the real training set data, and pass the source texts in these data to the large language model to obtain the translations for each source text, use these obtained translations as auxiliary machine translation, and merge them with the original triples to get quadruples for subsequent model training; Step 6: Upsample the real training set data and merge it with the pseudo-data to form the final training set data; Step 7: Perform preprocessing of word segmentation and sub-word segmentation on the training set data, then splice the processed source text, machine translation, and auxiliary machine translation as the input on the encoder side of the mBART architecture, use the processed post-edited translation as the input on the decoder side of the mBART architecture, and train the mBART model to complete the automatic post-editing task.

2. The method according to claim 1, characterized in that: The specific content of step 1 is as follows: Collect publicly available Chinese-English parallel corpora, including CCMT, UN, ParaCrawl, Wiki-Matrix, WikiTitles, and News Commentary. After data cleaning operations such as removing traditional Chinese characters, duplicate removal, and empty line removal, 20 million bilingual parallel corpora are obtained.

3. The method according to claim 2, characterized in that: The specific content of step 2 is as follows: Randomly sample 50,000 from the collected bilingual parallel corpora as negative examples, and use 5,000 training sets of the CCMT2023 automatic post-editing task as positive examples to train a classifier based on mBERT. Learn the field information of the real training set data through the mBERT classifier to determine whether the field to which a piece of data belongs is the same as the field of the training set data.

4. The method according to claim 3, characterized in that: The specific content of step 3 is as follows: Score the fields of all 20 million bilingual parallel corpus data on the interval [0, 1] according to the trained mBERT classifier. 0 points means completely irrelevant to the field of the training set data, and 1 point means completely the same as the field of the training set data. After sorting in descending order of scores, select the top 200,000 that are most similar to the positive examples of the training set as the basis for subsequent pseudo-data generation.

5. The method according to claim 4, characterized in that: The specific content of step 4 is as follows: All the Chinese in the selected 200,000 bilingual parallel corpora are input into the trained Transformer-based neural machine translation model. The English translations output by the model are used as machine translation translations, the Chinese in the original bilingual parallel corpora are used as the original texts, and the English is used as the post-edited translations. Finally, 200,000 pieces of high-quality machine translation automatic post-editing pseudo-data that belong to a similar field as the real training set data are obtained.

6. The method according to claim 5, characterized in that: The specific steps of step 5 are as follows: Combine the 200,000 pieces of pseudo-data obtained in step 4 with 5,000 pieces of real training set data, and input the original texts in these data into ChatGPT. Use "Translate "{src}" into English:" as a prompt to obtain the translations of ChatGPT for each original text. Use these obtained translations as auxiliary machine translation translations, and combine them with the original triple to obtain a quadruple for subsequent model training.

7. The method according to claim 6, characterized in that: The specific steps of step 6 are as follows: Upsample the 5,000 training sets of the CCMT2023 automatic post-editing task by 10 times, and combine them with the obtained 200,000 pieces of pseudo-data as the training set.

8. The method according to claim 7, characterized in that: The specific steps of step 7 are as follows: Use sentencepiece to tokenize and sub-word segment all the data, and then concatenate the processed original texts, machine translation translations, and auxiliary machine translation translations together as the input to the encoder side of the mBART architecture, and use the processed post-edited translations as the input to the decoder side of the mBART architecture; Before starting the formal training, it is necessary to first call fairseq for data preprocessing, build a dictionary and binarize the training data, and then use fairseq to call the parameters of the pre-trained mbart-cc25 to initialize the parameters of the mBART model. Then, perform the training of the automatic post-editing task on the above training set to complete the automatic post-editing.

9. A computer-readable storage medium, on which a computer program is stored, characterized in that, This program is executed by a processor to implement the method as claimed in claims 1-8.

10. A computer device, including a memory and a processor, the memory stores a computer program, characterized in that: When the processor executes the computer program, it implements the method as claimed in claims 1-8.

Citation Information

Patent Citations

  • Method and device for generating editing model corpus after machine translation

    CN111144137A

  • Machine translation post-editing method and system

    CN112836528A