Machine Translation Automatic Post-Editing Method and Device

Through the combination of term dictionary processing and automatic post-editing model, the problem of inaccurate translation of term words in machine translation is solved, and the accuracy of translation is improved.

CN116029310BActive Publication Date: 2025-07-29IOL WUHAN INFORMATION TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202111243079.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-10-25
Publication Date
2025-07-29
Estimated Expiration
2041-10-25

AI Technical Summary

Technical Problem

In the prior art, machine translations of term words are inaccurate, and automatic post-editing methods are difficult to correct, resulting in a decline in translation quality.

Method used

By obtaining the source language and machine-translated text, using the term dictionary for term tag processing, the third and fourth texts are generated, and input into a pre-trained automatic post-editing model for accurate translation.

Benefits of technology

Improves the accuracy of translation of texts in source languages containing terminology and enhances translation quality.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116029310B_ABST
    Figure CN116029310B_ABST
Patent Text Reader

Abstract

The present invention provides a method and apparatus for automatic post-editing of machine translation. The method includes: obtaining a first text and a second text; performing term tagging on the first text and the second text according to a term dictionary to obtain a third text and a fourth text; and inputting the third text and the fourth text into a pre-trained automatic post-editing model to obtain a target language text of the automatic post-editing. By performing term tagging on the first text and the second text and inputting the obtained third text and fourth text into the automatic post-editing model to obtain the target language text, the present invention further improves the accuracy of translating a source language text containing term words.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of machine translation, and in particular, to a method and device for automatic post-editing of machine translation. Background Art

[0002] Machine translation is a process of converting a natural language text into another natural language text by a computer. Automatic post-editing refers to further editing and modifying the machine translation text output by machine translation through a computer, so as to improve the quality of machine translation and make it closer to human translation. The combined use of machine translation and automatic post-editing can greatly reduce the workload of translators. Translators only need to perform a small amount of editing on the basis of the translation output by automatic post-editing to achieve fast and high-quality translation.

[0003] However, although current deep learning-based machine translation can achieve good translation results in most cases, it still cannot accurately translate named entity words such as term words. At the same time, the existing automatic post-editing methods are also difficult to correctly correct the translation errors of term words. Summary of the Invention

[0004] The present invention provides a method and device for automatic post-editing of machine translation, which is used to solve the defect in the prior art that named entity words such as term words cannot be accurately translated, and realizes the improvement of the translation accuracy of the source language text containing term words.

[0005] In a first aspect, the present invention provides a method for automatic post-editing of machine translation, including:

[0006] Obtain a first text and a second text; the first text is a source language text, and the second text is the text obtained by machine translating the first text;

[0007] Perform term tagging on the first text and the second text according to a term dictionary to obtain a third text and a fourth text; the term dictionary is a dictionary composed of pairs of source language term words and target language term words;

[0008] Input the third text and the fourth text into a pre-trained automatic post-editing model to obtain an automatically post-edited target language text; the automatic post-editing model is used to output the target language text corresponding to the first text according to the input third text and fourth text.

[0009] According to the method for automatic post-editing of machine translation provided by the present invention, the performing term tagging on the first text and the second text according to a term dictionary to obtain a third text and a fourth text includes:

[0010] Compare the term dictionary and the first text to obtain a first term; the first term is a source language term included in the first text;

[0011] Perform a first tagging process on the first term in the first text to obtain the third text;

[0012] Compare the term dictionary and the first term to obtain a second term; the second term is the target language term corresponding to the first term;

[0013] Perform a second tagging process on the second term and add it to the second text to obtain the fourth text.

[0014] According to a machine translation automatic post-editing method provided by the present invention, the automatic post-editing model is trained through the following steps:

[0015] Obtain a training corpus of the automatic post-editing model; the training corpus includes a plurality of training corpora, and each training corpus in the training corpus includes a source language training corpus, a machine translation training corpus, and a target language training corpus;

[0016] Train the automatic post-editing model according to the training corpus until the loss function of the automatic post-editing model no longer decreases, then stop training to obtain the automatic post-editing model.

[0017] According to a machine translation automatic post-editing method provided by the present invention, the obtaining of the training corpus of the automatic post-editing model includes:

[0018] Screen the parallel corpus according to the term dictionary to obtain a term parallel corpus; the term parallel corpus includes a plurality of term parallel corpora, and each term parallel corpus includes a source language text corpus and a corresponding target language text corpus;

[0019] The source language text corpus of the term parallel corpus is machine-translated to obtain a machine translation text corpus;

[0020] Perform a third tagging process on the source language text corpus to obtain the source language training corpus;

[0021] Perform a fourth tagging process on the machine translation text corpus to obtain the machine translation training corpus;

[0022] The target language text corpus is the target language training corpus.

[0023] A machine translation automatic post-editing method provided by the present invention, the automatic post-editing model includes a first encoder, a second encoder, a decoder and an output layer, and the decoder includes a first attention layer, a second attention layer and a self-attention layer;

[0024] Training the automatic post-editing model according to the training corpus includes:

[0025] Inputting the machine translation training corpus into the first encoder to obtain a machine translation context expression matrix; inputting the machine translation context expression matrix into the first attention layer;

[0026] Inputting the source language training corpus into the second encoder to obtain a source language context expression matrix; inputting the source language context expression matrix into the second attention layer;

[0027] Inputting the target language training corpus into the decoder;

[0028] Judging the change situation of the loss function of the automatic post-editing model.

[0029] A machine translation automatic post-editing method provided by the present invention, the loss function of the automatic post-editing model is:

[0030] L(θ) = -logP(T|S,M,term;θ)

[0031] Where P represents probability, the output of the automatic post-editing model is the conditional probability expressed by the loss function, T is the target language training corpus, S is the source language training corpus, M is the machine translation training corpus, term is a word pair composed of source language term words and target language term words included in the source language training corpus, and θ is the parameter of the automatic post-editing model.

[0032] In a second aspect, the present invention also provides a machine translation automatic post-editing device, including:

[0033] A text acquisition unit for acquiring a first text and a second text; the first text is a source language text, and the second text is a text obtained by machine translating the first text;

[0034] A label processing unit for performing term label processing on the first text and the second text according to a term dictionary to obtain a third text and a fourth text; the term dictionary is a dictionary composed of word pairs of source language term words and target language term words;

[0035] An editing model unit is configured to input the third text and the fourth text into a pre-trained automatic post-editing model to obtain a target language text after automatic post-editing; the automatic post-editing model is configured to output the target language text corresponding to the first text according to the input third text and fourth text.

[0036] In a third aspect, the present invention further provides an electronic device, including a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the program, the steps of any one of the machine translation automatic post-editing methods in the first aspect described above are implemented.

[0037] In a fourth aspect, the present invention further provides a non-transitory computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the steps of any one of the machine translation automatic post-editing methods in the first aspect described above are implemented.

[0038] In a fifth aspect, the present invention further provides a computer program product, including a computer program. When the computer program is executed by a processor, the steps of any one of the machine translation automatic post-editing methods described above are implemented.

[0039] The machine translation automatic post-editing method and device provided by the present invention perform term tag processing on the first text and the second text, input the obtained third text and fourth text into the automatic post-editing model to obtain the target language text, and further improve the translation accuracy of the source language text containing term words. BRIEF DESCRIPTION OF THE DRAWINGS

[0040] In order to more clearly illustrate the technical solutions in the present invention or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the drawings in the following description are some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.

[0041] Figure 1 is a schematic flowchart of the machine translation automatic post-editing method provided by the present invention;

[0042] Figure 2 is a schematic diagram of the specific source language sentence translation and editing of the machine translation automatic post-editing method provided by the present invention;

[0043] Figure 3 is a schematic diagram of the structure of the automatic post-editing model provided by the present invention;

[0044] Figure 4 is a schematic diagram of the structure of the machine translation automatic post-editing device provided by the present invention;

[0045] Figure 5 It is a schematic structural diagram of the electronic device provided by the present invention. Specific embodiments

[0046] In the present invention, the term "and / or" describes the association relationship of associated objects, indicating that there can be three relationships. For example, A and / or B can represent: A exists alone, A and B exist simultaneously, and B exists alone. The character " / " generally represents an "or" relationship between the associated objects before and after.

[0047] In the present invention, the term "plurality" refers to two or more, and other quantifiers are similar.

[0048] To make the objectives, technical solutions and advantages of the present invention clearer, the technical solutions in the present invention will be clearly and completely described below with reference to the accompanying drawings in the present invention. Obviously, the described embodiments are part of the embodiments of the present invention, rather than all of the embodiments. Based on the embodiments in the present invention, all other embodiments obtained by those of ordinary skill in the art without making creative efforts belong to the scope of protection of the present invention.

[0049] The following combines Figures 1 - 3 Describe the machine translation automatic post-editing method of the present invention.

[0050] Figure 1 It is a schematic flow diagram of the machine translation automatic post-editing method provided by the present invention. As Figure 1 shown, the method includes:

[0051] Step 101, obtain a first text and a second text; the first text is a source language text, and the second text is the text obtained by machine translating the first text.

[0052] Specifically, a term is a set of appellations used to represent concepts in a specific subject field. A term is a conventional language symbol that expresses or defines a scientific concept through voice or text, and is a tool for the exchange of ideas and knowledge.

[0053] When it is necessary to translate a source language text containing terms into a target language text in another natural language, it is necessary to obtain the first text. The first text is a source language text, and the source language text is a source language text containing terms.

[0054] Machine translation, also known as automatic translation, is a process of using a computer to convert one natural language (source language text) into another natural language (target language text).

[0055] The first text is machine-translated to obtain a machine translation text. The second text is the machine translation text obtained by machine translating the first text, and the second text is obtained.

[0056] The automatic post-editing method for machine translation provided by the present invention is not specifically limited to machine translation, and the machine translation may be any machine translation model or machine translation system in the prior art.

[0057] Step 102 : performing term labeling on the first text and the second text according to a term dictionary to obtain a third text and a fourth text; the term dictionary is a dictionary consisting of word pairs of source language term words and target language term words.

[0058] Specifically, a terminology dictionary is a specialized dictionary that includes terms within a specific scope and provides definitions and explanations. The terminology dictionary involved in the present invention can be any terminology dictionary or a combination of any number of terminology dictionaries. A terminology dictionary is a dictionary consisting of word pairs composed of source language terminology words and target language terminology words.

[0059] The term dictionary and the first text are compared, and term label processing is performed on the first text to obtain a third text.

[0060] The term dictionary and the second text are compared, and term label processing is performed on the second text to obtain a fourth text.

[0061] Term labeling can be a marking and attention processing method for term words, the purpose of which is to distinguish term words from other words.

[0062] Step 103: Input the third text and the fourth text into a pre-trained automatic post-editing model to obtain an automatically post-edited target language text; the automatic post-editing model is used to output a target language text corresponding to the first text based on the input third text and fourth text.

[0063] Specifically, the automatic post-editing model is a pre-trained automatic post-editing model, which is used to output a target language text corresponding to the first text based on the input third text and fourth text.

[0064] The obtained third text and fourth text are input into a pre-trained automatic post-editing model, and the output result of the automatic post-editing model is the target language text, that is, the target language text of another natural language.

[0065] for example: Figure 2 : is a schematic diagram of a specific source language sentence translation and editing method of the machine translation automatic post-editing method provided by the present invention, such as Figure 2 As shown, the source language sentence is the first text, and the machine translation result is the second text. The source language sentence and the machine translation result are labeled and input into the automatic post-editing model to obtain the automatic post-editing result, which is the target language text.

[0066] As can be seen from the above embodiments, the automatic post-editing method for machine translation provided by the present invention processes the first text and the second text with term tags, inputs the obtained third text and fourth text into the automatic post-editing model, and obtains the target language text, further improving the accuracy of the translation of the source language text containing term words.

[0067] Optionally, the processing the first text and the second text with term tags according to the term dictionary to obtain a third text and a fourth text includes:

[0068] Comparing the term dictionary with the first text to obtain first term words; the first term words are source language term words included in the first text;

[0069] Performing first tagging on the first term words in the first text to obtain the third text;

[0070] Comparing the term dictionary with the first term words to obtain second term words; the second term words are target language term words corresponding to the first term words;

[0071] Performing second tagging on the second term words and adding them to the second text to obtain the fourth text.

[0072] Specifically, comparing the term dictionary with the first text, finding all the term words included in the first text that appear in the term dictionary, and determining the first term words. The first term words are source language term words included in the first text, and the first term words may include one or more term words. Performing first tagging on the first term words in the first text, and taking the first text after the first tagging as the third text.

[0073] Comparing the term dictionary with the first term words to determine the second term words corresponding to the first term words. The second term words are target language term words corresponding to the first term words. Performing second tagging on the second term words, and adding the second term words after the second tagging to the second text as the fourth text.

[0074] For example: Denote the first text as S, and all the words included in S as S = {S1, S2, S3…SK}, where S contains K words. Comparing the term dictionary, the term words included in S that appear in the term dictionary are S2 and S6, and the first term words are the combination of S2 and S6. Performing first tagging on S2 and S6 in S. The first tagging can be adding tags " <tag>” and "< / tag> " before and after the term words, or adding other symbol tags, as long as S2 and S6 can be distinguished from other words in S. The first text after the first tagging is S = {S1, <tag1> S2< / tag1> ,S3,S4,S5,<tag2> S6< / tag2> …SK}, S = {S1, <tag1> S2< / tag1> , S3, S4, S5, <tag2> S6< / tag2> …SK} as the third text.

[0075] Compare the terminology dictionary and the first terminology words: the combination of S2 and S6, and determine that the second terminology words corresponding to the first terminology words are W2 and W6. Perform second tagging on W2 and W6. The first tagging and the second tagging methods can be the same or different.

[0076] Here, the first tagging and the second tagging methods adopt the same tagging method. Add the tag " <tag>” and "< / tag> " before and after W2 and W6. The second terminology words after the second tagging are <tag1> W2< / tag1> and <tag2> W6< / tag2> .

[0077] The second text is denoted as M = {M1, M2, M3…MH}. M = {M1, M2, M3…MH} and S = {S1, S2, S3…SK} are corresponding. H represents that there are H words in the second text. The number of words in the first text S and the second text H can be the same or different.

[0078] Add <tag1> W2< / tag1> and <tag2> W6< / tag2> to the second text, that is, M = { <tag1> W2< / tag1> , <tag2> W6< / tag2> , M1, M2, M3…MH}. Here, add <tag1> W2< / tag1> and <tag2> W3< / tag2> to the front of the second text, or it can also be added to other positions of the second text. Take M = { <tag1> W2< / tag1> , <tag2> W6< / tag2> , M1, M2, M3…MH} as the fourth text.

[0079] As can be seen from the above embodiments, the automatic post-editing method for machine translation provided by the present invention obtains the first terminology words by comparing the terminology dictionary and the first text; performs first tagging on the first terminology words in the first text to obtain the third text; compares the terminology dictionary and the first terminology words to obtain the second terminology words; performs second tagging on the second terminology words and adds them to the second text to obtain the fourth text, determining the third text and the fourth text input to the automatic post-editing model, and further improving the translation accuracy of the source language text containing terminology words.

[0080] Optionally, the automatic post-editing model is trained through the following steps:

[0081] Obtain the training corpus of the automatic post-editing model; the training corpus includes a plurality of training corpora, and each training corpus in the training corpus includes a source language training corpus, a machine translation training corpus, and a target language training corpus;

[0082] Train the automatic post-editing model according to the training corpus until the loss function of the automatic post-editing model no longer decreases, then stop training to obtain the automatic post-editing model.

[0083] Specifically, the training corpus includes a plurality of training corpora. Each training corpus in the training corpus includes a source language training corpus, a machine translation training corpus, and a target language training corpus. The source language training corpus, the machine translation training corpus, and the target language training corpus form a training corpus, and the plurality of training corpora constitute the training corpus.

[0084] Input the multiple training corpora in the obtained training corpus into the automatic post-editing model for training. When the loss function of the automatic post-editing model no longer decreases, stop training to obtain the automatic post-editing model.

[0085] As can be seen from the above embodiments, the machine translation automatic post-editing method provided by the present invention further improves the translation accuracy of the source language text containing term words by obtaining the training corpus of the automatic post-editing model and training the automatic post-editing model according to the training corpus to obtain the automatic post-editing model.

[0086] Optionally, the obtaining of the training corpus of the automatic post-editing model includes:

[0087] Screen the parallel corpus according to the term dictionary to obtain a term parallel corpus; the term parallel corpus includes a plurality of term parallel corpora, and the term parallel corpus includes a source language text corpus and a corresponding target language text corpus;

[0088] The source language text corpus of the term parallel corpus is machine-translated to obtain a machine translation text corpus;

[0089] Perform third tagging processing on the source language text corpus to obtain the source language training corpus;

[0090] Perform fourth tagging processing on the machine translation text corpus to obtain the machine translation training corpus;

[0091] The target language text corpus is the target language training corpus.

[0092] Specifically, the parallel corpus includes a large number of parallel corpora. Screen the parallel corpus according to the term dictionary to obtain a corpus that only includes term words as the term parallel corpus.

[0093] The term parallel corpus includes multiple term parallel corpora, and each term parallel corpus includes a source language text corpus and a corresponding target language text corpus.

[0094] Perform third tagging processing on the source language text corpus to obtain a source language training corpus. The third tagging processing includes:

[0095] Determine the term words in the source language text corpus according to the term dictionary;

[0096] Perform first tagging processing on the term words in the source language text corpus to obtain a source language training corpus.

[0097] The source language text corpus of the term parallel corpus is machine-translated to obtain a machine-translated text corpus.

[0098] Perform fourth tagging processing on the machine-translated text corpus to obtain a machine-translated training corpus. The fourth tagging processing includes:

[0099] According to the term dictionary and the term words in the source language text corpus, determine the target term words corresponding to the term words in the source language text corpus;

[0100] Perform second tagging processing on the target term words corresponding to the term words in the source language text corpus and add them to the machine-translated text corpus to obtain a machine-translated training corpus.

[0101] The target language text corpus is the target language training corpus.

[0102] As can be seen from the above embodiments, the machine translation automatic post-editing method provided by the present invention screens the parallel corpus according to the term dictionary to obtain a term parallel corpus; the term parallel corpus includes multiple term parallel corpora, and the term parallel corpus includes a source language text corpus and a corresponding target language text corpus; the source language text corpus of the term parallel corpus is machine-translated to obtain a machine-translated text corpus; perform third tagging processing on the source language text corpus to obtain a source language training corpus; perform fourth tagging processing on the machine-translated text corpus to obtain a machine-translated training corpus; the target language text corpus is the target language training corpus, which further improves the translation accuracy of the source language text containing term words.

[0103] Optionally, the automatic post-editing model includes a first encoder, a second encoder, a decoder, and an output layer, and the decoder includes a first attention layer, a second attention layer, and a self-attention layer;

[0104] The training of the automatic post-editing model according to the training corpus set includes:

[0105] Input the machine translation training corpus into the first encoder to obtain a machine translation context expression matrix; input the machine translation context expression matrix into the first attention layer;

[0106] Input the source language training corpus into the second encoder to obtain a source language context expression matrix; input the source language context expression matrix into the second attention layer;

[0107] Input the target language training corpus into the decoder;

[0108] Judge the change of the loss function of the automatic post-editing model.

[0109] Specifically, Figure 3 is a schematic diagram of the structure of the automatic post-editing model provided by the present invention. As Figure 3 shown, the automatic post-editing model includes a first encoder, a second encoder, a decoder and an output layer. The decoder includes multiple levels, and the number of levels can be 6 or 12. Each level includes a first attention layer, a second attention layer and a self-attention layer.

[0110] Train the automatic post-editing model according to the training corpus set. Input the machine translation training corpus into the first encoder to obtain a machine translation context expression matrix; input the machine translation context expression matrix into the first attention layer.

[0111] Input the source language training corpus into the second encoder to obtain a source language context expression matrix; input the source language context expression matrix into the second attention layer.

[0112] Input the target language training corpus into the decoder.

[0113] The first attention layer and the second attention layer respectively perform attention calculation on the machine translation context expression matrix and the source language context expression matrix output by the first encoder and the second encoder. The attention calculation formula is as follows:

[0114]

[0115] where Q is the query sequence. For the first attention layer, both K and V are the machine translation context expression matrix output by the first encoder; for the second attention layer, both K and V are the source language context expression matrix output by the second encoder, and d k is the dimension size of each vector of the context matrix.

[0116] Judge the change of the loss function of the automatic post-editing model. When the loss function of the automatic post-editing model no longer decreases, stop training and obtain the automatic post-editing model.

[0117] As can be seen from the above embodiments, the automatic post-editing method for machine translation provided by the present invention inputs the machine translation training corpus into the first encoder to obtain a machine translation context expression matrix; inputs the machine translation context expression matrix into the first attention layer; inputs the source language training corpus into the second encoder to obtain a source language context expression matrix; inputs the source language context expression matrix into the second attention layer; inputs the target language training corpus into the decoder, further improving the translation accuracy of the target automatic post-editing model for source language texts containing term words.

[0118] Optionally, the loss function of the automatic post-editing model is:

[0119] L(θ) = -logP(T|S,M,term;θ)

[0120] where P represents probability, the output of the automatic post-editing model is the conditional probability expressed by the loss function, T is the target language training corpus, S is the source language training corpus, M is the machine translation training corpus, term is a word pair composed of source language term words and target language term words included in the source language training corpus, and θ is the parameter of the automatic post-editing model.

[0121] Specifically, the loss function of the automatic post-editing model is:

[0122] L(θ) = -logP(T|S,M,term;θ)

[0123] where P represents probability, the output of the automatic post-editing model is the conditional probability expressed by the loss function, T is the target language training corpus, S is the source language training corpus, M is the machine translation training corpus, term is a word pair composed of source language term words and target language term words included in the source language training corpus, and θ is the parameter of the automatic post-editing model.

[0124] As can be seen from the above embodiments, for the automatic post-editing method for machine translation provided by the present invention, the loss function of the automatic post-editing model improves the training speed of the automatic post-editing model and further improves the translation accuracy of the automatic post-editing model for source language texts containing term words.

[0125] Next, the automatic post-editing device for machine translation provided by the present invention will be described. The automatic post-editing device for machine translation described below can be correspondingly referred to the automatic post-editing method for machine translation described above.

[0126] Figure 4 is a schematic structural diagram of the automatic post-editing device for machine translation provided by the present invention. As Figure 4 shown, the automatic post-editing device for machine translation of the present invention includes:

[0127] A text acquisition unit 401 is configured to acquire a first text and a second text; the first text is a source language text, and the second text is a text obtained by machine translating the first text.

[0128] A label processing unit 402 is configured to perform term labeling processing on the first text and the second text according to a term dictionary, so as to obtain a third text and a fourth text; the term dictionary is a dictionary composed of word pairs formed by source language term words and target language term words.

[0129] An editing model unit 403 is configured to input the third text and the fourth text into a pre-trained automatic post-editing model to obtain an automatically post-edited target language text; the automatic post-editing model is configured to output the target language text corresponding to the first text according to the input third text and fourth text.

[0130] Figure 5 It is a schematic structural diagram of an electronic device provided by the present invention. As Figure 5 shown, the electronic device may include: a processor 510, a communications interface 520, a memory 530, and a communication bus 540. Among them, the processor 510, the communications interface 520, and the memory 530 communicate with each other through the communication bus 540. The processor 510 may call logic instructions in the memory 530 to execute a machine translation automatic post-editing method, and the method includes:

[0131] Acquire a first text and a second text; the first text is a source language text, and the second text is a text obtained by machine translating the first text.

[0132] Perform term labeling processing on the first text and the second text according to a term dictionary, so as to obtain a third text and a fourth text; the term dictionary is a dictionary composed of word pairs formed by source language term words and target language term words.

[0133] Input the third text and the fourth text into a pre-trained automatic post-editing model to obtain an automatically post-edited target language text; the automatic post-editing model is configured to output the target language text corresponding to the first text according to the input third text and fourth text.

[0134] In addition, when the logical instructions in the above-mentioned memory 530 can be implemented in the form of software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on such an understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or a part of the technical solution, can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes several instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in various embodiments of the present invention. The aforementioned storage medium includes: various media such as USB flash drives, mobile hard disks, read-only memories (ROM, Read-Only Memory), random access memories (RAM, Random Access Memory), magnetic disks, or optical discs that can store program codes.

[0135] On the other hand, the present invention also provides a computer program product. The computer program product includes a computer program that can be stored on a non-transitory computer-readable storage medium. When the computer program is executed by a processor, the computer can execute the machine translation post-editing method provided by the above-mentioned various methods. The method includes:

[0136] Obtain a first text and a second text; the first text is a source language text, and the second text is the text obtained by machine translating the first text;

[0137] Perform term tagging processing on the first text and the second text according to a term dictionary to obtain a third text and a fourth text; the term dictionary is a dictionary composed of pairs of source language terms and target language terms;

[0138] Input the third text and the fourth text into a pre-trained post-editing model to obtain a post-edited target language text; the post-editing model is used to output the target language text corresponding to the first text according to the input third text and fourth text.

[0139] On another aspect, the present invention also provides a non-transitory computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, it realizes the machine translation post-editing method provided by the above-mentioned various methods. The method includes:

[0140] Obtain a first text and a second text; the first text is a source language text, and the second text is the text obtained by machine translating the first text;

[0141] Perform term tagging on the first text and the second text according to a term dictionary to obtain a third text and a fourth text; the term dictionary is a dictionary composed of word pairs of source language terms and target language terms.

[0142] Input the third text and the fourth text into a pre-trained automatic post-editing model to obtain an automatically post-edited target language text; the automatic post-editing model is used to output the target language text corresponding to the first text according to the input third text and fourth text.

[0143] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed to multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this embodiment. Those of ordinary skill in the art can understand and implement it without creative effort.

[0144] Through the description of the above embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by means of software plus a necessary general hardware platform, and of course also by hardware. Based on this understanding, the above technical solution, in essence, or the part that contributes to the prior art can be embodied in the form of a software product, which can be stored in a computer-readable storage medium, such as ROM / RAM, magnetic disk, optical disc, etc., including several instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute the methods described in each embodiment or some parts of the embodiments.

[0145] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, and are not intended to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements for some of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.

Claims

1. An automatic post-editing method for machine translation, characterized in that Including: Obtain a first text and a second text; the first text is a source language text, and the second text is a text obtained by machine translating the first text; Perform term tagging on the first text and the second text according to a term dictionary to obtain a third text and a fourth text; the term dictionary is a dictionary composed of word pairs of source language terms and target language terms; Input the third text and the fourth text into a pre-trained automatic post-editing model to obtain an automatically post-edited target language text; the automatic post-editing model is trained based on a term parallel corpus filtered according to the term dictionary, and the term parallel corpus includes a source language text corpus and a corresponding target language text corpus; the automatic post-editing model is used to output the target language text corresponding to the first text according to the input third text and fourth text.

2. The automatic post-editing method for machine translation according to claim 1, characterized in that The performing term tagging on the first text and the second text according to the term dictionary to obtain a third text and a fourth text includes: Compare the term dictionary with the first text to obtain a first term; the first term is a source language term included in the first text; Perform a first tagging process on the first term in the first text to obtain the third text; Compare the term dictionary with the first term to obtain a second term; the second term is the target language term corresponding to the first term; Perform a second tagging process on the second term and add it to the second text to obtain the fourth text.

3. The machine translation automatic post-editing method according to claim 1, wherein The automatic post-editing model is trained through the following steps: Obtain a training corpus set for the automatic post-editing model; the training corpus set includes a plurality of training corpora, and each training corpus in the training corpus set includes a source language training corpus, a machine translation training corpus, and a target language training corpus; Train the automatic post-editing model according to the training corpus set until the loss function of the automatic post-editing model no longer decreases, then stop training to obtain the automatic post-editing model.

4. The automatic post-editing method for machine translation according to claim 3, characterized in that The obtaining the training corpus set for the automatic post-editing model includes: Filter a parallel corpus according to the term dictionary to obtain a term parallel corpus; the term parallel corpus includes a plurality of term parallel corpora, and each term parallel corpus includes a source language text corpus and a corresponding target language text corpus; The source language text corpus of the term parallel corpus is machine translated to obtain a machine translation text corpus; Perform a third tagging process on the source language text corpus to obtain the source language training corpus; Perform a fourth tagging process on the machine translation text corpus to obtain the machine translation training corpus; The target language text corpus is the target language training corpus.

5. The automatic post-editing method for machine translation according to claim 3 or 4, characterized in that The automatic post-editing model includes a first encoder, a second encoder, a decoder, and an output layer, and the decoder includes a first attention layer, a second attention layer, and a self-attention layer; The training the automatic post-editing model according to the training corpus set includes: Input the machine translation training corpus into the first encoder to obtain a machine translation context expression matrix; input the machine translation context expression matrix into the first attention layer; Input the source language training corpus into the second encoder to obtain a source language context expression matrix; input the source language context expression matrix into the second attention layer; Input the target language training corpus into the decoder; Judge the change of the loss function of the automatic post-editing model.

6. The method for automatic post-editing of machine translation according to claim 3, wherein The loss function of the automatic post-editing model is: L(θ)=-logP(T|S,M,term;θ) where P represents probability, and the output of the automatic post-editing model is the conditional probability expressed by the loss function. T is the target language training corpus, S is the source language training corpus, M is the machine translation training corpus, term is a word pair composed of source language term words and target language term words included in the source language training corpus, and θ is the parameter of the automatic post-editing model.

7. An automatic post-editing device for machine translation, characterized in that, It includes: A text acquisition unit for acquiring a first text and a second text; the first text is a source language text, and the second text is the text obtained by machine translating the first text; A label processing unit for performing term label processing on the first text and the second text according to a term dictionary to obtain a third text and a fourth text; the term dictionary is a dictionary composed of word pairs of source language term words and target language term words; An editing model unit for inputting the third text and the fourth text into a pre-trained automatic post-editing model to obtain an automatically post-edited target language text; the automatic post-editing model is trained based on a term parallel corpus screened by the term dictionary, and the term parallel corpus includes a source language text corpus and a corresponding target language text corpus; the automatic post-editing model is used to output the target language text corresponding to the first text according to the input third text and fourth text.

8. An electronic device, comprising a memory, a processor, and a computer program stored on the memory and executable on the processor, characterized in that, When the processor executes the program, it implements the steps of the machine translation automatic post-editing method according to any one of claims 1 to 6.

9. A non-transitory computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by the processor, it implements the steps of the machine translation automatic post-editing method according to any one of claims 1 to 6.

10. A computer program product, comprising a computer program, characterized in that, When the computer program is executed by the processor, it implements the steps of the machine translation automatic post-editing method according to any one of claims 1 to 6.

Citation Information

Patent Citations

  • Text translation model training method and text translation method and device

    CN110555213A

  • Statistical machine translation apparatus and method

    US20100088085A1