Method and device for generating translation evaluation training data, equipment and storage medium

By generating sample post-edited sentences and vocabulary quality labels using multiple reference translation sentences, the problem of lack of post-edited data in the training of machine translation evaluation models is solved, achieving efficient training and improved accuracy of evaluation models.

CN114254658BActive Publication Date: 2026-03-17SHANGHAI LIULISHUO INFORMATION TECH CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-12-14
Publication Date
2026-03-17

AI Technical Summary

Technical Problem

In the training of machine translation evaluation models, the lack of publicly available post-edited data in existing technologies leads to low training efficiency and makes it impossible to effectively construct vocabulary-level quality labels.

Method used

By acquiring multiple reference translation sentences, generating sample sentences and editing them, and using these sentences to obtain the lexical quality labels of the original sample sentences, a training set is built to train the translation evaluation model.

Benefits of technology

It improves the training efficiency and accuracy of translation evaluation models without the need for manual post-editing or labeling, and ensures that the models can still be effectively trained without publicly available post-edited data.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114254658B_ABST
    Figure CN114254658B_ABST
Patent Text Reader

Abstract

A method, apparatus, device, and storage medium for generating translation evaluation training data are disclosed. The method includes: acquiring a sample original sentence to be translated; acquiring a sample translated sentence corresponding to the sample original sentence; acquiring a reference translation set corresponding to the sample original sentence, including multiple reference translated sentences; selecting the reference translated sentence with the highest similarity to the sample translated sentence as the sample post-edited sentence; using the sample post-edited sentence and the sample translated sentence to obtain sample translation quality labels for each word in the sample original sentence; and establishing a training set based on the sample translation quality labels. This invention utilizes the reference translated sentences of the original sentence to generate the sample post-edited sentence. Therefore, even without publicly available post-edited data, it eliminates the need for manual post-editing or labeling, and still allows the acquisition of post-edited sentences and sample translation quality labels. This ensures the training efficiency of the translation evaluation model while obtaining training data for training the model.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of machine translation technology, and in particular to a method, apparatus, device and storage medium for generating translation evaluation training data. Background Technology

[0002] Machine translation (MT) refers to the technology of using computers to translate a source text (generally called the source language) into a target language (generally called the target language). This translation process is completed by machine translation systems, thus offering higher efficiency compared to human translation.

[0003] In machine translation, ensuring both efficiency and accuracy is crucial. Therefore, evaluating the quality of machine translation is paramount (for example, providing feedback to optimize the machine translation system). Quality assessment typically involves either human or automated evaluation. Automated evaluation, utilizing translation evaluation models to intelligently assess the quality of machine translation output, offers advantages such as speed, timely feedback, and consistency of evaluation results. Consequently, automated machine evaluation is becoming the mainstream approach.

[0004] Currently, there are still certain limitations in training evaluation models. Summary of the Invention

[0005] The problem solved by the embodiments of the present invention is to provide a method, apparatus, device and storage medium for generating translation evaluation training data, which can obtain training data for training translation evaluation models while ensuring the training efficiency of translation evaluation models.

[0006] To address the aforementioned problems, this invention provides a method for generating translation evaluation training data, comprising: acquiring one or more sample original sentences to be translated; acquiring sample translated sentences corresponding to the sample original sentences, wherein the sample translated sentences are obtained by translating the sample original sentences; acquiring a reference translation set corresponding to the sample original sentences, wherein the reference translation set includes multiple reference translated sentences; selecting, from the reference translation set, the reference translated sentence with the highest similarity to the sample translated sentence as the sample post-edited sentence; using the sample post-edited sentence and the sample translated sentence, acquiring sample translation quality labels for each word in the sample original sentences; and establishing a training set based on the sample translation quality labels, wherein the training set includes the sample original sentences, sample translated sentences, and sample translation quality labels, and the training set is used to train a translation evaluation model.

[0007] Accordingly, embodiments of the present invention also provide a device for generating translation evaluation training data, comprising: a first sentence acquisition module, configured to acquire one or more sample original sentences to be translated; a sample translation sentence acquisition module, configured to acquire sample translation sentences corresponding to the sample original sentences, wherein the sample translation sentences are obtained by translating the sample original sentences; a reference translation set acquisition module, configured to acquire a reference translation set corresponding to the sample original sentences, wherein the reference translation set includes multiple reference translation sentences; a second sentence acquisition module, configured to select the reference translation sentence with the highest similarity to the sample translation sentence from the reference translation set as the sample post-editing sentence; a label acquisition module, configured to acquire sample translation quality labels for each word in the sample original sentences using the sample post-editing sentence and the sample translation sentences; and a training set establishment module, configured to establish a training set based on the sample translation quality labels, wherein the training set includes the sample original sentences, sample translation sentences, and sample translation quality labels, and the training set is used to train a translation evaluation model.

[0008] Accordingly, embodiments of the present invention also provide an apparatus, including at least one memory and at least one processor, wherein the memory stores one or more computer instructions, wherein the one or more computer instructions are executed by the processor to implement the translation evaluation training data generation method described in embodiments of the present invention.

[0009] Accordingly, embodiments of the present invention also provide a storage medium storing one or more computer instructions, which are used to implement the translation evaluation training data generation method described in embodiments of the present invention.

[0010] Compared with the prior art, the technical solution of the embodiments of the present invention has the following advantages:

[0011] In this embodiment of the invention, multiple reference translation statements of the original statement are used to generate sample post-edited statements, and the sample post-edited statements are used to generate sample translation quality labels for each word in the original sample statement, which are then used to train the translation evaluation model. Therefore, this embodiment of the invention can obtain post-edited statements and sample translation quality labels without manual post-editing or labeling, even without publicly available post-edited data. This ensures the training efficiency of the translation evaluation model while obtaining training data for training the model. Attached Figure Description

[0012] Figure 1 This is a flowchart of an embodiment of the method for generating translation evaluation training data according to the present invention;

[0013] Figure 2This is a functional block diagram of an embodiment of the translation evaluation training data generation device of the present invention;

[0014] Figure 3 This is a hardware structure diagram of a device provided in an embodiment of the present invention. Detailed Implementation

[0015] As can be seen from the background technology, there are still certain limitations in training the evaluation model.

[0016] Specifically, the evaluation of machine translation quality includes three levels: lexical, sentence, and document. In the lexical level, the acquisition of quality labels for words during the training set construction relies on human post-edit (PE) statements on the machine translation results.

[0017] However, in some translation evaluation tasks (e.g., Chinese-to-English translation evaluation), there is a lack of publicly available post-edited data, making it impossible to construct specific quality labels and thus impossible to train the model. In other words, it is necessary to rely on publicly available post-edited sentences. If post-editing is done manually, or if quality labels for specific words are directly labeled manually, it requires a lot of effort, resulting in low model training efficiency.

[0018] To address the aforementioned technical problems, embodiments of the present invention provide a method for generating translation evaluation training data. (See reference...) Figure 1 The diagram shows a flowchart of an embodiment of the method for generating translation evaluation training data according to the present invention.

[0019] In this embodiment of the invention, the method for generating translation evaluation training data includes the following basic steps:

[0020] Step S1: Obtain one or more sample original sentences to be translated;

[0021] Step S2: Obtain the sample translation statement corresponding to the original sample statement, wherein the sample translation statement is obtained by translating the original sample statement;

[0022] Step S3: Obtain the reference translation set corresponding to the original sample statement, wherein the reference translation set includes multiple reference translation statements;

[0023] Step S4: From the reference translation set, select the reference translation statement with the highest similarity to the sample translation statement as the sample post-edit statement;

[0024] Step S5: Using the sample post-edited statement and the sample translated statement, obtain the sample translation quality label of each word in the original sample statement;

[0025] Step S6: Establish a training set based on the sample translation quality labels. The training set includes the original sample sentences, the translated sample sentences, and the sample translation quality labels. The training set is used to train the translation evaluation model.

[0026] In this embodiment of the invention, multiple reference translation statements of the original statement are used to generate sample post-edited statements, and the sample post-edited statements are used to generate sample translation quality labels for each word in the original sample statement, which are then used to train the translation evaluation model. Therefore, this embodiment of the invention can obtain post-edited statements and sample translation quality labels without manual post-editing or labeling, even without publicly available post-edited data. This ensures the training efficiency of the translation evaluation model while obtaining training data for training the model.

[0027] To make the above-mentioned objects, features and advantages of the embodiments of the present invention more apparent and understandable, the specific embodiments of the present invention will be described in detail below with reference to the accompanying drawings.

[0028] refer to Figure 1 Execute step S1 to obtain one or more sample original sentences to be translated.

[0029] The sample original sentences refer to the sentences that need to be translated. Machine translation involves translating sentences from a source language to a specified target language; for example, in a Chinese-to-English translation task, the source language is Chinese and the target language is English. During the training of the translation evaluation model, these sample original sentences are used as part of the training set.

[0030] In this embodiment, the sample original statement is a source language statement. For example, the sample original statement is "I like cats because cats are elegant." It should be noted that the double quotation marks here are only used to limit the scope of the example content and are not an essential part of representing the content of the sample original statement. Those skilled in the art can use other less confusing symbols to limit the scope of the sample original statement content; the double quotation marks used thereafter are the same as described above. It should also be noted that "vocabulary" can be considered a special kind of vocabulary.

[0031] The number of sample original sentences can be one or more. In one embodiment, the sample original sentences can come from a source language corpus, which includes one or more sample original sentences. In another embodiment, the sample original sentences can also come from different source language corpora. In other embodiments, the sample original sentences can also come from multiple independent sentences from different sources. The corpus generally refers to a paragraph or article consisting of multiple sentences.

[0032] Continue to refer to Figure 1 Execute step S2 to obtain the sample translation statement corresponding to the original sample statement. The sample translation statement is obtained by translating the original sample statement.

[0033] The sample translated sentences correspond to the sample original sentences. During the training of the translation evaluation model, the sample original sentences are also used as part of the training set.

[0034] Furthermore, after obtaining the post-edited sample statements, the post-edited sample statements and the translated sample statements are used to obtain the sample translation quality labels for each word in the original sample statements. Post-editing typically refers to the process of manually correcting (e.g., modifying and polishing) the output of the machine translation system to make the output an acceptable translation, thereby improving translation quality.

[0035] It should be noted that in this embodiment, post-editing also includes a process of correcting the human translation results (for example, in the application scenario of students answering questions, correcting the students' translation results). Therefore, in this embodiment, the sample translated sentences are obtained by machine translation or human translation of the original sample sentences.

[0036] As an example, the original sample statement is "I like cats because they are elegant.", and the corresponding sample translation is "I lik cats because they are elegant."

[0037] Continue to refer to Figure 1 Step S3 is executed to obtain a reference translation set corresponding to the original sample statement. The reference translation set includes multiple reference translation statements.

[0038] The reference translation statement refers to the sentence used for comparison with the sample translation statement. The reference translation statement is generally a translation of higher quality, that is, the reference translation statement meets the confidence level requirements.

[0039] Each original sample statement corresponds one-to-one with a reference translation set. The reference translation set includes multiple reference translation statements. Multiple reference translation statements in the same reference translation set are used as candidate post-edited statements corresponding to the original sample statements, so that the post-edited statements corresponding to the original sample statements can be selected from the reference translation set later.

[0040] In vocabulary-level translation evaluation tasks, the training of translation evaluation models relies on post-edited sentences. However, in some translation evaluation tasks, there is no publicly available post-edited data, making it impossible to train the translation evaluation model. In this embodiment, multiple reference translation sentences of the original sentence are used to generate sample post-edited sentences. Even without publicly available post-edited data, post-edited sentences can still be generated (i.e., automatic post-editing is achieved without relying on publicly available post-edited data), and post-edited sentences do not need to be obtained through manual annotation. Thus, training data for training the translation evaluation model can be constructed without publicly available post-edited data, thereby ensuring the training efficiency of the translation evaluation model while obtaining training data for training the model.

[0041] It should be noted that each reference translation set includes multiple reference translation sentences, thereby increasing the number of reference translation sentences. This allows for the editing of sentences after obtaining more accurate samples, which in turn improves the evaluation accuracy of the subsequently trained translation evaluation model.

[0042] It is understandable that when a sentence in a source language is translated into a sentence in a specified target language, a single word in the source language sentence will typically have multiple translations in the target language. This allows for the generation of multiple reference translated sentences based on the original sample sentence. For example, these multiple reference translated sentences might be T1, T2, T3, T4, and T5, forming a reference translation set S, which is {T1, T2, T3, T4, T5}. As an example, if the original sample sentence is "I like cats because they are very elegant.", then the multiple reference translated sentences would include "I like cats because they are very elegant.", "I love cats because they are very graceful.", "I love cats because they are very elegant.", "I like cats because they are very graceful.", and "I love cats because of their elegance.", forming the reference translation set S.

[0043] In this embodiment, a reference translation set corresponding to the original sample sentence is obtained through one or both of machine translation and human translation.

[0044] Taking obtaining a reference translation set through human translation as an example, for some translation question scenarios, the question's metadata is usually stored, meaning that reference translation sentences are pre-stored, and these stored reference translation sentences can be used directly. It's understandable that obtaining a reference translation set through human translation is not limited to translation question scenarios. It should be noted that obtaining a reference translation set through human translation helps improve the accuracy of the reference translation sentences.

[0045] When obtaining a reference translation set corresponding to the original sample statement through machine translation, the steps for obtaining the reference translation set include: obtaining multiple candidate machine translation sets from different machine translation tools, each candidate machine translation set including one or more candidate machine translation statements; performing a first screening process on the multiple candidate machine translation sets to select multiple candidate machine translation statements that meet a first preset condition, the first preset condition including confidence level; obtaining reference translation statements from the multiple candidate machine translation statements that meet the first preset condition, the reference translation statements constituting the reference translation set.

[0046] Machine translation can automatically obtain reference translation sets, which is more efficient than human translation. It should be noted that when obtaining reference translation sets corresponding to the original sample sentences using machine translation, the machine translation tool needs to have a high accuracy rate.

[0047] In this embodiment, there is a one-to-one correspondence between the machine translation tool and the candidate machine translation set. Therefore, when multiple machine translation tools are used, the number of candidate machine translation sets will also be multiple.

[0048] In this embodiment, multiple reference translated sentences from the original sentence are subsequently used to generate sample edited sentences. Therefore, to improve the accuracy of the sample edited sentences, the first preset condition includes confidence level. That is, candidate machine-translated sentences with high confidence levels need to be selected as reference translated sentences, thereby improving the evaluation accuracy of the subsequently obtained trained translation evaluation model. Higher confidence levels indicate higher translation quality.

[0049] Therefore, one or more preset methods are used to perform a first screening process on the multiple candidate machine translation sets. In this embodiment, the preset methods include: selecting common candidate machine translation sentences from multiple candidate machine translation sets; or, selecting candidate machine translation sentences from machine translation tools that meet accuracy requirements; or, using the confidence score of each word output by the machine translation tool to obtain the average confidence score of all words in the candidate machine translation sentences, which is used as the translation result score; and selecting candidate machine translation sentences whose translation result score is greater than or equal to a preset score threshold.

[0050] When multiple candidate machine translation sets share a common candidate machine translation statement, it indicates that the accuracy of that common candidate machine translation statement is relatively high.

[0051] Machine translation tools that meet accuracy requirements are also considered highly reliable. Therefore, selecting candidate machine translation sentences from such tools helps ensure their accuracy. Examples of machine translation tools that meet accuracy requirements include Google Translate.

[0052] By using the confidence score of each word output by the machine translation tool, the accuracy of candidate machine-translated sentences can be evaluated more easily and directly. It should be noted that, in practice, the preset score threshold can be set according to actual needs.

[0053] In this embodiment, obtaining reference translation statements from a plurality of candidate machine translation statements that meet a first preset condition includes: determining whether the number of candidate machine translation statements that meet the first preset condition meets a first quantity threshold condition, wherein the first quantity threshold condition includes: the number of candidate machine translation statements that meet the first preset condition is greater than or equal to a first preset quantity; if the number of candidate machine translation statements that meet the first preset condition does not meet the quantity threshold condition, selecting all candidate machine translation statements that meet the first preset condition as reference translation statements; if the number of candidate machine translation statements that meet the first preset condition meets the quantity threshold condition, selecting the first preset quantity of candidate machine translation statements with the highest confidence from the candidate machine translation statements that meet the first preset condition as reference translation statements.

[0054] The number of candidate machine-translated statements that meet the first preset condition is judged by a first quantity threshold to ensure that a sufficient number of reference translation statements are obtained, and that the selected reference translation statements have high confidence levels. This allows the reference translation statements with the highest similarity to the sample translation statements to be selected as the post-edited statements. Therefore, when the number of candidate machine-translated statements that meet the first preset condition is less than the first preset number, all candidate machine-translated statements that meet the first preset condition are selected as reference translation statements to ensure a sufficient number of reference translation statements. When the number of candidate machine-translated statements that meet the first preset condition is greater than or equal to the first preset number, the first preset number of candidate machine-translated statements with the highest confidence levels are selected as reference translation statements, thus simultaneously satisfying the requirements for both quantity and confidence levels.

[0055] It should be noted that the first preset number should not be too small or too large. If the first preset number is too small, it will not provide enough reference translation sentences, which is not conducive to ensuring the accuracy of the edited sentences after the sample, and may easily lead to the inability to find the reference translation sentence with the highest similarity to the sample translation sentence. If the first preset number is too large, on the one hand, it may easily lead to excessive data processing, resulting in low efficiency of model training; on the other hand, it may easily lead to the use of candidate machine translation sentences with low accuracy as reference translation sentences. Therefore, in this embodiment, the first preset number is 5 to 20. For example, the second preset number is 10 or 15.

[0056] It should be noted that, depending on the actual situation, any combination of the above preset methods can be selected, so that after the first screening process, the number of candidate machine translation sentences that meet the first preset conditions is sufficient, thereby obtaining a sufficient number of reference translation sentences with high confidence.

[0057] In this embodiment, after obtaining the reference translation set corresponding to the original sample statement, before editing the statement after obtaining the sample in the reference translation set, the method further includes: performing synonym expansion processing on the reference translation statement to obtain synonym statements of the reference translation statement; adding the synonym statements as new reference translation statements to the reference translation set.

[0058] Here, a synonym refers to a statement that expresses the same meaning as the reference translation statement in a different way. In other words, it involves converting the reference translation statement into a different expression while maintaining its semantics. It's understandable that the synonym and the reference translation statement share a similar language. By obtaining more synonyms, the number of reference translation statements in the reference translation set is increased, thus expanding the reference translation set. This makes it easier to subsequently obtain the reference translation statement with the highest similarity to the sample translation statement, thereby further improving the accuracy of post-sample editing.

[0059] Specifically, the synonym expansion process includes: obtaining a set of candidate synonym statements for each reference translation statement, each set of candidate synonym statements including one or more candidate synonym statements; removing candidate synonym statements that are the same as any reference translation statement from the multiple sets of candidate synonym statements; after removing candidate synonym statements that are the same as any reference translation statement, performing a second filtering process on the remaining candidate synonym statements in the set of candidate synonym statements to select multiple candidate synonym statements that meet a second preset condition, the second preset condition including confidence level; and obtaining synonym statements from the multiple candidate synonym statements that meet the second preset condition.

[0060] Within the set of candidate synonyms for any given reference translation, there may be candidate synonyms that are identical to other reference translations. Synonym expansion is used to convert reference translations into different expressions; therefore, it is necessary to first remove candidate synonyms that are identical to any reference translation. Furthermore, a second filtering process is performed on the remaining candidate synonyms in the set to ensure that the selected synonyms have a high accuracy rate.

[0061] Specifically, the remaining candidate synonyms in the candidate synonym set undergo a second filtering process, selecting multiple candidate synonyms that meet the second preset condition. This includes selecting candidate synonyms with a repetition rate from the multiple candidate synonym sets. In other words, the repetition rate is used to characterize confidence. If there are repeated candidate synonyms in the multiple candidate synonym sets, it can be characterized that the repeated candidate synonyms have a high confidence level.

[0062] In this embodiment, obtaining synonyms from a plurality of candidate synonyms that meet a second preset condition includes: determining whether the number of candidate synonyms that meet the second preset condition meets a second quantity threshold condition, wherein the second quantity threshold condition includes: the number of candidate synonyms that meet the second preset condition is greater than or equal to a second preset quantity; if the number of candidate synonyms that meet the second preset condition does not meet the second quantity threshold condition, selecting all candidate synonyms that meet the second preset condition as synonyms; if the number of candidate synonyms that meet the second preset condition meets the second quantity threshold condition, selecting the top second preset quantity of candidate synonyms with the highest repetition rate from the candidate synonyms that meet the second preset condition as synonyms.

[0063] The number of candidate synonyms that meet the second preset condition is determined by a second quantity threshold condition to ensure a sufficient number of synonyms are obtained, and that the selected synonyms have high confidence levels, thereby improving the accuracy of the reference translation. This allows for the selection of the reference translation with the highest similarity to the sample translation as the post-edited sample. Therefore, when the number of candidate synonyms that meet the second preset condition is less than the second preset number, all candidate synonyms that meet the second preset condition are selected as synonyms to ensure a sufficient number. When the number of candidate synonyms that meet the second preset condition is greater than or equal to the second preset number, the top second preset number of candidate synonyms with the highest confidence levels are selected as synonyms, thus simultaneously satisfying the requirements for both quantity and confidence.

[0064] It should be noted that the number of second presets should not be too few or too many. If the number of second presets is too few, not enough synonyms will be provided, resulting in poor performance in synonym expansion of the reference translation. If the number of second presets is too many, on the one hand, it can easily lead to excessive data processing, resulting in low model training efficiency; on the other hand, it can easily lead to the selection of candidate synonyms with low accuracy as synonyms. Therefore, in this embodiment, the number of second presets is 5 to 20. For example, the number of second presets is 10 or 15.

[0065] In this embodiment, a synonym transcription system is used to obtain a synonymous translation for each reference translation. The synonym transcription system has a transcription model that outputs a statement with the same or similar meaning after receiving the input statement. This system improves the efficiency of obtaining synonymous translations.

[0066] Continue to refer to Figure 1 Then, in step S4, the reference translation statement with the highest similarity to the sample translation statement is selected from the reference translation set as the sample post-edit statement.

[0067] By selecting the reference translation with the highest similarity as the sample post-editing statement, the accuracy of the sample post-editing statement is improved.

[0068] In this embodiment, the reference translation statement with the smallest edit distance from the sample translation statement is selected as the post-sample edit statement. A smaller edit distance indicates a higher similarity between the two sentences. Edit distance quantifies the similarity between the sample translation statement and the reference translation statement, making it easier to select the reference translation statement with the highest similarity from multiple reference translation statements.

[0069] Specifically, the edit distance between the sample translated statement and each reference translated statement is obtained. After obtaining multiple edit distances, the reference translated statement corresponding to the smallest edit distance is selected as the sample translated statement for editing. The edit distance refers to the number of operations required for the sample translated statement to be identical to the reference translated statement. One operation includes inserting a word, deleting a word, or replacing a word.

[0070] For example, if the sample translation is "I lik cats because they are very elegant.", then the edited version of the sample selected from the multiple reference translations is "I like cats because they are very elegant." Only one operation is needed: replacing "lik" with "like".

[0071] Continue to refer to Figure 1Then, in step S5, the sample translation quality labels of each word in the original sample sentence are obtained using the sample post-edited sentence and the sample translated sentence.

[0072] Translation quality labels are also used as part of the training set during the training of the translation evaluation model.

[0073] The sample post-edited sentences are obtained using multiple reference translation sentences of the original sentences. Therefore, by using the sample post-edited sentences to generate sample translation quality labels for each word in the original sample sentences, this embodiment can obtain the post-edited sentences and sample translation quality labels without the need for manual post-editing or labeling, even without publicly available post-edited data. This ensures the training efficiency of the translation evaluation model while obtaining training data for training the model.

[0074] In this embodiment, obtaining sample translation quality labels for each word in the original sample sentence using post-edited and translated sample sentences includes: performing a matching degree detection on the translated and post-edited sample sentences, and adding confidence labels to each word in the post-edited sample sentence based on the matching degree detection results. The confidence labels are used to indicate whether the translation quality is acceptable or unacceptable. The words in the original and post-edited sample sentences are then aligned, and confidence labels are added to the corresponding words in the original sample sentence. For words in the original sample sentence that do not have a corresponding relationship, confidence labels indicating acceptable translation quality are added. The confidence labels added to the original sample sentence serve as the sample translation quality labels.

[0075] By performing a matching degree test and adding confidence labels to each word, the translation quality of each word can be accurately determined, thus achieving translation evaluation at the word level. For example, the confidence label "OK" indicates that the translation quality is acceptable, and the confidence label "BAD" indicates that the translation quality is unacceptable.

[0076] Specifically, the matching degree detection of the sample translated statement and the sample post-edited statement includes: obtaining the correspondence between words in the sample translated statement and the sample post-edited statement; and performing matching degree detection on words with corresponding relationships.

[0077] In this embodiment, the principle of minimum edit distance is used to obtain the correspondence between words in the sample translated statement and the sample edited statement. For example, the sample translated statement is "I lik cats because they are very elegant.", and the sample edited statement selected from the above multiple reference translated statements is "I like cats because they are very elegant." According to the principle of minimum edit distance, only one operation is needed to replace "lik" with "like". Therefore, "lik" and "like" have a correspondence.

[0078] Accordingly, in the step of adding confidence labels to each word in the post-edited sentence, the confidence labels for each word (including punctuation) in the post-edited sentence are OK BAD OK OK OK OK OK OK. It should be noted that punctuation can be considered a special type of word.

[0079] Word alignment is a natural language processing technique used to identify the correspondence between words in two languages. In other words, when given a set of sentences to be translated, word alignment is automatically generated to obtain the correspondence between the words. Specifically, a common representation is i→j, which maps the target word at position i to the source word at position j. Here, the target word is the word in the edited version of the sample sentence, and the source word is the word in the original sample sentence.

[0080] It should be noted that in the original sample sentences, there are certain words that do not need to be translated directly. In such cases, during word alignment, there may be words in the original sample sentences that do not have a direct correspondence. This lack of correspondence is not caused by poor translation quality. Therefore, confidence labels indicating acceptable translation quality are added to these uncorresponding words in the original sample sentences. For example, the confidence label "OK" is added to these uncorresponding words.

[0081] If there is a correspondence between words in the original sample statement and words in the edited sample statement, then the words in the original sample statement are given the same confidence label as the corresponding words in the edited sample statement. For example, if the confidence label of any word in the edited sample statement is "OK", then the corresponding word in the original sample statement is also given the confidence label "OK". Similarly, if the confidence label of any word in the edited sample statement is "BAD", then the corresponding word in the original sample statement is also given the confidence label "BAD".

[0082] For example, the original sample sentence is "I like cats because they are very elegant.", the translated sample sentence is "I likcats because they are very elegant.", and the edited sample sentence is "I like cats because they are very elegant.". After obtaining the sample translation quality labels for each word in the original sample sentence, the confidence labels for each word (including punctuation) in the edited sample sentence are OK BAD OK OK OK OK OK OK. Correspondingly, the word correspondences between the original and edited sample sentences are: I → I, like → like, cats → cats, because → because, they → they, very → very, elegant → elegant, . → .. Furthermore, the commas in the original sample sentence have no corresponding correspondence; therefore, the confidence labels for each word (including punctuation) in the original sample sentence are: OK BAD OK OK OK OK OK OK OK.

[0083] It should be noted that the confidence labels are not limited to using "OK" and "BAD" for differentiation. In other embodiments, other labeling methods can also be used, for example, using the number "1" to indicate that the translation quality is acceptable and using the number "0" to indicate that the translation quality is unacceptable.

[0084] Continue to refer to Figure 1 Execute step S6 to build a training set based on the sample translation quality labels. The training set includes the original sample sentences, the translated sample sentences, and the sample translation quality labels. The training set is used to train the translation evaluation model.

[0085] Based on the foregoing description, even without publicly available post-edited data, this embodiment can generate sample post-edited sentences using multiple reference translated sentences of the original sentence. This allows the generation of sample translation quality labels for each word in the original sample sentence using the sample post-edited sentences, thereby establishing a training set for training the translation evaluation model. This ensures the training efficiency of the translation evaluation model while obtaining training data for training the model.

[0086] Accordingly, embodiments of the present invention also provide a device for generating translation evaluation training data. Figure 2 This is a functional block diagram of an embodiment of the translation evaluation training data generation device of the present invention.

[0087] The device for generating translation evaluation training data includes: a first sentence acquisition module 10, used to acquire one or more sample original sentences to be translated; a sample translation sentence acquisition module 20, used to acquire sample translation sentences corresponding to the sample original sentences, wherein the sample translation sentences are obtained by translating the sample original sentences; a reference translation set acquisition module 30, used to acquire a reference translation set corresponding to the sample original sentences, wherein the reference translation set includes multiple reference translation sentences; a second sentence acquisition module 40, used to select the reference translation sentence with the highest similarity to the sample translation sentence from the reference translation set as the sample post-edited sentence; a label acquisition module 50, used to acquire sample translation quality labels for each word in the sample original sentences using the sample post-edited sentences and sample translation sentences; and a training set establishment module 60, used to establish a training set based on the sample translation quality labels, wherein the training set includes sample original sentences, sample translation sentences, and sample translation quality labels, and the training set is used to train the translation evaluation model.

[0088] The device for generating translation evaluation training data includes a second sentence acquisition module 40 and a label acquisition module 50. It uses multiple reference translation sentences of the original sentence to generate sample post-edited sentences, and uses the sample post-edited sentences to generate sample translation quality labels for each word in the original sample sentence, which are then used to train the translation evaluation model. Therefore, this embodiment can obtain post-edited sentences and sample translation quality labels without manual post-editing or labeling, even without publicly available post-edited data. This ensures the training efficiency of the translation evaluation model while obtaining training data for training the model.

[0089] The original sample sentences refer to the sentences that need to be translated; that is, the original sample sentences are sentences in the source language. During the training of the translation evaluation model, the original sample sentences are used as part of the training set. For a detailed description of the original sample sentences, please refer to the corresponding descriptions in the foregoing embodiments, which will not be repeated here.

[0090] The sample translated sentences are obtained by translating the original sample sentences. The sample translated sentences correspond to the original sample sentences and are also used as part of the training set during the training of the translation evaluation model. Furthermore, the sample translated sentences are used as input to the label acquisition module 50, thereby enabling the label acquisition module 50 to obtain sample translation quality labels for each word in the original sample sentences.

[0091] In this embodiment, the sample translated statements are obtained through machine translation or human translation. That is, the sample translated statements can be the result of human translation or machine translation.

[0092] The reference translation set acquisition module 30 is used to obtain multiple reference translation statements corresponding to the original sample statement. Reference translation statements are sentences used for comparison with the sample translation statement; they are generally high-quality translations, meaning they meet the confidence level requirements.

[0093] The original sample statement corresponds one-to-one with the reference translation set. The reference translation set includes multiple reference translation statements. Multiple reference translation statements in the same reference translation set are used as candidate sample post-editing statements corresponding to the original sample statement, so that a suitable reference translation statement is selected from multiple reference translation statements as the sample post-editing statement.

[0094] In vocabulary-level translation evaluation tasks, the training of translation evaluation models relies on post-edited sentences. However, in some translation evaluation tasks, there is no publicly available post-edited data, making it impossible to train the translation evaluation model. In this embodiment, multiple reference translation sentences of the original sentence are used to generate sample post-edited sentences. Even without publicly available post-edited data, post-edited sentences can still be automatically generated without manual intervention. This ensures the training efficiency of the translation evaluation model while obtaining training data for training the model.

[0095] It should be noted that each reference translation set includes multiple reference translation sentences, thereby increasing the number of reference translation sentences. This allows for the editing of sentences after obtaining more accurate samples, which in turn improves the evaluation accuracy of the subsequently trained translation evaluation model.

[0096] It is understandable that when a sentence in a source language is translated into a sentence in another specified target language, a word in the source language sentence will usually have multiple translation results in the target language, thus enabling the acquisition of multiple reference translation sentences based on the original sample sentence.

[0097] In this embodiment, the reference translation statements in the reference translation set are one or both of machine translation results and human translation results.

[0098] Taking the example of a human-translated reference sentence, for some translation question scenarios, the question's metadata is usually stored, meaning reference sentences are pre-stored, and multiple stored reference sentences can be used directly. It's understandable that the situation of human-translated reference sentences is not limited to translation question scenarios. It should be noted that human-translated reference sentences help improve their accuracy.

[0099] When the reference translation statement is a machine translation result, the reference translation set acquisition module 30 includes: a candidate machine translation set acquisition unit, used to acquire multiple candidate machine translation sets from different machine translation tools, each candidate machine translation set including one or more candidate machine translation statements; a first filtering unit, used to perform a first filtering process on the multiple candidate machine translation sets, selecting multiple candidate machine translation statements that meet a first preset condition, the first preset condition including confidence level; and a reference translation set acquisition unit, used to acquire reference translation statements from the multiple candidate machine translation statements that meet the first preset condition, the reference translation statements constituting a reference translation set.

[0100] When the reference translation is a machine translation result, the reference translation set acquisition module can automatically acquire the reference translation set, which helps to improve the efficiency of acquiring the reference translation set.

[0101] In this embodiment, there is a one-to-one correspondence between the machine translation tool and the candidate machine translation set. Therefore, when multiple machine translation tools are used, the number of candidate machine translation sets will also be multiple.

[0102] In this embodiment, to improve the accuracy of the edited sentences after sampling, the first preset condition includes confidence level. That is, candidate machine-translated sentences with high confidence levels need to be selected as reference translated sentences, thereby improving the evaluation accuracy of the subsequently obtained trained translation evaluation model. Higher confidence levels indicate higher translation quality.

[0103] Therefore, the first screening unit uses one or more preset methods to perform a first screening process on multiple candidate machine translation sets. In this embodiment, the preset methods include: selecting common candidate machine translation statements from multiple candidate machine translation sets as reference translation statements; or, selecting candidate machine translation statements from machine translation tools that meet accuracy requirements as reference translation statements; or, using the confidence score of each word output by the machine translation tool to obtain the average confidence score of all words in the candidate machine translation statements, using it as the translation result score, and selecting candidate machine translation statements whose translation result score is greater than or equal to a preset score threshold as reference translation statements.

[0104] When multiple candidate machine translation sets share a common candidate machine translation statement, it indicates that the accuracy of that common candidate machine translation statement is relatively high.

[0105] Machine translation tools that meet accuracy requirements are also considered highly reliable. Therefore, selecting candidate machine translation sentences from such tools helps ensure their accuracy. Examples of machine translation tools that meet accuracy requirements include Google Translate.

[0106] By using the confidence score of each word output by the machine translation tool, the accuracy of candidate machine-translated sentences can be evaluated more easily and directly. It should be noted that, in practice, the preset score threshold can be set according to actual needs.

[0107] In this embodiment, the reference translation set acquisition unit includes: a first judgment subunit, used to judge whether the number of candidate machine translation statements that meet the first preset condition meets the first quantity threshold condition, the first quantity threshold condition including: the number of candidate machine translation statements that meet the first preset condition is greater than or equal to the first preset quantity; and a first selection subunit, used to select all candidate machine translation statements that meet the first preset condition as reference translation statements when the number of candidate machine translation statements that meet the first preset condition does not meet the quantity threshold condition, and to select the first preset quantity of candidate machine translation statements with the highest confidence from the candidate machine translation statements that meet the first preset condition as reference translation statements when the number of candidate machine translation statements that meet the first preset condition meets the quantity threshold condition.

[0108] In this embodiment, the first preset quantity is 5 to 20.

[0109] It should be noted that, depending on the actual situation, the first screening unit can also select any combination of the above preset methods, so that after the first screening process, the number of candidate machine translation sentences that meet the first preset conditions is sufficient, thereby obtaining a sufficient number of reference translation sentences with high confidence.

[0110] The device for generating translation evaluation training data further includes a synonym expansion module 35 disposed between the reference translation set acquisition module 30 and the second sentence acquisition module 40. This module is used to perform synonym expansion processing on the reference translation sentences, acquire synonym sentences of the reference translation sentences, and add the synonym sentences as new reference translation sentences to the reference translation set.

[0111] By obtaining more synonyms, the number of reference translations in the reference translation set can be increased, thereby expanding the reference translation set. This makes it easier to obtain the reference translation with the highest similarity to the sample translation, which in turn helps to further improve the accuracy of the post-sample editing.

[0112] In this embodiment, the synonym expansion module 35 includes: a candidate synonym statement set acquisition unit, used to acquire a candidate synonym statement set for each reference translation statement, wherein each candidate synonym statement set includes one or more candidate synonym statements; a second filtering unit, used to remove candidate synonym statements that are the same as any reference translation statement from multiple candidate synonym statement sets; a third filtering unit, used to perform a second filtering process on the remaining candidate synonym statements in the candidate synonym statement set after removing candidate synonym statements that are the same as any reference translation statement, and select multiple candidate synonym statements that meet a second preset condition, wherein the second preset condition includes confidence level; and a synonym statement acquisition unit, used to acquire synonym statements from multiple candidate synonym statements that meet the second preset condition.

[0113] Specifically, the third filtering unit is used to select candidate synonyms with a repetition rate from multiple sets of candidate synonyms. In other words, the repetition rate is used to represent the confidence level.

[0114] In this embodiment, the third filtering unit includes: a first judgment subunit, used to judge whether the number of candidate synonyms that meet the second preset condition meets the second quantity threshold condition, the second quantity threshold condition including: the number of candidate synonyms that meet the second preset condition is greater than or equal to the second preset quantity; and a first selection subunit, used to select all candidate synonyms that meet the second preset condition as synonyms when the number of candidate synonyms that meet the second preset condition does not meet the second quantity threshold condition, and to select the top second preset quantity of candidate synonyms with the highest repetition rate from the candidate synonyms that meet the second preset condition as synonyms when the number of candidate synonyms that meet the second preset condition meets the second quantity threshold condition.

[0115] In this embodiment, the second preset quantity is 5 to 20.

[0116] In this embodiment, the candidate synonym set acquisition unit is a synonym transcription system. The synonym transcription system has a transcription model used to output statements with the same or similar meanings after obtaining the input statement. The synonym transcription system improves the efficiency of acquiring candidate synonym statements.

[0117] The second sentence acquisition module 40 selects the reference translation sentence with the highest similarity as the sample post-editing sentence, thereby improving the accuracy of the sample post-editing sentence.

[0118] In this embodiment, the second sentence acquisition module 40 is used to select the reference translation sentence with the smallest edit distance from the sample translation sentence as the sample post-edit sentence. The smaller the edit distance, the higher the similarity between the two sentences. By using the edit distance, the similarity between the sample translation sentence and the reference translation sentence can be quantified, making it easier to select the reference translation sentence with the highest similarity from multiple reference translation sentences.

[0119] Specifically, the second statement acquisition module 40 includes: an edit distance acquisition unit, used to acquire the edit distance between the sample translated statement and each reference translated statement; and a fourth filtering unit, used to select the reference translated statement corresponding to the smallest edit distance from the plurality of edit distances as the sample post-edit statement. Here, edit distance refers to: how many operations the sample translated statement needs to perform to be the same as the reference translated statement, wherein one operation includes: inserting a word, deleting a word, or replacing a word.

[0120] The label acquisition module 50 is used to obtain the sample translation quality labels of each word in the original sample sentence using the sample post-edited sentence and the sample translated sentence. During the training of the translation evaluation model, the translation quality labels are also used as part of the training set.

[0121] The sample post-edited sentences are obtained using multiple reference translation sentences of the original sentences. Therefore, by using the sample post-edited sentences to generate sample translation quality labels for each word in the original sample sentences, this embodiment can obtain the post-edited sentences and automatically obtain the sample translation quality labels even without publicly available post-edited data, without the need for manual post-editing or labeling. This ensures the training efficiency of the translation evaluation model while obtaining training data for training the model.

[0122] In this embodiment, the tag acquisition module 50 includes: a matching degree detection unit, used to perform matching degree detection on the sample translated sentence and the sample post-edited sentence, and add confidence tags to each word in the sample post-edited sentence according to the matching degree detection result, the confidence tags being used to indicate whether the translation quality is qualified or unqualified; and a labeling unit, used to align the words in the sample original sentence and the sample post-edited sentence, add confidence tags to the corresponding words in the sample original sentence, and add confidence tags to words in the sample original sentence that have no corresponding relationship to indicate that the translation quality is qualified, the confidence tags added to the sample original sentence serving as sample translation quality tags.

[0123] By performing a matching degree test and adding confidence labels to each word, the translation quality of each word can be accurately determined, thus achieving translation evaluation at the word level. For example, the confidence label "OK" indicates that the translation quality is acceptable, and the confidence label "BAD" indicates that the translation quality is unacceptable.

[0124] In this embodiment, the matching degree detection unit is used to obtain the correspondence between words in the sample translated statement and the sample post-edited statement, and to perform matching degree detection on words with corresponding relationships. Specifically, the matching degree detection unit uses the minimum edit distance principle to obtain the correspondence between words in the sample translated statement and the sample post-edited statement. That is, the words with the smallest edit distance in the sample translated statement and the sample post-edited statement have a corresponding relationship.

[0125] Word alignment is a natural language processing technique used to identify the correspondence between words in two languages. In other words, when given a set of sentences to be translated, word alignment is automatically generated to obtain the correspondence between the words. Specifically, a common representation is i→j, which maps the target word at position i to the source word at position j. Here, the target word is the word in the edited version of the sample sentence, and the source word is the word in the original sample sentence.

[0126] It should be noted that in the original sample sentences, there are certain words that do not need to be translated directly. In such cases, during word alignment, there may be words in the original sample sentences that do not have a direct correspondence. This lack of correspondence is not caused by poor translation quality. Therefore, confidence labels are added to these words in the original sample sentences to indicate that the translation quality is acceptable. For example, the confidence label "OK" is added to these words.

[0127] If there is a correspondence between words in the original sample statement and words in the edited sample statement, then the words in the original sample statement are given the same confidence label as the corresponding words in the edited sample statement. For example, if the confidence label of any word in the edited sample statement is "OK", then the corresponding word in the original sample statement is also given the confidence label "OK". Similarly, if the confidence label of any word in the edited sample statement is "BAD", then the corresponding word in the original sample statement is also given the confidence label "BAD".

[0128] It should be noted that the confidence labels are not limited to using "OK" and "BAD" for differentiation. In other embodiments, other labeling methods can also be used, for example, using the number "1" to indicate that the translation quality is acceptable and using the number "0" to indicate that the translation quality is unacceptable.

[0129] The training set building module 60 builds a training set based on the sample translation quality labels. Based on the foregoing description, even without publicly available post-edited data, this embodiment can generate sample post-edited sentences using multiple reference translated sentences of the original sentence. These sample post-edited sentences are then used to generate sample translation quality labels for each word in the original sample sentence, thereby building a training set for training the translation evaluation model. This ensures both the acquisition of training data for the translation evaluation model and the efficiency of its training.

[0130] This invention also provides a device that can implement the translation evaluation training data generation method provided in this invention by loading the above-described translation evaluation training data generation method in the form of a program.

[0131] refer to Figure 3 The diagram illustrates the hardware structure of a device according to an embodiment of the present invention. The device in this embodiment includes: at least one processor 01, at least one communication interface 02, at least one memory 03, and at least one communication bus 04.

[0132] In this embodiment, the number of processor 01, communication interface 02, memory 03 and communication bus 04 is at least one, and the processor 01, communication interface 02 and memory 03 communicate with each other through the communication bus 04.

[0133] The communication interface 02 can be an interface of a communication module used for network communication, such as the interface of a GSM module.

[0134] The processor 01 may be a central processing unit (CPU), an application-specific integrated circuit (ASIC), or one or more integrated circuits configured to implement the detection method described in this embodiment.

[0135] The memory 03 may include high-speed RAM memory, and may also include non-volatile memory, such as at least one disk storage device.

[0136] The memory 03 stores one or more computer instructions, which are executed by the processor 01 to implement the translation evaluation training data generation method provided in the foregoing embodiments.

[0137] It should be noted that the aforementioned terminal device may also include other devices (not shown) that may not be essential to understanding the content disclosed in the embodiments of the present invention; given that these other devices may not be essential for understanding the content disclosed in the embodiments of the present invention, the embodiments of the present invention will not describe them one by one.

[0138] This invention also provides a storage medium storing one or more computer instructions, which are used to implement the translation evaluation training data generation method provided in the foregoing embodiments.

[0139] In the translation evaluation training data generation method provided in this embodiment, multiple reference translation sentences of the original sentence are used to generate sample post-edited sentences, and the sample post-edited sentences are used to generate sample translation quality labels for each word in the original sample sentence, which are then used to train the translation evaluation model. Therefore, this embodiment of the invention can obtain post-edited sentences and sample translation quality labels without manual post-editing or labeling, even without publicly available post-edited data, thereby ensuring the training efficiency of the translation evaluation model while obtaining training data for training the model.

[0140] The embodiments of the present invention described above are combinations of elements and features of the present invention. Unless otherwise stated, the elements or features described are optional. Individual elements or features may be practiced without combination with other elements or features. Furthermore, embodiments of the present invention may be constructed by combining some elements and / or features. The order of operations described in the embodiments of the present invention may be rearranged. Some constructions of any embodiment may be included in another embodiment and may be replaced by corresponding constructions of another embodiment. It will be apparent to those skilled in the art that claims in the appended claims that are not expressly referenced to each other may be combined to form embodiments of the present invention, or may be included as new claims in amendments made after the filing of this application.

[0141] Embodiments of the present invention can be implemented by various means, such as hardware, firmware, software, or combinations thereof. In a hardware configuration, the method according to an exemplary embodiment of the present invention can be implemented by one or more application-specific integrated circuits (ASICs), digital signal processors (DSPs), digital signal processing devices (DSPDs), programmable logic devices (PLDs), field-programmable gate arrays (FPGAs), processors, controllers, microcontrollers, microprocessors, etc.

[0142] In firmware or software configuration, embodiments of the present invention can be implemented in the form of modules, processes, functions, etc. Software code can be stored in a memory unit and executed by a processor. The memory unit is located inside or outside the processor and can send data to and receive data from the processor via various known means.

[0143] The above description of the disclosed embodiments enables those skilled in the art to make or use the invention. Various modifications to these embodiments will be readily apparent to those skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of the invention. Therefore, the invention is not to be limited to the embodiments shown herein, but is accorded the widest scope consistent with the principles and novel features disclosed herein.

[0144] While the embodiments of the present invention have been disclosed above, the present invention is not limited thereto. Any person skilled in the art can make various modifications and alterations without departing from the spirit and scope of the present invention; therefore, the scope of protection of the present invention should be determined by the scope defined in the claims.

Claims

1. A method for generating training data for translation evaluation, characterized in that, The method comprises: obtaining one or more sample original sentences to be translated; obtaining sample translated sentences corresponding to the sample original sentences, the sample translated sentences being obtained by translating the sample original sentences; obtaining a reference translation set corresponding to the sample original sentences, the reference translation set comprising a plurality of reference translated sentences; wherein the reference translation set corresponding to the sample original sentences is obtained by machine translation; and wherein obtaining the reference translation set comprises: obtaining a plurality of candidate machine translation sets from different machine translation tools, each of the candidate machine translation sets comprising one or more candidate machine translated sentences; performing first screening processing on the plurality of candidate machine translation sets to select a plurality of candidate machine translated sentences that satisfy a first preset condition, the first preset condition comprising a confidence level; obtaining reference translated sentences from the plurality of candidate machine translated sentences that satisfy the first preset condition, the reference translated sentences constituting the reference translation set; selecting, from the reference translation set, a reference translated sentence that has the highest similarity to the sample translated sentence as a sample post-editing sentence; using the sample post-editing sentence and the sample translated sentence to obtain sample translation quality labels for each word in the sample original sentence; establishing a training set according to the sample translation quality labels, the training set comprising the sample original sentence, the sample translated sentence, and the sample translation quality labels, the training set being used to train a translation evaluation model.

2. The method of claim 1, wherein, The first screening processing on the plurality of candidate machine translation sets is performed using one or more of the following preset methods: selecting common candidate machine translated sentences from the plurality of candidate machine translation sets; or, selecting candidate machine translated sentences from machine translation tools that satisfy an accuracy requirement; or, using a confidence score of each word output by the machine translation tool to obtain an average value of confidence scores of all words in the candidate machine translated sentence as a translation result score; and selecting the candidate machine translated sentence whose translation result score is greater than or equal to a preset score threshold.

3. The method of claim 1, wherein the training data is generated by: Obtaining reference translated sentences from the plurality of candidate machine translated sentences that satisfy the first preset condition comprises: determining whether the number of candidate machine translated sentences that satisfy the first preset condition satisfies a first quantity threshold condition, the first quantity threshold condition comprising: the number of candidate machine translated sentences that satisfy the first preset condition being greater than or equal to a first preset number; in the case where the number of candidate machine translated sentences that satisfy the first preset condition does not satisfy the quantity threshold condition, selecting all candidate machine translated sentences that satisfy the first preset condition as reference translated sentences; in the case where the number of candidate machine translated sentences that satisfy the first preset condition satisfies the quantity threshold condition, selecting the first preset number of candidate machine translated sentences with the highest confidence level from the candidate machine translated sentences that satisfy the first preset condition as reference translated sentences.

4. The method of claim 3, wherein, The first preset number is 5 to 20.

5. The method of claim 1, wherein, Before selecting the reference translation sentence with the highest similarity to the sample translation sentence from the reference translation set as the sample post-editing sentence, the method further includes: performing synonym expansion processing on the reference translation sentence to obtain a synonym sentence of the reference translation sentence, and adding the synonym sentence as a new reference translation sentence to the reference translation set.

6. The method of claim 5, wherein, The synonym expansion processing on the reference translation sentence to obtain a synonym sentence of the reference translation sentence includes: obtaining a candidate synonym sentence set of each reference translation sentence, each candidate synonym sentence set including one or more candidate synonym sentences; removing any candidate synonym sentence that is the same as any reference translation sentence from the plurality of candidate synonym sentence sets; after removing any candidate synonym sentence that is the same as any reference translation sentence, performing second screening processing on the remaining candidate synonym sentences in the candidate synonym sentence set to select a plurality of candidate synonym sentences that satisfy a second preset condition, the second preset condition including a confidence level; obtaining a synonym sentence from the plurality of candidate synonym sentences that satisfy the second preset condition.

7. The method of claim 6, wherein the translation evaluation training data is generated by: The second screening processing on the remaining candidate synonym sentences in the candidate synonym sentence set to select a plurality of candidate synonym sentences that satisfy a second preset condition includes: selecting a candidate synonym sentence with a repetition rate from the plurality of candidate synonym sentence sets.

8. The method of claim 6, wherein the training data is generated by, The obtaining of a synonym sentence from the plurality of candidate synonym sentences that satisfy the second preset condition includes: determining whether the number of candidate synonym sentences that satisfy the second preset condition satisfies a second quantity threshold condition, the second quantity threshold condition including: the number of candidate synonym sentences that satisfy the second preset condition being greater than or equal to a second preset number; in the case where the number of candidate synonym sentences that satisfy the second preset condition does not satisfy the second quantity threshold condition, selecting all candidate synonym sentences that satisfy the second preset condition as synonym sentences; in the case where the number of candidate synonym sentences that satisfy the second preset condition satisfies the second quantity threshold condition, selecting the first second preset number of candidate synonym sentences with the highest repetition rate from the candidate synonym sentences that satisfy the second preset condition as synonym sentences.

9. The method of claim 8, wherein, The second preset number is 5 to 20.

10. The method of claim 6, wherein, The synonym sentence of the reference translation sentence is obtained through a synonym transcription system.

11. The method of claim 1, wherein, The selecting of the reference translation sentence with the highest similarity to the sample translation sentence from the reference translation set as the sample post-editing sentence includes: selecting the reference translation sentence with the smallest edit distance to the sample translation sentence as the sample post-editing sentence.

12. The method of claim 1, wherein, The sample translation sentence is obtained by machine translation or manual translation of the sample original sentence.

13. The method of claim 1, wherein, The obtaining of the sample translation quality label of each vocabulary in the sample original sentence using the sample post-editing sentence and the sample translation sentence includes: The matching degree of the sample translated sentence and the sample post-editing sentence is detected, and according to the matching degree detection result, a confidence label is added to each vocabulary in the sample post-editing sentence, and the confidence label is used to represent that the translation quality is qualified or unqualified; The vocabulary in the sample original sentence and the sample post-editing sentence is aligned, the confidence label is added to the corresponding vocabulary in the sample original sentence, and a confidence label used to represent that the translation quality is qualified is added to the vocabulary without a corresponding relationship in the sample original sentence, and the confidence label added to the sample original sentence is used as a sample translation quality label.

14. The method of claim 13, wherein the translation evaluation training data is generated by: The matching degree detection on the sample translated sentence and the sample post-editing sentence comprises: obtaining the corresponding relationship of the vocabulary in the sample translated sentence and the sample post-editing sentence; The matching degree detection is performed on the vocabulary with the corresponding relationship.

15. The method of claim 14, wherein the training data is generated by: The corresponding relationship of the vocabulary in the sample translated sentence and the sample post-editing sentence is obtained by using the minimum edit distance.

16. A device for generating translation evaluation training data, characterized in that, It comprises: A first sentence acquisition module is configured to acquire one or more sample original sentences to be translated; A sample translated sentence acquisition module is configured to acquire a sample translated sentence corresponding to the sample original sentence, wherein the sample translated sentence is obtained by translating the sample original sentence; A reference translation set acquisition module is configured to acquire a reference translation set corresponding to the sample original sentence, wherein the reference translation set comprises a plurality of reference translated sentences; wherein the reference translation set corresponding to the sample original sentence is obtained by machine translation; wherein the reference translation set is obtained by: Obtaining a plurality of candidate machine translation sets from different machine translation tools, wherein each candidate machine translation set comprises one or more candidate machine translated sentences; Performing a first screening process on the plurality of candidate machine translation sets to select a plurality of candidate machine translated sentences satisfying a first preset condition, wherein the first preset condition comprises a confidence degree; Obtaining a reference translated sentence from the plurality of candidate machine translated sentences satisfying the first preset condition, wherein the reference translated sentence constitutes the reference translation set; A second sentence acquisition module is configured to select, from the reference translation set, a reference translated sentence with the highest similarity to the sample translated sentence as a sample post-editing sentence; A label acquisition module is configured to obtain a sample translation quality label of each vocabulary in the sample original sentence by using the sample post-editing sentence and the sample translated sentence; A training set establishment module is configured to establish a training set according to the sample translation quality label, wherein the training set comprises the sample original sentence, the sample translated sentence, and the sample translation quality label, and the training set is used to train a translation evaluation model.

17. An apparatus, comprising: The device comprises at least one memory and at least one processor, wherein the memory stores one or more computer instructions, and the one or more computer instructions are executed by the processor to implement the method for generating translation evaluation training data according to any one of claims 1 to 15. The device comprises at least one memory and at least one processor, wherein the memory stores one or more computer instructions, and the one or more computer instructions are executed by the processor to implement the method for generating translation evaluation training data according to any one of claims 1 to 15.

18. A storage medium, characterized by The storage medium stores one or more computer instructions for implementing the generation method of the translation evaluation training data according to any one of claims 1 to 15.

Citation Information

Patent Citations

  • Corpus evaluation model training method and device, storage medium and computer equipment

    CN110263349A