A Method and System for Generating Pseudo Data for Translation Quality Estimation Based on ELECTRA
Through the ELECTRA-based translation quality estimation pseudo-data generation method, manual editing of translation and machine-translated translations are used to generate pseudo-data, which solves the problem of data scarcity in machine translation quality estimation technology, and improves the performance and accuracy of the translation quality evaluation model.
Patent Information
- Application Number
- CN202111470031.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-12-03
- Publication Date
- 2025-07-11
- Estimated Expiration
- 2041-12-03
AI Technical Summary
Existing machine translation quality estimation techniques have problems with small data set size and scarce data, which leads to the model being easily overfitted during training.
The pseudo-data generation method based on ELECTRA is used to generate pseudo-data by manually editing the translation and machine-translated translation. The ELECTRA model is used for training. First, use manual and then edit the translation for initial training, and then use machine-translated translation and the original data set for secondary training to generate a translation quality evaluation model at the sentence or word level.
It improves the performance of the translation quality estimation model, solves the problem of data scarcity, reduces the phenomenon of overfitting, and improves the accuracy of translation quality evaluation.
Smart Images

Figure CN114330373B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of translation, and specifically to a method and system for generating pseudo-data for translation quality estimation based on ELECTRA. Background Art
[0002] Machine translation technology is a research direction in the field of NLP, and the automatic evaluation of the quality of machine translation translations is very important for the research of machine translation. Machine translation quality estimation technology is a technology for evaluating the effect of machine translation translations without reference translations. As a machine translation automatic evaluation technology that does not require reference translations, it can evaluate the quality of translations only using the source language and machine translation translations, and can still be widely applicable without reference translations.
[0003] The dataset for the machine translation quality estimation task consists of four parts: source language sentences, machine translations, human post-edited translations, and quality annotations. Among them, the human post-edited translations are obtained by translation practitioners post-editing the machine translation translations according to the semantics of the source language. Since the data annotation process for translation quality estimation is very complex and the cost of human post-edited translations is high, this has led to a generally small scale of the translation quality estimation dataset, resulting in the problem of data scarcity. Currently, the translation quality estimation model of the pre-trained language model enables the model to learn knowledge from a large-scale unsupervised corpus during the pre-training stage and transfer it to the downstream translation quality estimation task, relatively alleviating the problem of the lack of the translation quality estimation task dataset. However, from the perspective of data scale, the scale of the translation quality estimation dataset has not expanded, and the problem of overfitting still occurs when the model is trained using a small dataset in the translation estimation stage.
[0004] Machine translation quality estimation technology is a machine translation automatic evaluation technology that does not require reference translations. It can evaluate the quality of translations only using the source language and machine translation translations, so it is widely applicable without reference translations. Although translation quality estimation technology has many advantages, the corpus used for training translation quality estimation models often requires professional translators to post-edit machine translations manually, resulting in the general problems of small dataset scale and data scarcity in translation quality estimation tasks. Summary of the Invention
[0005] The present invention provides a method for generating pseudo-data for translation quality estimation based on ELECTRA, aiming at the problem of scarce translation quality estimation data.
[0006] The present invention provides a system for generating pseudo-data for translation quality estimation based on ELECTRA, aiming at the problem of scarce translation quality estimation data.
[0007] The present invention is achieved through the following technical solutions:
[0008] A method for generating pseudo-data for translation quality estimation based on ELECTRA, which performs translation sentence by sentence. The generation method includes the following steps:
[0009] Step J1: Generate pseudo-data using the human post-edited translations in the target language direction of the QE dataset to be augmented or generate pseudo-data using machine translation translations.
[0010] Step J2: Perform the first training on the pseudo-data generated based on the human post-edited translations in Step J1 using the trained ELECTRA model.
[0011] Step J3: Mix the pseudo-data generated based on the machine translation translations in Step J1 with the original dataset and perform the second training on the ELECTRA model obtained in Step J2.
[0012] Step J4: Obtain a sentence-level translation quality assessment model by performing two rounds of ELECTRA model training in Step J3.
[0013] Step J5: Verify the performance of the model in Step J4.
[0014] A method for generating pseudo-data for translation quality estimation based on ELECTRA. Specifically, the generation of pseudo-data from manually edited data in Step 1 is as follows:
[0015] Step J1.1: Use the human post-edited translations or machine translation translations in the target language direction of the QE dataset to be augmented.
[0016] Step J1.2: Input the human post-edited translation or machine translation translation in Step J1.1 as the mother text into the ELECTRA generator.
[0017] Step J1.3: Generate a rewritten new human post-edited translation or new machine translation translation based on the ELECTRA generator.
[0018] Step J1.4: Calculate the corresponding HTER score of the human post-edited translation or new machine translation translation in Step J1.3 through the TERCOM toolkit.
[0019] Step J1.5: The HTER score in Step J1.4, along with the final source text, the new machine translation rewritten by the generator, and the human post-edited translation, form a new sentence-level translation quality estimation pseudo-data quadruple.
[0020] A method for generating pseudo-data for translation quality estimation based on ELECTRA. Specifically, the generation of the rewritten new translation in Step J1.2 is as follows:
[0021] The generator selects words in the sentence for prediction and replaces the original words with the predicted ones, canceling the masking operation on the input. Therefore, the generator predicts and replaces each word in the sentence once, rather than only predicting and replacing some words;
[0022] The formal representation of the above process is as follows. A sentence p = [p1,... p n , is encoded by the trained generator G to obtain For the prediction of all words at position t, the probability is:
[0023] p G (p t ∣p) = softmax(exp(p t )·h G (p) t ) (1)
[0024] Finally, select the word p with the highest probability and replace it with the word at this position:
[0025]
[0026] A sentence-level translation quality estimation pseudo-data generation system based on ELECTRA, the generation system includes
[0027] A pseudo-data generation unit for generating pseudo-data from the human post-edited translations in the target language direction of the QE dataset to be augmented or generating pseudo-data from machine translation translations;
[0028] An ELECTRA training unit for training the pseudo-data generated from the human post-edited translations and the pseudo-data generated from machine translation translations;
[0029] A sentence quality evaluation unit for evaluating the translation quality at the sentence level after training by the ELECTRA training unit.
[0030] A method for generating pseudo-data for translation quality estimation based on ELECTRA, which performs translation on a word-by-word basis. The generation method specifically includes the following steps:
[0031] Step C1: Use the words in the target language direction of the QE dataset to be augmented for manual editing data. Finally, a quadruple of pseudo-data for word-level translation quality estimation consisting of the source text, the new machine translation rewritten by the generator, the human post-edited translation, and the OK / BAD annotation file is obtained;
[0032] Step C2: Based on the quadruple of pseudo-data generated in Step C1, use the ELECTRA model for the first training;
[0033] Step C3: Use the original dataset to perform a second training based on the ELECTRA model obtained in Step C2;
[0034] Step C4: Obtain a word-level translation quality evaluation model through the second ELECTRA model training in Step C3;
[0035] Step C5: Verify the performance of the evaluation model in Step C4.
[0036] A method for generating pseudo-data for translation quality estimation based on ELECTRA. Specifically, Step C1 is as follows:
[0037] Step C1.1: Use the human post-edited translations in the target language direction of the QE dataset to be augmented;
[0038] Step C1.2: Select the human post-edited translations in Step 1 as the mother copy for generating pseudo-data and input them into the ELECTRA generator. Before input, generate the corresponding number of OK tags according to the number of words in the human post-edited translations;
[0039] Step C1.3: Compare the rewritten sentences output by the generator with the original sentences word by word, find the corresponding positions of the replaced words in the sentences, and change the tags at the corresponding positions from OK to BAD;
[0040] Step C1.4: Obtain pseudo-data from the human post-edited translations based on the new tags in Step C1.3;
[0041] Step C1.5: Finally, the original text, the new machine translation rewritten by the generator, the human post-edited translation, and the OK / BAD annotation file obtained by the above method constitute a quadruple of pseudo-data for word-level translation quality estimation.
[0042] A system for generating pseudo-data for translation quality estimation based on ELECTRA. The generation system includes
[0043] A pseudo-data generation unit for generating pseudo-data from human post-edited translations or machine translation translations of the words in the target language direction of the QE dataset to be augmented;
[0044] An ELECTRA training unit for training the pseudo-data generated from human post-edited translations and the pseudo-data generated from machine translation translations;
[0045] A word quality evaluation unit for evaluating the word-level translation quality obtained after the training of the ELECTRA training unit.
[0046] The beneficial effects of the present invention are:
[0047] The present invention uses the generator of ELECTRA as a means to generate QE pseudo-data, and uses the generated pseudo-data to participate in training to further improve the model performance. Description of the Drawings
[0048] Figure 1 Schematic diagram of specific word generation at the sentence level of the present invention.
[0049] Figure 2 Schematic diagram of specific word generation at the word level of the present invention.
[0050] Figure 3 Flowchart of the training method at the sentence level of the present invention, where (a) is the flowchart of training with manually post-edited translation pseudo-data; (b) is the flowchart of training with machine translation pseudo-data; (c) is to generate a sentence-level quality assessment model.
[0051] Figure 4 Flowchart of the training method at the word level of the present invention, where (a) is the flowchart of training with manually post-edited translation pseudo-data; (b) is to generate a word-level quality assessment model. Detailed Embodiments
[0052] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.
[0053] A method for generating translation quality estimation pseudo-data based on ELECTRA, which performs translation in units of sentences, and the generation method includes the following steps:
[0054] Step J1: Generate pseudo-data using manually post-edited translation in the target language direction in the QE dataset to be expanded or machine translation translation to generate pseudo-data;
[0055] Step J2: Perform the first training on the pseudo-data generated from the manually post-edited translation in Step J1 using the trained ELECTRA model;
[0056] Step J3: Mix the pseudo-data generated from the machine translation translation in Step J1 with the original dataset and perform the second training on the ELECTRA model obtained in Step J2;
[0057] Step J4: Obtain a sentence-level translation quality assessment model by performing two ELECTRA model trainings in Step J3;
[0058] Step J5: Verify the performance of the model in Step J4.
[0059] Further, the specific process of generating pseudo data by manually editing data in step 1 is as follows:
[0060] Step J1.1: Use the human post-edited translation or machine translation of the target language direction in the QE dataset to be augmented;
[0061] Step J1.2: Input the human post-edited translation or machine translation in step J1.1 into the ELECTRA generator as the mother text;
[0062] Step J1.3: Generate a rewritten new human post-edited translation or new machine translation based on the ELECTRA generator;
[0063] Step J1.4: Calculate the corresponding HTER score of the human post-edited translation or new machine translation in step J1.3 through the TERCOM toolkit;
[0064] Step J1.5: The HTER score in step J1.4, along with the final source text, the new machine translation rewritten by the generator, and the human post-edited translation, these four parts form a new sentence-level translation quality estimation pseudo data quadruple.
[0065] Further, the specific process of generating a rewritten new translation in step J1.2 is as follows:
[0066] The generator will select words in the sentence for prediction and use the predicted words to replace the original words, canceling the mask operation on the input. Therefore, the generator will predict and replace each word in the sentence once, rather than only predicting and replacing some words;
[0067] The formal representation of the above process is as follows. A sentence p = [p1,... p n , after being encoded by the trained generator G, obtains For the probability of all words predicted at position t:
[0068] p G (p t ∣p) = softmax(exp(p t ) · h G (p) t ) (1)
[0069] Finally, select the word p with the highest probability and replace it with the word at this position:
[0070]
[0071] Therefore, the sentences output after being processed by the ELECTRA generator can be regarded as newly generated machine translations with partial translation errors compared to the original sentences. During the actual operation process, two texts from different sources were first selected as the mother texts for pseudo-data generation and input into the ELECTRA generator. The human post-edited translations and machine translations in the target language direction of the QE dataset to be augmented were used as the mother texts for pseudo-data generation respectively, and pseudo-data generation operations were carried out according to the proposed method, obtaining two pseudo-datasets for the translation quality estimation task.
[0072] In the training process, the method of first using the pseudo-data generated from human-edited translations to initially train the model and then using the dataset after mixing the pseudo-data generated from machine translations with the original data for secondary training was adopted.
[0073] A pseudo-data generation system for translation quality estimation based on ELECTRA
[0074] A pseudo-data generation unit for generating pseudo-data from human post-edited translations or machine translations in the target language direction of the QE dataset to be augmented;
[0075] An ELECTRA training unit for training the pseudo-data generated from human post-edited translations and the pseudo-data generated from machine translations;
[0076] A sentence quality evaluation unit for evaluating the translation quality at the sentence level after training by the ELECTRA training unit.
[0077] Furthermore, the translation is carried out on a word-by-word basis, and the generation method specifically includes the following steps:
[0078] Step C1: Use the words in the target language direction of the QE dataset to be augmented for human-edited data, and finally obtain a quadruple of pseudo-data for word-level translation quality estimation consisting of the source text, the new machine translation rewritten by the generator, the human post-edited translation, and the OK / BAD annotation file obtained by the above method;
[0079] Step C2: Based on the pseudo-data quadruple generated in Step C1, use the ELECTRA model for the first training;
[0080] Step C3: Use the original dataset and conduct secondary training based on the ELECTRA model obtained in Step C2;
[0081] Step C4: Obtain a word-level translation quality evaluation model after the ELECTRA model in Step C3 has been trained twice;
[0082] Step C5: Verify the performance of the evaluation model in Step C4.
[0083] Further, step C1 is specifically as follows:
[0084] Step C1.1: Use the human post-edited translation in the target language direction of the QE dataset to be augmented.
[0085] Step C1.2: Select the human post-edited translation in step 1 as the mother copy for generating pseudo-data and input it into the ELECTRA generator, and generate the corresponding number of OK tags according to the number of words in the human post-edited translation before input; the reason for fully annotating as OK is that the human post-edited translation does not contain translation errors.
[0086] Step C1.3: Compare the rewritten sentences output by the generator with the original sentences word by word, find the corresponding positions of the replaced words in the sentences, and change the tags at the corresponding positions from OK to BAD; thus, the corresponding correct annotation file is obtained.
[0087] Step C1.4: Generate pseudo-data from the human post-edited translation based on the new tags obtained in step C1.3.
[0088] Step C1.5: Finally, the source text, the new machine translation rewritten by the generator, the human post-edited translation, and the OK / BAD annotation file obtained by the above method constitute a quadruple of pseudo-data for word-level translation quality estimation.
[0089] A pseudo-data generation system for translation quality estimation based on ELECTRA, the generation system includes
[0090] A pseudo-data generation unit for generating pseudo-data from the human post-edited translation in the target language direction of the QE dataset to be augmented or generating pseudo-data from machine translation.
[0091] An ELECTRA training unit for training the pseudo-data generated from the human post-edited translation and the pseudo-data generated from machine translation.
[0092] A word quality evaluation unit for evaluating the translation quality at the word level after training by the ELECTRA training unit.
[0093] For sentence-level QE pseudo-data, two types of pseudo-data with different data distributions are generated by using machine translation as the input mother copy to generate pseudo-data and using human post-edited translation to generate pseudo-data, and a method is proposed to first use the pseudo-data generated from the human post-edited translation to initially train the model and then use the dataset obtained by mixing the pseudo-data generated from machine translation with the original data for secondary training. For word-level pseudo-data, aiming at the problem of unbalanced label distribution of training data, pseudo-data with a more reasonable distribution is generated, and a method is adopted to first use the obtained pseudo-data to train the model and then use the original dataset for secondary training.
[0094] Example 2
[0095] To verify the proposed method for generating pseudo data for sentence-level translation quality estimation based on ELECTRA, experiments were conducted on the CCMT2019 EN-ZH dataset when generating pseudo data for translation quality estimation from different input sources. To explore whether the constructed pseudo data for translation quality estimation can improve the model's performance when participating in training, this section selected the CCMT2019 EN-ZH sentence-level QE task dataset for experimental verification. First, the machine translations and human post-edited translations in this dataset were used as inputs respectively, and two different pseudo datasets were obtained using the proposed pseudo data generation method. In addition, there is a certain probability that all words in the sentences input to the ELCTRA generator will not be replaced when output, so the sentences described above were excluded. In the translation quality estimation training stage, the generated additional pseudo data was added to the original QE dataset for training. Regarding the specific method of training with the additional pseudo data, the following training strategy was adopted: first, the model was initially trained with the pseudo data generated from the human post-edited translations, and then the model was secondarily trained with the dataset obtained by mixing the pseudo data generated from the machine translations and the original data.
[0096] To verify the proposed method for generating pseudo data for word-level translation quality estimation based on ELECTRA, relevant experiments were conducted on the WMT2017 DE-EN word-level QE task dataset. To prove the effectiveness of the QE pseudo dataset generated by the proposed method in improving the translation quality estimation results, a pseudo dataset with the same scale and the same OK and BAD label ratios was generated by means of random replacement. The specific operation was as follows: by comparing the human post-edited translations and the pseudo dataset generated from the post-edited translations, the positions of all words predicted to be replaced by the ELECTRA generator were found, and the words at these positions were replaced with other random words. After generating the pseudo data, the model was trained first with the obtained pseudo data and then secondarily trained with the original dataset.
[0097] Experimental parameter settings
[0098] The overall experiment was conducted in a Pytorch-based environment. In the pre-training stage, first, the ELECTRA pre-trained language model related to the experiment and the corresponding vocabulary file were loaded from Hugging face
[10] . The base version of the model was used, that is, a 12-layer Transformer-encoder with a hidden layer size of 768 and 12 attention heads. Next, the prepared corpus was used for secondary pre-training. AdamW was used as the optimizer, the sequence_length was uniformly set to 128, and the batch size was set to 32. When using a decaying learning rate with an initial learning rate of 2e-4, it took 5-7 days to train different pre-trained language models until convergence using 4 2080Ti graphics cards.
[0099] In the translation quality estimation stage, the pre-trained model continued to participate in the training. AdamW was still selected as the optimizer in this paper. Under the conditions of a batch size of 16 and a learning rate of 5e-5, the model could be trained in less than 1 hour using 2 2080Ti graphics cards. Among them, the model termination condition was that the loss of the validation set did not decrease within 2 rounds.
[0100] Experimental results of the pseudo-data generation and training method for sentence-level translation quality estimation
[0101] The scale of the pseudo-data generated by two methods is shown in Table 1
[0102] Table 1 Scale of the pseudo-data for sentence-level translation quality estimation generated
[0103]
[0104] The two obtained pseudo-datasets were respectively experimented with the above two training strategies, and the experimental results are shown in Table 2.
[0105] Table 2 Experimental results of different pseudo-corpus training strategies in the CCMT2019 EN-ZH sentence-level task
[0106]
[0107] The experimental results show that the method of first using the pseudo-data generated by manually post-edited translations to initially train the model and then using the dataset mixed with the pseudo-data generated by machine translations and the original data for secondary training finally achieved the best translation quality estimation results shown in the last row of Table 2.
[0108] Experimental results of the pseudo-data generation and training method for word-level translation quality estimation
[0109] The word-level QE pseudo-data with the scale shown in Table 3 was obtained:
[0110] Pseudo-data scale of word-level translation quality estimation generated in Table 3
[0111]
[0112] The experimental results of the constructed word-level QE pseudo-dataset under different strategies are shown in Table 4.
[0113] Experimental results of different pseudo-corpus training strategies in the WMT2017 DE-EN word-level task
[0114]
[0115] The experimental results show that for the pseudo-dataset obtained by using the proposed word-level pseudo-data generation method, the training strategy of first training the model with the pseudo-data obtained by this method and then performing secondary training with the original dataset can enable the model to obtain the best performance.
Claims
1. A method for generating pseudo-data for translation quality estimation based on ELECTRA, characterized in that Translate sentence by sentence. The generation method includes the following steps: Step J1: Generate pseudo data using the human post-edited translations in the target language direction of the QE dataset to be augmented or generate pseudo data using machine translation translations. Step J2: Use the pseudo data generated from the human post-edited translations in Step J1 to perform the first training using the trained ELECTRA model. Step J3: Mix the pseudo data generated from the machine translation translations in Step J1 with the original dataset and perform the second training on the ELECTRA model obtained in Step J2. Step J4: Obtain a sentence-level translation quality assessment model by performing two rounds of ELECTRA model training in Step J3. Step J5: Verify the performance of the model in Step J4. Specifically for generating pseudo data from the human post-edited data in Step J1, Step J1.1: Use the human post-edited translations or machine translation translations in the target language direction of the QE dataset to be augmented. Step J1.2: Use the human post-edited translation or machine translation translation in Step J1.1 as the parent input into the ELECTRA generator. Step J1.3: Generate a rewritten new human post-edited translation or new machine translation translation based on the ELECTRA generator. Step J1.4: Calculate the corresponding HTER score for the human post-edited translation or new machine translation translation in Step J1.3 through the TERCOM toolkit. Step J1.5: The HTER score in Step J1.4, along with the final source text, the new machine translation rewritten by the generator, and the human post-edited translation, these four parts form a new sentence-level translation quality estimation pseudo data quadruple. Specifically for generating the rewritten new human post-edited translation or new machine translation translation in Step J1.3, The generator will select words in the sentence for prediction and use the predicted words to replace the original words, canceling the mask operation on the input. Therefore, the generator will predict and replace each word in the sentence once, rather than only predicting and replacing some words. The formal representation of the above process is as follows. A sentence p = [p1, … p n , is encoded by the trained generator G to obtain The probabilities of all words are predicted for the position t: p G (p t |p) = softmax(exp(p t )·h G (p) t ) (1) Finally, select the word p with the highest probability and replace it with the word at that position:
2. The method for generating pseudo data for translation quality estimation based on ELECTRA according to claim 1, wherein Translate word by word. The generation method specifically includes the following steps: Step C1: Use the words in the target language direction of the QE dataset to be augmented for manual editing of data. Finally, a word-level translation quality estimation pseudo data quadruple is formed by the source text, the new machine translation rewritten by the generator, the human post-edited translation, and the OK / BAD annotation file. Step C2: Based on the pseudo data quadruple generated in Step C1, perform the first training using the ELECTRA model. Step C3: Use the original dataset to perform the second training on the ELECTRA model obtained in Step C2. Step C4: Obtain a word-level translation quality assessment model by performing two rounds of ELECTRA model training in Step C3. Step C5: Verify the performance of the assessment model in Step C4.
3. The method for generating pseudo data for translation quality estimation based on ELECTRA according to claim 2, wherein, Specifically, Step C1 is as follows: Step C1.1: Use the human post-edited translations in the target language direction of the QE dataset to be augmented. Step C1.2: Select the human post-edited translation in Step C1.1 as the master copy for generating pseudo-data and input it into the ELECTRA generator. Before input, generate OK tags with the same number as the number of words in the human post-edited translation. Step C1.3: Compare the rewritten sentences output by the generator with the original sentences word by word, find the positions of the replaced words in the sentences, and change the tags at those positions from OK to BAD. Step C1.4: Obtain pseudo-data from the human post-edited translation based on the new tags in Step C1.
3. Step C1.5: Finally, the source text, the new machine translation rewritten by the generator, the human post-edited translation, and the OK / BAD annotation file obtained by the above method constitute a quadruple of pseudo-data for word-level translation quality estimation.
4. A pseudo-data generation system for translation quality estimation based on ELECTRA, characterized in that, The generation system uses a method for generating pseudo-data for translation quality estimation based on ELECTRA as described in Claim 1. The system includes: A pseudo-data generation unit for generating pseudo-data from the human post-edited translations in the target language direction in the QE dataset to be augmented or generating pseudo-data from machine translation translations. An ELECTRA training unit for training the pseudo-data generated from the human post-edited translations and the pseudo-data generated from machine translation translations. A sentence quality evaluation unit for evaluating the translation quality at the sentence level after training by the ELECTRA training unit.
5. A pseudo-data generation system for translation quality estimation based on ELECTRA, characterized in that, The generation system uses a method for generating pseudo-data for translation quality estimation based on ELECTRA as described in Claim 3. The generation system includes: A pseudo-data generation unit for generating pseudo-data from the human post-edited translations of the words in the target language direction in the QE dataset to be augmented or generating pseudo-data from machine translation translations. An ELECTRA training unit for training the pseudo-data generated from the human post-edited translations and the pseudo-data generated from machine translation translations. A word quality evaluation unit for evaluating the translation quality at the word level after training by the ELECTRA training unit.
Citation Information
Patent Citations
Sentence-level machine translation quality estimation model training method based on mixed granularity
CN110472253A
Machine translation quality evaluation method, device, equipment and medium
CN112347795A
Semantic checking method and system for multi-language mixed text
CN113158695A