Translation quality evaluation model training method, evaluation method and related equipment
By using batch divergence loss value in the translation quality evaluation model training, the problem that mean square error cannot fit the translation sequence correlation of paragraphs is solved, which improves the accuracy of the evaluation results and the reliability of model training.
Patent Information
- Application Number
- CN202411967206.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-26
- Publication Date
- 2025-05-06
- Estimated Expiration
- 2044-12-26
AI Technical Summary
When training the translation quality evaluation model in the prior art, mean square error cannot effectively fit the sequential correlation between paragraph translations, resulting in low accuracy of the evaluation results.
The translation quality evaluation model is trained by using the batch divergence loss value, and by calculating the batch divergence loss value between the translation to be evaluated and the label translation, fitting the score values and sequential correlations of the manual evaluation.
The training reliability and accuracy of the translation quality evaluation model are improved, and the evaluation accuracy of the evaluation results are avoided due to different evaluation samples.
Smart Images

Figure CN119940374A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of data processing technology, and in particular to a translation quality assessment model training method, an assessment method and related equipment. Background Art
[0002] Translation quality evaluation can help us understand the quality of manual or machine translations. It has important applications in translation annotation, translation engine effect comparison, machine translation model development and other fields. Automatic evaluation has lower costs and higher efficiency than manual evaluation.
[0003] In related technologies, in order to improve the reliability of automatic translation quality evaluation, the mean square error is usually used to calculate the error loss between the manual evaluation score and the model evaluation score during the training process of the quality evaluation model to train the quality evaluation model multiple times, so that the model evaluation score is more similar to the manual evaluation score. However, for regression translation evaluation tasks in which there is a paragraph order correlation in batch translations, the mean square error can only fit the similar performance of the translation quality evaluation model for the evaluation score of a single paragraph translation, but cannot fit the order between paragraph translations. As a result, after the translation quality evaluation model for processing such regression tasks is trained according to the mean square error, the evaluation results obtained by using the translation quality evaluation model to perform quality evaluation on the actual target translation are less accurate. Summary of the invention
[0004] The embodiments of the present application provide a translation quality assessment model training method, an assessment method and related equipment, which can improve the training reliability of the translation quality assessment model, thereby improving the accuracy of the assessment results obtained by using the translation quality assessment model to perform quality assessment on the actual target translation.
[0005] To achieve the above-mentioned purpose, a first aspect of an embodiment of the present application proposes a translation quality assessment model training method, the method comprising:
[0006] Acquire a translation quality assessment dataset, wherein the translation quality assessment dataset includes at least one training batch, wherein the training batch includes a plurality of sample text pairs and first translation quality scores of the sample text pairs, wherein the sample text pairs include at least a translation to be assessed and a label translation;
[0007] For each of the training batches, scoring the translation to be evaluated based on the labeled translation in the sample text pair using a translation quality evaluation model to obtain a second translation quality score;
[0008] Calculating a batch divergence loss value corresponding to the training batch according to the first translation quality score and the second translation quality score;
[0009] The translation quality assessment model is trained according to the batch divergence loss value to obtain the trained translation quality assessment model.
[0010] In some embodiments, calculating the batch divergence loss value corresponding to the training batch according to the first translation quality score and the second translation quality score includes:
[0011] Normalizing each of the first translation quality scores one by one to obtain a first normalized score;
[0012] Normalizing each of the second translation quality scores one by one to obtain a second normalized score;
[0013] The batch divergence loss value is obtained based on a ratio of the first normalized score to the second normalized score.
[0014] In some embodiments, obtaining the batch divergence loss value based on the ratio of the first normalized score to the second normalized score includes:
[0015] Based on the ratio of each first normalized score to the corresponding second normalized score, logarithmic processing is performed to obtain a relative ratio;
[0016] Obtaining a relative batch divergence loss value based on the product of the relative ratio and the corresponding second normalized score;
[0017] All the relative batch divergence loss values are averaged to obtain the batch divergence loss value.
[0018] In some embodiments, normalizing each of the first translation quality scores one by one to obtain a first normalized score includes:
[0019] Accumulating all the first translation quality scores to obtain a first accumulated score;
[0020] Each of the first translation quality scores is taken as a first target score one by one, and the first normalized score is obtained based on the ratio of the first target score to the first accumulated score.
[0021] In some embodiments, the training of the translation quality assessment model according to the batch divergence loss value includes:
[0022] Obtain a mean square loss weight value, and obtain a batch divergence loss weight value based on one minus the mean square loss weight value;
[0023] Obtaining a mean square loss value based on the first translation quality score and all the second translation quality scores;
[0024] Based on the product of the mean square loss weight value and the mean square loss value, plus the product of the batch divergence loss weight value and the batch divergence loss value, a composite loss value is obtained, and model training is performed on the translation quality assessment model based on the composite loss value.
[0025] In some embodiments, the step of using the translation quality assessment model to score the translation to be assessed based on the label translation in the sample text pair to obtain a second translation quality score includes:
[0026] When the labeled translation only includes the reference translation, the translation quality assessment model is obtained based on the BLEURT model, and the translation to be assessed is scored based on the reference translation in the sample text pair using the translation quality assessment model to obtain the second translation quality score;
[0027] When the labeled translation includes only the source text, the translation quality assessment model is obtained based on the COMETkiwi model, and the translation to be assessed is scored based on the source text in the sample text pair using the translation quality assessment model to obtain the second translation quality score;
[0028] When the labeled translation includes the reference translation and the source text, the translation quality assessment model is obtained based on the COMET model, and the translation to be assessed is scored using the translation quality assessment model based on the source text and the reference translation in the sample text pair to obtain the second translation quality score.
[0029] To achieve the above purpose, a second aspect of the embodiments of the present application proposes a translation quality assessment method, the method comprising:
[0030] Obtain target translation data;
[0031] The target translation data is input into the translation quality assessment model obtained after training as described in the first aspect to perform data processing to obtain a target quality score corresponding to the target translation data.
[0032] To achieve the above-mentioned purpose, a third aspect of the embodiments of the present application proposes a translation quality assessment model training device, the device comprising:
[0033] A data set acquisition module, used to acquire a translation quality assessment data set, wherein the translation quality assessment data set includes at least one training batch, wherein the training batch includes a plurality of sample text pairs and first translation quality scores of the sample text pairs, wherein the sample text pairs include at least a translation to be assessed and a labeled translation;
[0034] A scoring module, configured to score the translation to be evaluated based on the labeled translation in the sample text pair using a translation quality evaluation model for each of the training batches, to obtain a second translation quality score;
[0035] A divergence loss calculation module, configured to calculate a batch divergence loss value corresponding to the training batch according to the first translation quality score and the second translation quality score;
[0036] A training module is used to train the translation quality assessment model according to the batch divergence loss value to obtain the trained translation quality assessment model.
[0037] To achieve the above-mentioned purpose, the fourth aspect of an embodiment of the present application proposes an electronic device, which includes a memory and a processor, the memory stores a computer program, and the processor implements the translation quality assessment model training method as described in the first aspect when executing the computer program.
[0038] To achieve the above-mentioned purpose, the fifth aspect of an embodiment of the present application proposes a storage medium, which is a computer-readable storage medium, and the storage medium stores a computer program. When the computer program is executed by a processor, the translation quality assessment model training method described in the first aspect is implemented.
[0039] The translation quality assessment model training method, assessment method and related equipment proposed in the embodiment of the present application include: first, obtaining a translation quality assessment data set, the translation quality assessment data set includes at least one training batch, the training batch includes multiple sample text pairs and first translation quality scores of the sample text pairs, and the sample text pairs at least include a translation to be assessed and a labeled translation; then, for each training batch, using the translation quality assessment model to score the translation to be assessed based on the labeled translation in the sample text pair to obtain a second translation quality score; next, calculating a batch divergence loss value corresponding to the training batch according to the first translation quality score and the second translation quality score; finally, training the translation quality assessment model according to the batch divergence loss value to obtain a trained translation quality assessment model. In the training of a translation quality assessment model for quality assessment corresponding to a plurality of translation texts in a related order in a training batch, the embodiment of the present application uses a batch divergence loss value obtained by calculating the batch divergence loss corresponding to a regression task to fit the score value association and order association between a first translation quality score corresponding to a manual evaluation and a second translation quality score obtained by model evaluation, so as to perform model training, thereby improving the training reliability of the translation quality assessment model, avoiding the problem of inconsistent evaluation accuracy of the translation quality assessment model caused by different evaluation samples during the training process and the actual application process, and further improving the accuracy of the evaluation result obtained by using the translation quality assessment model to perform quality evaluation on the actual target translation.
[0040] Other features and advantages of the present application will be described in the following description, and partly become apparent from the description, or understood by practicing the present application. The purpose and other advantages of the present application can be realized and obtained by the structures specifically pointed out in the description, claims and drawings. BRIEF DESCRIPTION OF THE DRAWINGS
[0041] Figure 1 It is a flowchart of a translation quality assessment model training method provided in one embodiment of the present application.
[0042] Figure 2 yes Figure 1 Flow chart of step 102 in FIG.
[0043] Figure 3 yes Figure 1 Flow chart of step 103 in FIG.
[0044] Figure 4 yes Figure 3 Flow chart of step 301 in FIG.
[0045] Figure 5 yes Figure 3 Flow chart of step 303 in FIG.
[0046] Figure 6 yes Figure 1 Flow chart of step 104 in FIG.
[0047] Figure 7 This is a flowchart of a translation quality assessment method provided by another embodiment of the present application.
[0048] Figure 8 It is a flowchart of a translation quality assessment model training and assessment method provided by another embodiment of the present application.
[0049] Fig. 9 It is a structural diagram of a translation quality assessment model training device provided in one embodiment of the present application.
[0050] Fig.10 It is a schematic diagram of the hardware structure of an electronic device provided in one embodiment of the present application. DETAILED DESCRIPTION
[0051] In order to make the purpose, technical solution and advantages of the present application more clearly understood, the present application is further described in detail below in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application and are not used to limit the present application.
[0052] It should be noted that although the functional modules are divided in the device schematic and the logical order is shown in the flowchart, in some cases, the steps shown or described may be performed in a different order than the module division in the device or the order in the flowchart.
[0053] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as those commonly understood by those skilled in the art to which this application belongs. The terms used herein are only for the purpose of describing the embodiments of this application and are not intended to limit this application.
[0054] Automatic translation quality evaluation refers to the automatic qualitative or quantitative evaluation of the translation quality of the text from the source language to the target language without manual intervention. Translation quality evaluation can help us understand the quality of manual or machine translation. It has important applications in translation annotation, translation engine effect comparison, machine translation model development and other fields. Automatic evaluation has lower cost and higher efficiency than manual evaluation. Traditional automatic translation quality evaluation methods usually use calculation methods based on vocabulary matching or semantic matching. The vocabulary matching method segments the translation and the standard reference translation respectively, and then calculates the vocabulary overlap between the two. The higher the vocabulary overlap between the translation and the reference translation, the better the translation quality. Some other methods count the number of steps required to edit the translation to the reference translation, namely the edit distance, as the judgment standard of vocabulary overlap. Commonly used vocabulary matching methods include BLEU[1], chrF, sacreBLEU, TER, WER, etc. The semantic matching-based method uses an evaluation model to calculate the semantic similarity between the source text and the translation or between the reference translation and the translation. The higher the semantic similarity, the closer the semantics of the translation are to the source text, and the higher the translation quality. The evaluation model includes unsupervised and supervised models. The unsupervised model uses a pre-trained neural network to represent the semantics of the source text, and then calculates the cosine value of the semantic vectors of the two as the semantic similarity; the supervised model uses the existing manually annotated data of translation quality scores for supervised training, and fits the manual evaluation scores of translation quality by training the neural network. The trained translation quality evaluation model can directly predict the translation quality score for a given source text and translation pair, reference translation and translation pair, or a triple input of source text, translation and reference translation.
[0055] Manual and automatic evaluation of translation quality are collectively referred to as translation quality evaluation indicators (metrics). The correlation with human evaluation is the main criterion for measuring automatic evaluation indicators of translation quality, that is, for a given source language to target language translation corpus {s i ,h i} i=1,2,...,M , where s i Indicates the source text, h i represents the corresponding translation, M represents the number of translation corpora, and the task of automatic translation quality evaluation index is to evaluate all translation sentence pairs. i ,h i >The translation quality is scored, and the score is p i The interval is usually between [0, 1] or [0, 100]. The manual evaluation index is used for translation sentences. i ,h i >The translation quality score is q i , human evaluation relevance is to calculate {q i } i=1,2,...,M With {p i} i=1,2,...,M The correlation coefficient is used to measure the correlation between the automatic translation quality evaluation index and the manual evaluation. The higher the correlation, the more consistent the automatic translation quality evaluation index is with the human evaluation. The calculation methods of the correlation coefficient are generally Pearson, Spearman, Kendall or Kendall-tau, etc. The correlation with the human evaluation is a floating point number with a value between 0 and 1.
[0056] Traditional supervised translation quality automatic evaluation models generally use mean-square error (MSE) as the loss function during training. When training a neural network, all training samples are usually batched. The mean square error between the model fitting score and the manual evaluation score of the N samples in each batch can be used as the training loss of the model, thereby performing model training and iterative optimization. The calculation formula is shown in the following formula (1).
[0057]
[0058] In the related art, in order to improve the reliability of automatic evaluation of translation quality, the error loss between the manual evaluation score and the model evaluation score is usually calculated using the mean square error during the training process of the quality assessment model to train the quality assessment model multiple times, so that the model evaluation score is more similar to the manual evaluation score.
[0059] However, the quality assessment model uses mean square error (MSE) as the loss function during training. Although it can theoretically accurately fit the quality score of manual evaluation, it cannot be associated with the evaluation indicator of correlation with manual evaluation, which leads to the problem of inconsistent goals during model training and evaluation. That is, mean square error is used for measurement during training, while Pearson correlation is used for measurement during evaluation. This leads to the trained model fitting the quality score more accurately, but its evaluation of good and bad translations may not be consistent with human evaluation. For example, for two translated sentence pairs, the quality scores evaluated by humans are 0.8 and 0.9, while the quality scores predicted by model 1 are 0.9 and 0.8 respectively, and the quality scores predicted by model 2 are 0.7 and 0.8 respectively. Although the mean square error predicted by the two models is the same, which is 0.01, the relative ranking of the quality scores predicted by model 2 is consistent with human evaluation, and its Pearson coefficient is 1.0 (model 1 is -1.0), which has a better correlation with human evaluation. That is, for regression translation evaluation tasks in which there is a correlation between paragraph order in batch translations, the mean square error can only fit the similar performance of the translation quality evaluation model for the evaluation scores of a single paragraph translation, but cannot fit the order between paragraph translations. As a result, after the translation quality evaluation model for this regression task is trained according to the mean square error, the evaluation results obtained by using the translation quality evaluation model to evaluate the quality of the actual target translation are less accurate.
[0060] In order to improve the training reliability of the translation quality assessment model, and thus improve the accuracy of the evaluation results obtained by using the translation quality assessment model to perform quality assessment on an actual target translation, the embodiment of the present application is directed to the training of a translation quality assessment model for quality assessment corresponding to a plurality of translation texts in a related order in a training batch, and uses a batch divergence loss value obtained by calculating the batch divergence loss corresponding to a regression task to fit the score value association and order association between a first translation quality score corresponding to a manual evaluation and a second translation quality score obtained by model evaluation, so as to perform model training, thereby improving the training reliability of the translation quality assessment model, avoiding the problem of inconsistent evaluation accuracy of the translation quality assessment model during the training process and the actual application process due to different evaluation samples, and thereby improving the accuracy of the evaluation results obtained by using the translation quality assessment model to perform quality assessment on an actual target translation.
[0061] The following will further describe the translation quality assessment model training method, assessment method and related devices provided in the embodiments of the present application. The translation quality assessment model training method and assessment method provided in the embodiments of the present application can be applied to any intelligent terminal equipped with a translation quality assessment model.
[0062] The following will first describe in detail the translation quality assessment model training method in the embodiment of the present application. Figure 1 , which is an optional flowchart of the translation quality assessment model training method provided in the embodiment of the present application, Figure 1 The method in the embodiment may include but is not limited to steps 101 to 104. Figure 1 The order of step 101 to step 104 is not specifically limited, and the order of steps can be adjusted or some steps can be reduced or added according to actual needs.
[0063] Step 101: Obtain a translation quality assessment dataset.
[0064] The following is a detailed description of step 101.
[0065] In some embodiments, in order to achieve accurate prediction of the quality score of manual evaluation and high consistency of the quality of the manual evaluation (i.e., correlation with human evaluation) and alleviate the problem of inconsistent goals in model training and testing in the current method, it is necessary to obtain a translation quality evaluation dataset including a large number of training samples in advance. The translation quality evaluation dataset includes sample data of at least one training batch, each training batch includes a plurality of sample text pairs with an associated order and a first translation quality score q associated with the sample text pair. i , the sample text pair includes the translation h to be evaluated i and the translation to be evaluated i The corresponding label translation, which includes the source text s i and reference translation i At least one of the following. The first translation quality score q i Indicates that the translation h to be evaluated is manually i The scoring standards or scales of the scoring corpus (i.e. translation target texts) are not uniform, and the scores need to be normalized first so that they are between 0 and 1, with 0 representing the lowest score (i.e. low translation quality) and 1 representing the highest score (i.e. high translation quality).
[0066] Based on this, the basic sample data corresponding to the training batch includes three cases, namely two types of triples <reference translation r i , translation to be evaluated h i , the first translation quality score q i >、<Translation to be evaluated i , source texts i , the first translation quality score q i > and a four-tuple <source text s i , translation to be evaluated h i , reference translation i , the first translation quality score qi >.
[0067] Step 102: For each training batch, the translation quality assessment model is used to score the translation to be assessed based on the labeled translation in the sample text to obtain a second translation quality score.
[0068] The following is a detailed description of step 102.
[0069] In some embodiments, after obtaining sample data of each training batch, a suitable translation quality assessment model is selected for each training batch to assess each translation h to be assessed in the training batch. i Scoring is performed to obtain the model for each translation h to be evaluated i The second translation quality score p i .
[0070] The following will further describe how to select an appropriate translation quality assessment model.
[0071] Reference Figure 2 , using the translation quality assessment model to score the translation to be assessed based on the labeled translation in the sample text to obtain a second translation quality score, including the following steps 201 to 203.
[0072] Step 201: When the labeled translation only includes the reference translation, a translation quality assessment model is obtained based on the BLEURT model, and the translation quality assessment model is used to score the translation to be assessed based on the reference translation in the sample text to obtain a second translation quality score.
[0073] Step 202: When the labeled translation only includes the source text, a translation quality assessment model is obtained based on the COMETkiwi model, and the translation quality assessment model is used to score the translation to be assessed based on the source text in the sample text to obtain a second translation quality score.
[0074] Step 203: When the labeled translation includes a reference translation and a source text, a translation quality assessment model is obtained based on the COMET model, and the translation quality assessment model is used to score the translation to be assessed based on the source text and the reference translation in the sample text to obtain a second translation quality score.
[0075] Steps 201 to 203 are described in detail below.
[0076] In some embodiments, when the label translation only includes the translation to be evaluated h i Corresponding reference translation i When the basic sample data corresponding to the training batch is <reference translation r i , translation to be evaluated h i , the first translation quality score qi >, the model only needs to directly compare the reference translation r i and the translation to be evaluated i The difference between the two is that BLEURT is selected as the model architecture of the translation quality assessment model. The BLEURT architecture provides a more robust and consistent machine translation evaluation indicator, which evaluates the translation quality by training the learned bilingual (source language and target language) representation. Thus, the translation quality assessment model based on the BLEURT architecture is used to evaluate the translation quality based on the reference translation r in the sample text pair. i To be evaluated translation i Score and get the translation h to be evaluated i The corresponding second translation quality score p i .
[0077] In contrast, when the label translation only includes the translation to be evaluated h i The corresponding source text s i When the basic sample data corresponding to the training batch is < the translation to be evaluated h i , source texts i , the first translation quality score q i >, the model needs to capture the source text s i and the translation to be evaluated i , and then convert these semantic representations into quality scores. Therefore, COMETkiwi is selected as the model architecture of the translation quality assessment model. COMETkiwi can handle the semantic differences between the source language and the target language. COMETkiwi uses a multilingual pre-trained model to capture the semantic similarity between the source text and the target text to evaluate the translation quality. This model is suitable for scenarios where there is no reference translation, but the quality between the source text and the translation to be evaluated needs to be evaluated. Therefore, the translation quality assessment model based on the COMETkiwi architecture is used to evaluate the quality of the source text in the sample text pair. i To be evaluated translation i Score and get the translation h to be evaluated i The corresponding second translation quality score p i .
[0078] And, when the label translation includes the translation to be evaluated h i Corresponding source texts i and reference translation i When the basic sample data corresponding to the training batch is < source text s i , translation to be evaluated h i , reference translation i , the first translation quality score q i >, the model needs to process the source text s at the same time i、Translation to be evaluated i and reference translation i Therefore, COMET is selected as the model framework of the translation quality assessment model. The COMET model framework compares the source text s i and the translation to be evaluated i The semantic similarity between them, and the translation h to be evaluated i and reference translation i The semantic similarity between them is used to evaluate the translation h i This model is suitable for scenarios that require comprehensive consideration of the source text, the translation to be evaluated, and the reference translation information. i and source texts i To be evaluated translation i Score and get the translation h to be evaluated i The corresponding second translation quality score p i .
[0079] Through the above steps 201 to 203, for different label translation data situations, the correlation between the corresponding label translation and the translation to be evaluated is utilized, and the characteristics of different model frameworks are combined to select the corresponding appropriate model framework to generate a translation quality evaluation model, thereby improving the accuracy of scoring the translation to be evaluated, and improving the reliability of the translation quality evaluation model training.
[0080] Step 103: Calculate a batch divergence loss value corresponding to the training batch according to the first translation quality score and the second translation quality score.
[0081] The following is a detailed description of step 103.
[0082] In some embodiments, after obtaining the translation to be evaluated h i Corresponding to the first translation quality score q evaluated by humans i and the second translation quality score p corresponding to the model evaluation i After that, based on all the associated translations to be evaluated h i The first translation quality score q i and the second translation quality score p i Calculate the batch divergence loss corresponding to the training batch KL , as described below.
[0083] Reference Figure 3 , calculating the batch divergence loss value corresponding to the training batch according to the first translation quality score and the second translation quality score, including the following steps 301 to 303.
[0084] Step 301: normalize each first translation quality score one by one to obtain a first normalized score.
[0085] Step 302: normalize each second translation quality score one by one to obtain a second normalized score.
[0086] Steps 301 to 302 are described in detail below.
[0087] In some embodiments, after obtaining all the translations to be evaluated h i The first translation quality score q i and the second translation quality score p i Afterwards, in order to avoid different translations to be evaluated i In case that the evaluation scores are of different orders of magnitude due to inconsistent scoring standards or scales, it is necessary to normalize each first translation quality score one by one to obtain a first normalized score, and to normalize each second translation quality score one by one to obtain a second normalized score. The normalization process is described in detail as follows.
[0088] Reference Figure 4 , normalizing each first translation quality score to obtain a first normalized score, including the following steps 401 to 402.
[0089] Step 401: Accumulate all first translation quality scores to obtain a first accumulated score.
[0090] Step 402: taking each first translation quality score as a first target score one by one, and obtaining a first normalized score based on the ratio of the first target score to the first accumulated score.
[0091] Steps 401 to 402 are described in detail below.
[0092] In some embodiments, after obtaining all the translations to be evaluated h i The first translation quality score q i After that, first accumulate all the first translation quality scores q i To get the first cumulative score Where N is the translation h to be evaluated in the training batch. i The total number of translations.
[0093] Next, each first translation quality score is taken as the first target score q i , and based on the first target score q i and the first cumulative score The ratio of , get the first normalized score Q i As shown in the following formula (2).
[0094]
[0095] Similarly, each second translation quality p i The second normalized score P corresponding to the score i As shown in the following formula (3).
[0096]
[0097] Step 303: Based on the ratio of the first normalized score to the second normalized score, a batch divergence loss value is obtained.
[0098] Steps 301 to 303 are described in detail below.
[0099] In some embodiments, after obtaining all the translations to be evaluated h i The corresponding first normalized score Q i and the second normalized score P i Afterwards, based on each first normalized score Q i and the corresponding second normalized score P i to obtain the batch divergence loss value loss corresponding to the training batch KL , as described below.
[0100] Reference Figure 5 , based on the ratio of the first normalized score to the second normalized score, a batch divergence loss value is obtained, including the following steps 501 to 503.
[0101] Step 501: Based on the ratio of each first normalized score to the corresponding second normalized score, logarithmic processing is performed to obtain a relative ratio.
[0102] Step 502: Based on the product of the relative ratio and the corresponding second normalized score, a relative batch divergence loss value is obtained.
[0103] Step 503: average all relative batch divergence loss values to obtain a batch divergence loss value.
[0104] Steps 501 to 503 are described in detail below.
[0105] In some embodiments, after obtaining all the translations to be evaluated h i The corresponding first normalized score Q i and the second normalized score P i After that, each translation to be evaluated will be evaluated one by one i The ratio of the corresponding first normalized score to the corresponding second normalized score is processed logarithmically to obtain h for each translation to be evaluated. i The corresponding relative ratio log Qi / P i ). Then based on the relative ratio log(Q i / P i ) and the corresponding second normalized score P i The product of each translation h to be evaluated is obtained i The corresponding relative batch divergence loss value P i log(Q i / P i ). Next, for all the translations to be evaluated h i The corresponding relative batch divergence loss values are averaged to obtain the batch divergence loss value loss KL As shown in the following formula (4).
[0106]
[0107] Through the above steps 301 to 303, 401 to 402, and 501 to 503, the evaluation score is converted into a normalized probability distribution by using normalization processing to eliminate the influence between different dimensions, and then the KL divergence between the true probability distribution and the predicted probability distribution is calculated, so as to improve the consistency of the quality score prediction score and solve the problem that the mean square error loss only considers the difference in the quality score prediction score. The regression calculation characteristics of the KL divergence loss are used to consider the sequential correlation between the translations to be evaluated, thereby improving the training reliability of the translation quality evaluation model; in addition, the flexible compatibility of the KL divergence loss is also used, which can be applicable to model networks of different architectures (including the above-mentioned BLEURT, COMET, COMETkiwi, etc.), thereby improving the flexibility and versatility of the translation quality evaluation model training.
[0108] Step 104: Train the translation quality assessment model according to the batch divergence loss value to obtain a trained translation quality assessment model.
[0109] Step 104 is described in detail below.
[0110] In some embodiments, after obtaining the batch divergence loss value loss corresponding to the training batch KL After that, the batch divergence loss value corresponding to the training batch will be KL The selected translation quality assessment model is trained, thereby introducing the KL divergence loss in the model training to utilize the regression calculation characteristics of the KL divergence loss, solving the problem that the mean square error loss only considers the difference in quality score prediction values.
[0111] Using batch divergence loss KL During model training, you can directly use the batch divergence loss value loss KLThe translation quality assessment model is trained, or the batch divergence loss value loss KL The mean square error (MSE) loss mentioned in the above formula (1) is combined with the feature advantages of the two to further train the model, as described below.
[0112] Reference Figure 6 , training the translation quality assessment model according to the batch divergence loss value, including the following steps 601 to 603.
[0113] Step 601: Obtain a mean square loss weight value, and obtain a batch divergence loss weight value based on one minus the mean square loss weight value.
[0114] Step 602: Obtain a mean square loss value based on the first translation quality score and all second translation quality scores.
[0115] Step 603: A composite loss value is obtained based on the product of the mean square loss weight value and the mean square loss value, plus the product of the batch divergence loss weight value and the batch divergence loss value, and a translation quality assessment model is trained based on the composite loss value.
[0116] Steps 601 to 603 are described in detail below.
[0117] In some embodiments, after obtaining all the translations to be evaluated h in the training batch i The corresponding first translation quality score q i and the second translation quality score p i Then, the above formula (1) is used to calculate the mean square loss value loss of the training batch MSE .
[0118] Then get the mean square loss value loss MSE The corresponding mean square loss weight value α is obtained by subtracting the mean square loss weight value from one to get the batch divergence loss value loss KL The corresponding batch divergence loss weight value is (1-α). Among them, α represents the weight of the MSE loss in the overall loss function, and its value ranges from 0 to 1. It can be set according to actual needs or selected based on cross-validation of the validation set to select the optimal weight.
[0119] Next, based on the product of the mean square loss weight value and the mean square loss value αloss MSE , plus the product of the batch divergence loss weight and the batch divergence loss value (1-α)loss KL , the composite loss value loss is obtained as shown in the following formula (5).
[0120] loss = α * loss MSE +(1-α)*loss KL(5)
[0121] Then, the translation quality assessment model is trained based on the composite loss value.
[0122] Through the above steps 601 to 603, batch KL divergence loss is used to optimize the relative order of model predictions to make it closer to the relative order of manual evaluation, which can improve the order and numerical correlation between the model prediction scores and the manual evaluation scores, and use the mean square error to further improve the numerical correlation between the model prediction scores and the manual evaluation scores, thereby better improving the accuracy and reliability of translation quality evaluation model training.
[0123] In addition, the present application also provides a translation quality assessment method, referring to Figure 7 , is an optional flow chart of a translation quality assessment method provided in an embodiment of the present application. The translation quality assessment method includes the following steps 701 to 702.
[0124] Step 701: Obtain target translation data.
[0125] Step 702: Input the target translation data into the translation quality assessment model obtained after training to perform data processing to obtain a target quality score corresponding to the target translation data.
[0126] Steps 701 to 702 are described in detail below.
[0127] In some embodiments, after the translation quality assessment model is trained multiple times using sample data of at least one training batch, when the intelligent terminal equipped with the translation quality assessment model responds to a quality score calculation request for target translation data, the obtained target translation data is input into the trained translation quality assessment model for data processing, thereby accurately obtaining the target quality score corresponding to the target translation data.
[0128] Reference Figure 8 , is a flowchart of a translation quality assessment model training and assessment method provided in an embodiment of the present application. Figure 8 As shown in , firstly, sample data of multiple training batches (including translation data and corresponding translation quality scoring data) are collected, and then based on the data situation of the sample data, the corresponding deep learning model network framework is selected to generate a translation quality assessment model, and then the batch KL divergence loss corresponding to the training batch is calculated, and a training loss value is generated based on the batch KL divergence loss, and the translation quality assessment model is trained based on the training loss value; and the trained translation quality assessment model is used to calculate the translation evaluation score of the target translation data.
[0129] In some embodiments, in order to achieve accurate prediction of the quality score of manual evaluation, and also to achieve high consistency of the quality of the manual evaluation of the translation (i.e., the correlation with human evaluation), and to alleviate the problem of inconsistent goals in the current method during model training and testing, the present invention proposes a method for automatic evaluation of translation quality based on batch KL divergence (Kullback-Leiblerdivergence) loss, and introduces correlation loss on the basis of the existing mean square error loss. The correlation loss takes the normalized model prediction score and human evaluation score in a single batch as two probability distributions, the former is the predicted probability distribution, and the latter is the real probability distribution. The difference between the two distributions is calculated using KL divergence or JS divergence, and the difference can effectively perceive the relative order in the probability distribution, thereby optimizing the correlation of manual evaluation. Reasonable adjustment of the weights of mean square error loss and correlation loss to construct a new loss during model training can effectively improve the correlation between the trained automatic translation quality evaluation model and human evaluation, and the weight can be used as a hyperparameter to be verified on the validation set to achieve the optimal correlation improvement. The present invention pluggably introduces batch KL divergence loss without affecting the network structure and parameters of the original automatic translation quality evaluation model, is applicable to all currently commonly used network structures for automatic translation quality evaluation, and has good versatility and flexibility.
[0130] The translation quality assessment model training method, assessment method and related equipment proposed in the embodiment of the present application include: first, obtaining a translation quality assessment data set, the translation quality assessment data set includes at least one training batch, the training batch includes multiple sample text pairs and first translation quality scores of the sample text pairs, and the sample text pairs at least include a translation to be evaluated and a labeled translation; then, for each training batch, when the labeled translation only includes a reference translation, a translation quality assessment model is obtained based on the BLEURT model, and the translation quality assessment model is used to score the translation to be evaluated based on the reference translation in the sample text pair to obtain a second translation quality score; when the labeled translation only includes a source text, a translation quality assessment model is obtained based on the COMETkiwi model, and the translation quality assessment model is used to score the translation to be evaluated based on the source text in the sample text pair to obtain a second translation quality score; when the labeled translation includes a reference translation and a source text, a translation quality assessment model is obtained based on the COMET model, and the translation quality assessment model is used to score the translation to be evaluated based on the source text in the sample text pair and the reference translation to obtain a first translation quality score. Next, all the first translation quality scores are accumulated to obtain a first accumulated score, each first translation quality score is used as a first target score one by one, and a first normalized score is obtained based on the ratio of the first target score to the first accumulated score, each second translation quality score is normalized one by one to obtain a second normalized score, a relative ratio is obtained based on the ratio of each first normalized score to the corresponding second normalized score, and a relative batch divergence loss is obtained based on the product of the relative ratio and the corresponding second normalized score. Loss value, average all relative batch divergence loss values to obtain the batch divergence loss value; finally, obtain the mean square loss weight value, and obtain the batch divergence loss weight value based on one minus the mean square loss weight value, obtain the mean square loss value based on the first translation quality score and all second translation quality scores, and obtain the composite loss value based on the product of the mean square loss weight value and the mean square loss value, plus the product of the batch divergence loss weight value and the batch divergence loss value, and train the translation quality assessment model based on the composite loss value to obtain the trained translation quality assessment model.
[0131] In the training of a translation quality assessment model for quality assessment corresponding to a plurality of translation texts having an associated order in a training batch, the embodiment of the present application utilizes a batch divergence loss value obtained by calculating the batch divergence loss corresponding to the regression task to fit the score value association and the order association between a first translation quality score corresponding to the manual evaluation and a second translation quality score obtained by the model evaluation, so as to perform model training, thereby improving the training reliability of the translation quality assessment model, avoiding the problem of inconsistent evaluation accuracy of the translation quality assessment model caused by different evaluation samples during the training process and the actual application process, and further improving the accuracy of the evaluation result obtained by using the translation quality assessment model to perform quality assessment on the actual target translation; wherein, for different labeled translation data situations, utilizing the correlation between the corresponding labeled translation and the translation to be evaluated, and combining the characteristics of different model frameworks, selecting a corresponding appropriate model framework to generate a translation quality assessment model, thereby improving the accuracy of scoring the translation to be evaluated, so as to improve the reliability of the translation quality assessment model training; and utilizing the normalized The evaluation scores are converted into normalized probability distributions to eliminate the influence of different dimensions, and then the KL divergence between the true probability distribution and the predicted probability distribution is calculated, which improves the consistency of the quality score prediction scores and solves the problem that the mean square error loss only considers the difference in the quality score prediction scores. The regression calculation characteristics of the KL divergence loss are used to consider the order correlation between the translations to be evaluated, thereby improving the training reliability of the translation quality evaluation model; in addition, the flexible compatibility of the KL divergence loss is also used, which can be applied to model networks of different architectures (including the above-mentioned BLEURT, COMET, COMETkiwi, etc.), thereby improving the flexibility and versatility of the translation quality evaluation model training; and the batch KL divergence loss is used to optimize the relative order of model predictions to make them closer to the relative order of manual evaluation, which can improve the order and numerical correlation between the model prediction scores and the manual evaluation scores, and the mean square error is used to further improve the numerical correlation between the model prediction scores and the manual evaluation scores, thereby better improving the accuracy and reliability of the translation quality evaluation model training.
[0132] The present application also provides a translation quality assessment model training device, which can implement the above translation quality assessment model training method, referring to Fig. 9 , the device 900 comprises:
[0133] A data set acquisition module 910 is used to acquire a translation quality assessment data set, where the translation quality assessment data set includes at least one training batch, where the training batch includes a plurality of sample text pairs and first translation quality scores of the sample text pairs, where the sample text pairs include at least a translation to be assessed and a labeled translation;
[0134] Scoring module 920, for scoring the translation to be evaluated based on the labeled translation in the sample text using the translation quality evaluation model for each training batch, to obtain a second translation quality score;
[0135] A divergence loss calculation module 930, configured to calculate a batch divergence loss value corresponding to a training batch according to the first translation quality score and the second translation quality score;
[0136] The training module 940 is used to train the translation quality assessment model according to the batch divergence loss value to obtain a trained translation quality assessment model.
[0137] In some embodiments, the divergence loss calculation module 930 is further configured to:
[0138] Normalizing each first translation quality score one by one to obtain a first normalized score;
[0139] Normalizing each second translation quality score one by one to obtain a second normalized score;
[0140] Based on the ratio of the first normalized score and the second normalized score, the batch divergence loss value is obtained.
[0141] In some embodiments, the divergence loss calculation module 930 is further configured to:
[0142] Based on the ratio of each first normalized score to the corresponding second normalized score, a relative ratio is obtained by logarithmic processing;
[0143] Based on the product of the relative ratio and the corresponding second normalized score, the relative batch divergence loss value is obtained;
[0144] All relative batch divergence loss values are averaged to get the batch divergence loss value.
[0145] In some embodiments, the divergence loss calculation module 930 is further configured to:
[0146] Accumulating all first translation quality scores to obtain a first cumulative score;
[0147] Each first translation quality score is taken as a first target score one by one, and a first normalized score is obtained based on a ratio of the first target score to the first accumulated score.
[0148] In some embodiments, the training module 940 is further configured to:
[0149] Get the mean square loss weight value, and get the batch divergence loss weight value based on one minus the mean square loss weight value;
[0150] Obtaining a mean square loss value based on the first translation quality score and all second translation quality scores;
[0151] Based on the product of the mean square loss weight value and the mean square loss value, plus the product of the batch divergence loss weight value and the batch divergence loss value, a composite loss value is obtained, and the translation quality assessment model is trained based on the composite loss value.
[0152] In some embodiments, the scoring module 920 is further configured to:
[0153] When the labeled translation only includes the reference translation, a translation quality assessment model is obtained based on the BLEURT model, and the translation quality assessment model is used to score the translation to be assessed based on the reference translation in the sample text to obtain a second translation quality score;
[0154] When the labeled translation only includes the source text, a translation quality assessment model is obtained based on the COMETkiwi model, and the translation quality assessment model is used to score the translation to be assessed based on the source text in the sample text to obtain a second translation quality score;
[0155] When the labeled translation includes a reference translation and a source text, a translation quality assessment model is obtained based on the COMET model, and the translation quality assessment model is used to score the translation to be assessed based on the source text and the reference translation in the sample text to obtain a second translation quality score.
[0156] In the above embodiments, the description of each embodiment has its own emphasis. For the part that is not described in detail in a certain embodiment, the specific implementation of the translation quality assessment model training device is basically the same as the specific implementation of the above translation quality assessment model training method, and will not be repeated here.
[0157] In an embodiment of the present application, a translation quality assessment model training device uses a batch divergence loss value obtained by calculating the batch divergence loss corresponding to the regression task to fit the score value association and order association between the first translation quality score corresponding to the manual evaluation and the second translation quality score obtained by the model evaluation in the training of the translation quality assessment model corresponding to the training batch including multiple translation texts with an associated order, so as to perform model training, thereby improving the training reliability of the translation quality assessment model, avoiding the problem of inconsistent evaluation accuracy of the translation quality assessment model caused by different evaluation samples during the training process and the actual application process, and thus improving the accuracy of the evaluation result obtained by using the translation quality assessment model to perform quality evaluation on the actual target translation; wherein, for different label translation data situations, the correlation between the corresponding label translation and the translation to be evaluated is used, and the corresponding appropriate model framework is selected in combination with the characteristics of different model frameworks to generate a translation quality assessment model, thereby improving the accuracy of scoring the translation to be evaluated, so as to improve the reliability of the translation quality assessment model training; Furthermore, the evaluation scores are converted into normalized probability distributions by normalization processing to eliminate the influence between different dimensions, and then the KL divergence between the true probability distribution and the predicted probability distribution is calculated, thereby improving the consistency of the quality score prediction scores and solving the problem that the mean square error loss only considers the difference in the quality score prediction scores, so as to utilize the regression calculation characteristics of the KL divergence loss and consider the sequential correlation between the translations to be evaluated, thereby improving the training reliability of the translation quality evaluation model; in addition, the flexible compatibility of the KL divergence loss is also utilized, which can be applied to model networks of different architectures (including the above-mentioned BLEURT, COMET, COMETkiwi, etc.), thereby improving the flexibility and versatility of the translation quality evaluation model training; and, the batch KL divergence loss is used to optimize the relative order of the model predictions to make them closer to the relative order of the manual evaluation, which can improve the order and numerical correlation between the model prediction scores and the manual evaluation scores, and the mean square error is used to further improve the numerical correlation between the model prediction scores and the manual evaluation scores, thereby better improving the accuracy and reliability of the translation quality evaluation model training.
[0158] The present application also provides an electronic device, including:
[0159] at least one memory;
[0160] at least one processor;
[0161] at least one program;
[0162] The program is stored in the memory, and the processor executes the at least one program to implement the above-mentioned translation quality assessment model training method implemented in this application. The electronic device can be any intelligent terminal including a mobile phone, a tablet computer, a personal digital assistant (PDA), a car computer, etc.
[0163] See also Fig.10 , Fig.10 The hardware structure of an electronic device of another embodiment is illustrated, and the electronic device includes:
[0164] The processor 1001 may be implemented by a general-purpose CPU (Central Processing Unit), a microprocessor, an application-specific integrated circuit (Application Specific Integrated Circuit, ASIC), or one or more integrated circuits, and is used to execute relevant programs to implement the technical solutions provided in the embodiments of the present application;
[0165] The memory 1002 can be implemented in the form of ROM (Read Only Memory), static storage device, dynamic storage device or RAM (Random Access Memory). The memory 1002 can store operating systems and other applications. When the technical solutions provided in the embodiments of this specification are implemented by software or firmware, the relevant program codes are stored in the memory 1002, and the processor 1001 calls and executes the translation quality assessment model training method of the embodiments of this application;
[0166] Input / output interface 1003, used to implement information input and output;
[0167] The communication interface 1004 is used to realize the communication interaction between the device and other devices. The communication can be realized through a wired manner (such as USB, network cable, etc.) or a wireless manner (such as mobile network, WIFI, Bluetooth, etc.);
[0168] A bus 1005 , which transmits information between various components of the device (e.g., the processor 1001 , the memory 1002 , the input / output interface 1003 , and the communication interface 1004 );
[0169] The processor 1001 , the memory 1002 , the input / output interface 1003 and the communication interface 1004 are connected to each other in communication within the device via the bus 1005 .
[0170] An embodiment of the present application further provides a storage medium, which is a computer-readable storage medium and stores a computer program. When the computer program is executed by a processor, the above-mentioned translation quality assessment model training method is implemented.
[0171] The memory, as a non-transient computer-readable storage medium, can be used to store non-transient software programs and non-transient computer executable programs. In addition, the memory may include a high-speed random access memory, and may also include a non-transient memory, such as at least one disk storage device, a flash memory device, or other non-transient solid-state storage device. In some embodiments, the memory may optionally include a memory remotely disposed relative to the processor, and these remote memories may be connected to the processor via a network. Examples of the above-mentioned network include, but are not limited to, the Internet, an intranet, a local area network, a mobile communication network, and combinations thereof.
[0172] The embodiments described in the embodiments of the present application are intended to more clearly illustrate the technical solutions of the embodiments of the present application and do not constitute a limitation on the technical solutions provided in the embodiments of the present application. Those skilled in the art will appreciate that with the evolution of technology and the emergence of new application scenarios, the technical solutions provided in the embodiments of the present application are also applicable to similar technical problems.
[0173] Those skilled in the art will appreciate that the technical solutions shown in the figures do not constitute a limitation on the embodiments of the present application, and may include more or fewer steps than shown in the figures, or a combination of certain steps, or different steps.
[0174] The device embodiments described above are merely illustrative, and the units described as separate components may or may not be physically separated, that is, they may be located in one place or distributed on multiple network units. Some or all of the modules may be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0175] Those skilled in the art will appreciate that all or some of the steps in the methods disclosed above, and the functional modules / units in the systems and devices may be implemented as software, firmware, hardware, or a suitable combination thereof.
[0176] The terms "first", "second", "third", "fourth", etc. (if any) in the specification of the present application and the above-mentioned drawings are used to distinguish similar objects, and are not necessarily used to describe a specific order or sequence. It should be understood that the data used in this way can be interchangeable where appropriate, so that the embodiments of the present application described herein can be implemented in an order other than those illustrated or described herein. In addition, the terms "including" and "having" and any of their variations are intended to cover non-exclusive inclusions, for example, a process, method, system, product or device comprising a series of steps or units is not necessarily limited to those steps or units clearly listed, but may include other steps or units that are not clearly listed or inherent to these processes, methods, products or devices.
[0177] It should be understood that in the present application, "at least one (item)" means one or more, and "plurality" means two or more. "And / or" is used to describe the association relationship of associated objects, indicating that three relationships may exist. For example, "A and / or B" can mean: only A exists, only B exists, and A and B exist at the same time, where A and B can be singular or plural. The character " / " generally indicates that the objects associated before and after are in an "or" relationship. "At least one of the following" or similar expressions refers to any combination of these items, including any combination of single or plural items. For example, at least one of a, b or c can mean: a, b, c, "a and b", "a and c", "b and c", or "a and b and c", where a, b, c can be single or multiple.
[0178] In the several embodiments provided in the present application, it should be understood that the disclosed devices and methods can be implemented in other ways. For example, the device embodiments described above are only schematic. For example, the division of the above units is only a logical function division. There may be other division methods in actual implementation, such as multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. The mutual coupling or direct coupling or communication connection shown or discussed can be through some interfaces, indirect coupling or communication connection of devices or units, which can be electrical, mechanical or other forms.
[0179] The units described above as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed on multiple network units. Some or all of the units may be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0180] In addition, each functional unit in each embodiment of the present application may be integrated into one processing unit, or each unit may exist physically separately, or two or more units may be integrated into one unit. The above-mentioned integrated unit may be implemented in the form of hardware or in the form of software functional units.
[0181] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present application, or the part that contributes to the prior art, or all or part of the technical solution can be embodied in the form of a software product, and the computer software product is stored in a storage medium, including multiple instructions to enable a computer device (which can be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods of various embodiments of the present application. The aforementioned storage medium includes: U disk, mobile hard disk, read-only memory (Read-Only Memory, referred to as ROM), random access memory (Random Access Memory, referred to as RAM), disk or optical disk and other media that can store programs.
[0182] The preferred embodiments of the present invention are described above with reference to the accompanying drawings, but the scope of the rights of the present invention is not limited thereto. Any modification, equivalent substitution and improvement made by a person skilled in the art without departing from the scope and essence of the present invention should be within the scope of the rights of the present invention.
Claims
1. A translation quality assessment model training method, characterized in that: The method comprises: Acquire a translation quality assessment dataset, wherein the translation quality assessment dataset includes at least one training batch, wherein the training batch includes a plurality of sample text pairs and first translation quality scores of the sample text pairs, wherein the sample text pairs include at least a translation to be assessed and a label translation; For each of the training batches, scoring the translation to be evaluated based on the labeled translation in the sample text pair using a translation quality evaluation model to obtain a second translation quality score; Calculating a batch divergence loss value corresponding to the training batch according to the first translation quality score and the second translation quality score; The translation quality assessment model is trained according to the batch divergence loss value to obtain the trained translation quality assessment model.
2. The translation quality assessment model training method according to claim 1, characterized in that: The calculating the batch divergence loss value corresponding to the training batch according to the first translation quality score and the second translation quality score includes: Normalizing each of the first translation quality scores one by one to obtain a first normalized score; Normalizing each of the second translation quality scores one by one to obtain a second normalized score; The batch divergence loss value is obtained based on a ratio of the first normalized score to the second normalized score.
3. The translation quality assessment model training method according to claim 2, characterized in that: The obtaining the batch divergence loss value based on the ratio of the first normalized score to the second normalized score comprises: Based on the ratio of each first normalized score to the corresponding second normalized score, logarithmic processing is performed to obtain a relative ratio; Obtaining a relative batch divergence loss value based on the product of the relative ratio and the corresponding second normalized score; All the relative batch divergence loss values are averaged to obtain the batch divergence loss value.
4. The translation quality assessment model training method according to claim 2, characterized in that: The step of normalizing each of the first translation quality scores one by one to obtain a first normalized score includes: Accumulating all the first translation quality scores to obtain a first accumulated score; Each of the first translation quality scores is taken as a first target score one by one, and the first normalized score is obtained based on the ratio of the first target score to the first accumulated score.
5. The translation quality assessment model training method according to claim 1, characterized in that: The training of the translation quality assessment model according to the batch divergence loss value includes: Obtain a mean square loss weight value, and obtain a batch divergence loss weight value based on one minus the mean square loss weight value; Obtaining a mean square loss value based on the first translation quality score and all the second translation quality scores; Based on the product of the mean square loss weight value and the mean square loss value, plus the product of the batch divergence loss weight value and the batch divergence loss value, a composite loss value is obtained, and model training is performed on the translation quality assessment model based on the composite loss value.
6. The translation quality assessment model training method according to claim 1, characterized in that: The step of using the translation quality assessment model to score the translation to be assessed based on the label translation in the sample text pair to obtain a second translation quality score includes: When the labeled translation only includes the reference translation, the translation quality assessment model is obtained based on the BLEURT model, and the translation to be assessed is scored based on the reference translation in the sample text pair using the translation quality assessment model to obtain the second translation quality score; When the labeled translation includes only the source text, the translation quality assessment model is obtained based on the COMETkiwi model, and the translation to be assessed is scored based on the source text in the sample text pair using the translation quality assessment model to obtain the second translation quality score; When the labeled translation includes the reference translation and the source text, the translation quality assessment model is obtained based on the COMET model, and the translation to be assessed is scored using the translation quality assessment model based on the source text and the reference translation in the sample text pair to obtain the second translation quality score.
7. A translation quality assessment method, characterized in that: The method comprises: Obtain target translation data; The target translation data is input into the translation quality assessment model obtained after training as claimed in claim 1 for data processing to obtain a target quality score corresponding to the target translation data.
8. A translation quality assessment model training device, characterized in that: The device comprises: A data set acquisition module, used to acquire a translation quality assessment data set, wherein the translation quality assessment data set includes at least one training batch, wherein the training batch includes a plurality of sample text pairs and first translation quality scores of the sample text pairs, wherein the sample text pairs include at least a translation to be assessed and a labeled translation; A scoring module, configured to score the translation to be evaluated based on the labeled translation in the sample text pair using a translation quality evaluation model for each of the training batches, to obtain a second translation quality score; A divergence loss calculation module, configured to calculate a batch divergence loss value corresponding to the training batch according to the first translation quality score and the second translation quality score; A training module is used to train the translation quality assessment model according to the batch divergence loss value to obtain the trained translation quality assessment model.
9. An electronic device comprising a memory and a processor, wherein the memory stores a computer program, wherein: When the processor executes the computer program, the translation quality assessment model training method according to any one of claims 1 to 6 or the translation quality assessment method according to claim 7 is implemented.
10. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the translation quality assessment model training method according to any one of claims 1 to 6 or the translation quality assessment method according to claim 7 is implemented.
Citation Information
Patent Citations
Machine translation quality evaluation method, device, equipment and medium
CN112347795A
Machine translation quality evaluation method and device, equipment and storage medium
CN115310460A
Multilingual machine translation quality evaluation method, model, equipment and storage medium
CN115952808A
Text translation method and device, computer equipment and readable storage medium
CN118839705A
Machine translation quality estimation method based on multi-semantic space
CN119005214A