A translation quality analysis method and system
By employing a multi-dimensional evaluation approach, combining paragraph unit decomposition, syntactic tree analysis, automatic question-answering model, and paraphrase coding model, the translation is comprehensively scored. This solves the problem of inaccurate evaluation caused by a single reference translation in existing technologies, and achieves a more objective and reliable assessment of translation quality.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- SHENZHEN BILIN SOFTWARE CO LTD
- Filing Date
- 2022-06-08
- Publication Date
- 2026-04-17
AI Technical Summary
Existing translation quality analysis methods rely on a single reference translation, resulting in evaluation results that are not objective and accurate enough, and fail to reflect the differences between various translation styles and approaches.
The translation is comprehensively scored using a multi-dimensional evaluation method, including paragraph unit decomposition, syntax tree analysis, automatic question answering model and paraphrase coding model, combined with weight calculation, taking into account the differences in different translation styles and methods.
It enables a more objective and reliable assessment of translation quality, reflects the differences in various translation styles and methods, and improves the accuracy and reliability of the evaluation.
Smart Images

Figure CN115204191B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of computer technology, and specifically to a method and system for analyzing translation quality. Background Technology
[0002] With societal progress and the increasing exchange of different languages, cultures, and ideas, the demand for translation services for diverse languages and scripts is rapidly growing. Some translation is done manually, while others are automated by robots. Human translation relies on the translator's experience and language skills, resulting in varying levels of quality; while automated robot translation, due to differences in its translation logic, produces inconsistent results.
[0003] Regardless of whether the translation is obtained by human translators or automated translation robots, it needs to be analyzed to determine its quality level. Currently, some researchers have proposed automated translation analysis methods that use a reference translation as an evaluation standard to score the translation being analyzed and derive a final evaluation result, which is then used as the quality level of the translation. This evaluation method has certain limitations: First, it requires prior access to the reference translation, the method of obtaining which is unknown; if it is translated by others or by automated translation, the quality level of the reference translation cannot be determined. Second, there are various ways of expressing a translation, and different expressions do not have a clear hierarchy; therefore, using a single reference translation as an evaluation standard is too one-sided, and the resulting evaluation result is not objective enough.
[0004] Therefore, it is necessary to propose a more reasonable technical solution to solve the technical problems existing in the current technology. Summary of the Invention
[0005] To overcome at least one of the aforementioned defects, this invention proposes a translation quality analysis method and system. By referencing different styles of translation expression, the method evaluates and analyzes the translation through a multi-dimensional evaluation approach. The semantics and syntax of the translation are broken down separately for comprehensive scoring, resulting in a more objective evaluation of translation quality.
[0006] To achieve the above objectives, the present invention may adopt the following technical solution:
[0007] A translation quality analysis method includes:
[0008] The original text to be translated is divided into several paragraph units. Multiple translation reference paragraphs are generated for each paragraph unit. The paragraphs are scored according to their specific expression components. When evaluating each translation reference paragraph and assigning a level value, a syntax tree for each translation reference paragraph is generated. An evaluation base value is assigned based on different syntax trees. The evaluation base value is multiplied by the weight value of the translation reference paragraph to obtain the level value. Each translation reference paragraph is rated and assigned a level value. The translation reference paragraphs are arranged and combined to generate several reference documents. The sum of the level values of all translation reference paragraphs in a reference document is used as the rating score of that reference document.
[0009] The translation to be evaluated is divided into the same number of paragraph units to be analyzed. The degree of similarity of the translation is checked for each paragraph, and the closest reference document is determined. The rating score of the reference document is used as the initial score of the translation to be evaluated.
[0010] The paragraph unit to be analyzed is matched with the translation reference paragraph unit. The level value of the translation reference paragraph unit is used as the paragraph unit value of the paragraph unit to be analyzed. Positive and negative values are obtained by comparing the consistency of question and answer results through an automatic question answering model and comparing the difference in the number of sentence expression types through a paraphrase coding model. The sum of the paragraph unit value, positive value and negative value is used as the evaluation score of the paragraph. The sum of the evaluation scores of all paragraph units to be analyzed is used as the evaluation score of the translation.
[0011] The aforementioned publicly disclosed translation quality analysis method first generates several source text translation reference documents and rates each reference document to determine the quality of different translation methods. Then, it matches the translation to be evaluated to a specific level, obtaining an initial score for that level. During translation within that level, based on the differences between the translation to be evaluated and the reference translations, it adjusts the score of the translation to be analyzed using positive and negative values to obtain the final translation quality score.
[0012] The quality assessment method does not rely on a single reference translation as the standard, but takes into account different translation styles and methods. Multiple translation styles can be objectively and reasonably evaluated and identified, making the method more reliable.
[0013] Furthermore, when processing the original text to be translated, this invention processes it into paragraph units and obtains their translation reference paragraphs. The paragraph units to be evaluated are then analyzed in correspondence with the translation reference paragraphs. Simultaneously, the language of the translation reference paragraphs is the same as the language of the translation to be evaluated. The language of the translation reference paragraphs is determined after identifying the language of the translation to be evaluated. Specifically, this invention employs the following method: before breaking down the original text into multiple paragraph units and generating translation reference paragraphs, the language of the translation to be evaluated is identified. When the translation to be evaluated includes two or more languages, the weight of each language is determined by the number of characters, and the language with the highest weight is designated as the target language, with the remaining languages designated as supplementary languages. Using this approach, the target language is used as the language of the reference translation when translating the original text.
[0014] Furthermore, the entire text is divided into multiple paragraph units for individual scoring. Since there are certain differences between these paragraph units, a weighting system needs to be established for scoring, increasing the weight of more important paragraphs and decreasing the weight of relatively less important ones. The method of setting weights is not unique; this invention proposes the following optimized approach: determine the weight based on the number of words in each paragraph unit, using the percentage of each paragraph unit's words in the total number of words as its weight value. With this approach, paragraph units with more content have a higher weight, and those with less content have a lower weight.
[0015] Furthermore, in this invention, when specifically evaluating the level value of each paragraph unit, scoring is performed according to the specific paragraph expression composition. Here, an optimization is made, and the following feasible option is proposed: When evaluating each translation reference paragraph and assigning a level value, a syntax tree for each translation reference paragraph is generated. An evaluation base value is assigned based on different syntax trees. The level value is obtained by multiplying the evaluation base value with the weight value of the translation reference paragraph. Using this approach, several syntax trees are preset, and corresponding base values are assigned. Therefore, once the syntax tree of a translation reference paragraph is determined, the base value for that translation reference paragraph is determined, thereby determining its level value. Simultaneously, the maximum rating score for a reference document is 100 points, meaning the sum of the standard scores for each translation reference paragraph is 100 points. However, depending on different translation methods, the scores of multiple translation reference paragraphs obtained for a single paragraph unit may differ. Therefore, the rating scores of reference documents vary, with most reference documents having a rating score below 100 points.
[0016] Furthermore, to compare the translation comprehensibility between the reference document and the translation to be evaluated, the substantive translation content of the translation to be evaluated is examined through a question-and-answer format, thereby obtaining the substantive content in the translation. Specifically, in this invention, one of the following options can be adopted: An automatic question-and-answer model is trained, and the paragraph unit to be analyzed and the translation reference paragraph unit are respectively input into the automatic question-and-answer model to obtain multiple sets of question-and-answer results and compare the consistency of the question-and-answer results; when the question-and-answer results of the paragraph unit to be analyzed differ from those of the translation reference paragraph unit, a negative value is recorded for the paragraph unit to be analyzed, and when the question-and-answer results of the paragraph unit to be analyzed are consistent with those of the reference paragraph unit, a positive value is recorded for the paragraph unit to be analyzed; when the number of question-and-answer result sets of the paragraph unit to be analyzed differs from the number of question-and-answer result sets of the translation reference paragraph unit, a negative value is recorded for the paragraph unit to be analyzed, and when the number of question-and-answer result sets of the paragraph unit to be analyzed is equal to the number of question-and-answer result sets of the translation reference paragraph unit, a positive value is recorded for the paragraph unit to be analyzed. When using this approach, it indicates the degree of similarity between the substantive content of the translation to be analyzed and the content of the reference document. The higher the similarity, the higher the consistency of the question and answer results, and the more accurate the scoring results of the translation to be evaluated.
[0017] Furthermore, to compare the consistency of expression between the reference document and the document to be evaluated, and to determine the descriptive richness of the translation to be evaluated, several feasible solutions can be adopted. This invention optimizes and provides one feasible option: train a paraphrasing encoding model, input the paragraph unit to be analyzed and the translation reference paragraph unit into the paraphrasing encoding model respectively, and count the number of sentence expression types in each model. When the number of sentence expression types in the paragraph unit to be analyzed is equal to that in the translation reference paragraph unit...
[0018] When the difference in the number of expression types among the units is within a preset range, a positive value is recorded for the unit to be analyzed; otherwise, a negative value is recorded. Using this approach, each sentence in the reference translation can be analyzed to determine its linguistic expression type and statistically analyze it. Simultaneously, the linguistic expression type of each unit in the translation to be evaluated is determined sentence-by-sentence and statistically analyzed. Therefore, the degree of similarity between the linguistic expression of the units in the translation to be evaluated and those in the reference translation can be determined; a higher degree of similarity results in a higher score, and vice versa.
[0019] Furthermore, among the various translation methods corresponding to paragraph units, when generating multiple translation reference paragraphs for each paragraph unit, different translation modes are referenced to achieve the generation of translation reference units. In this invention, optimization and improvement are made, and the following feasible option is proposed: When generating multiple translation reference paragraphs for each paragraph unit, according to the distinction of using syntax trees as translation reference paragraphs, multiple different syntax trees are selected, and a translation reference paragraph is generated according to each syntax tree. When adopting such a scheme, the syntax tree is used as the overall framework of a paragraph unit, thereby distinguishing different paragraph unit structures.
[0020] Furthermore, when matching the paragraph unit to be analyzed to a close translation reference paragraph, one feasible option is to determine the syntax tree of the paragraph unit to be analyzed and compare it with the syntax tree of the translation reference paragraph. The translation reference paragraph with the closest syntax tree is then selected as the closest translation reference paragraph. Using this approach, the translation reference paragraph can be quickly determined based on the syntax tree results, thereby quickly determining the paragraph unit value of the translated text paragraph unit to be analyzed.
[0021] Furthermore, when determining the closest reference document, the paragraph unit is used as the smallest unit. The closest translation reference paragraph is determined for each paragraph unit, and the resulting document, after arranging these closest reference paragraphs, is considered the closest reference document. Using this approach allows for the evaluation of the translation based on a relatively similar translation method and style.
[0022] The above content describes the method for evaluating translation quality. This invention also discloses a system for evaluating translation quality, which will be explained below.
[0023] A translation quality analysis system, comprising:
[0024] The acquisition unit is used to acquire translation reference paragraphs. When the original text to be translated is divided into several paragraph units, the acquisition unit reads and translates each paragraph unit and obtains the level value of the translation reference paragraph according to multiple syntax trees.
[0025] The paragraph identification and segmentation unit is used to identify paragraphs in the original text to be translated, the reference translation, and the translation to be evaluated, and to segment them into several paragraph units respectively.
[0026] The weight calculation unit is used to analyze the original text to be translated to determine the weight of each paragraph unit;
[0027] The language recognition unit is used to identify paragraph units of the translation to be analyzed and determine the language type within the paragraph units.
[0028] The language analysis unit is used to identify and statistically analyze the expression patterns of sentences in each paragraph unit;
[0029] The syntactic analysis and evaluation unit is used to analyze each paragraph unit and form a syntactic tree, while analyzing and comparing the similarity of the syntactic trees of the corresponding paragraph units of the reference translation and the translation to be evaluated;
[0030] The calculation unit is used to calculate the paragraph unit value, positive and negative values of the translated paragraph unit to be evaluated by obtaining positive and negative values.
[0031] The aforementioned publicly available translation quality analysis system can be used to break down the original text to be translated and generate several translation reference documents for each paragraph unit. When analyzing the translation to be analyzed, each paragraph unit of the translation to be analyzed is matched with a translation reference document. Finally, a complete reference document is formed based on the combination of translation reference paragraphs matched with the paragraph units. The initial score of the reference document is then adjusted to obtain the overall score of the translation to be evaluated.
[0032] Compared with the prior art, some of the beneficial effects of the technical solution disclosed in this invention include:
[0033] This invention provides translation reference paragraphs with various translation methods and styles, each with different ratings and scores. The scores of each paragraph are weighted and factored into the overall evaluation of the translation reference document to determine the final score of the translation being evaluated. The method provided by this invention offers more references, a more objective evaluation, and more reliable evaluation results. Attached Figure Description
[0034] To more clearly illustrate the technical solutions of the embodiments of the present invention, the accompanying drawings used in the embodiments will be briefly introduced below. It should be understood that the following drawings only show some embodiments of the present invention and should not be regarded as a limitation on the scope. For those skilled in the art, other related drawings can be obtained based on these drawings without creative effort.
[0035] Figure 1 This is a schematic diagram illustrating the overall logic of the evaluation method.
[0036] Figure 2 This is a schematic diagram of the components of the evaluation system.
[0037] Figure 3 This is a schematic diagram illustrating the question-answering result analysis process of an automatic question-answering model.
[0038] Figure 4 This is a schematic diagram illustrating the training process of the encoding model. Detailed Implementation
[0039] The present invention will be further explained below with reference to the accompanying drawings and specific embodiments.
[0040] In view of the fact that existing translation quality analysis methods use a single reference standard, the evaluation results are not objective and accurate enough, resulting in large evaluation errors. This embodiment optimizes and improves the method to overcome the shortcomings of the existing technology.
[0041] Example 1
[0042] This embodiment discloses a translation quality analysis method, which obtains corresponding reference content from multiple translation reference paragraphs to evaluate the translation to be evaluated, specifically including:
[0043] S1: Divide the original text to be translated into several paragraph units, generate multiple translation reference paragraphs for each paragraph unit, rate each translation reference paragraph and assign a grade value; arrange and combine the translation reference paragraphs to generate several reference documents, and use the sum of the grade values of all translation reference paragraphs in the reference document as the rating score of the reference document.
[0044] S2: Divide the translation to be evaluated into the same number of paragraph units to be analyzed, and check the degree of similarity of the translation for each paragraph to determine the closest reference document. Use the rating score of the reference document as the initial score of the translation to be evaluated.
[0045] S3: Match the paragraph unit to be analyzed with the translation reference paragraph unit, use the level value of the translation reference paragraph unit as the paragraph unit value of the paragraph unit to be analyzed, and obtain the positive and negative values of the paragraph unit to be analyzed after comparing with the translation reference paragraph unit. Use the sum of the paragraph unit value, positive value and negative value as the evaluation score of the paragraph, and use the sum of the evaluation scores of all paragraph units to be analyzed as the evaluation score of the translation.
[0046] The aforementioned publicly available translation quality analysis method first generates several source text translation reference documents and rates each reference document to determine the quality of different translation styles. Then, it matches the translation to be evaluated to a specific level, obtaining an initial score for that level. During translation within that level, the score is adjusted based on the differences between the translation to be evaluated and the reference translations, using positive and negative values to arrive at the final translation quality score. This quality assessment method does not rely on a single reference translation as the standard but considers different translation styles and methods, ensuring that multiple translation styles can be objectively and reasonably evaluated, thus enhancing reliability.
[0047] In this embodiment, when processing the original text to be translated, it processes each paragraph unit individually to obtain its translation reference paragraph. The paragraph unit to be evaluated is then analyzed in correspondence with the translation reference paragraph. Furthermore, the language of the translation reference paragraph is the same as the language of the text to be evaluated. The language of the translation reference paragraph is determined after identifying the language of the text to be evaluated. Specifically, this...
[0048] The implementation method is as follows: before splitting the original text to be translated into multiple paragraph units and generating translation reference paragraphs, the system identifies...
[0049] When evaluating a translation in two or more languages, determine the weight of each language based on its word count, designating the language with the highest weight as the target language and the others as supplementary languages. In this approach, use the target language as the reference language for the translation and translate the original text accordingly.
[0050] The entire text is divided into multiple paragraph units for individual scoring. Since there are inherent differences between these paragraph units, a weighting system needs to be established for scoring, increasing the weight of more important paragraphs and decreasing the weight of less important ones. The method for setting these weights is not unique; this embodiment proposes an optimized approach: determining the weight based on the number of words in each paragraph unit, using the percentage of each paragraph unit's words in the total word count as its weight value. With this approach, paragraph units with more content have a higher weight, and those with less content have a lower weight.
[0051] In this embodiment, when evaluating the level value of each paragraph unit, scoring is performed according to the specific paragraph expression composition. An optimization is made here, employing the following feasible option: When evaluating each translation reference paragraph and assigning a level value, a syntax tree for each translation reference paragraph is generated. An evaluation base value is assigned based on different syntax trees. The level value is obtained by multiplying the evaluation base value with the weight value of the translation reference paragraph. With this approach, several syntax trees are pre-defined, and corresponding base values are assigned. Therefore, once the syntax tree of a translation reference paragraph is determined, the base value for that translation reference paragraph is determined, thereby determining its level value. Simultaneously, the maximum rating score for a reference document is 100 points, meaning the sum of the standard scores for each translation reference paragraph is 100 points. However, depending on different translation methods, the scores of multiple translation reference paragraphs for a single paragraph unit may differ, resulting in variations in the rating scores of reference documents, with most reference documents having rating scores below 100 points.
[0052] To compare the translation comprehensibility between the reference document and the translation to be evaluated, a question-and-answer format is used to examine the substantive translation content of the translation to be evaluated, thereby obtaining the substantive content in the translation. Specifically, in this embodiment, one of the following options can be adopted: train an automatic question-and-answer model, input the paragraph unit to be analyzed and the translation reference paragraph unit into the automatic question-and-answer model respectively, obtain multiple sets of question-and-answer results, and compare the consistency of the question-and-answer results; when the question-and-answer results of the paragraph unit to be analyzed are different from those of the translation reference paragraph unit, a negative value is recorded for the paragraph unit to be analyzed, and when the question-and-answer results of the paragraph unit to be analyzed are consistent with those of the reference paragraph unit, a positive value is recorded for the paragraph unit to be analyzed; when the number of question-and-answer result sets of the paragraph unit to be analyzed is different from the number of question-and-answer result sets of the translation reference paragraph unit, a negative value is recorded for the paragraph unit to be analyzed, and when the number of question-and-answer result sets of the paragraph unit to be analyzed is equal to the number of question-and-answer result sets of the translation reference paragraph unit, a positive value is recorded for the paragraph unit to be analyzed. When using this approach, it indicates the degree of similarity between the substantive content of the translation to be analyzed and the content of the reference document. The higher the similarity, the higher the consistency of the question and answer results, and the more accurate the scoring results of the translation to be evaluated.
[0053] When analyzing question-answering results using an automatic question-answering model, the following process is included:
[0054] S01: Input the translation to be evaluated into a pre-trained automatic question answering model to obtain at least one set of questions and answers, wherein the automatic question answering model is a neural network used to extract questions and answers from the text;
[0055] S02: Determine the percentage of correct answers in the at least one set of questions and answers as the question-and-answer score;
[0056] S03: Obtain the standard score obtained by analyzing the answer results of the automatic question-answering model on the reference translation;
[0057] S04: The question-and-answer score is corrected using the standard score to obtain the comprehensibility score of the translation to be evaluated.
[0058] To compare the consistency of expression between the reference document and the document to be evaluated, and to determine the descriptive richness of the translation to be evaluated, several feasible solutions can be adopted. This embodiment optimizes and adopts one of the feasible options: training to obtain...
[0059] The paraphrasing coding model takes the paragraph unit to be analyzed and the translation reference paragraph unit as inputs, and then performs statistical analysis on each.
[0060] The number of sentence expression types is calculated as follows: if the difference between the number of sentence expression types in the paragraph unit being analyzed and the number of sentence expression types in the translation reference paragraph unit is within a preset range, a positive value is recorded for the paragraph unit being analyzed; otherwise, a negative value is recorded. Using this approach, each sentence in the translation reference paragraph can be analyzed to determine its language expression type and statistically analyzed. Simultaneously, the language expression type used in the paragraph unit of the translation to be evaluated is determined sentence-by-sentence and statistically analyzed. Therefore, the degree of similarity between the paragraph unit of the translation to be evaluated and the paragraph unit of the translation reference in terms of expression type can be determined. A higher degree of similarity results in a higher score, and vice versa.
[0061] Preferably, the restatement coding model is trained using the following method:
[0062] S01: Obtain a set of original sentences in the first language that are in the same language as the translation to be evaluated;
[0063] S02: For each original sentence in the first language in the set of original sentences in the first language, the original sentence in the first language is translated into a translation in the second language through the first translation model, and then the translation in the second language is translated into a restatement sentence in the first language through the second translation model. The original sentence and the restatement sentence in the first language are combined into a restatement sentence pair, and a sentence is randomly selected and combined with the original sentence in the first language to form a non-restatement sentence pair.
[0064] S03: Using the set of restatement pairs as positive examples and the non-restatement pairs as negative examples, a classifier is trained using machine learning methods to obtain the restatement encoding model.
[0065] Among the various translation methods corresponding to paragraph units, when generating multiple translation reference paragraphs for each paragraph unit, different translation modes are referenced to achieve the generation of translation reference units. In this embodiment, optimization and improvement are made, and the following feasible option is adopted: when generating multiple translation reference paragraphs for each paragraph unit, according to the distinction between syntax trees as translation reference paragraphs, multiple different syntax trees are selected, and a translation reference paragraph is generated according to each syntax tree. When adopting this scheme, the syntax tree is used as the overall framework of a paragraph unit, thereby distinguishing different paragraph unit structures.
[0066] When matching the paragraph unit to be analyzed to a close translation reference paragraph, one feasible option is to determine the syntax tree of the paragraph unit to be analyzed and compare it with the syntax tree of the translation reference paragraph. The translation reference paragraph with the closest syntax tree is then selected as the closest translation reference paragraph. Using this approach, the translation reference paragraph can be quickly determined based on the syntax tree results, thereby quickly determining the paragraph unit value of the translated text paragraph unit to be analyzed.
[0067] When determining the closest reference document, the paragraph unit is used as the smallest unit. The closest translation reference paragraph is identified for each paragraph unit, and the resulting document, arranged in order of these closest reference paragraphs, is considered the closest reference document. This approach allows for the evaluation of the translation based on a relatively similar translation method and style.
[0068] Example 2
[0069] The above embodiments describe the methods for evaluating translation quality. This embodiment also discloses a translation quality evaluation system, which will be explained below.
[0070] A translation quality analysis system, comprising:
[0071] The acquisition unit is used to obtain translation reference paragraphs. When the original text to be translated is divided into several paragraph units, the acquisition unit reads and translates each paragraph unit and obtains translation reference paragraphs according to multiple syntax trees.
[0072] The paragraph identification and segmentation unit is used to identify paragraphs in the original text to be translated, the reference translation, and the translation to be evaluated, and to segment them into several paragraph units respectively.
[0073] The weight calculation unit is used to analyze the original text to be translated to determine the weight of each paragraph unit;
[0074] The language recognition unit is used to identify paragraph units of the translation to be analyzed and determine the language type within the paragraph units.
[0075] The language analysis unit is used to identify and statistically analyze the expression patterns of sentences in each paragraph unit;
[0076] The syntactic analysis and evaluation unit is used to analyze each paragraph unit and form a syntactic tree, while analyzing and comparing the similarity of the syntactic trees of the corresponding paragraph units of the reference translation and the translation to be evaluated;
[0077] The calculation unit calculates the paragraph unit value, positive and negative values of the translated paragraph unit to be evaluated by obtaining positive and negative values.
[0078] The aforementioned publicly available translation quality analysis system can be used to break down the original text to be translated and generate several translation reference documents for each paragraph unit. When analyzing the translation to be analyzed, each paragraph unit of the translation to be analyzed is matched with a translation reference document. Finally, a complete reference document is formed based on the combination of translation reference paragraphs matched with the paragraph units. The initial score of the reference document is then adjusted to obtain the overall score of the translation to be evaluated.
[0079] The above are the embodiments listed in this example. However, this example is not limited to the optional embodiments described above. Those skilled in the art can arbitrarily combine the above methods to obtain other various embodiments. Anyone can derive other various forms of embodiments under the guidance of this example. The above specific embodiments should not be construed as limiting the scope of protection of this example. The scope of protection of this example should be defined in the claims.
Claims
1. A method for analyzing the quality of a translation, characterized in that, include: The original text to be translated is divided into several paragraph units. Multiple translation reference paragraphs are generated for each paragraph unit. The paragraphs are scored according to their specific expression components. When evaluating each translation reference paragraph and assigning a level value, a syntax tree for each translation reference paragraph is generated. An evaluation base value is assigned based on the different syntax trees. The evaluation base value is multiplied by the weight value of the translation reference paragraph to obtain the level value. Each translation reference paragraph is rated and assigned a level value. The translation reference paragraphs are arranged and combined to generate several reference documents. The sum of the level values of all translation reference paragraphs in a reference document is used as the rating score of that reference document. The weight of each paragraph unit is determined based on the number of words in each paragraph unit, and the weight value of each paragraph unit is the percentage of the number of words in the total number of words. When evaluating each translation reference paragraph and assigning a level value, a syntax tree for each translation reference paragraph is generated. An evaluation base value is assigned based on the different syntax trees. The level value is obtained by multiplying the evaluation base value with the weight value of the translation reference paragraph. The translation to be evaluated is divided into the same number of paragraph units to be analyzed, and the degree of similarity of the translation is checked for each paragraph. The closest reference document is determined, and the rating score of the reference document is used as the initial score of the translation to be evaluated. The paragraph unit to be analyzed is matched with the translation reference paragraph unit. The level value of the translation reference paragraph unit is used as the paragraph unit value of the paragraph unit to be analyzed. Positive and negative values are obtained by comparing the consistency of question and answer results through an automatic question answering model and comparing the difference in the number of sentence expression types through a paraphrase coding model. The sum of the paragraph unit value, positive value and negative value is used as the evaluation score of the paragraph. The sum of the evaluation scores of all paragraph units to be analyzed is used as the evaluation score of the translation. An automatic question-answering model is trained. The paragraph unit to be analyzed and the translation reference paragraph unit are input into the automatic question-answering model respectively. Multiple sets of question-answering results are obtained and their consistency is compared. When the question-answering result of the paragraph unit to be analyzed is different from that of the translation reference paragraph unit, a negative value is recorded for the paragraph unit to be analyzed. When the question-answering result of the paragraph unit to be analyzed is consistent with that of the reference paragraph unit, a positive value is recorded for the paragraph unit to be analyzed. When the number of question-answering result sets of the paragraph unit to be analyzed is different from that of the translation reference paragraph unit, a negative value is recorded for the paragraph unit to be analyzed. When the number of question-answering result sets of the paragraph unit to be analyzed is equal to that of the translation reference paragraph unit, a positive value is recorded for the paragraph unit to be analyzed.
2. The translation quality analysis method according to claim 1, characterized in that: Before breaking down the original text to be translated into multiple paragraph units and generating translation reference paragraphs, the language of the translation to be evaluated is identified. When the translation to be evaluated includes more than two languages, the weight of each language is determined by the number of words in each language, and the language with the highest weight is taken as the target language, while the remaining languages are taken as supplementary languages.
3. The translation quality analysis method according to claim 1, characterized in that: The paraphrasing encoding model is trained. The paragraph unit to be analyzed and the translation reference paragraph unit are input into the paraphrasing encoding model respectively. The number of sentence expression types is counted for each. When the difference between the number of sentence expression types of the paragraph unit to be analyzed and the number of sentence expression types of the translation reference paragraph unit is within a preset range, a positive value is recorded for the paragraph unit to be analyzed; otherwise, a negative value is recorded.
4. The translation quality analysis method according to claim 1, characterized in that: When generating multiple translation reference paragraphs for each paragraph unit, different syntax trees are selected based on the distinction between them as translation reference paragraphs, and a translation reference paragraph is generated for each syntax tree.
5. The translation quality analysis method according to claim 1, characterized in that: The syntax tree of the paragraph unit to be analyzed is determined and compared with the syntax tree of the translation reference paragraph. The translation reference paragraph with the closest syntax tree is selected as the closest translation reference paragraph.
6. The translation quality analysis method according to claim 4, characterized in that: When determining the closest reference document, the paragraph unit is used as the smallest unit. The closest translation reference paragraph is determined for each paragraph unit. The reference document determined by arranging the closest translation reference paragraphs is taken as the closest reference document.
7. A translation quality analysis system, characterized in that, include: The paragraph identification and segmentation unit is used to split the original text to be translated into several paragraph units, and to split the translation to be evaluated into the same number of paragraph units to be analyzed. The syntactic analysis and evaluation unit is used to generate multiple translation reference paragraphs for each paragraph unit, score them according to the specific paragraph expression composition, generate the syntactic tree for each translation reference paragraph when evaluating each translation reference paragraph and assigning a level value, and check the similarity of the translation for each paragraph to determine the closest reference document, and use the rating score of the reference document as the initial score of the translation to be evaluated. The weight calculation unit is used to determine the weight of each paragraph unit based on the number of words in each paragraph unit. The weight value of each paragraph unit is the proportion of the number of words in the total number of words. The acquisition unit is used to assign an evaluation base value based on different syntax trees. The evaluation base value is multiplied by the weight value of the translation reference paragraph to obtain the level value. Each translation reference paragraph is rated and assigned a level value. When evaluating each translation reference paragraph and assigning a level value, the syntax tree of each translation reference paragraph is generated accordingly. An evaluation base value is assigned based on different syntax trees. The evaluation base value is multiplied by the weight value of the translation reference paragraph to obtain the level value. The language recognition unit is used to identify the language of the translation to be evaluated before splitting the original text to be translated into multiple paragraph units and generating translation reference paragraphs. When the translation to be evaluated includes two or more languages, the weight of each language is determined by the number of words in each language, and the language with the highest weight is taken as the target language, and the remaining languages are taken as additional languages. The language analysis unit is used to arrange and combine the translation reference paragraphs to generate several reference documents. The sum of the level values of all translation reference paragraphs in the reference document is used as the rating score of the reference document. The calculation unit matches the paragraph unit to be analyzed with the translation reference paragraph unit. The level value of the translation reference paragraph unit is used as the paragraph unit value of the paragraph unit to be analyzed. Positive and negative values are obtained by comparing the consistency of question-and-answer results using an automatic question-and-answer model and comparing the difference in the number of sentence expression types using a paraphrasing encoding model. The sum of the paragraph unit value, positive values, and negative values is used as the paragraph's evaluation score. The sum of the evaluation scores of all paragraph units to be analyzed is used as the translation's evaluation score. An automatic question-and-answer model is trained, and the paragraph unit to be analyzed and the translation reference paragraph unit are input into the automatic question-and-answer model respectively. In the answer model, multiple sets of question-and-answer results are obtained and their consistency is compared. When the question-and-answer results of the paragraph unit to be analyzed differ from those of the translation reference paragraph unit, a negative value is recorded for the paragraph unit to be analyzed, and when the question-and-answer results of the paragraph unit to be analyzed are consistent with those of the reference paragraph unit, a positive value is recorded for the paragraph unit to be analyzed. When the number of question-and-answer result sets of the paragraph unit to be analyzed differs from the number of question-and-answer result sets of the translation reference paragraph unit, a negative value is recorded for the paragraph unit to be analyzed, and when the number of question-and-answer result sets of the paragraph unit to be analyzed is equal to the number of question-and-answer result sets of the translation reference paragraph unit, a positive value is recorded for the paragraph unit to be analyzed.
Citation Information
Patent Citations
Method and device for evaluating translation quality
CN111027331A
Method and apparatus for evaluating translation quality
US20210174033A1