Evaluation method and device based on document-level machine translation, storage medium and equipment
Patent Information
- Application Number
- CN202610707154.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2026-05-21
- Publication Date
- 2026-09-18
- Estimated Expiration
- 2046-05-21
AI Technical Summary
[0011]本申请提供了一种基于文档级机器翻译的评测方法、装置、存储介质及设备,用于解决文档级机器翻译的评测结果不准确的问题
通过构建实体一致性、时态一致性和篇章连贯性三个维度的综合评测框架,相比于传统测评指标仅能评估词汇重合,在篇章层面失效的技术缺陷,本申请能真实反映译文在长文档中的可读性和逻辑严密性,为文档级翻译提供了可靠的量化标准。同时,本申请仅依赖于源文文档、机器译文和参考译文的文本内容,无需访问翻译模型的内部结构,可广泛应用于各类黑盒翻译系统的质量测试,具有极高的工程落地价值。
Smart Images

Figure CN122389895B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of machine translation technology, and in particular to an evaluation method, apparatus, storage medium, and device based on document-level machine translation. Background Technology
[0002] Document-level machine translation refers to machine translation that targets complete documents. During the translation process, it not only handles dependencies within sentences but also considers discourse phenomena between sentences (such as reference, cohesion, and coherence). We need to calculate the accuracy of document-level machine translation to evaluate its effectiveness.
[0003] Currently, the evaluation of document-level machine translation mainly adopts the following technical solutions: (1) Evaluation method based on n-gram matching These technical solutions use the surface-level lexical matching degree between the machine translation and the reference translation as the core evaluation criterion. Their core logic is to calculate the translation quality score by utilizing the matching ratio of n-grams between the machine translation and the reference translation. Typical indicators include Bilingual Evaluation Understudy (BLEU) and Metric for Evaluation of Translation with Explicit Ordering (METEOR). BLEU primarily uses the exact matching rate, combined with a length penalty factor, to quantitatively evaluate the overall quality of the machine translation. METEOR, on the other hand, introduces mechanisms such as word form restoration and synonym matching to improve tolerance for lexical variations.
[0004] However, n-gram-based evaluation methods mainly focus on lexical similarity at the sentence level, and have limited ability to reflect cross-sentence semantic relationships, contextual consistency and overall document readability. They cannot reflect the overall quality of the document, which means that machine translations with high BLEU scores may still have obvious entity confusion, tense jumps and logical breaks at the document level.
[0005] (2) Evaluation method for semantic representation based on pre-trained language model These technical solutions introduce deep learning models to model semantic vectors of machine translations and reference translations, and evaluate translation quality based on distance or similarity in the semantic space. Typical methods, such as COMET, utilize pre-trained language models to jointly encode the source language text, machine translation, and reference translation, thereby capturing deeper semantic correspondences. This can alleviate errors caused by surface-level lexical mismatches to some extent and has better robustness to semantically equivalent but differently expressed translations.
[0006] Evaluation based on pre-trained language models is computationally expensive, and the evaluation results are significantly affected by model structure, training corpus, and parameter settings. Furthermore, these methods primarily focus on overall semantic consistency, lacking specific explicit modeling for fine-grained linguistic phenomena unique to document-level machine translation, such as entity consistency, tense inference, and discourse coherence.
[0007] (3) Local evaluation methods for specific language phenomena These technical solutions target specific linguistic phenomena in machine translation for specialized evaluation. For example, they assess pronoun translation accuracy (APT). They identify the correspondence between source language text and machine-translated text through word alignment algorithms or coreference resolution tools; or they analyze and evaluate the handling of certain linguistic phenomena during translation by accessing the model's internal representation. These solutions are highly targeted in analyzing the model's ability to handle specific linguistic phenomena and are suitable for research-based evaluation or model diagnostic tasks.
[0008] Local evaluation methods for specific linguistic phenomena often rely on complex third-party tools (such as word alignment tools like FastAlign or parsers) or require access to the internal attention weights of the translation model. This "white-box" or heavily reliant evaluation approach has significant limitations when evaluating commercial black-box systems (such as Google Translate and DeepL).
[0009] (4) Manual evaluation method Human evaluation methods typically involve subjective assessments of machine translation results by experts with backgrounds in linguistics or translation. Evaluation criteria generally include accuracy, fluency, contextual consistency, and overall readability. Human evaluation comprehensively considers various linguistic phenomena, providing a relatively complete judgment on document-level translation quality, and is considered one of the reference standards for machine translation evaluation.
[0010] Manual evaluation methods are costly, time-consuming, and difficult to scale. Furthermore, subjective differences exist between different evaluators, making it difficult to achieve large-scale, repeatable, automated model iteration. Summary of the Invention
[0011] This application provides a method, apparatus, storage medium, and device for evaluating document-level machine translation, to address the problem of inaccurate evaluation results for document-level machine translation. The technical solution is as follows: According to a first aspect of this application, a document-level machine translation evaluation method is provided, the method comprising: Obtain the source document and its corresponding machine translation and reference translation; Generate coreference chains for entities in the source document, find the set of translated items corresponding to each coreference chain in the machine translation, and the translated items in the translated item set correspond to the mentions in the coreference chain; generate an entity consistency score based on the comprehensive similarity between every two translated items in the translated item set, and the entity consistency score represents the consistency of nouns of the same entity in the machine translation. The machine translation and the reference translation are labeled with tense to obtain a first tense tag sequence and a second tense tag sequence. A tense consistency score is generated based on the matching degree of the first tense tag sequence and the second tense tag sequence corresponding to the same sentence in the machine translation and the reference translation. The tense consistency score represents the consistency of the machine translation in terms of time expression. According to a preset mapping table, the related words in the machine translation and the reference translation are mapped to the corresponding semantic relationship categories respectively; a discourse coherence score is generated based on the matching degree of the semantic relationship categories corresponding to the related words at the same semantic position in the machine translation and the reference translation, and the discourse coherence score represents the consistency of the machine translation in the discourse logical relationship. A comprehensive score for document-level machine translation is generated based on the entity consistency score, the temporal consistency score, and the discourse coherence score.
[0012] In one possible implementation, generating an entity consistency score based on the comprehensive similarity between every two translated items in the translated item set includes: For each set of translated items corresponding to a coreference chain, calculate the comprehensive similarity between every two translated items in the set, and calculate the entity consistency score of the set based on all comprehensive similarities. The final entity consistency score is calculated based on the entity consistency scores of all translated item sets.
[0013] In one possible implementation, calculating the comprehensive similarity between every two translated items in the set of translated items includes: For every two translated items in the set of translated items, calculate the semantic similarity, edit distance similarity, and speech similarity between the two translated items; The semantic similarity, the edit distance similarity, and the speech similarity are weighted and fused to obtain the comprehensive similarity between the two translated items.
[0014] In one possible implementation, calculating the semantic similarity, edit distance similarity, and speech similarity between the two translated items includes: The semantic similarity is obtained by calculating the cosine similarity between the word vectors of the two translated terms. Calculate the edit distance between the two translated items to obtain the edit distance similarity. The phonetic similarity between the two translated items is calculated based on a string encoding algorithm or a phonetic transcription conversion algorithm.
[0015] In one possible implementation, generating a temporal consistency score based on the matching degree of the first and second temporal tag sequences corresponding to the same statement in the machine translation and the reference translation includes: The machine-translated text is divided into a first explicit set and a first implicit set, and the reference translation is divided into a second explicit set and a second implicit set. The first explicit set and the second explicit set are sets of statements that contain time clues, and the first implicit set and the second implicit set are sets of statements that do not contain the time clues. For the same statement in the first explicit set and the second explicit set, determine whether the first tense label sequence and the second tense label sequence corresponding to the statement are consistent; calculate the proportion of statements with consistent tense label sequences to obtain the explicit matching degree. For the same statement in the first implicit set and the second implicit set, determine whether the first tense label sequence and the second tense label sequence corresponding to the statement are consistent; calculate the proportion of statements with consistent tense label sequences to obtain the implicit matching degree. The explicit matching degree and the implicit matching degree are weighted and fused to obtain the temporal consistency score.
[0016] In one possible implementation, the step of performing temporal tagging on the statements in the machine translation and the reference translation to obtain a first temporal tag sequence and a second temporal tag sequence includes: For each sentence in the machine translation, a tense tagger is used to tag each word in the sentence with a tense category label to obtain a first tense tag sequence; For each sentence in the reference translation, a tense labeler is used to label each word in the sentence with a tense category label, resulting in a second tense label sequence.
[0017] In one possible implementation, generating a discourse coherence score based on the matching degree of semantic relationship categories corresponding to related words at the same semantic position in the machine translation and the reference translation includes: Identify the conjunctions in each sentence of the machine translation and the reference translation; For the same semantic position, query the first semantic relationship category corresponding to the associated word at the semantic position in the machine translation, query the second semantic relationship category corresponding to the associated word at the semantic position in the reference translation, and determine whether the first semantic relationship category and the second semantic relationship category are consistent; The coherence score of the passage is obtained by calculating the proportion of related words with consistent semantic relationships.
[0018] According to a second aspect of this application, an evaluation apparatus based on document-level machine translation is provided, the apparatus comprising: The acquisition module is used to acquire the source document and its corresponding machine translation and reference translation; The entity consistency assessment module is used to generate coreference chains for entities in the source document, find the set of translated items corresponding to each coreference chain in the machine translation, and the translated items in the translated item set correspond to the mentions in the coreference chain; and generate an entity consistency score based on the comprehensive similarity between every two translated items in the translated item set, the entity consistency score representing the consistency of nouns of the same entity in the machine translation. The tense consistency evaluation module is used to perform tense annotation on the statements in the machine translation and the reference translation respectively to obtain a first tense label sequence and a second tense label sequence; and to generate a tense consistency score based on the matching degree of the first tense label sequence and the second tense label sequence corresponding to the same statement in the machine translation and the reference translation, wherein the tense consistency score represents the consistency of the machine translation in terms of time expression. The discourse coherence assessment module is used to map the related words in the machine translation and the reference translation to the corresponding semantic relationship categories according to a preset mapping table; and to generate a discourse coherence score based on the matching degree of the semantic relationship categories corresponding to the related words at the same semantic position in the machine translation and the reference translation. The discourse coherence score represents the consistency of the machine translation in the discourse logical relationship. The comprehensive evaluation module is used to generate a comprehensive score for document-level machine translation based on the entity consistency score, the temporal consistency score, and the discourse coherence score.
[0019] According to a third aspect of this application, a computer-readable storage medium is provided, wherein at least one instruction is stored therein, the at least one instruction being loaded and executed by a processor to implement the document-level machine translation-based evaluation method as described above.
[0020] According to a fourth aspect of this application, an electronic device is provided, the electronic device including the above-described document-level machine translation-based evaluation device.
[0021] The beneficial effects of the technical solution provided in this application include at least the following: By constructing a comprehensive evaluation framework encompassing entity consistency, tense consistency, and discourse coherence, this application overcomes the limitations of traditional evaluation metrics that only assess lexical overlap and fail at the discourse level. It accurately reflects the readability and logical rigor of translations in long documents, providing a reliable quantitative standard for document-level translation. Furthermore, this application relies solely on the text content of the source document, machine translation, and reference translation, without requiring access to the internal structure of the translation model. This makes it widely applicable to quality testing of various black-box translation systems, demonstrating significant engineering value.
[0022] By identifying the differences in the translation of the same entity in different sentences at the coreference chain level, and then combining semantic similarity, edit distance similarity, and speech similarity to comprehensively evaluate the entity coherence of the entire machine translation, it is possible to accurately locate the frequency of changes in the name of the same entity during the translation process, providing a fine-grained analytical basis for the diagnosis of translation models.
[0023] By subdividing tense evaluation into explicit tense and implicit tense, explicit tense evaluation examines the translation model's accuracy in translating sentences with explicit time cues, while implicit tense evaluation examines the translation model's ability to infer tense from context when there are no explicit time cues. This classification evaluation mechanism can accurately point out the shortcomings of the translation model's ability to use context for tense reasoning when there are no explicit time words. This is a refined diagnostic capability that traditional evaluation indicators do not possess.
[0024] By constructing a semantic relation mapping mechanism, a tolerant assessment of the logical coherence of a text is achieved. Even if the machine translation uses different conjunctions than the reference translation, as long as the semantic relation categories expressed by the two are the same, it can still be judged as correct. It can effectively identify problems where the semantics are fluent but the logical conjunctions are used incorrectly, which significantly improves the accuracy of the evaluation. Attached Figure Description
[0025] To more clearly illustrate the technical solutions in the embodiments of this application, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0026] Figure 1 This is a schematic diagram of an evaluation method based on document-level machine translation provided in one embodiment of this application; Figure 2 This is a flowchart of an evaluation method based on document-level machine translation provided in one embodiment of this application; Figure 3This is a flowchart of an evaluation method based on document-level machine translation provided in one embodiment of this application; Figure 4 This is a structural block diagram of an evaluation device based on document-level machine translation provided in one embodiment of this application; Figure 5 This is a structural block diagram of an electronic device provided in one embodiment of this application. Detailed Implementation
[0027] To make the objectives, technical solutions, and advantages of the embodiments of this application clearer, the implementation methods of this application will be further described in detail below with reference to the accompanying drawings.
[0028] The following is a description of the terms used in this application.
[0029] (1) Document-Level Machine Translation Document-level machine translation refers to machine translation methods that target complete documents. In the translation process, it is necessary not only to handle the dependencies within sentences, but also to consider discourse phenomena between sentences (such as reference, cohesion, and coherence).
[0030] (2) Entity Consistency Entity consistency refers to maintaining a unified, stable, and coherent vocabulary when nouns representing the same real-world entity are mentioned multiple times in machine translation, avoiding ambiguity caused by the use of different words.
[0031] (3) Tense Consistency Tense consistency refers to the logical consistency of time expression in the sentences of a translation. That is, it is able to accurately and coherently express the linguistic features of events changing over time according to the context, especially when translating from a language without explicit tense (such as Chinese) to a language with explicit tense (such as English).
[0032] (4) Coherence Discourse coherence refers to the consistency and rationality between sentences in a translation in terms of logical relationships, semantic connections, and discourse structure. It is an important indicator for measuring the readability of a translation.
[0033] (5) Co-reference Chain Coreference chains are semantic association chains in text where multiple words or phrases point to the same entity. They include sets of nouns or pronouns that point to the same entity and are used to determine the scope of entity consistency assessment.
[0034] (6) Black-box Evaluation Black-box evaluation refers to a method of evaluating a model based solely on its input and output results without obtaining its internal parameters and structure.
[0035] like Figure 1 As shown, this application relates to a comprehensive evaluation framework for consistency and coherence in document-level machine translation, comprising three parts: an entity consistency evaluation module, a tense consistency evaluation module, and a discourse coherence evaluation module. After inputting the source document, machine translation, and reference translation, the entity consistency evaluation module uses coreference chain constraints to evaluate the entity consistency score of the source document and machine translation, avoiding the sacrifice of translation diversity by simply emphasizing lexical consistency. The tense consistency evaluation module uses tense tag sequences to evaluate the tense consistency score of explicit (intra-sentence information) and implicit (contextual information) aspects in the machine translation and reference translation. The discourse coherence evaluation module constructs a semantic coherence mapping table and evaluates the discourse coherence score by calculating the hit rate between the machine translation and the reference translation on logical relationships (rather than simple lexical relationships). Finally, the entity consistency score, tense consistency score, and discourse coherence score are weighted and fused to obtain the final comprehensive score.
[0036] like Figure 2 The diagram illustrates a flowchart of a document-level machine translation-based evaluation method according to an embodiment of this application. This document-level machine translation-based evaluation method can be applied to electronic devices. The document-level machine translation-based evaluation method may include: Step 201: Obtain the source document and its corresponding machine translation and reference translation.
[0037] The source document is the original language document to be translated, such as a Chinese document in a Chinese-to-English translation scenario; the machine translation is the target language document output by a document-level machine translation model after translating the source document, such as an English document in a Chinese-to-English translation scenario; the reference translation is the standard translation result corresponding to the source document, usually obtained by human translation by experts in the field of translation, and is used as the evaluation standard.
[0038] Step 202: Generate coreference chains for entities in the source document, and search for the set of translated items corresponding to each coreference chain in the machine translation. The translated items in the translated item set correspond to the mentions in the coreference chain. Generate an entity consistency score based on the comprehensive similarity between every two translated items in the translated item set. The entity consistency score represents the consistency of nouns of the same entity in the machine translation.
[0039] This application can generate an entity consistency score based on the source document and the machine translation, which indicates the consistency of the nouns of the same entity in the machine translation.
[0040] Specifically, coreference resolution tools (such as AllenNLP, Stanford CoreNLP, etc.) are used to resolve coreference in the source document, identifying all entities and their coreference chains in the source document. Each coreference chain corresponds to an entity in the source document, and each element in the coreference chain is called a mention. Assume the set of coreference chains in the source document is C = {c1, c2, c3, ..., c...} n}, where n is the total number of core-pointing chains, and for the i-th core-pointing chain c i ={e i1 e i2 e i3 , ..., e im}, where m is the total number of mentions, e ij This represents the j-th mention in the i-th coreference chain.
[0041] To improve the accuracy and efficiency of the evaluation, the extracted coreference chains can be filtered, retaining only those with a length greater than 1 and containing nouns, in order to eliminate meaningless pronoun interference.
[0042] For each cofinite chain c i Search the machine translation for each mention e in the coreference chain. ij The corresponding translation item t ij All the found translation terms are combined into a set T of translation terms corresponding to the coreference chain. i ={t i1 t i2 t i3 , ..., t im}, where t ij For reference to item e ij The corresponding translation item.
[0043] For each set of translated items, the comprehensive similarity between every two translated items in the set is calculated, and the entity consistency score corresponding to the set is calculated based on all comprehensive similarities. Then, the entity consistency scores of all translated item sets are aggregated to obtain the final entity consistency score. The entity consistency score ranges from 0 to 1, with a higher score indicating that the machine translation is more consistent with the reference translation in terms of entity reference.
[0044] Step 203: Tense annotation is performed on the statements in the machine translation and the reference translation to obtain the first tense label sequence and the second tense label sequence; a tense consistency score is generated based on the matching degree of the first tense label sequence and the second tense label sequence corresponding to the same statement in the machine translation and the reference translation. The tense consistency score represents the consistency of the machine translation in terms of time expression.
[0045] This application can generate a tense consistency score based on machine translation and reference translation, which indicates the consistency of the machine translation in terms of temporal expression.
[0046] Specifically, for each sentence in the machine translation, each word in the sentence is tense-tagged, resulting in a first tense tag sequence corresponding to the sentence length. For each sentence in the reference translation, each word in the sentence is tense-tagged, resulting in a second tense tag sequence corresponding to the sentence length. Each tag in the tense tag sequence represents the tense category of the corresponding word, and the tense tags include past tense, present tense, future tense, modal verbs, verb infinitives, and non-finite verbs.
[0047] For the same statement, the first and second tense tag sequences are used to calculate the matching degree between them. Then, the proportion of statements whose tense tag sequences match is calculated across all statement pairs to obtain a tense consistency score. The tense consistency score ranges from 0 to 1, with a higher score indicating that the machine translation is more consistent with the reference translation in terms of temporal expression.
[0048] Step 204: According to the preset mapping table, the related words in the machine translation and the reference translation are mapped to the corresponding semantic relationship categories respectively; a discourse coherence score is generated based on the matching degree of the semantic relationship categories corresponding to the related words at the same semantic position in the machine translation and the reference translation. The discourse coherence score represents the consistency of the machine translation in the discourse logical relationship.
[0049] This application can generate a discourse coherence score based on machine translation and reference translation, which indicates the consistency of the machine translation in terms of discourse logic.
[0050] Specifically, a semantic relationship mapping table is pre-constructed, mapping multiple related words to their corresponding semantic relationship categories. These categories include contrastive, causal, conditional, progressive, adversative, parallel, temporal, and purposive relationships, and each semantic relationship category can correspond to multiple related words to reflect the diversity of language expression. Then, related words in each sentence of the machine translation and the reference translation are identified. Related words can appear at the beginning, middle, or end of a sentence; this embodiment does not impose restrictions on their position, as long as the sentence contains related words, they are identified.
[0051] For conjunctions at the same semantic position in both the machine translation and the reference translation, the semantic relationship mapping table is queried to obtain the semantic relationship category corresponding to the conjunction, and it is determined whether the mapped semantic relationship categories are consistent. Here, "same semantic position" refers to a corresponding position in the statement that performs the same logical function. If both the machine translation and the corresponding reference translation contain conjunctions at the corresponding position, and their mapped semantic relationship categories are the same, then the logical relationship of the statement pair is determined to match; otherwise, the logical relationship of the statement pair is determined to be inconsistent.
[0052] The coherence score is calculated by determining the proportion of sentences with matching logical relationships among all sentence pairs containing conjunctions. The coherence score ranges from 0 to 1, with a higher score indicating greater consistency between the machine translation and the reference translation in terms of logical relationships.
[0053] Step 205: Generate a comprehensive score for document-level machine translation based on entity consistency score, temporal consistency score, and discourse coherence score.
[0054] Specifically, entity consistency score, temporal consistency score and discourse coherence score can be weighted and fused to obtain the final comprehensive score.
[0055] In summary, the document-level machine translation evaluation method provided in this application constructs a comprehensive evaluation framework encompassing entity consistency, tense consistency, and discourse coherence. Compared to traditional evaluation metrics that only assess lexical overlap and fail at the discourse level, this application accurately reflects the readability and logical rigor of the translation in long documents, providing a reliable quantitative standard for document-level translation. Furthermore, this application relies solely on the text content of the source document, the machine translation, and the reference translation, without requiring access to the internal structure of the translation model. It can be widely applied to quality testing of various black-box translation systems, demonstrating significant engineering practical value.
[0056] like Figure 3 The diagram illustrates a flowchart of a document-level machine translation-based evaluation method according to an embodiment of this application. This document-level machine translation-based evaluation method can be applied to electronic devices. The document-level machine translation-based evaluation method may include: Step 301: Obtain the source document and its corresponding machine translation and reference translation.
[0057] The source document is the original language document to be translated, the machine translation is the target language document output by the document-level machine translation model after translating the source document, and the reference translation is the standard translation result corresponding to the source document.
[0058] Step 302: Generate coreference chains for entities in the source document, and search for the set of translated items corresponding to each coreference chain in the machine translation. The translated items in the set of translated items correspond to the mentions in the coreference chains.
[0059] The process for generating coreference chains and translation item sets is detailed in step 202 and will not be repeated here.
[0060] Step 303: For each coreference chain, calculate the comprehensive similarity between every two translated items in the translated item set, and calculate the entity consistency score of the translated item set based on all comprehensive similarities; calculate the final entity consistency score based on the entity consistency scores of all translated item sets, where the entity consistency score represents the consistency of nouns of the same entity in the machine translation.
[0061] Specifically, calculating the comprehensive similarity between any two translated items in the translation item set can include: for each pair of translated items in the translation item set, calculating the semantic similarity, edit distance similarity, and speech similarity between the two translated items; and weighting and fusing the semantic similarity, edit distance similarity, and speech similarity to obtain the comprehensive similarity between the two translated items.
[0062] Assume semantic similarity is sim sem The edit distance similarity is sim lev The semantic similarity is sim pho Then the overall similarity S sim The calculation formula is: S sim =λ1sim sem +λ2sim lev +λ3sim pho λ1, λ2 and λ3 are weighting coefficients that satisfy λ1+λ2+λ3=1. The specific values can be configured according to the actual application scenario.
[0063] In this embodiment, calculating the semantic similarity, edit distance similarity, and speech similarity between two translated items may include: (1) Calculate the cosine similarity between the word vectors of the two translation terms to obtain the semantic similarity.
[0064] Word vectors are extracted for each of the two translated terms. A cosine similarity is calculated between the two word vectors, and this cosine similarity is taken as the semantic similarity between the two translated terms. This embodiment does not limit the method used to extract the word vectors or the algorithm for calculating the cosine similarity.
[0065] The purpose of calculating semantic similarity is to capture the semantic equivalence between translated items and avoid misjudging two translated items as inconsistent due to differences in their superficial forms.
[0066] (2) Calculate the edit distance between the two translation items to obtain the edit distance similarity.
[0067] Edit distance, also known as Levenshtein distance, refers to the minimum number of edit operations required to transform one string into another.
[0068] The string lengths of the two translated items are extracted, and the edit distance similarity is calculated using the formula for calculating the edit distance similarity between the string lengths and the translated items.
[0069] The purpose of calculating edit distance similarity is to tolerate spelling differences and morphological variations, which is important for evaluating translation consistency between different English variants.
[0070] (3) Calculate the phonetic similarity between two translation items based on string encoding algorithm or phonetic conversion algorithm.
[0071] The Soundex algorithm is a method that encodes English words into alphanumeric forms based on their pronunciation. Words with similar pronunciations share the same Soundex code. The similarity score between two translated words is calculated using the formula of the Soundex algorithm.
[0072] Phonetic transcription conversion algorithms involve converting translated terms into International Phonetic Alphabet (IPA) representations and then calculating the similarity between phonetic transcription sequences. First, a tool is used to convert the translated terms into phonetic transcription strings; then, the edit distance similarity or longest common subsequence similarity between two phonetic transcription strings is calculated to obtain the speech similarity.
[0073] The purpose of calculating speech similarity is to prevent misjudgment caused by homophones and avoid underestimating entity consistency due to differences in appearance.
[0074] After obtaining the comprehensive similarity between every two translated items in the translated item set, the entity consistency score corresponding to each coreference chain can be calculated, and then the arithmetic mean of all entity consistency scores can be calculated to obtain the final entity consistency score.
[0075] Among them, the entity consistency score corresponds to each corefinite chain. Where c represents a coreference chain, Indicates the translation term t i and translation item t j The overall similarity.
[0076] Step 304: Tense annotation is performed on the statements in the machine translation and the reference translation to obtain the first tense tag sequence and the second tense tag sequence.
[0077] The process involves performing tense tagging on sentences in both the machine translation and the reference translation to obtain a first tense label sequence and a second tense label sequence. This can be achieved by: for each sentence in the machine translation, using a tense tagger to label each word in the sentence with a tense category label, resulting in the first tense label sequence; and for each sentence in the reference translation, using the same tagger to label each word in the sentence with a tense category label, resulting in the second tense label sequence. The tense tagger can be based on the Viterbi algorithm or a pre-trained tagger.
[0078] For each statement, firstly, it is segmented into multiple independent words; then, the segmented words are tagged with parts of speech to obtain the part-of-speech category of each word; then, based on the form of the word itself and the part-of-speech tagging results, the tense category of each word is determined; finally, the tense category tags of each word are combined in the original word order to obtain the tense tag sequence corresponding to the statement.
[0079] Step 305: Generate a tense consistency score based on the matching degree of the first tense label sequence and the second tense label sequence corresponding to the same sentence in the machine translation and the reference translation. The tense consistency score represents the consistency of the machine translation in terms of temporal expression.
[0080] Specifically, a tense consistency score is generated based on the matching degree between the first tense label sequence and the second tense label sequence corresponding to the same statement in the machine translation and the reference translation. This can include: (1) The statements in the machine translation are divided into a first explicit set and a first implicit set, and the statements in the reference translation are divided into a second explicit set and a second implicit set. The first explicit set and the second explicit set are sets of statements containing time clues, and the first implicit set and the second implicit set are sets of statements not containing time clues.
[0081] The basis for classifying statements into explicit and implicit sets is whether they contain time clues. Time clues refer to words or structures that can directly or indirectly indicate when an action occurred, including but not limited to: time adverbs (such as yesterday, today, tomorrow, etc.), time adverbs (such as last year, next week, etc.), and time conjunctions (such as after, before, etc.).
[0082] For each sentence in the machine translation, all words in the sentence are scanned to determine if the aforementioned time clues exist. If at least one time clue exists, the sentence is assigned to the first explicit set; otherwise, it is assigned to the first implicit set. Similarly, for each sentence in the reference translation, the same method is used to divide them into a second explicit set and a second implicit set.
[0083] (2) For the same statement in the first explicit set and the second explicit set, determine whether the first tense label sequence and the second tense label sequence corresponding to the statement are consistent; count the proportion of statements with consistent tense label sequences to obtain the explicit matching degree.
[0084] For the first and second tense label sequences corresponding to the same statement, determine whether the tense labels at all or some positions in the sequence are the same. If they are the same, determine that the tense label sequences of the statement are consistent; otherwise, determine that the tense label sequences of the statement are inconsistent.
[0085] The formula for calculating the explicit matching degree is as follows: Where 1() represents the indicator function, D e Tense represents an explicit set, s represents the statements in the explicit set, and Tense represents the set. MT (s) represents the sequence of first tense labels for statements in the first explicit set. Ref (s) represents the second tense label sequence of statements in the second explicit set.
[0086] (3) For the same statement in the first implicit set and the second implicit set, determine whether the first tense label sequence and the second tense label sequence corresponding to the statement are consistent; count the proportion of statements with consistent tense label sequences to obtain the implicit matching degree.
[0087] The formula for calculating the implicit matching degree is as follows: Where 1() represents the indicator function, D f Represents an implicit set. Tense represents statements in an implicit set. MT ( () represents the sequence of first tense labels for statements in the first implicit set. Ref ( ) represents the second tense label sequence of statements in the second implicit set.
[0088] (4) The explicit matching degree and the implicit matching degree are weighted and fused to obtain the temporal consistency score.
[0089] Final temporal consistency score (Tense) score =αTense score (D) e )+βTense score (D) f ), where α and β are weighting coefficients that satisfy α+β=1, and the specific values can be configured according to the actual application scenario.
[0090] Step 306: Based on the preset mapping table, map the related words in the machine translation and the reference translation to the corresponding semantic relationship categories.
[0091] A mapping table is a predefined set of mapping relationships used to map multiple related words to corresponding semantic relationship categories. Semantic relationship categories are abstract classifications of the semantic functions of related words. In this embodiment, 16 categories are defined, such as contrastive relationships, causal relationships, conditional relationships, progressive relationships, adversative relationships, and parallel relationships.
[0092] After obtaining the pre-defined semantic relationship mapping table, for each related word in the machine translation, the mapping table is queried to obtain its corresponding semantic relationship category.
[0093] Step 307: Identify the conjunctions in each sentence of the machine translation and the reference translation.
[0094] By identifying related words in a sentence, candidate words for semantic relationship category mapping are determined.
[0095] Step 308: For the same semantic position, query the first semantic relationship category corresponding to the related word at the semantic position in the machine translation, query the second semantic relationship category corresponding to the related word at the semantic position in the reference translation, and determine whether the first semantic relationship category and the second semantic relationship category are consistent.
[0096] After determining the same semantic location, query the semantic relationship category corresponding to the related words at that semantic location in the machine translation and the reference translation respectively, and determine whether the two are consistent. If they are the same, it is determined that the logical relationship at that semantic location is matched; otherwise, it is determined that they are not matched.
[0097] The formula for calculating the matching degree of related words in the logical space is as follows: , where P MT P represents conjunctions in machine translation. Ref M represents the conjunctions in the reference translation, and M() represents the function that maps the conjunctions to the corresponding semantic relation categories.
[0098] Step 309: Calculate the proportion of conjunctions with consistent semantic relationship categories to obtain the discourse coherence score. The discourse coherence score indicates the consistency of machine translation in discourse logical relationships.
[0099] Passage coherence score Where z represents the total number of conjunctions, Coh(b) j ) represents the matching degree of the associated word at the j-th semantic position in the logical space.
[0100] Step 310: Generate a comprehensive score for document-level machine translation based on entity consistency score, temporal consistency score, and discourse coherence score.
[0101] Specifically, entity consistency score, temporal consistency score and discourse coherence score can be weighted and fused to obtain the final comprehensive score.
[0102] In summary, the document-level machine translation evaluation method provided in this application constructs a comprehensive evaluation framework encompassing entity consistency, tense consistency, and discourse coherence. Compared to traditional evaluation metrics that only assess lexical overlap and fail at the discourse level, this application accurately reflects the readability and logical rigor of the translation in long documents, providing a reliable quantitative standard for document-level translation. Furthermore, this application relies solely on the text content of the source document, the machine translation, and the reference translation, without requiring access to the internal structure of the translation model. It can be widely applied to quality testing of various black-box translation systems, demonstrating significant engineering practical value.
[0103] By identifying the differences in the translation of the same entity in different sentences at the coreference chain level, and then combining semantic similarity, edit distance similarity, and speech similarity to comprehensively evaluate the entity coherence of the entire machine translation, it is possible to accurately locate the frequency of changes in the name of the same entity during the translation process, providing a fine-grained analytical basis for the diagnosis of translation models.
[0104] By subdividing tense evaluation into explicit tense and implicit tense, explicit tense evaluation examines the translation model's accuracy in translating sentences with explicit time cues, while implicit tense evaluation examines the translation model's ability to infer tense from context when there are no explicit time cues. This classification evaluation mechanism can accurately point out the shortcomings of the translation model's ability to use context for tense reasoning when there are no explicit time words. This is a refined diagnostic capability that traditional evaluation indicators do not possess.
[0105] By constructing a semantic relation mapping mechanism, a tolerant assessment of the logical coherence of a text is achieved. Even if the machine translation uses different conjunctions than the reference translation, as long as the semantic relation categories expressed by the two are the same, it can still be judged as correct. It can effectively identify problems where the semantics are fluent but the logical conjunctions are used incorrectly, which significantly improves the accuracy of the evaluation.
[0106] like Figure 4 The diagram illustrates a structural block diagram of a document-level machine translation-based evaluation apparatus according to an embodiment of this application. This apparatus can be applied to electronic devices and includes: Module 410 is used to acquire the source document and its corresponding machine translation and reference translation; The entity consistency assessment module 420 is used to generate coreference chains for entities in the source document, find the set of translated items corresponding to each coreference chain in the machine translation, and the translated items in the translated item set correspond to the mentions in the coreference chain; and generate an entity consistency score based on the comprehensive similarity between every two translated items in the translated item set. The entity consistency score represents the consistency of the nouns of the same entity in the machine translation. The tense consistency evaluation module 430 is used to perform tense annotation on the statements in the machine translation and the reference translation respectively to obtain a first tense label sequence and a second tense label sequence; a tense consistency score is generated based on the matching degree of the first tense label sequence and the second tense label sequence corresponding to the same statement in the machine translation and the reference translation. The tense consistency score represents the consistency of the machine translation in terms of time expression. The discourse coherence assessment module 440 is used to map the related words in the machine translation and the reference translation to the corresponding semantic relationship categories according to the preset mapping table; and to generate a discourse coherence score based on the matching degree of the semantic relationship categories corresponding to the related words at the same semantic position in the machine translation and the reference translation. The discourse coherence score represents the consistency of the machine translation in the discourse logical relationship. The comprehensive evaluation module 450 is used to generate a comprehensive score for document-level machine translation based on entity consistency score, temporal consistency score, and discourse coherence score.
[0107] In an optional embodiment, the entity consistency assessment module 420 is further configured to: For each coreference chain, calculate the comprehensive similarity between every two translated items in the translated item set, and calculate the entity consistency score of the translated item set based on all comprehensive similarities. The final entity consistency score is calculated based on the entity consistency scores of all translated item sets.
[0108] In an optional embodiment, the entity consistency assessment module 420 is further configured to: For every two translation items in the set of translation items, calculate the semantic similarity, edit distance similarity, and speech similarity between the two translation items; We perform weighted fusion of semantic similarity, edit distance similarity, and speech similarity to obtain the comprehensive similarity between the two translated items.
[0109] In an optional embodiment, the entity consistency assessment module 420 is further configured to: Calculate the cosine similarity between the word vectors of two translated terms to obtain semantic similarity; Calculate the edit distance between two translated items to obtain the edit distance similarity; The phonetic similarity between two translated items is calculated based on string encoding algorithms or phonetic transcription conversion algorithms.
[0110] In an optional embodiment, the temporal consistency evaluation module 430 is further configured to: The machine-translated text is divided into a first explicit set and a first implicit set, and the reference translation is divided into a second explicit set and a second implicit set. The first explicit set and the second explicit set are sets of sentences containing time clues, and the first implicit set and the second implicit set are sets of sentences not containing time clues. For the same statement in the first explicit set and the second explicit set, determine whether the first tense label sequence and the second tense label sequence corresponding to the statement are consistent; count the proportion of statements with consistent tense label sequences to obtain the explicit matching degree. For the same statement in the first implicit set and the second implicit set, determine whether the first tense label sequence and the second tense label sequence corresponding to the statement are consistent; count the proportion of statements with consistent tense label sequences to obtain the implicit matching degree; The explicit and implicit matching scores are weighted and fused to obtain a temporal consistency score.
[0111] In an optional embodiment, the temporal consistency evaluation module 430 is further configured to: For each sentence in the machine translation, a tense tagger is used to tag each word in the sentence with a tense category label, resulting in the first tense label sequence; For each sentence in the reference translation, a tense labeler is used to label each word in the sentence with a tense category label, resulting in a second tense label sequence.
[0112] In an optional embodiment, the discourse coherence assessment module 440 is further configured to: Identify conjunctions in each sentence of the machine translation and the reference translation; For the same semantic position, query the first semantic relationship category corresponding to the related words at the semantic position in the machine translation, query the second semantic relationship category corresponding to the related words at the semantic position in the reference translation, and determine whether the first semantic relationship category and the second semantic relationship category are consistent; The proportion of conjunctions with consistent semantic relationships is used to obtain a score for the coherence of the text.
[0113] In summary, the document-level machine translation evaluation device provided in this application, by constructing a comprehensive evaluation framework encompassing entity consistency, tense consistency, and discourse coherence, overcomes the limitations of traditional evaluation metrics that can only assess lexical overlap and fail at the discourse level. This application can accurately reflect the readability and logical rigor of the translation in long documents, providing a reliable quantitative standard for document-level translation. Furthermore, this application relies solely on the text content of the source document, the machine translation, and the reference translation, without requiring access to the internal structure of the translation model. It can be widely applied to quality testing of various black-box translation systems, demonstrating significant engineering practical value.
[0114] By identifying the differences in the translation of the same entity in different sentences at the coreference chain level, and then combining semantic similarity, edit distance similarity, and speech similarity to comprehensively evaluate the entity coherence of the entire machine translation, it is possible to accurately locate the frequency of changes in the name of the same entity during the translation process, providing a fine-grained analytical basis for the diagnosis of translation models.
[0115] By subdividing tense evaluation into explicit tense and implicit tense, explicit tense evaluation examines the translation model's accuracy in translating sentences with explicit time cues, while implicit tense evaluation examines the translation model's ability to infer tense from context when there are no explicit time cues. This classification evaluation mechanism can accurately point out the shortcomings of the translation model's ability to use context for tense reasoning when there are no explicit time words. This is a refined diagnostic capability that traditional evaluation indicators do not possess.
[0116] By constructing a semantic relation mapping mechanism, a tolerant assessment of the logical coherence of a text is achieved. Even if the machine translation uses different conjunctions than the reference translation, as long as the semantic relation categories expressed by the two are the same, it can still be judged as correct. It can effectively identify problems where the semantics are fluent but the logical conjunctions are used incorrectly, which significantly improves the accuracy of the evaluation.
[0117] One embodiment of this application provides a computer-readable storage medium storing at least one instruction, which is loaded and executed by a processor to implement the document-level machine translation-based evaluation method described above.
[0118] Figure 5A schematic block diagram of an example electronic device 500 that can be used to implement embodiments of this application is shown. The electronic device is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device may also represent various forms of mobile devices, such as personal digital processors, cellular phones, smartphones, wearable devices, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely illustrative and are not intended to limit the implementation of the application described and / or claimed herein.
[0119] like Figure 5 As shown, the electronic device 500 includes a computing unit 501, which can perform various appropriate actions and processes according to a computer program stored in a read-only memory (ROM) 502 or a computer program loaded from a storage unit 508 into a random access memory (RAM) 503. The RAM 503 may also store various programs and data required for the operation of the electronic device 500. The computing unit 501, ROM 502, and RAM 503 are interconnected via a bus 504. An input / output (I / O) interface 505 is also connected to the bus 504.
[0120] Multiple components in electronic device 500 are connected to I / O interface 505, including: input unit 506, such as keyboard, mouse, etc.; output unit 507, such as various types of monitors, speakers, etc.; storage unit 508, such as disk, optical disk, etc.; and communication unit 509, such as network card, modem, wireless transceiver, etc. Communication unit 509 allows electronic device 500 to exchange information / data with other devices through computer networks such as the Internet and / or various telecommunications networks.
[0121] The computing unit 501 can be a variety of general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of the computing unit 501 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various special-purpose artificial intelligence (AI) computing chips, various computing units running machine learning model algorithms, a digital signal processor (DSP), and any suitable processor, controller, microcontroller, etc. The computing unit 501 performs the various methods and processes described above, such as data processing methods. For example, in some embodiments, the data processing methods may be implemented as computer software programs tangibly contained in a machine-readable medium, such as storage unit 508. In some embodiments, part or all of the computer program may be loaded and / or installed on the electronic device 500 via ROM 502 and / or communication unit 509. When the computer program is loaded into RAM 503 and executed by the computing unit 501, one or more steps of the data processing methods described above may be performed. Alternatively, in other embodiments, the computing unit 501 may be configured to perform data processing methods by any other suitable means (e.g., by means of firmware).
[0122] Various embodiments of the systems and techniques described above herein can be implemented in digital electronic circuit systems, integrated circuit systems, field-programmable gate arrays (FPGAs), application-specific integrated circuits (ASICs), application-specific standard products (ASSPs), systems-on-a-chip (SoCs), complex programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments may include implementations in one or more computer programs that can be executed and / or interpreted on a programmable system including at least one programmable processor, which may be a dedicated or general-purpose programmable processor, capable of receiving data and instructions from a storage system, at least one input device, and at least one output device, and transmitting data and instructions to the storage system, the at least one input device, and the at least one output device.
[0123] The program code used to implement the methods of this application may be written in any combination of one or more programming languages. This program code may be provided to a processor or controller of a general-purpose computer, special-purpose computer, or other programmable data processing device, such that when executed by the processor or controller, the functions / operations specified in the flowcharts and / or block diagrams are implemented. The program code may be executed entirely on a machine, partially on a machine, as a standalone software package partially on a machine and partially on a remote machine, or entirely on a remote machine or server.
[0124] In the context of this application, a machine-readable medium can be a tangible medium that may contain or store a program for use by or in conjunction with an instruction execution system, apparatus, or device. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. Machine-readable media can be, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of machine-readable storage media include electrical connections based on one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fibers, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination of the foregoing.
[0125] To provide interaction with a user, the systems and techniques described herein can be implemented on a computer having: a display device for displaying information to the user (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor); and a keyboard and pointing device (e.g., a mouse or trackball) through which the user provides input to the computer. Other types of devices can also be used to provide interaction with the user; for example, feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including sound input, voice input, or tactile input).
[0126] The systems and technologies described herein can be implemented in computing systems that include backend components (e.g., as a data server), or computing systems that include middleware components (e.g., an application server), or computing systems that include frontend components (e.g., a user computer with a graphical user interface or web browser through which a user can interact with implementations of the systems and technologies described herein), or any combination of such backend, middleware, or frontend components. The components of the system can be interconnected via digital data communication of any form or medium (e.g., a communication network). Examples of communication networks include local area networks (LANs), wide area networks (WANs), and the Internet.
[0127] Computer systems can include clients and servers. Clients and servers are generally located far apart and typically interact via communication networks. Client-server relationships are created by computer programs running on the respective computers and having a client-server relationship with each other. Servers can be cloud servers, servers in distributed systems, or servers incorporating blockchain technology.
[0128] It should be understood that the various forms of processes shown above can be used to rearrange, add, or delete steps. For example, the steps described in this disclosure can be executed in parallel, sequentially, or in different orders, as long as the desired result of the technical solution disclosed in this application can be achieved, and this is not limited herein.
[0129] The specific embodiments described above do not constitute a limitation on the scope of protection of this application. Those skilled in the art should understand that various modifications, combinations, sub-combinations, and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of this application should be included within the scope of protection of this application.
Claims
1. An evaluation method based on document-level machine translation, characterized in that, The method includes: Obtain the source document and its corresponding machine translation and reference translation; Generate coreference chains for entities in the source document, find the set of translated items corresponding to each coreference chain in the machine translation, and the translated items in the translated item set correspond to the mentions in the coreference chain; generate an entity consistency score based on the comprehensive similarity between every two translated items in the translated item set, and the entity consistency score represents the consistency of nouns of the same entity in the machine translation. The machine translation and the reference translation are labeled with tense to obtain a first tense tag sequence and a second tense tag sequence. A tense consistency score is generated based on the matching degree of the first tense tag sequence and the second tense tag sequence corresponding to the same sentence in the machine translation and the reference translation. The tense consistency score represents the consistency of the machine translation in terms of time expression. According to a preset mapping table, the related words in the machine translation and the reference translation are mapped to the corresponding semantic relationship categories respectively; a discourse coherence score is generated based on the matching degree of the semantic relationship categories corresponding to the related words at the same semantic position in the machine translation and the reference translation, and the discourse coherence score represents the consistency of the machine translation in the discourse logical relationship. A comprehensive score for document-level machine translation is generated based on the entity consistency score, the temporal consistency score, and the discourse coherence score. The step of generating a tense consistency score based on the matching degree of the first tense tag sequence and the second tense tag sequence corresponding to the same statement in the machine translation and the reference translation includes: dividing the statements in the machine translation into a first explicit set and a first implicit set, and dividing the statements in the reference translation into a second explicit set and a second implicit set, wherein the first explicit set and the second explicit set are sets of statements containing time clues, and the first implicit set and the second implicit set are sets of statements not containing the time clues; for the same statement in the first explicit set and the second explicit set, determining whether the first tense tag sequence and the second tense tag sequence corresponding to the statement are consistent; calculating the proportion of statements with consistent tense tag sequences to obtain an explicit matching degree; for the same statement in the first implicit set and the second implicit set, determining whether the first tense tag sequence and the second tense tag sequence corresponding to the statement are consistent; calculating the proportion of statements with consistent tense tag sequences to obtain an implicit matching degree; and performing a weighted fusion of the explicit matching degree and the implicit matching degree to obtain the tense consistency score. The step of generating a discourse coherence score based on the matching degree of semantic relationship categories corresponding to the related words at the same semantic position in the machine translation and the reference translation includes: identifying related words in each sentence of the machine translation and the reference translation; for the same semantic position, querying the first semantic relationship category corresponding to the related word at the semantic position in the machine translation, querying the second semantic relationship category corresponding to the related word at the semantic position in the reference translation, and determining whether the first semantic relationship category and the second semantic relationship category are consistent; and calculating the proportion of related words with consistent semantic relationship categories to obtain the discourse coherence score.
2. The evaluation method based on document-level machine translation according to claim 1, characterized in that, The step of generating an entity consistency score based on the comprehensive similarity between every two translated items in the translated item set includes: For each set of translated items corresponding to a coreference chain, calculate the comprehensive similarity between every two translated items in the set, and calculate the entity consistency score of the set based on all comprehensive similarities. The final entity consistency score is calculated based on the entity consistency scores of all translated item sets.
3. The evaluation method based on document-level machine translation according to claim 2, characterized in that, The calculation of the overall similarity between every two translated items in the set of translated items includes: For every two translated items in the set of translated items, calculate the semantic similarity, edit distance similarity, and speech similarity between the two translated items; The semantic similarity, the edit distance similarity, and the speech similarity are weighted and fused to obtain the comprehensive similarity between the two translated items.
4. The evaluation method based on document-level machine translation according to claim 3, characterized in that, The calculation of semantic similarity, edit distance similarity, and speech similarity between the two translated items includes: The semantic similarity is obtained by calculating the cosine similarity between the word vectors of the two translated terms. Calculate the edit distance between the two translated items to obtain the edit distance similarity. The phonetic similarity between the two translated items is calculated based on a string encoding algorithm or a phonetic transcription conversion algorithm.
5. The evaluation method based on document-level machine translation according to claim 1, characterized in that, The step of performing temporal tagging on the statements in the machine translation and the reference translation to obtain a first temporal tag sequence and a second temporal tag sequence includes: For each sentence in the machine translation, a tense tagger is used to tag each word in the sentence with a tense category label to obtain a first tense tag sequence; For each sentence in the reference translation, a tense labeler is used to label each word in the sentence with a tense category label, resulting in a second tense label sequence.
6. An evaluation device based on document-level machine translation, characterized in that, The device includes: The acquisition module is used to acquire the source document and its corresponding machine translation and reference translation; The entity consistency assessment module is used to generate coreference chains for entities in the source document, find the set of translated items corresponding to each coreference chain in the machine translation, and the translated items in the translated item set correspond to the mentions in the coreference chain; and generate an entity consistency score based on the comprehensive similarity between every two translated items in the translated item set, the entity consistency score representing the consistency of nouns of the same entity in the machine translation. The tense consistency evaluation module is used to perform tense annotation on the statements in the machine translation and the reference translation respectively to obtain a first tense label sequence and a second tense label sequence; and to generate a tense consistency score based on the matching degree of the first tense label sequence and the second tense label sequence corresponding to the same statement in the machine translation and the reference translation, wherein the tense consistency score represents the consistency of the machine translation in terms of time expression. The discourse coherence assessment module is used to map the related words in the machine translation and the reference translation to the corresponding semantic relationship categories according to a preset mapping table; and to generate a discourse coherence score based on the matching degree of the semantic relationship categories corresponding to the related words at the same semantic position in the machine translation and the reference translation. The discourse coherence score represents the consistency of the machine translation in the discourse logical relationship. The comprehensive evaluation module is used to generate a comprehensive score for document-level machine translation based on the entity consistency score, the temporal consistency score, and the discourse coherence score. The tense consistency assessment module is further configured to: divide the statements in the machine translation into a first explicit set and a first implicit set, and divide the statements in the reference translation into a second explicit set and a second implicit set, wherein the first explicit set and the second explicit set are sets of statements containing time clues, and the first implicit set and the second implicit set are sets of statements not containing the time clues; for the same statement in the first explicit set and the second explicit set, determine whether the first tense label sequence and the second tense label sequence corresponding to the statement are consistent; calculate the proportion of statements with consistent tense label sequences to obtain the explicit matching degree; for the same statement in the first implicit set and the second implicit set, determine whether the first tense label sequence and the second tense label sequence corresponding to the statement are consistent; calculate the proportion of statements with consistent tense label sequences to obtain the implicit matching degree; and perform a weighted fusion of the explicit matching degree and the implicit matching degree to obtain the tense consistency score. The discourse coherence assessment module is further configured to: identify the conjunctions in each sentence of the machine translation and the reference translation; for the same semantic position, query the first semantic relationship category corresponding to the conjunction at the semantic position in the machine translation, query the second semantic relationship category corresponding to the conjunction at the semantic position in the reference translation, determine whether the first semantic relationship category and the second semantic relationship category are consistent; and calculate the proportion of conjunctions with consistent semantic relationship categories to obtain the discourse coherence score.
7. A computer-readable storage medium, characterized in that, The storage medium stores at least one instruction, which is loaded and executed by a processor to implement the document-level machine translation evaluation method as described in any one of claims 1 to 5.
8. An electronic device, characterized in that, The electronic device includes the evaluation apparatus based on document-level machine translation as described in claim 6.
Citation Information
Patent Citations
Translation quality evaluation method and device, electronic equipment and storage medium
CN120579558A
Method for computing similarity between text spans using factored word sequence kernels
US20090175545A1