Translation omission detection method based on directional fuzzy test
By adopting a directed fuzz testing method in a machine translation system, using mutation operators to generate test inputs and using word alignment tools to detect translation omissions, the problem of lack of targetedness in detecting translation omissions in existing technologies is solved, and detection efficiency and user experience are improved.
Patent Information
- Application Number
- CN202510784275.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-12
- Publication Date
- 2025-09-12
AI Technical Summary
Existing machine translation systems lack specificity in detecting translation omissions, which affects translation quality and user experience.
A translation omission detection method based on directed fuzz testing is adopted. Three mutation operators (character-level mutation, word-level mutation, and punctuation mutation) are designed to modify the original seed input, generate new test input, and use word alignment tools to detect translation omissions.
It improves the pertinence and efficiency of translation omission detection, enables deeper exploration of the behavioral patterns of machine translation systems, reduces the workload of manual verification, and enhances user trust.
Smart Images

Figure CN120633679A_ABST
Abstract
Description
Technical Field
[0001] The invention belongs to the field of machine translation testing and relates to a translation omission detection method based on directional fuzzy testing. Background Art
[0002] In recent years, with the rapid development of deep learning and natural language processing technologies, machine translation systems have made great progress. Research shows that today's machine translation systems can already reach a level comparable to human translation. However, the robustness and security of such systems still have certain shortcomings. In actual applications, erroneous translations occasionally occur, which may lead to economic losses or even safety risks. Among various translation errors, translation omissions are considered to be a particularly serious problem. Compared with minor flaws that may exist in the translation, users are generally more intolerant of missing content, and such errors can significantly undermine their trust in machine translation systems. Therefore, translation omissions not only affect translation quality, but may also hinder the widespread application of machine translation systems. However, the current detection methods for translation omissions are still very limited, and existing machine translation testing methods lack specific targeting for this problem, which makes it challenging to further improve the reliability of the system and user experience.
[0003] Currently, the main approach to studying the robustness of machine translation systems relies on metamorphic testing. A variety of testing methods based on metamorphic testing have been developed in academia and have successfully detected numerous translation errors. The basic principle of these methods is to modify the original input using mutation operators to generate new test inputs. The expected output for these new inputs is then derived based on the metamorphic relationship. If the actual output of a test input is inconsistent with the expected output, the system is deemed to have an error. However, these methods typically employ a black-box testing strategy, meaning that the test generation process is completely independent of the machine translation system's internal execution feedback. While this approach ensures a certain degree of test versatility, it may limit the ability to deeply explore the behavioral patterns of the machine translation system. Furthermore, metamorphic testing methods often construct test oracles based on the output of the original input. However, if the translation result of the original input itself contains errors, the test oracles may fail, resulting in a decrease in the method's error detection capabilities.
[0004] In contrast, in the field of traditional software testing, directed grey-box fuzz testing is considered to be an extremely efficient method for detecting specific types of vulnerabilities. This method helps the test generate more accurate inputs by collecting test feedback that is highly relevant to the target vulnerability, thereby effectively exploring the internals of the software and gradually triggering specific vulnerabilities. In view of the great difficulties faced by machine translation systems in the automated detection of translation omissions, the present invention proposes a translation omission detection method based on directed fuzz testing, NMTFuzz (Fuzzing for Neural Machine Translation). The core concept of NMTFuzz is to introduce test guides that are highly relevant to translation omissions, screen and retain seed inputs as the basis for subsequent generation of test inputs. On this basis, the seed inputs are mutated through heuristic strategies to generate test inputs that are more likely to trigger translation omissions. This process is iterative, aiming to gradually and deeply explore the behavioral patterns of the machine translation system and accurately locate its translation omission problems. Summary of the Invention
[0005] Machine translation systems face significant difficulties in automatically detecting translation omissions. Existing techniques based on metamorphic testing require further manual verification and struggle to identify translation omissions when the test predictions themselves are incorrect. Furthermore, most existing methods generate a single round of test inputs, resulting in insufficient testing depth. This paper proposes a translation omission detection method based on targeted fuzzy testing to address these issues.
[0006] The present invention provides a translation omission detection method based on directed fuzzy testing, comprising:
[0007] Step 1: Create a seed queue and modify the original seed input to generate new test input using three designed mutation operators (character-level mutation, word-level mutation, and punctuation mutation);
[0008] Step 2: Calculate the ROUGE-1 value of the original seed input as the test input quality evaluation indicator, discard unqualified test inputs, execute qualified test inputs and collect test feedback;
[0009] Step 3: Use a word alignment tool to map the input text and the translated output text, count the unaligned non-stop words in the input text, and detect translation omissions;
[0010] Step 4: Calculate whether the generated test input has a greater probability of triggering a translation omission than the original seed input. If so, add the test input to the seed queue; otherwise, discard the test input.
[0011] First, the specific steps of step 1 above are as follows:
[0012] In step 1.1, we first sample from the corpus to construct a seed set for testing. Each original seed is used to initialize a dedicated seed queue for subsequent test input generation.
[0013] In step 1.2, the original seed is modified by applying different mutation operators (including character-level mutation, word-level mutation, and punctuation mutation) to generate new test inputs. During the mutation process, a maximum number of mutations must be set for each seed to prevent over-exploration of a single seed.
[0014] Secondly, the specific steps of step 2 above are as follows:
[0015] Step 2.1: Measure the deviation of the new test input from the original seed by calculating the ROUGE-1 value between the generated test input and the corresponding original seed input;
[0016] In step 2.2, the generated test inputs are screened based on their ROUGE-1 values. If a test input has a low ROUGE-1 value, this indicates that it has deviated significantly from the original seed and is likely of poor quality, so it is discarded. However, if the ROUGE-1 value of a test input is high enough, it is used as input to the machine translation system to generate translations and collect test feedback.
[0017] Thirdly, the specific steps of step 3 above are as follows:
[0018] In step 3.1, the generated test input and translation text are segmented using a word segmentation tool, breaking the sentences into independent words or phrases. The two word sequences after segmentation are then fed into a word alignment model, which is responsible for establishing a word-level mapping relationship between the input text and the translation result.
[0019] In step 3.2, the error detection module performs a matching analysis based on the output of the word alignment model, counting the number of unaligned non-stop words in the input text. If the number of unaligned non-stop words exceeds a preset threshold, it indicates a significant omission in the translated text, and the system reports a translation omission error.
[0020] Fourthly, the specific steps of the above step 4 are as follows:
[0021] Step 4.1: Collect the probability distribution of each generation step when translating the original seed and the newly generated test input, and calculate its test bootstrap value by feedback of the selected text probability and EOS probability;
[0022] In step 4.2, the seed queue is updated according to the test bootstrap value. If the newly generated test input has a larger test bootstrap value than the existing test input, it indicates that it is more likely to trigger translation omission and is added to the seed queue. Otherwise, it is directly discarded.
[0023] Fifthly, the above-mentioned translation omission test guidance specifically includes:
[0024] Collect the probability distribution of each generation step of the translation text;
[0025] Calculate the probability difference between the EOS token selected in each generation step and the text token with the maximum probability when the translation is not completed, and take the maximum difference as the translation omission test guide, which reflects the possibility of triggering translation omission.
[0026] The specific calculation method of the test guide is shown in formula (1).
[0027]
[0028] Among them, represents the probability of selecting a text token in the i-th generation step of machine translation, It represents the probability of selecting EOS token in this generation step, and n represents the maximum generation step length of the unfinished translation.
[0029] Compared with the prior art, the present invention has the following beneficial effects:
[0030] 1. The proposed invention has significant advantages in detecting translation omissions. Its core lies in the iterative generation mechanism of the directed gray-box fuzz testing framework, which avoids the traditional method of randomly generating all test inputs at once. Instead, it gradually generates test inputs that are more likely to trigger errors through a progressively more complex mutation strategy. This "cumulative detection" method can more deeply explore potential vulnerabilities in the translation system, significantly improving the targetedness and efficiency of detection. In addition, the invention pays special attention to the impact of punctuation on translation results. By introducing a punctuation mutation operator, it can simulate the interference of punctuation mutation on translation, thereby discovering types of translation omissions that have not been fully studied before.
[0031] 2. The proposed invention can detect translation omissions independently of the output results of the original test input, overcoming the limitation of the technology based on metamorphic testing that relies on the original test input as a test oracle. When the machine translation system generates an erroneous output for the original test input, the error detection ability of the existing methods tends to drop significantly. The present invention circumvents this problem through a word alignment method. It uses automated word alignment technology to locate translation omissions, which can focus more on the integrity of the translation content itself without relying on the mapping relationship between the original test input and the generated test input. In addition, the solution directly reports translation omissions instead of reporting suspected illegal test input pairs like the metamorphic testing method, which reduces the workload of manual inspection. BRIEF DESCRIPTION OF THE DRAWINGS
[0032] Figure 1 This is a flow chart of a translation omission detection method based on directed fuzz testing.
[0033] Figure 2 This is an overall flow chart of a translation omission detection method based on directed fuzz testing.
[0034] Figure 3 This is the information of the machine translation system under test selected in the experimental link of the present invention.
[0035] Figure 4 This is the data set selected for the experimental test phase of this invention.
[0036] Figure 5 is the number of translation omissions detected by NMTFuzz and the metamorphic testing-based baseline method when generating the same number of test inputs.
[0037] Figure 6 This is a comparison of the diversity of NMTFuzz and a baseline method based on metamorphic testing in detecting translation omissions.
[0038] Figure 7 This is a case study of NMTFuzz and a baseline method based on metamorphic testing in detecting translation omissions.
[0039] Figure 8 It is an analysis of the number and diversity of translation omissions detected by NMTFuzz with and without test guidance.
[0040] Figure 9 is the number of translation omissions reported by NMTFuzz during the test generation process with and without test guidance. DETAILED DESCRIPTION
[0041] The present invention will be further described below with reference to the accompanying drawings and implementation examples. It should be noted that the implementation examples described are only intended to facilitate understanding of the present invention and do not have any limiting effect on the present invention.
[0042] This paper aims to propose a translation omission detection method based on a targeted fuzz testing framework for machine translation systems. The invention provides a comprehensive fuzz testing framework and conducts extensive experiments to demonstrate the feasibility and effectiveness of the method.
[0043] like Figure 1 As shown, a translation omission detection method based on directed fuzzy testing of the present invention includes:
[0044] Step 201: Create a seed queue and use the three designed mutation operators to modify the original seed input to generate new test input.
[0045] The mutation operators used in this invention include token-level and punctuation-level mutation operators. Token-level mutation operators are further divided into character mutation operators and word replacement mutation operators based on the BERT pre-trained model (Bidirectional Encoder Representations from Transformers).
[0046] In step 2011, the character mutation operator first selects a random word in the seed input, and then modifies a character at a random position in the word to simulate keyboard input errors that may occur in daily life, including: 1) replacing the selected character with a random character; 2) adding a random character at a random position in the word; 3) deleting the character at a random position in the word.
[0047] In step 212, a word replacement mutation based on the BERT pre-trained model replaces randomly selected words in the seed sentence with the symbol [MASK], and then uses the BERT model to predict the possible tokens for the position of [MASK]. To increase the diversity of the generated test inputs, the mutation operator randomly selects a token from the first k possible tokens provided by BERT to replace [MASK].
[0048] In step 2013, the punctuation mutation operator modifies the punctuation in the seed input, including: 1) replacing the randomly selected punctuation with other punctuation; 2) adding new punctuation before or after the randomly selected punctuation; 3) inserting a space before or after the punctuation; 4) deleting the punctuation.
[0049] Step 202 : Calculate the ROUGE-1 value of the mutated input and the original seed input as a test input quality assessment indicator, and discard unqualified test inputs.
[0050] Step 2021: For each generated variant input, calculate its ROUGE-1 value compared to the original seed input. ROUGE-1 measures the consistency between the two by counting the ratio of the number of tokens in the variant input that match the seed input to the total number of tokens in the seed input. The calculation method is as follows:
[0051]
[0052] Among them, s is the original input, s' is the mutated input, Count match (token n ) counts the number of identical tokens in s and s', Count(token n )Count the number of all tokens in s.
[0053] In step 2022, the generated test inputs are screened based on the calculated ROUGE-1 value. Low ROUGE-1 values indicate that the mutated input deviates significantly from the seed input and is of low quality, so such inputs are discarded. Only mutated inputs with high ROUGE-1 values are retained as candidate inputs for further testing.
[0054] Step 203 : Translate the test input and map the input text and the translated output text using a word alignment tool, count the unaligned non-stop words in the input text, and detect translation omissions.
[0055] In step 2031, the present invention uses AWESOME, the most effective tool in the field of word alignment, to identify translation omissions. The main function of AWESOME is to generate the alignment relationship of tokens between two languages through semantic analysis. It takes the source language sentence and the translation sentence generated by machine translation as input, and outputs multiple ij pairs, each ij pair indicating that the i-th token in the source language sentence corresponds to the j-th token in the translation sentence. Through this tool, the present invention can directly locate which source language tokens have not found corresponding content in the translation, and judge whether translation omissions have occurred accordingly. In actual operation, in order to avoid excessive interference with unimportant content in the translation result, this study excludes stop words (such as "the", "is", etc., which are functional words with lower importance in semantic expression) during analysis. Non-stop words (such as nouns, verbs, adjectives, etc., which have actual semantic meaning) are the core of the sentence content and are usually not omitted. In the translation omission detection of the present invention, only the non-stop words that are not aligned are counted to capture the actual translation omission phenomenon more accurately and intuitively.
[0056] In step 2032, the present invention reduces false positives for translation omissions by setting a threshold σ. In practical applications, AWESOME may have certain alignment errors, such as missing alignment of certain non-stop words or aligning to the wrong token. Such missed alignments may be caused by sentence complexity or ambiguity in the language expression. The present invention sets a threshold σ based on actual test performance of translation omissions. A translation omission is reported only when the number of misaligned non-stop words exceeds this threshold.
[0057] Step 204 : Calculate whether the generated test input has a greater probability of triggering a translation omission than the original seed input. If so, add the test input to the seed queue; otherwise, discard the test input.
[0058] In step 2041, the present invention proposes two key evaluation criteria for triggering translation omissions: first, the test input should have a higher probability of the system selecting EOS token when the translation is not completed; second, the system should have a lower probability of selecting text token in the same generation step. The test input needs to be more inclined to select EOS token rather than the text token that continues to be generated when the translation system has not yet completed the translation. In a certain generation step, if the probability of the system selecting EOS token is higher than the probability of selecting text token, it can be considered that the risk of translation omission has increased significantly. This comparison of probabilities provides a theoretical basis for the screening of test inputs. The present invention defines the translation omission test guide as follows based on the probability distribution during machine translation generation:
[0059]
[0060] in, represents the probability of selecting a text token in the i-th generation step of machine translation, It represents the probability of selecting EOS token in this generation step, and n represents the maximum generation step length of the unfinished translation.
[0061] In step 2402, the smaller the probability difference Δ defined in step 2401, the closer the test input is to the condition for triggering a translation omission. The present invention uses this formula to evaluate each generated test input. Compared to the original seed input, if the Δ value of the generated input is smaller, it indicates that it has a greater probability of triggering a translation omission. These inputs will be retained and used as new seeds for the next round of test generation. Through multiple rounds of iteration, the present invention can gradually optimize the generated input to ensure that the final test input can fully reveal the translation omission problem of the system. When the seed queue is empty or other termination conditions are met, the fuzzy test generation process terminates.
[0062] This paper mainly uses a directed fuzzy testing framework to detect translation omissions in machine translation. We use an open source machine translation system in the real world to test the effect. Figure 3 The information of the machine translation system selected in the experimental phase of the present invention is displayed. Figure 4 We present the test datasets we selected for the experiments and compare our developed NMTFuzz with the current state-of-the-art baseline methods.
[0063] In order to verify the effect of NMTFuzz, a translation omission detection method based on directional fuzz testing, we Figure 3 The three open source machine translation systems shown are H-NLP, mBART50 and M2M100. Figure 4 The data set was tested. At the same time, the present invention also verified the mainstream mutation testing methods SIT, CAT, and NMTFuzz using a single mutation operator strategy for testing machine translation systems under the condition of generating the same test input as the control. The experimental results are shown in Figure 2. Figure 5 and Figure 6 As shown. By observing Figure 5 and Figure 6 It is found that the NMTFuzz method using three mutation operators detects more translation omissions than other baseline methods. In addition, the NMTFuzz method using three mutation operators is also better than the baseline method in terms of the diversity of generated test inputs. This experimental result illustrates the effectiveness of the present invention in detecting translation omissions. In addition, the experiment also sampled some test cases such as Figure 7 As shown in the figure, it is used to illustrate the difference between NMTFuzz and the metamorphic testing method in detecting translation omissions, as well as the intuitiveness and effectiveness of NMTFuzz error detection. In addition to the comparison of the final detection effect, the present invention also verifies the effectiveness of the proposed test guide. The experimental results are shown in the figure. Figure 8 As shown. Figure 8 Experimental results demonstrate that the proposed test guides for guiding omission detection can effectively generate test guides that trigger omissions. Compared to random guides, NMTFuzz using test guides detected more omission errors and exhibited greater diversity, achieving better testing results.
Claims
1. A translation omission detection method based on directed fuzzy testing, characterized in that: The steps include: Step 1: Create a seed queue and generate new test inputs by modifying the original seed input using three designed mutation operators: character-level mutation, word-level mutation, and punctuation mutation. Step 2: Calculate the ROUGE-1 value of the mutated input and the original seed input as the test input quality evaluation indicator, discard unqualified test inputs, execute qualified test inputs and collect test feedback; Step 3: Use a word alignment tool to map the input text and the translated output text, count the unaligned non-stop words in the input text, and detect translation omissions; Step 4: Calculate whether the generated test input has a greater probability of triggering a translation omission than the original seed input. If so, add the test input to the seed queue; otherwise, discard the test input.
2. A method according to claim 1, characterized in that The specific implementation of step 1 includes the following steps: Step 1.1, construct a seed set by sampling the corpus and generate a seed queue for each original seed; In step 1.2, the original seed is modified by the mutation operator to obtain a new test input. In this process, the maximum number of mutations for each seed is set to prevent excessive exploration of the seed from causing a decrease in the ability to detect errors.
3. A method according to claim 1, characterized in that The specific implementation of step 2 includes the following steps: Step 2.1, calculate the ROUGE-1 value by comparing the generated test input and the original seed input. This value indicates the degree of deviation between the generated test input and the original seed. In step 2.2, a lower ROUGE-1 value indicates that the generated test input is of low quality and is discarded directly; Otherwise, the test input is used as the input of the machine translation model to obtain the translated text and test feedback.
4. A method according to claim 1, characterized in that The specific implementation of step 3 includes the following steps: Step 3.1: Use the word segmentation tool to segment the test input text and the generated translation text, and then input the two groups of words into the word alignment model for word alignment; In step 3.2, the error detection module matches the word mapping output by the word alignment model and counts the non-stop words that are not aligned in the original input text. When the number of non-stop words that are not aligned exceeds the threshold, a translation omission is reported.
5. A method according to claim 1, characterized in that The specific implementation of step 4 includes the following steps: Step 4.1: Collect the probability distribution of each generation step when translating the original seed and the newly generated test input, and calculate its test bootstrap value by feedback from the probability of the selected text and the probability of the special token EOS (End of Sentence) that ends the translation; Step 4.2, update the seed queue according to the original seed and the test boot value of the newly generated test input. If the newly generated test input has a larger test boot value than the previously generated test input, add the test input to the seed queue, otherwise discard the test input.
6. A method according to claim 5, characterized in that The translation omission guidance specifically includes: Collect the probability distribution of each generation step of the translation text; Calculate the probability difference between the EOS token selected at each generation step and the text token with the highest probability when the translation is incomplete. Take the maximum difference as the translation omission test guide, which reflects the probability of triggering translation omission. The specific calculation method of the test guide Δ is as follows; in, represents the probability of selecting a text token in the i-th generation step of machine translation, It represents the probability of selecting EOS token in this generation step, and n represents the maximum generation step length of the unfinished translation.