Fuzzy test mutation scheduling optimization method for question-answering system based on firefly algorithm
Through the fuzzy testing mutation scheduling optimization method of the Q&A system based on the firefly algorithm, the problem of insufficient selection strategy of mutation operators in the QA system is solved, and more efficient testing resource utilization and error discovery are achieved.
Patent Information
- Application Number
- CN202510515382.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-23
- Publication Date
- 2025-07-18
AI Technical Summary
The lack of effective mutation operator selection strategies in the fuzz testing of existing QA systems leads to wasted testing resources and inefficiency, and traditional methods cannot adapt to the diversity and dynamic changes of QA systems.
The fuzzy test mutation scheduling optimization method based on the firefly algorithm is used to update the seed's preference for mutation operators through a custom firefly algorithm, and generate fireflies using the confusion feedback information, automatically identify the efficiency differences of the mutation operators and optimize resource allocation.
The efficiency and effectiveness of fuzz testing of QA system are improved, the efficiency differences of mutant operators are automatically identified, resource allocation is optimized, and the testing performance and error discovery rate are improved.
Smart Images

Figure CN120336188A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of question - answering system detection, and relates to an optimization method for fuzzy - testing mutation scheduling of a question - answering system based on a firefly algorithm. Background Art
[0002] In modern software engineering, with the rapid development of artificial intelligence and natural language processing technologies, question - answering (QA) systems, as a type of software based on natural language processing (NLP) technologies, aim to answer questions raised by users in a natural - language manner. It has been widely applied in various industries and has become an important tool for information acquisition and knowledge management. However, with the expansion of its application scope, the complexity and potential risks of QA systems are also increasing continuously. Although these QA systems have achieved remarkable success in performing tasks, as software programs, they also have potential software defects, which may lead to significant economic losses or life - safety problems. The existence of these problems highlights the urgency of quality control and bug - fixing for QA systems to ensure their reliability and usability in various application fields.
[0003] In recent years, scholars often detect errors in QA systems when answering user questions by constructing datasets and designing metamorphic relations. Specifically, the former mainly relies on static datasets for testing. By constructing standardized datasets, it directly evaluates the accuracy and effectiveness of the system in generating answers given questions and contexts. However, the method of static datasets has obvious limitations: First, the creation of datasets requires a large amount of time and human resources; second, the language phenomena and task forms covered by datasets are limited, making it difficult to fully reveal the defects of the system; finally, such testing methods cannot dynamically generate new test samples or quickly adapt to system updates and changes, restricting the improvement of testing efficiency. Therefore, this type of method is more used for preliminary verification of the basic functions of the system. The latter, metamorphic testing, has become an effective means to verify the functional integrity and logical consistency of QA systems because of its characteristic of not requiring a clear test Oracle. Since the answers of QA systems are often diverse and uncertain, it makes it difficult for traditional testing methods that simply rely on comparison with correct answers to effectively cover all scenarios. Metamorphic testing indirectly evaluates whether the system behavior meets expectations when the correctness of the output cannot be directly verified by defining the expected logical relationship (i.e., metamorphic relation) between the input and the output, thus effectively solving this problem. However, these testing methods based on metamorphic testing focus on the definition of metamorphic relations and the design of mutation methods. These contributions are beneficial to the generation of high-quality test cases, but they do not provide systematic testing guidance and selection strategies, ignoring the efficiency and effectiveness of the test generation process, which limits it to a certain extent when generating large-scale test cases. Therefore, more efficient generation methods are needed to enhance the testing efficiency.
[0004] Based on this, some researchers experiment with highly automated testing methods such as fuzz testing to further improve testing efficiency. The rise of fuzz testing methods has brought new solutions to the testing of QA software. As a classic automated software testing technology, fuzz testing is effectively applicable to the black-box characteristics of QA systems. Without concerning the internal structure and principles of the software, it can generate a large number of diverse test cases and discover existing problems by detecting incorrect answers of the QA system. However, the application of fuzz testing in QA systems is still in the initial development stage. This is because many fuzz testing optimization methods under traditional software are mostly not applicable to improving the fuzz testing framework of QA systems. Therefore, there is still great room for development in the current research on fuzz testing of QA systems.
[0005] Mutation scheduling is one of the most critical links in the fuzz testing process, which determines the priority of each mutation operator (MO) being called in the mutation link. However, existing testing tools for QA systems randomly select mutation operators during the testing process, lacking inventions regarding the selection of mutation operators.
[0006] Specifically, the necessity of mutation scheduling includes the following three points: First, due to the differences in the effects of mutation operators, randomly and fairly selecting mutation operators may waste unnecessary test resources on inefficient operators, thus reducing the overall fuzzing efficiency; Second, different QA systems have different preferences for mutation operators. There are obvious differences in aspects such as the technical principles, designs, and training data among different QA systems, which leads to different error rates when QA systems handle test cases generated by different mutation operators; Finally, due to the differences in text in terms of structure, grammar, semantics, and even punctuation features, different seeds also have differences in their preferences for mutation operators.
[0007] Therefore, based on the above motivations, the present invention proposes an optimized method for mutation scheduling in fuzz testing of question - answering systems, MuQA, based on the firefly algorithm, which fully solves the problems brought by the above motivations. This method defines the efficiency of mutation operators on seeds as preferences and represents the preferences of seeds for all mutation operators through a vector. Through a customized firefly algorithm, considering the brightness, position, and distance of fireflies generated by seeds in each iteration to affect the preference update of seeds for mutation operators. Among them, each firefly is determined by the mutation information in the seed iteration and is used to reflect the efficiency of all operators in the current iteration. Comparative experiments show that MuQA has better test effects compared to the baseline test methods, and relevant ablation experiments further fully verify the effectiveness of the customized firefly algorithm in mutation scheduling in this paper. This method provides a quantifiable basis for operator selection in mutation testing of deep - learning systems, breaking through the limitations of traditional fixed - weight scheduling strategies. Summary of the Invention
[0008] The object of the present invention is to solve the problem of waste of test resources caused by neglecting the mutation operator selection strategy in the traditional QA system test method, which in turn leads to low test efficiency. For this purpose, the present invention proposes an optimized method for mutation scheduling in fuzz testing of question - answering systems based on the firefly algorithm.
[0009] The present invention provides an optimized method for mutation scheduling in fuzz testing of question - answering systems based on the firefly algorithm, including:
[0010] Step 1, according to the current preference of the seed, change it into a probability distribution for guiding the seed to select a mutation operator during mutation through the method of sum normalization, and control the mutation of the seed at a specified energy according to this distribution;
[0011] Step 2, statistically collect the perplexity feedback information of the test cases generated after each mutation of the seed when input into the QA system, and generate fireflies representing the efficiency of each mutation operator in this iteration through customized firefly features;
[0012] Step 3: Add the new fireflies to the swarm corresponding to the seeds, comprehensively consider the weights of each firefly according to the optimized firefly algorithm, and update the preference of the seeds for each mutation operator.
[0013] In the first aspect, the specific steps of the above Step 1 are as follows:
[0014] Step 1.1: Normalize the preference vector representing the preference degree of the seeds for each mutation operator to obtain the probability distribution for guiding the specific mutation link, and select the mutation operator to perform mutation according to the distribution.
[0015] Step 1.2: Perform the mutation operation, input the generated test case into the QA system, obtain the corresponding answer and the score tensor of each token in the vocabulary in the answer, use the softmax function to process it to obtain the probability value of the token, and obtain the change in perplexity caused by this mutation according to the perplexity calculation formula, and reflect the influence of the mutation operator on the confidence degree of the QA system through this change to reflect the efficiency of the mutation operator.
[0016] In the second aspect, the specific steps of the above Step 2 are as follows:
[0017] Step 2.1: Statistically record the changes in perplexity caused by all mutations. First, compare each change value with the average value of the changes, regard the test case with a higher value as locally interesting, and then compare this change with a fixed preset value, and regard the test case with a higher value as globally interesting.
[0018] Step 2.2: Start generating fireflies for this iteration. First, calculate the proportion of locally interesting test cases for each mutation operator, save all proportion values in a vector with the same length as the preference vector, and use it as the position of the firefly to reflect the relative efficiency between mutation operators. Then calculate the proportion of all globally interesting test cases in this iteration as the brightness of the firefly to reflect the efficiency of this iteration relative to other iterations in history.
[0019] In the third aspect, the specific steps of the above Step 3 are as follows:
[0020] Step 3.1: Determine the attractiveness of each firefly in the swarm corresponding to the seeds. Regard the difference between the current iteration and the iteration number to which the firefly belongs as the distance to reflect the timeliness of the firefly's preference for update, and calculate the distance and the brightness of this firefly as the attractiveness of this firefly according to formula (1):
[0021]
[0022] Among them, γ is the attractiveness decay factor used to control the speed of the decay of the attractiveness between fireflies with distance, β0 is the brightness of the firefly, and r is the distance.
[0023] Step 3.2: Update the preference of the seeds according to the firefly algorithm customized in the present invention. The attraction I will be used as a weight to affect the amplitude of the movement of the preference vector towards the position of the firefly. After the movement is completed, the next firefly will be considered until all fireflies have been considered.
[0024] Compared with the prior art, the beneficial effects of the present invention are as follows:
[0025] 1. In the test of the QA system for the fuzzy test mutation scheduling optimization method based on the firefly algorithm proposed, the perplexity is used to reflect the uncertainty degree of the QA system about its own answer, so as to reflect the efficiency difference between different mutation operators. It can automatically identify the efficiency of mutation operators, allocate more test resources to high-quality mutation operators, and make the fuzzy test work of the QA system get rid of the traditional and simple random selection strategy. From a macroscopic level, the present invention not only distinguishes the efficiency difference between mutation operators, but also eliminates the prior work of operator efficiency that the tester needs to carry out according to different QA systems, and can automatically and dynamically identify the operator efficiency, filling the relevant research gap and solving the problem of waste of test resources.
[0026] 2. Due to the particularity of the seeds of the QA system in terms of its own structure, semantics, etc., there are obvious differences in the preferences of different seeds for mutation operators. The fuzzy test mutation scheduling optimization method for the QA system based on the firefly algorithm proposed in this method designs a unique mutation guidance distribution for each seed in a more microscopic level. The efficiency of the mutation operator on the seed is defined as the preference. Through the customized firefly algorithm, considering the brightness, position, and distance of the fireflies generated by the seeds in each iteration to affect the preference update of the seeds for the mutation operators. Each firefly and the mutation information in the seed iteration are determined to reflect the efficiency of all operators in the current iteration. The invention further explores the preference of each seed, makes more full use of test resources, and achieves more outstanding test performance. BRIEF DESCRIPTION OF THE DRAWINGS
[0027] Figure 1 It is the overall flowchart of the fuzzy test mutation scheduling optimization method for the QA system based on the firefly algorithm.
[0028] Figure 2 It is the detailed flowchart of the fuzzy test mutation scheduling optimization method MuQA for the QA system based on the firefly algorithm.
[0029] Figure 3 It is the information of the QA system used in the experimental part of the present invention.
[0030] Figure 4 It is the information of the QA system dataset used in the experimental part of the present invention.
[0031] Figure 5 These are the test results of MuQA proposed by the present invention and other QA system testing methods on each QA system and its corresponding dataset. The values in the cells represent the false discovery rate under the same test budget.
[0032] Figure 6 These are the test results of MuQA proposed by the present invention on each QA system and its corresponding dataset before and after disabling the mutation scheduling method when the loop feedback is enabled.
[0033] Figure 7 These are the test results of MuQA proposed by the present invention on each QA system and its corresponding dataset before and after disabling the mutation scheduling method when the loop feedback is disabled.
[0034] Figure 8 These are the dynamic changes of the false discovery rate and the number of errors with the consumption of test resources (energy) before and after disabling the mutation scheduling method when the loop feedback is enabled for MuQA proposed by the present invention.
[0035] Figure 9 These are the dynamic changes of the false discovery rate and the number of errors with the consumption of test resources (energy) before and after disabling the mutation scheduling method when the loop feedback is disabled for MuQA proposed by the present invention. Detailed implementation manners
[0036] The present invention will be further described below in conjunction with the accompanying drawings and implementation cases. It should be noted that the described implementation cases are only for facilitating the understanding of the present invention and do not impose any limitations on it.
[0037] The present invention aims at the testing of QA systems and proposes an optimization method for fuzzy testing mutation scheduling of question answering systems based on the firefly algorithm to more efficiently utilize limited test resources to select appropriate mutation operators for mutation. The invention provides a complete fuzzy testing framework for QA systems and conducts sufficient experiments to prove the feasibility and effectiveness of the method.
[0038] As Figure 1 shown, an optimization method for fuzzy testing mutation scheduling of a question answering system based on the firefly algorithm of the present invention includes:
[0039] Step 201 changes, according to the current preference of the seed, into a probability distribution for guiding the seed to select a mutation operator during mutation through the method of sum normalization, and controls the mutation of the seed at the specified energy according to this distribution.
[0040] The objective of the present invention is to determine the optimal probability distribution of each mutation operator at the current iteration depth for each seed. The present invention regards the efficiency of an operator as the preference of a seed for it, and represents the preference of a seed for all mutation operators by a vector with the same length as the number of operators. Each element will be initialized to 0.5 at the beginning and will move within the space of [0, 1]. The preference vector indirectly reflects the efficiency of the mutation operator, and the preference of the seed for different mutation operators determines the probability distribution used to guide the mutation.
[0041] Step 201: Establish a seed queue, and modify the original seed input through three designed mutation operators to generate new test inputs.
[0042] In step 2011, the preference vector of the seed is used to indirectly reflect the efficiency of the mutation operator. The preference of the seed for different mutation operators determines the probability distribution used to guide the mutation. By normalizing the preference vector through formula (2), the probability distribution X used to guide the seed to mutate can be obtained:
[0043]
[0044] where x i represents the probability of selecting mutation operator i, and p i represents the preference of the seed for mutation operator i. In the subsequent mutations of the seed, each time the energy is consumed, the operator that performs the mutation will be selected according to the probability distribution;
[0045] In step 2012, perform the selected mutation operation, input the generated test case into the QA system, obtain the corresponding answer and the score tensor of each token in the answer in the vocabulary, and use the softmax function to process it to obtain the probability value of the token. According to the perplexity calculation formula shown in formula (3), obtain the perplexity ppl caused by this mutation:
[0046]
[0047] where p(A) is the generation probability of the answer A of the QA system. This method reflects the influence of the mutation operator on the confidence level of the QA system by the difference from the original perplexity to reflect the efficiency of the mutation operator. Test cases with a higher difference will be regarded as new seeds and added to the seed set.
[0048] Step 202: Statistically collect the perplexity feedback information of the test cases generated after each mutation of the seeds. Generate fireflies representing the efficiency of each mutation operator in this iteration through the custom firefly features.
[0049] Step 2021 statistically records the perplexity changes caused by all mutations, and obtains the average value τ1 of all perplexity changes in this iteration. Compare each change value with it. Test cases with a value higher than τ1 will be regarded as locally interesting, and then compare this change with a fixed preset value τ2. Test cases with a higher value will be regarded as globally interesting. It should be noted that the value of τ2 is usually greater than τ1 to prevent the over-expansion of the seed set.
[0050] Step 2022 generates fireflies for the mutation of this iteration of this subset according to the information statistically recorded in the previous step. Fireflies are obtained by calculating the brightness and position of the fireflies.
[0051] The position of the firefly is similar to the preference vector of the seed and is represented by a vector with exactly the same structure. The position of the firefly determines the moving direction when the preference vector is updated, and is used to reflect the performance of each mutation operator in the iteration to which the firefly belongs. However, it should be noted that the position of firefly i only reflects the efficiency of each mutation operator in the i-th iteration to a certain extent, and is not equivalent to the preference vector of the seed, because the preference vector is jointly determined by multiple fireflies through multiple iterations. The position of the firefly is obtained through calculation after statistically recording the perplexity changes during each mutation. Each element of the position is relatively independent and only reflects the effect of one mutation operator. If the perplexity change caused by the mutation exceeds a certain threshold, then the generated test case will be regarded as interesting, and the position of the firefly on a certain operator is expressed as the ratio of the number of interesting test cases to the number of times the mutation operator is called. Let the number of interesting test cases obtained by the seed through mutation operator j in energy j be c i , then L i 's position l j in the j direction is:
[0052]
[0053] Brightness is used to reflect the overall effectiveness of the firefly in this iteration. The algorithm will represent this index by the total number of all interesting test cases and the total energy obtained by the seed. Compared with the position, brightness is more inclined to reflect the effect of mutation globally to reflect whether the firefly is worth considering. The brightness β0 of the firefly can be calculated by formula (5):
[0054]
[0055] Step 203 adds the new fireflies to the swarm corresponding to the seeds, comprehensively considers the weights of each firefly according to the optimized firefly algorithm, and updates the preference of the seeds for each mutation operator.
[0056] Step 2031 adds each generated firefly to the swarm corresponding to the seed. Each seed has an independent swarm, and the number of fireflies in it is equal to the number of iterations the seed has experienced. The attractiveness of each firefly is defined as the weight affecting the update of the seed preference vector, and the calculation of the attractiveness will be obtained according to the brightness and distance of the firefly. Among them, the distance refers to the difference between the current iteration and the number of iterations to which the firefly belongs, which is used to reflect the timeliness of the firefly's preference for update. The attractiveness of each firefly can be obtained through formula (1):
[0057]
[0058] Among them, γ is the attractiveness decay factor, which is used to control the speed of the attractiveness decay between fireflies with distance, β0 is the brightness of the firefly, and r is the distance;
[0059] Step 2032 will consider all the fireflies in the swarm at the same time, and move the preference vector to its position according to the attractiveness of each firefly to the current iteration. Specifically, at the end of the (t + 1)-th iteration, the preference vector P of the seed t+1 will be updated according to formula (6):
[0060]
[0061] Among them, n is the total number of fireflies in the swarm since the start of the test, and I i , L i respectively represent the attractiveness and position of firefly i.
[0062] The present invention is mainly used to detect errors when the QA system answers user questions. We select two common QA systems in the real world as the objects to be tested, and conduct experiments on three datasets, namely BoolQ, QuAD2, and NarrativeQA respectively. Figure 3 Shows the QA system information selected in the experimental session of the present invention, etc. Figure 4 Shows the test datasets we selected in the experiment.
[0063] To verify the effect of the MuQA, a fuzzy test mutation scheduling optimization method for the question-answering system based on the firefly algorithm of the invention, the present invention will use QATest and QAQA as the comparison baselines of MuQA, and compare the discovery rates of the error test inputs detected by them. To ensure the fairness of the experiment, we set the same upper limit of the energy budget for QATest. It should be noted that since QAQA is not based on the fuzzy test technology and has the defect that it cannot guarantee the non-repetition of test cases with this relatively large budget upper limit, the experiment will follow the experimental settings in QAQA and show its historical highest error rate. From Figure 5From the experimental results, it can be seen that MuQA has the highest number of errors and error rate on all tested programs and datasets. Specifically, QATest and MuQA, both based on fuzz testing, far exceed the highest error rate of QAQA; and MuQA has achieved a significant improvement on the two datasets of SQuAD2 and NarrativeQA.
[0064] Meanwhile, the present invention compares with itself through ablation experiments to examine whether preference update helps to correctly identify the preference of seeds for mutation operators and improve the generation quantity of error cases. We disable the mutation scheduling in MuQA and further examine the true contribution of the method to the test effect by comparing the differences in test results before and after, so as to eliminate the improvement of the test effect caused by factors such as mutation operators or fuzz testing frameworks. Figure 6 The relevant experimental results are shown, including the number of errors, the error rate, and the improvement information of the number of errors. Figure 8 It shows the dynamic change trends of the error rate and the number of errors with energy consumption. It can be observed from the figure that the test effect after disabling mutation scheduling has decreased significantly on all datasets. Generally speaking, enabling the mutation scheduling module is significantly helpful for error discovery, which further proves the key role of this module in testing the performance of QA systems.
[0065] In addition, to better explore the effect of the scheduling algorithm itself, we also conducted a similar experiment while turning off the fuzz testing loop feedback of MuQA. Specifically, after turning off the feedback, no new seeds will be added to the seed set, and the initial seed set will remain unchanged until the test results. Under this setting, we can more intuitively explore the role of the scheduling algorithm. Figure 7 The relevant experimental results are shown. Figure 9 It shows the dynamic change trends of the error rate and the number of errors with energy consumption. The results show that the mutation scheduling algorithm makes a more prominent contribution to the test results, and the growth of the error rate is more significant. This is because when the feedback is turned on, the vulnerability of many new seeds added to the seed set is greater, and they are more likely to trigger errors after any mutation operation, which causes the gap between the preferences of these seeds for mutation operators to be "smoothed out" to a certain extent. This reason makes the role of the scheduling algorithm not so obvious. However, obviously, the scheduling algorithm can still distinguish the efficiency differences between different operators when facing these new seeds and achieve an improvement in the test effect.
Claims
1. A mutation scheduling optimization method for fuzz testing of a question-and-answer system based on the firefly algorithm, characterized in that, It includes the following steps: Step 1: According to the current preference of the seed, it is changed into a probability distribution for guiding the seed to select a mutation operator during mutation through the sum normalization method, and the mutation of the seed at the specified energy is controlled according to this distribution; Step 2: Statistically collect the perplexity feedback information of the test cases generated after each mutation of the seed when input into the QA system, and generate fireflies representing the efficiency of each mutation operator in this iteration through the custom firefly characteristics; Step 3: Add the new fireflies to the swarm corresponding to the seed, comprehensively consider the weight of each firefly according to the optimized firefly algorithm, and update the preference of the seed for each mutation operator.
2. The method according to claim 1, characterized in that, The specific implementation of the above Step 1 includes the following steps: Step 1.1: Normalize the preference vector representing the preference degree of the seed for each mutation operator to obtain a probability distribution for guiding the specific mutation link, and select the mutation operator for performing mutation according to the distribution; Step 1.2: Perform the mutation operation, input the generated test case into the QA system, obtain the corresponding answer and the score tensor of each token in the vocabulary in the answer, process it with the softmax function to obtain the probability value of the token, and obtain the perplexity change caused by this mutation according to the perplexity calculation formula, and reflect the influence of the mutation operator on the confidence of the QA system through this change to reflect the efficiency of the mutation operator.
3. The method according to claim 1, wherein The specific implementation of the above Step 2 includes the following steps: Step 2.1: Statistically record the perplexity changes caused by all mutations. First, compare each change value with the average value of the changes, and regard the test case with a higher value as locally interesting. Then, compare this change with a fixed preset value, and regard the test case with a higher value as globally interesting; Step 2.2: Start generating fireflies for this iteration. First, calculate the proportion of locally interesting test cases for each mutation operator, and save all the proportion values in a vector with the same length as the preference vector, which is used as the position of the firefly to reflect the relative efficiency between each mutation operator. Then, calculate the proportion of all globally interesting test cases in this iteration, which is used as the brightness of the firefly to reflect the efficiency of this iteration relative to other iterations in history.
4. The method according to claim 1, characterized in that, The specific implementation of the above Step 3 includes the following steps: Update the preference of the seed according to the custom firefly algorithm of the present invention, consider each firefly generated by the seed so far one by one and move towards it, regard the difference between the current iteration and the iteration number to which the firefly belongs as the distance to reflect the timeliness of the firefly for updating the preference, and calculate the distance and the brightness of this firefly as the attraction of this firefly according to formula (1): Where γ is the attraction attenuation factor used to control the speed of the attraction decay between fireflies, β0 is the brightness of the firefly, and r is the distance. The attraction I will be used as a weight to affect the movement of the preference vector towards the position where the firefly is located. After the movement is completed, the next firefly will continue to be considered until all fireflies are considered.