A method for seed selection and anti-disturbance mutation in neural network fuzz testing
Through the seed selection and mutation strategy of the DeepCUFuzz framework, combined with coverage and uncertainty evaluation, redundancy is eliminated and gradient perturbation and adversarial attacks are introduced, which solves the limitations of existing fuzz testing methods and achieves more efficient model testing and defect discovery.
Patent Information
- Application Number
- CN202510935748.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-08
- Publication Date
- 2025-09-26
- Estimated Expiration
- 2045-07-08
AI Technical Summary
Existing coverage-guided deep neural network fuzz testing methods have limitations in seed selection and mutation strategies. It is difficult to balance test efficiency, coverage potential, uncertainty and input diversity in high-reliability testing scenarios, which makes the model prone to misjudgment when faced with small perturbations.
Using the DeepCUFuzz framework, a seed selection model guided by coverage and uncertainty is combined with Euclidean distance to eliminate redundancy. A mutation strategy that coordinates gradient perturbation with adversarial attacks is designed to generate test inputs with balanced diversity and coverage, thereby improving the robustness of the model.
The model's testing efficiency and defect detection capabilities have been significantly improved. The test adequacy of various neuron coverage indicators has increased by 4.56% to 18.53%. In particular, the SNAC indicator on the LeNet-1 model has increased by 22.95%, enhancing the robustness of the model.
Smart Images

Figure CN120430349B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of artificial intelligence testing technology, and in particular relates to a neural network fuzzy test seed selection and anti-disturbance mutation method. Background Art
[0002] With the widespread application of deep neural networks (DNNs) in safety-critical fields such as autonomous driving, medical diagnosis, and intelligent security, their complex internal structures and black-box nature pose significant challenges to system reliability and safety. Traditional static evaluation methods based on manually labeled data can reflect the model's accuracy in common scenarios, but they struggle to systematically and comprehensively identify potential flaws in models with marginal or anomalous inputs. Even minor perturbations, such as noise, occlusion, or adversarial attacks, can cause deep learning models to misjudge, leading to serious safety incidents.
[0003] In recent years, Coverage-Guided Fuzzing (CGF) has been introduced as a dynamic testing paradigm for DNN reliability assessment. Its core idea is to define a "coverage" metric for the internal behavior of a neural network and use this as a feedback signal to dynamically generate or mutate test inputs to maximize the activation of uncovered neurons or states in the network, further identifying potential misclassifications.
[0004] Existing coverage-guided fuzz testing (CGF) methods still face significant bottlenecks, primarily due to the limitations of seed input selection and mutation strategies. Current seed selection methods are generally single or unbalanced. For example, DeepXplore and DeepTest use random selection, which, while highly efficient, is highly blind and difficult to effectively reach the model's decision boundary or low-probability failure region. Methods such as DLRegion select based on classification uncertainty or error probability. Although they can focus on low-confidence samples, these samples are often over-concentrated in the feature space, resulting in test data redundancy. Methods such as DeepSmartFuzzer attempt to reduce redundancy through clustering, but performing clustering calculations in high-dimensional space is complex and difficult to meet the real-time requirements of actual testing. In addition, most existing strategies often ignore the comprehensive consideration of test sample diversity and the degree of deviation from the training data distribution (i.e., uncertainty), limiting their applicability in high-reliability testing. In terms of mutation strategies, while traditional image transformation operations such as translation, rotation, and cropping can generate a variety of test inputs, they are disconnected from neuron coverage feedback, lack specificity, and have limited improvement effects. Adversarial attack techniques such as the fast gradient sign method (FGSM) and projected gradient descent (PGD) perform well in triggering misclassifications, but due to their over-reliance on local perturbation optimization, they often only activate a small number of neurons, making it difficult to achieve systematic coverage improvement. Current methods often focus on a single metric between coverage and misclassification rate, lacking a mechanism for collaborative optimization of the two. More critically, how to balance the collaborative optimization of test efficiency, coverage potential, uncertainty, and input diversity in complex test scenarios with massive candidate seeds, high-dimensional feature spaces, and diverse coverage criteria remains a research challenge that urgently needs to be overcome, seriously restricting the practical promotion of CGF technology in high-reliability application scenarios. Summary of the Invention
[0005] To address the limitations of existing coverage-guided deep neural network (DNN) fuzz testing methods in seed selection and mutation strategies, this paper presents a neural network fuzz testing seed selection and perturbation-resistant mutation method based on coverage and uncertainty-guided deep neural network fuzz testing (Deep CoverageUncertainty Fuzz, DeepCUFuzz). This paper utilizes the DeepCUFuzz framework and implements a seed selection model that prioritizes coverage and uncertainty (the degree to which the test input deviates from the training data in feature space). This model combines initial neuron coverage with Euclidean distance to eliminate redundancy, achieving a balanced optimization of diversity and coverage. A mutation strategy is designed that synergizes gradient perturbation with adversarial attacks, utilizing the activation gradients of target neurons to generate targeted perturbations to cover low-activation regions. The Fast Gradient Signed Method (FGSM) is also embedded in the adversarial attack to enhance the challenge of the test input, improve the model's robustness, and significantly improve test efficiency and defect detection capabilities.
[0006] To achieve the above objectives, the present invention provides the following technical solutions:
[0007] A neural network fuzzy test seed selection and anti-disturbance mutation method includes the following steps:
[0008] (1) Seed selection stage: For candidate test inputs, the coverage and uncertainty are jointly evaluated. By evaluating the degree of deviation of the test input from the training data in the feature space, and combining the Euclidean distance and the initial coverage score to eliminate redundancy, the initial fuzz test seeds that can activate new neurons and reveal the blind spots of the model are selected;
[0009] (2) Mutation phase: Based on a mutation strategy combining gradient perturbation and adversarial attack, test inputs that can enhance coverage or expose model misbehavior are generated from the selected initial fuzz test seeds;
[0010] (3) Overall process integration stage: The DeepCUFuzz fuzz testing framework builds a closed-loop testing process through collaborative seed selection and mutation, continuously improving the testing effectiveness and coverage capabilities of deep neural networks.
[0011] Preferably, the seed selection stage specifically includes the following steps:
[0012] (1-1) Uncertainty scoring stage: Through feature space modeling and category support estimation, the uncertainty of model prediction is quantified to identify test inputs with potential high fault trigger rates;
[0013] (1-2) Redundancy elimination stage: Calculate the similarity between input samples based on Euclidean distance, thereby eliminating redundant test samples and improving the diversity of test data;
[0014] (1-3) Seed screening stage: Combining the initial neuron coverage and uncertainty score, inputs with strong coverage and high prediction uncertainty are preferentially selected as initial fuzzy seeds.
[0015] Preferably, step (1-1) is specifically as follows:
[0016] (1-1-1) The penultimate layer of the deep neural network is used as the latent feature space to represent the semantic embedding of the input;
[0017] (1-1-2) For each test input, calculate the distance between it and the training sample in the feature space and select K nearest neighbor training samples;
[0018] (1-1-3) Based on the normalized exponential distance and constructing the support score for each category :
[0019]
[0020] in, is the representation vector of the test input in the feature space, is with The representation vector of the sample closest to the jth in the feature space; is the label of the jth training sample, is the indicator function (when the label of the training sample Output 1 if it is equal to category c, otherwise output 0). It is an adjustable parameter used to control the distance weight;
[0021] (1-1-4) Using the Contradiction Uncertainty Scoring Index , the degree of deviation between the predicted class support of the quantitative model and the optimal class support:
[0022]
[0023] in, is the deep neural network's response to the test input predictions, is the maximum support among all categories. This formula quantifies the deviation between the predicted category support and the optimal support: Much smaller than When , the score approaches 1, indicating that the model predictions significantly conflict with the categories supported by the training data. That is, there is a contradiction between the model decision logic and the potential distribution pattern of the training data. This may be due to a decision blind spot caused by insufficient learning or generalization bias, which is consistent with the cognitive uncertainty caused by the difference in the posterior distribution in the Bayesian framework, indicating that this test input is more likely to reveal the failure of the DNN.
[0024] Preferably, steps (1-2) are specifically as follows:
[0025] (1-2-1) For each test input, calculate the Euclidean distance between it and its K nearest neighbor test samples based on its representation in the feature space :
[0026]
[0027] in, is with The representation vector of the sample closest to the jth in the feature space;
[0028] (1-2-2) Using the above Euclidean distance as a redundancy measure between input samples, candidate samples with a small distance to other selected samples are removed in the subsequent seed selection to preserve the diversity of the test set.
[0029] Preferably, steps (1-3) are specifically as follows:
[0030] (1-3-1) Calculate the initial neuron coverage for each test input as a measure of the input’s ability to activate different regions of the model;
[0031] (1-3-2) Combine neuron coverage and uncertainty score, perform redundancy elimination based on the distance between samples, construct a comprehensive scoring function, and assign a single weight to each sample - uncertainty weight is assigned in descending order according to the uncertainty score, feature distance weight is assigned in descending order according to the Euclidean distance, and coverage weight is assigned in descending order according to the initial neuron coverage value; then calculate the comprehensive weight value according to the formula composite weight = (1-α) × uncertainty weight + α × distance weight + β × coverage weight, where α represents the distance weight coefficient and β represents the coverage weight coefficient; according to the weight size, from large to small, filter out the top N seeds as the candidate seed set (seed pool) for the initial fuzz test seed.
[0032] Preferably, the mutation stage specifically includes the following steps:
[0033] (2-1) Select the seed input currently used for mutation from the candidate seed set, then calculate its gradient information relative to the input for the preset target neuron; based on the gradient direction, construct a perturbation with limited amplitude and apply it to the original input image to generate a new test sample as the mutant of this round;
[0034] (2-2) If the mutant successfully induces abnormal behavior of the deep neural network model (such as prediction errors), it is judged as an adversarial example and included in the adversarial test set; if the mutant does not trigger incorrect behavior but improves the coverage of neurons, the example will be further selected as a seed and participate in the subsequent mutation process, thereby achieving continuous mining of the internal behavioral state of the model;
[0035] (2-3) To enhance diversity and improve exploration efficiency, while generating mutants through each round of gradient perturbation, an adversarial attack is performed on the same seed input to generate additional adversarial samples with boundary exploration capabilities. The Fast Gradient Sign Method (FGSM) is used as an adversarial attack method, and the random initialization method in the particle swarm algorithm is introduced to avoid falling into local optimality, thereby improving mutation diversity and global exploration capabilities.
[0036] As a preferred embodiment, the overall process stage specifically includes the following steps:
[0037] (3-1) Initial seed generation: DeepCUFuzz first analyzes the input test set, extracts the feature embedding of each test sample based on the intermediate layer (such as the penultimate layer) of the deep neural network model, constructs the support distribution based on the training data, and uses the contradiction uncertainty score to identify samples with deviations in prediction;
[0038] (3-2) Coverage evaluation and redundancy elimination: For candidate test inputs, calculate their initial neuron coverage to measure their activation ability on the model structure. At the same time, use the Euclidean distance metric to test the similarity between input samples and eliminate redundant samples to improve seed diversity.
[0039] (3-3) Seed screening: Based on the uncertainty score and the initial coverage score, a seed priority ranking is constructed to select inputs with high uncertainty, novel coverage areas, and different from other seeds as the initial fuzz test seeds;
[0040] (3-4) Mutant generation: For the screened seed input, a gradient-guided perturbation calculation is performed. That is, the neurons to be activated are selected and their gradients with respect to the input are obtained, and small perturbations are constructed to generate mutants. At the same time, the FGSM adversarial attack algorithm is introduced to generate adversarial samples that can significantly challenge the model's discrimination boundary.
[0041] (3-5) Result evaluation and feedback iteration: The mutant is input into the deep neural network model. If its predicted label changes (i.e., triggers an erroneous behavior), it is determined to be an adversarial example and collected into the adversarial set. If no error is triggered but the model coverage is improved, the mutant is added back to the seed pool as a candidate input for subsequent mutations.
[0042] (3-6) Iterative execution until the seed pool is exhausted: The fuzz testing process is carried out in a loop, continuously selecting inputs from the seed pool, generating mutants, evaluating the results, and updating the seed pool until all initial fuzz testing seeds are fully processed.
[0043] Preferably, the value of N is 30.
[0044] The beneficial effects of the present invention are:
[0045] (1) The present invention can quantify the potential coverage value of test samples based on the deviation of feature space, and combine it with the redundancy elimination strategy based on Euclidean distance to improve the diversity and effectiveness of seeds. In addition, the designed gradient-adversarial hybrid mutation mechanism organically combines the gradient guidance of target neuron activation with the boundary exploration capability of adversarial attack, which can achieve in-depth exploration of low coverage areas and high uncertainty areas of the model. Through the above mechanism, more representative adversarial test samples can be generated, the model vulnerability discovery capability can be improved, and the model robustness can be effectively enhanced after retraining;
[0046] (2) Compared with the existing technology, the test adequacy of the present invention on various neuron coverage indicators (such as neuron coverage NC, k-multi-segment neuron coverage KMNC, neuron boundary coverage NBC, strong neuron activation coverage SNAC and top k neuron coverage TKNC) is improved by an average of 4.56% to 18.53%, among which the SNAC indicator on the LeNet-1 model is improved by as much as 22.95%, significantly enhancing the ability to mine potential defects in deep neural networks. BRIEF DESCRIPTION OF THE DRAWINGS
[0047] Figure 1 It is a schematic flow diagram of the present invention;
[0048] Figure 2 This is a comparison chart of the coverage of various coverage-guided fuzz testing frameworks running 10 times under the LeNet-1 model; the histograms from left to right are the present invention (DeepCUFuzz), DLRegion, DLFuzz, DeepXplore, and random selection (Random);
[0049] Figure 3 This is a comparison chart of the coverage of various coverage-guided fuzz testing frameworks running 10 times under the LeNet-5 model; the histograms from left to right are the present invention (DeepCUFuzz), DLRegion, DLFuzz, DeepXplore, and Random;
[0050] Figure 4This is a coverage comparison chart of various coverage-guided fuzz testing frameworks running 10 times on the ResNet-20 model; the histograms from left to right are the present invention (DeepCUFuzz), DLRegion, DLFuzz, DeepXplore, and Random;
[0051] Figure 5 This is a coverage comparison chart of various coverage-guided fuzz testing frameworks running 10 times under the VGG-16 model; the bar charts from left to right are the present invention (DeepCUFuzz), DLRegion, DLFuzz, DeepXplore, and Random. DETAILED DESCRIPTION
[0052] In order to better understand the above technical solution, the above technical solution will be described in detail below with reference to the accompanying drawings and specific embodiments.
[0053] Reference Figure 1 A neural network fuzzy test seed selection and anti-disturbance mutation method includes the following steps:
[0054] (1) Seed selection stage (seed selection strategy): For candidate test inputs, the coverage and uncertainty are jointly evaluated. By evaluating the degree of deviation of the test input from the training data in the feature space, and combining the Euclidean distance and the initial coverage score to eliminate redundancy, the initial fuzz test seeds (initial seeds) that can both activate new neurons and reveal the blind spots of the model are selected;
[0055] (2) Mutation phase: A mutation strategy based on a combination of gradient perturbation and adversarial attack aims to generate test inputs from the selected fuzz test seeds that can enhance coverage or expose model misbehavior;
[0056] (3) Overall process integration stage: The DeepCUFuzz fuzz testing framework builds a closed-loop testing process through collaborative seed selection strategy and mutation strategy to continuously improve the testing effectiveness and coverage capabilities of deep neural networks.
[0057] Furthermore, the seed selection stage specifically includes the following steps:
[0058] (1-1) Uncertainty Scoring (Uncertainty Assessment) Phase: Through feature space modeling and category support estimation, the uncertainty of model prediction is quantified to identify test inputs with potential high fault trigger rates. The specific design is as follows:
[0059] (1-1-1) The penultimate layer of the deep neural network is used as the latent feature space to represent the semantic embedding of the input;
[0060] (1-1-2) For each test input, calculate the distance between it and the training sample in the feature space and select K nearest neighbor training samples;
[0061] (1-1-3) Based on the normalized exponential distance and the support score of each category, the following formula is used for calculation:
[0062]
[0063] in, is the representation vector of the test input in feature space (extracted from the penultimate layer of the DNN), is with The representation vector of the sample closest to the jth in the feature space; is the label of the jth training sample; is the indicator function (when the label of the training sample Output 1 if it is equal to category c, otherwise output 0); It is an adjustable parameter used to control the distance weight;
[0064] (1-1-4) Define the "contradiction uncertainty score" indicator to quantify the degree of deviation between the model's predicted category support and the optimal category support. Its calculation formula is as follows:
[0065]
[0066] in, is the DNN's response to the test input predictions, is the maximum support among all categories. This formula quantifies the deviation between the predicted category support and the optimal support: Much smaller than When , the score approaches 1, indicating that the model predictions significantly conflict with the categories supported by the training data. That is, there is a contradiction between the model decision logic and the potential distribution pattern of the training data. This may be due to a decision blind spot caused by insufficient learning or generalization bias, which is consistent with the cognitive uncertainty caused by the difference in the posterior distribution in the Bayesian framework, indicating that this test input is more likely to reveal the failure of the DNN.
[0067] (1-2) Redundancy elimination stage: The similarity between input samples is calculated based on the Euclidean distance, thereby eliminating redundant test samples and improving the diversity of test data. The specific design is as follows:
[0068] (1-2-1) For each test input, based on its representation in the feature space, the Euclidean distance between it and its K nearest neighbor test samples is calculated as shown in the following formula:
[0069]
[0070] in, is with The representation vector of the sample closest to the jth in the feature space;
[0071] (1-2-2) The above distance is used as a redundancy measure between input samples, and candidate samples with a small distance to other selected samples are removed in subsequent seed selection to preserve the diversity of the test set.
[0072] (1-3) Seed screening (coverage guidance) stage: Combining the initial neuron coverage and uncertainty score, we prioritize inputs with strong coverage and high prediction uncertainty as the initial fuzzy seeds. The specific design is as follows:
[0073] (1-3-1) Calculate the initial neuron coverage of each test input (neuron coverage criteria include neuron coverage (NC), k-multi-segment neuron coverage (KMNC), neuron boundary coverage (NBC), strong neuron activation coverage (SNAC) and top k neuron coverage (TKNC)) as an indicator of the input's ability to activate different regions of the model;
[0074] (1-3-2) Combining neuron coverage and uncertainty scores, while performing redundancy elimination based on the distance between samples, a comprehensive scoring function is constructed to assign individual weights to each sample: uncertainty weights are assigned in descending order of uncertainty scores, feature distance weights are assigned in descending order of Euclidean distance, and coverage weights are assigned in descending order of the initial neuron coverage value; then, the comprehensive weight value is calculated according to the formula composite weight = (1-α) × uncertainty weight + α × distance weight + β × coverage weight, where α represents the distance weight coefficient and β represents the coverage weight coefficient, specifically α is 0.2 and β is 0.5; based on the weights, the top N seeds are selected from large to small as the candidate seed set for the initial fuzzy seed;
[0075] (1-3-3) In the candidate seed set, redundancy elimination is further performed based on the distance between samples (the specific implementation is shown in step (1-2)), and the top N seeds are selected as the final seed set to achieve a balance between diversity and uncertainty.
[0076] The seed selection strategy of the present invention aims to eliminate redundancy from the test set of a given dataset based on the uncertainty score and the distance of the seed in the sample space, and select the initial seeds through coverage guidance.
[0077] The present invention first calculates the initial neuron coverage value and coverage ranking of each sample in the test set; secondly, the sample uncertainty score is calculated by quantifying the model prediction confidence. Specifically, the potential feature vector of the test sample in the penultimate layer of the deep neural network is extracted, and the class support is calculated based on the K nearest neighbor samples of the same category in the training set (calculated by Euclidean distance). The uncertainty score is defined as 1 minus the (ratio of the predicted label support to the highest class support). A score close to 1 indicates low model decision confidence. Subsequently, redundancy is calculated based on the feature space distribution between samples: for each sample, the sum of its Euclidean distance to the K nearest neighbor samples is calculated. The smaller the distance, the higher the sample redundancy. Finally, the three indicators (high uncertainty priority, high initial neuron coverage, and high feature space uniqueness) are integrated through a weighted composite ranking mechanism: first, a single weight is assigned to each sample - the uncertainty weight is assigned in descending order of the uncertainty score (the higher the score, the greater the weight), the feature distance weight is assigned in descending order of the Euclidean distance (the greater the distance, the greater the weight), and the coverage weight is assigned in descending order of the initial neuron coverage value; then, the composite weight is assigned according to the formula = The composite weight value is calculated by adding (1-0.2)×uncertainty weight + 0.2×distance weight + 0.5×coverage weight. Finally, the test set samples are sorted in descending order by the composite weight, and the first N samples are selected as the initial seeds (N=30 in this article) to ensure that the selected seeds have both decision boundary exploration capabilities and neuron coverage potential.
[0078] Furthermore, the mutation stage specifically includes the following steps:
[0079] (2-1) Based on the coverage-uncertainty dual-criteria seed selection strategy, the seed input currently used for mutation is selected from the candidate seed set; then, for the preset target neuron, its gradient information relative to the input is calculated; based on the gradient direction, a perturbation with limited amplitude is constructed and applied to the original input image (original image) to generate a new test sample as the mutant (adversarial image) of this round;
[0080] (2-2) If the mutant successfully induces abnormal behavior of the model (such as prediction errors), it is judged as an adversarial example and included in the adversarial test set; if the mutant does not trigger incorrect behavior but significantly improves the coverage of neurons, the example will be further selected as a seed and participate in the subsequent mutation process, thereby achieving continuous mining of the internal behavioral state of the deep neural network model;
[0081] (2-3) To enhance diversity and improve exploration efficiency, while generating mutants through each round of gradient perturbation, the present invention also performs an adversarial attack on the same seed input to generate additional adversarial samples with boundary exploration capabilities. Given the limitations of computing resources and time overhead, the Fast Gradient Sign Method (FGSM) is preferred as the adversarial attack method, and the random initialization method in the particle swarm algorithm is introduced to avoid falling into local optimality, thereby improving mutation diversity and global exploration capabilities.
[0082] This paper proposes a deep neural network fuzz testing method based on coverage and uncertainty guidance, named DeepCUFuzz. By constructing a priority-driven seed selection mechanism, a mutation strategy combining gradient perturbation and adversarial attacks, and a dynamic update process based on coverage feedback, it achieves efficient and automated testing of deep learning models.
[0083] During the specific testing process, the DeepCUFuzz execution process (overall process stage) includes the following steps:
[0084] (3-1) Initial seed generation:
[0085] DeepCUFuzz first analyzes the input test set, extracts the feature embedding of each test sample based on the model's intermediate layers (such as the penultimate layer), constructs a support distribution based on the training data, and uses the contradiction uncertainty score to identify samples with deviations in prediction.
[0086] (3-2) Coverage evaluation and redundancy elimination:
[0087] For candidate test inputs, calculate their initial neuron coverage to measure their ability to activate the model structure. Use the Euclidean distance metric to measure the similarity between test inputs and remove redundant samples to improve seed diversity.
[0088] (3-3) Seed screening:
[0089] Based on the combination of uncertainty score and initial coverage score, a seed priority ranking is constructed to select inputs with high uncertainty, novel coverage area and different from other seeds as the initial seeds for fuzz testing;
[0090] (3-4) Mutant generation:
[0091] For the screened seed input, a gradient-guided perturbation calculation is performed. This involves selecting neurons to be activated and obtaining their gradients with respect to the input, constructing small perturbations to generate mutants. At the same time, the FGSM adversarial attack algorithm is introduced to generate adversarial samples that can significantly challenge the model's discrimination boundaries.
[0092] (3-5) Result evaluation and feedback iteration:
[0093] The mutant is input into the model. If its predicted label changes (i.e., triggers an erroneous behavior), it is considered an adversarial example and collected into the adversarial set. If no error is triggered but the model coverage is improved, the mutant is added back to the seed pool as a candidate input for subsequent mutations.
[0094] (3-6) Iterate until the seed pool is exhausted:
[0095] The above fuzz testing process is performed in a loop, continuously selecting inputs from the seed pool, generating mutants, evaluating the results, and updating the seed pool until all initial seeds are fully processed.
[0096] Result Analysis
[0097] Comparison of seed selection strategies
[0098] To reduce the impact of randomness in the neuron selection process, we repeated each method 10 times and averaged the results. Table 1 shows the comparison (%) results of the seed selection strategy of our invention (coverage-uncertainty-based seed selection method) with two other seed selection strategies (model confidence-based seed selection method and random selection method) on five metrics (neuron coverage criteria): neuron coverage (NC), k-multisegment neuron coverage (KMNC), neuron boundary coverage (NBC), strong neuron activation coverage (SNAC), and top k-neuron coverage (TKNC).
[0099] Table 1
[0100]
[0101] As can be seen from Table 1, the seed selection strategy of the present invention achieved optimality in 90% of cases (18 out of 20). However, in the two small models, Lenet1 and Lenet5, due to the shallowness of the network, the selected seeds with high uncertainty may cause the neuron activation values to be concentrated in a few intervals, which makes the coverage effect of KMNC inferior to the seed selection method based on the maximum prediction probability of the model.
[0102] Compared with the existing coverage-guided fuzz testing framework, the coverage improvement effect of the present invention (DeepCUFuzz) on various DNNs and datasets under different coverage indicators is compared.
[0103] In order to reduce the influence of randomness in the neuron selection process, we repeated each method 10 times and took the average result. The coverage comparison results of DeepCUFuzz and various coverage-guided fuzz testing frameworks running 10 times are shown in Table 2 (using different methods, including DeepCUFuzz of the present invention, DLRegion, The coverage comparison results of each coverage-guided fuzz testing framework running 10 times on LeNet-1, LeNet-5, ResNet-20, and VGG-16 models are shown in the figure. Figures 2 to 5 shown.
[0104] Table 2
[0105]
[0106] from Figures 2 to 5 It can be seen intuitively that the present invention (DeepCUFuzz) achieves the best performance among all methods. Compared with the second most effective technique, the coverage of the present invention increases by 8.90% to 19.03% on NC, 13.05% to 15.09% on KMNC, 2.21% to 21.48% on NBC, 2.56% to 22.95% on SNAC, and 0.30% to 7.05% on TKNC.
[0107] Table 2 shows that DeepCUFuzz significantly improves test coverage through the coordinated optimization of seed selection and mutation strategies. Experimental data shows that on the MNIST dataset (LeNet-5 model), DeepCUFuzz achieves a NC of 67.44%, significantly exceeding DLRegion (61.24%) and DeepXplore (60.85%). This advantage stems from its dual-priority screening strategy: coverage-priority prioritizes seeds that cover new paths, while uncertainty-priority prioritizes the degree to which test samples deviate from the training distribution, with regions of high uncertainty more likely to expose model flaws. This combination effectively explores inactive neurons and decision boundary weaknesses. Furthermore, the mutation strategy generates targeted inputs through gradient perturbations and introduces adversarial attacks to strengthen boundary condition testing. For example, on the CIFAR-10 dataset (ResNet-20 model), DeepCUFuzz achieves significantly higher KMNC (13.69%) and SNAC (7.91%) than other methods, demonstrating its deep ability to mine distribution intervals and activation boundaries.
[0108] Results of DeepCUFuzz model robustness improvement on different validation sets
[0109] We generate adversarial examples using a mutation strategy that combines gradient perturbations with adversarial attacks under the coverage criterion. We then use these to retrain the DNN model and calculate the model's adversarial accuracy to evaluate its robustness. Table 3 shows the results of improved model robustness on different validation sets.
[0110] Table 3
[0111]
[0112] Results show that on the MNIST dataset, the 1207 adversarial examples generated by the LeNet-1 model improved the accuracy of post-adversarial training by 1.36%, while the 1759 adversarial examples generated by the LeNet-5 model resulted in a 0.37% improvement in robustness. Since the model's initial accuracy was high on the original test set, retraining with the generated adversarial examples further improved its adversarial accuracy. This demonstrates that adversarial examples generated by the mutation strategy, which combines gradient perturbation with adversarial attacks, can improve the model's robustness to a certain extent.
[0113] The present invention solves the problems of lack of directionality in seed selection, excessive redundant samples, and disconnection between mutation strategy and coverage target in existing coverage-guided fuzz testing. The present invention includes four stages: uncertainty assessment, redundancy elimination, hybrid mutation, and dynamic sorting. In the uncertainty assessment stage, the present invention proposes a cognitive uncertainty score based on the deviation between the model prediction confidence and the training data support, and quantitatively scores the candidate samples to measure their potential coverage value; then in the redundancy elimination stage, the similarity of samples in the feature space is measured by Euclidean distance, and highly clustered samples are eliminated in real time, thereby achieving a balance between coverage potential and diversity; when entering the hybrid mutation stage, the target neuron activation gradient perturbation is synergistically integrated with the FGSM adversarial attack, which not only uses gradient information to directionally activate low coverage areas, but also uses adversarial perturbations to deeply explore decision boundaries; finally, in the dynamic sorting stage, the latest coverage feedback and uncertainty score are combined to prioritize the seed queue to ensure the joint optimal scheduling of coverage and diversity. The present invention significantly improves the efficiency and defect detection capability of deep neural network fuzz testing through a seed selection model that jointly optimizes coverage and uncertainty and a gradient-adversarial hybrid mutation strategy. Extensive experimental results on MNIST and CIFAR10 demonstrate that compared to existing state-of-the-art methods, this method achieves an average improvement of 4.56% to 18.53% in coverage across five metrics: neuron coverage (NC), k-multisegment neuron coverage (KMNC), neuron boundary coverage (NBC), strong neuron activation coverage (SNAC), and top k-neuron coverage (TKNC). SNAC improves by 22.95% on the LeNet1 model, revealing more decision boundary violations. This technological breakthrough enables the generated test inputs to more comprehensively explore the model's decision boundary, effectively exposing violations that are difficult to detect with traditional methods, and providing stronger support for robustness verification of deep learning models.
[0114] The above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit the same. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or make equivalent replacements for some of the technical features therein. However, these modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the various embodiments of the present invention.
Claims
1. A neural network fuzzy test seed selection and anti-disturbance mutation method, characterized in that: The steps include: (1) Seed selection stage: Through the joint evaluation of coverage and uncertainty, the initial fuzz test seeds are screened by evaluating the degree of deviation of the test input from the training data in the feature space, and combining the Euclidean distance and the initial coverage score to eliminate redundancy; The process of screening the initial fuzz test seeds is as follows: first, the initial neuron coverage of each test input is calculated as an indicator of the input's ability to activate different regions of the model; then, the neuron coverage and uncertainty score are combined, and redundancy elimination is performed based on the distance between samples. A comprehensive scoring function is constructed, and a single weight is assigned to each sample. Subsequently, a composite weight is obtained, and the top N seeds are screened from large to small according to the weight size as the candidate seed set for the initial fuzz test seeds; (2) Mutation phase: Based on a mutation strategy combining gradient perturbation and adversarial attack, test inputs that can enhance coverage or expose model misbehavior are generated from the selected initial fuzz test seeds; (3) Overall process integration phase: Build a closed-loop testing process through collaborative seed selection and mutation to continuously improve the testing effectiveness and coverage of deep neural networks; The mutation stage specifically includes the following steps: (2-1) Select the seed input currently used for mutation from the candidate seed set, then calculate its gradient information relative to the input for the preset target neuron; based on the gradient direction, construct a perturbation with limited amplitude and apply it to the original input image to generate a new test sample as the mutant of this round; (2-2) If the mutant successfully induces abnormal behavior of the deep neural network model, it is judged as an adversarial sample and included in the adversarial test set; if the mutant does not trigger incorrect behavior but improves the coverage of neurons, the sample will be further selected as a seed to participate in the subsequent mutation process; (2-3) While generating mutants through each round of gradient perturbation, an adversarial attack is performed on the same seed input to generate additional adversarial samples with boundary exploration capabilities; the fast gradient sign method is used as the adversarial attack method, and the random initialization method in the particle swarm algorithm is introduced to avoid falling into the local optimum.
2. The neural network fuzzy test seed selection and anti-disturbance mutation method according to claim 1 is characterized in that: The seed selection stage specifically includes the following steps: (1-1) Uncertainty scoring stage: Through feature space modeling and category support estimation, the uncertainty of model prediction is quantified to identify test inputs with potential high fault trigger rates; (1-2) Redundancy elimination stage: Calculate the similarity between input samples based on Euclidean distance, thereby eliminating redundant test samples and improving the diversity of test data; (1-3) Seed screening stage: Combine the initial neuron coverage and uncertainty score to screen out the initial fuzz test seeds.
3. The neural network fuzzy test seed selection and anti-disturbance mutation method according to claim 2 is characterized in that: Step (1-1) is as follows: (1-1-1) The penultimate layer of the deep neural network is used as the latent feature space to represent the semantic embedding of the input; (1-1-2) For each test input, calculate the distance between it and the training sample in the feature space and select K nearest neighbor training samples; (1-1-3) Based on the normalized exponential distance and constructing the support score for each category: (1-1-4) The contradiction uncertainty score indicator is used to quantify the degree of deviation between the model's predicted category support and the optimal category support.
4. The neural network fuzzy test seed selection and anti-disturbance mutation method according to claim 2 is characterized in that: Steps (1-2) are as follows: (1-2-1) For each test input, calculate the Euclidean distance between it and its K nearest neighbor test samples based on its representation in the feature space; (1-2-2) The above Euclidean distance is used as a redundancy measure between input samples and is added to the composite weight calculation as a distance weight in the subsequent seed selection.
5. The neural network fuzzy test seed selection and anti-disturbance mutation method according to claim 1 is characterized in that: The overall process stage specifically includes the following steps: (3-1) Initial seed generation: First, the input test set is analyzed, and the feature embedding of each test sample is extracted based on the middle layer of the deep neural network model. The support distribution is constructed based on the training data, and the contradictory uncertainty score is used to identify samples with deviations in the prediction; (3-2) Coverage evaluation and redundancy elimination: For candidate test inputs, calculate their initial neuron coverage to measure their activation ability on the model structure. At the same time, use the Euclidean distance metric to test the similarity between input samples and eliminate redundant samples. (3-3) Seed screening: Based on the uncertainty score and the initial coverage score, a seed priority ranking is constructed to screen the initial fuzz test seeds; (3-4) Mutant generation: For the screened seed input, perform gradient-guided perturbation calculations; at the same time, introduce the FGSM adversarial attack algorithm to generate adversarial samples that can significantly challenge the model's discrimination boundary; (3-5) Result evaluation and feedback iteration: The mutant is input into the deep neural network model. If its predicted label changes, it is determined to be an adversarial example and collected into the adversarial set. If no error is triggered but the model coverage is improved, the mutant is added back to the seed pool as a candidate input for subsequent mutations. (3-6) Iterative execution until the seed pool is exhausted: The fuzz testing process is carried out in a loop, continuously selecting inputs from the seed pool, generating mutants, evaluating the results, and updating the seed pool until all initial fuzz testing seeds are fully processed.
6. The neural network fuzzy test seed selection and anti-disturbance mutation method according to claim 1 is characterized in that: The value of N is 30.
Citation Information
Patent Citations
Intelligent system test data generation method based on uncertainty
CN113762335A
Fuzzy test method adopting region-based neuron selection strategy and terminal
CN113986717A