Fuzzy test seed generation method based on diffusion generation model

Generating diversified fuzz test seeds through diffusion generative model solves the problem of easy crash in generative adversarial networks and insufficient diversity of seeds generated by variational autoencoders, and improves the efficiency and coverage of parallel fuzz tests.

CN120371709APending Publication Date: 2025-07-25XIAN UNIV OF TECH
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510504233.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-22
Publication Date
2025-07-25

AI Technical Summary

Technical Problem

Existing fuzz testing technologies are prone to crash in the generative adversarial network, and insufficient diversity in generative seeds, resulting in inefficient parallel fuzz testing and difficult to meet the need for efficient vulnerabilities.

Method used

Diffusion generation model is used to generate diverse fuzz test seeds, and data preprocessing is performed by collecting parallel fuzz test data sets, training diffusion generation models, and generating high-quality seed files for parallel fuzz testing.

Benefits of technology

The path coverage and crash triggering efficiency of parallel fuzz testing are improved, and the overall efficiency of fuzz testing is improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120371709A_ABST
    Figure CN120371709A_ABST
Patent Text Reader

Abstract

The invention discloses a fuzz test seed generation method based on a diffusion generation model, and aims to solve the problem that a traditional seed selection strategy is often low in efficiency and unstable in effect during fuzz test. According to the method, valuable input files are collected through parallel fuzzy testing to serve as a training data set, and after data preprocessing, the input files are input into a diffusion generation model to learn data features, and diversified seed files are generated. The seed files can quickly trigger the path of the target application program, so that the efficiency and coverage rate of vulnerability detection are remarkably improved. Experiments show that the seed file newly generated by the model improves the branch coverage rate and the path coverage rate of the fuzz test, and the efficiency of the parallel fuzz test is improved to a certain extent.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of fuzz testing, and specifically to a method for generating fuzz testing seeds based on a diffusion generative model. Background Art

[0002] As an automated software testing technology, fuzz testing has become a core means of discovering software vulnerabilities by injecting a large number of unexpected inputs into the target program to trigger abnormal behaviors. According to the differences in the strategies for generating test cases, existing fuzz testing technologies can be divided into the following two categories: Mutation-based fuzz testing, which is based on an initial seed file and generates new test cases by randomly flipping bytes, block replacement, etc. Representative tools such as AFL, LibFuzzer, etc. Its advantage lies in its simple implementation and the lack of prior knowledge of the input format of the target program, but it highly depends on the quality of the initial seeds and has limited ability to cover the deep logic of complex structured inputs. Generation-based fuzz testing constructs compliant test cases according to a predefined input format grammar, and typical tools such as Peach, Sulley. This method can generate syntactically valid inputs, but it requires manual writing of format description files and is difficult to adapt to rapidly changing protocols or test scenarios with unknown formats.

[0003] With the continuous growth of the scale and complexity of software systems, traditional single-threaded fuzz testing methods face efficiency bottlenecks and are difficult to meet the requirements of efficiently discovering vulnerabilities. Parallel fuzz testing can significantly improve the test efficiency and coverage by introducing multi-threading or distributed architectures to execute multiple test instances simultaneously. This method can not only generate and execute test cases faster but also make full use of the multi-core processing capabilities and distributed resources of modern computing platforms. In addition, parallel fuzz testing also involves key technologies such as load balancing, state synchronization, and result merging to ensure the efficiency and accuracy of testing. In recent years, the emergence of tools such as AFL++, PAFL, and FastAFL has further promoted the research and application of parallel fuzz testing.

[0004] Fuzz testing tools based on genetic algorithms do not consider the input format of the target application. The tool mutates the initial input seeds by methods such as byte flipping and crossover according to the genetic algorithm. Then, to discover crashes, the tool uses the initial input seed set and the mutated input files as inputs to the application. Since they have little prior knowledge requirements for the target application, mutation-based fuzz testing tools work efficiently. However, sometimes they get stuck due to simple crash detection strategies. To improve the efficiency of mutation-based fuzz tools, a good method is to use a better seed set.

[0005] In recent years, the combination of deep learning technology and fuzz testing has become a research hotspot, providing new ideas and tools for traditional fuzz testing. In particular, diffusion generative models, as an emerging alternative to generative adversarial networks (GANs) and variational autoencoders (VAEs), have attracted attention due to their excellent performance in generating high-quality and diverse data. The mode collapse of GANs: Due to their adversarial training characteristics, GANs are prone to mode collapse, that is, the diversity of generated samples drops sharply. The long-range dependence modeling defect of RNNs: Sequence generation models based on recurrent neural networks have difficulty maintaining long-distance syntactic constraints when processing nested structure inputs, resulting in poor semantic consistency of generated test cases. In fuzz testing, the diversity and quality of seeds directly affect test coverage and vulnerability discovery capabilities. Traditional methods for generating seeds have a relatively high randomness, which may lead to redundant test cases and reduced efficiency. By introducing diffusion generative models, their powerful generation capabilities can be utilized to generate diverse seeds based on specific characteristics of the target program, improving the coverage and diversity of the initial input set.

[0006] In parallel fuzz testing, the application of diffusion generative models has more advantages. Diverse initial seeds provide a rich input space for parallel task allocation, helping to avoid duplicate testing between different tasks and improving overall efficiency. This combination of deep learning and fuzz testing represents an important direction for the intelligent development of testing tools. Summary of the Invention

[0007] The object of the present invention is to provide a fuzz testing seed generation method based on a diffusion generative model. Currently, mainstream methods rely on using generative models such as generative adversarial networks or variational autoencoders to generate new seeds for fuzz testing. However, due to their adversarial training characteristics, generative adversarial networks are prone to mode collapse, affecting the ability of the generated seeds to cover paths. Due to the constraint conditions of latent space modeling in variational autoencoders, the generated seeds lack diversity. The diffusion generative model used in the present invention, as an emerging alternative to generative adversarial networks and variational autoencoders, can quickly generate high-quality and diverse fuzz testing seeds, improving the efficiency of parallel fuzz testing. The innovation of the present invention lies in (1) using a diffusion generative model to generate diverse seed files. (2) The generated new seeds can improve the efficiency of parallel fuzz testing. The advantage of the present invention is that it does not require knowledge of the detailed structure of the program under test. Only need to construct a training data set and input it into the model for training. After training, the diverse files generated by the model are used as the initial seed set for parallel fuzz testing. Experimental results show that both the path coverage rate and the crash trigger efficiency of the final fuzz testing results of the present invention have been improved.

[0008] To achieve the above object, the present invention provides the following technical solution: A fuzz testing seed generation method based on a diffusion generative model, comprising the following steps:

[0009] Step 1, collect the training set: First, perform parallel fuzz testing on the target program, and use all files that trigger new paths and unique crashes as the training data set;

[0010] Step 2, data preprocessing: Preprocess the training data set collected in Step 1 to generate data samples that conform to model training;

[0011] Step 3, build a diffusion generation model: The model includes a training process and a sampling process. In the training process, noise is added to the data samples step by step, and in the sampling process, the denoising is performed in reverse on the noisy data and iterated forward;

[0012] Step 4, generate an initial seed: Continuously train through the model in Step 3. After the training is completed, a weight file will be generated. Use the weight file to generate a new binary file as the initial seed for fuzz testing;

[0013] Step 5, perform fuzz testing using the newly generated seed: Use the binary file generated in Step 4 as the initial seed to start fuzz testing.

[0014] Further, the parallel fuzz testing of the target program in Step 1 is as follows: Use the existing seeds as the input for fuzz testing, perform parallel fuzz testing on the target program for a predetermined time. The existing seeds are conventional input files (such as.txt files, binary files, etc.).

[0015] Further, the specific steps of the data preprocessing in Step 2 are as follows:

[0016] Step 2.1, read files in binary: Read all files in the training data set in binary form to obtain binary strings;

[0017] Step 2.2, Base64 encoding: Use Base64 to encode the binary strings and convert the encoding results into values within the digital range (0 - 64) to form the transformed samples;

[0018] Step 2.3, normalize the samples: Perform normalization operations on the transformed samples to make the sample data at the same order of magnitude;

[0019] Step 2.4, convert to a matrix: Store the normalized data in a matrix and input it into the model for training.

[0020] Further, the training process in Step 3 is as follows: For the data samples preprocessed in Step 2, starting from the real data distribution, gradually add noise to each data sample; add a little noise at each step until the data becomes a pure noise distribution;

[0021] The sampling process is as follows: Train a neural network to simulate the reverse denoising process. Through multiple-step iteration, gradually recover from the noise sample to the noise-free sample.

[0022] Furthermore, through continuous training with the model in step 4, a weight file will be generated after the training is completed. Specifically: During the entire training process, the model continuously learns the denoising ability, and after the training is completed, the learned ability is saved to the final weight file for generating new data in the subsequent steps.

[0023] Furthermore, in step 4, use the weight file to generate a new binary file as the initial seed for fuzz testing. Specifically: Randomly sample a pure noise data, use the denoising ability learned from the weight file to denoise the randomly sampled data to generate new matrix data, and then convert the matrix data into a binary file as the initial seed for fuzz testing.

[0024] Furthermore, in step 2.3, the normalization formula is:

[0025]

[0026] In the formula, X represents the normalized data, x is the value of a single data, min is the minimum value of the column where the data is located, and max is the maximum value of the column where the data is located.

[0027] Furthermore, the neural network model formula is:

[0028]

[0029] where α t is a coefficient related to the time step, ∈ θ (x t ,t) is the noise prediction network obtained through training, x t is the data at the t-th step in this process, x t-1 is the data of the previous step, that is, after denoising the data at the t-th step.

[0030] Compared with the prior art, the beneficial effects of the present invention are as follows:

[0031] A fuzz testing seed generation method based on a diffusion generation model according to the present invention, through continuous training of the diffusion generation model, learns the characteristics of files that can trigger new paths and crashes, and finally generates diverse seed files for parallel fuzz testing. Experimental results show that the seed files generated by the present invention can effectively improve the path coverage rate of parallel fuzz testing and the probability of triggering unique crashes. The model of the present invention improves the efficiency of parallel fuzz testing to a certain extent. Description of the Drawings

[0032] Figure 1 It is a flowchart of a method for generating fuzz testing seeds based on a diffusion generative model according to the present invention;

[0033] Figure 2 It is a schematic diagram of a diffusion generative model;

[0034] Figure 3 It is an overall architecture diagram of a method for generating fuzz testing seeds based on a diffusion generative model according to the present invention;

[0035] Figure 4 It is a comparison diagram of the effects of path coverage and branch coverage of fuzz testing an open-source program libxml and normal fuzz testing using the method of the present invention. Specific embodiments

[0036] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all of the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.

[0037] Embodiment

[0038] A method for generating fuzz testing seeds based on a diffusion generative model according to the present invention, as Figure 1 shown, is specifically implemented according to the following steps:

[0039] Step 1, collect a training set; specifically as follows:

[0040] Step 1.1, use the existing seeds as the input for fuzz testing, and perform parallel fuzz testing on the target program for a predetermined time, and the predetermined time is 24h.

[0041] Step 1.2, after the fuzz testing ends, collect all the files that trigger new paths and unique crashes to form a training data set.

[0042] Step 2, data preprocessing: perform data preprocessing on the training data set collected in Step 1 to generate data samples that meet model training; specifically as follows:

[0043] The fuzz testing program only focuses on whether the input data can trigger abnormal behaviors or vulnerabilities of the target program, rather than the format or content of the files. Therefore, the files in the training data set include various formats and even damaged files. If the training set is directly used, the model will not be able to run.

[0044] Step 2.1, read files in binary: read all the files in the training set in binary form to obtain binary strings.

[0045] Step 2.2, Base64 Encoding: Use Base64 to encode the binary string and convert it into numbers (0 - 64).

[0046] Step 2.3, Normalize the Samples: To eliminate the influence of dimensions between different evaluation metrics, perform a normalization operation on the converted samples so that the sample data is on the same order of magnitude. The normalization formula is as follows:

[0047]

[0048] In the formula, X represents the normalized data, x is the value of a single data point, min is the minimum value of the data in the column, and max is the maximum value of the data in the column.

[0049] Step 2.4, Conversion Matrix: The model is expected to receive structured data for training in order to effectively learn data features. Therefore, store the normalized data in a matrix, and the model receives the matrix data for training;

[0050] Step 3, Build a Diffusion Generation Model: The model includes a training process and a sampling process. The training process adds noise step by step to the initial data, and the sampling process iteratively denoises the noisy data backward. The implementation details of the training process and the sampling process, as well as the specific loss function in the training process, are as follows:

[0051] Step 3.1, Training Process: The core of the training process is the diffusion process, which gradually adds noise to the data until the data becomes pure noise. Assume the input data is x0. The diffusion process gradually adds noise through multiple steps (T steps) to generate a series of noisy data x1, x2, …, x T , until the data at the T-th step becomes close to complete noise. Each step of adding noise is achieved by adding Gaussian noise, and the formula is usually expressed as:

[0052]

[0053] where α t is a coefficient related to the time step, and ∈ t is the standard normal distribution noise.

[0054] The goal of the model is to learn how to reverse this process, that is, to recover the original sample from the noisy data. To achieve this goal, the model is trained by predicting the noise in each step of the diffusion process. The loss function is usually the mean squared error of denoising in each step, that is, the difference between the predicted noise ∈ t and the real noise at each time step. The form of the loss function:

[0055]

[0056] where, ∈θ (x t , t) is a noise prediction network obtained through training, ∈ t is the true noise, and the loss function minimizes the error between the model-predicted noise and the true noise.

[0057] In the training stage, what is learned is how to reverse the diffusion process, that is, given the noise data x t , predict the original data x0. The goal of the reverse process is to recover from the noise sample x T to the sample x0 in the data distribution. This step is crucial for the model to learn, and the denoising network needs to be optimized during the training process.

[0058] Step 3.2, Sampling process: In the sampling stage, the trained model will be used to generate new samples. The core of sampling is the reverse process, that is, gradually recovering the sample from pure noise. In the sampling stage, it starts with a completely random noise sample x T (usually standard Gaussian noise), and then gradually denoises through the reverse process to obtain a realistic sample x0. The reverse process can gradually denoise through multiple iterations, and the formula is usually expressed as:

[0059]

[0060] where ∈ θ (x t , t) is a noise prediction network obtained through training.

[0061] The sampling process gradually recovers the data at each time step and finally obtains the generated sample at step 0. At each step, the model predicts the noise ∈ t , and then denoises according to the formula. The prediction and correction at each step depend on the parameters learned by the model in the training stage. The generated sample has a certain degree of randomness because the noise sample x T is different each time. This randomness helps to generate diverse samples.

[0062] The training stage includes the diffusion process and the denoising process. The goal is to learn to recover the original data from noise by optimizing the denoising model. The sampling stage starts from pure noise and gradually denoises through the reverse process to generate new samples. The denoising at each step depends on the denoising network learned in the training stage. The advantages of the diffusion model include that the generated data quality is usually high, it can generate more delicate details, and it has a more stable training process. However, it requires multiple steps of reverse denoising, which results in a large time overhead for training and generation. At each step, the model prediction needs to be calculated to gradually recover the data from the noise.

[0063] Step 4, Generate the initial seed, specifically as follows:

[0064] Step 4.1, Denormalization: The data input into the model is in matrix form, so the data generated by the model is also in matrix form. It cannot be directly used as the seed for fuzz testing. The matrix needs to be converted into the form of a binary file. First, extract the data from the matrix and perform denormalization on the data.

[0065] Step 4.2, Base64 Decoding: Perform Base64 decoding on the denormalized data and restore it to a binary file to form a new seed for fuzz testing;

[0066] Step 5, Use the newly generated seed for fuzz testing: Use the binary file generated in Step 4 as the initial seed to start fuzz testing; specifically as follows:

[0067] Perform parallel fuzz testing: After training through the model constructed in Step 3, the generated matrix data is converted into a binary file through Step 4 and can be used as the seed for fuzz testing. Use the new seed for parallel fuzz testing, and evaluate the effectiveness of this method according to the branch coverage and the number of triggered paths after the test is completed.

[0068] As the number of training rounds increases, the model will gradually generate higher-quality samples, the loss function continuously decreases, the data quality improves and tends to the distribution of real data. Through continuous training of the diffusion generation model of the present invention, the learning can trigger the characteristics of new paths and crashing files, and finally generate diverse seed files for parallel fuzz testing. The experimental results show that the seed files generated by the present invention can effectively improve the path coverage rate and the probability of triggering crashes in parallel fuzz testing. The model of the present invention improves the efficiency of parallel fuzz testing to a certain extent.

[0069] Figure 3 It is the overall architecture diagram of a fuzz testing seed generation method based on a diffusion generation model of the present invention; it shows the overall process of this method, including using PAFL to collect valuable files as the training data set, and inputting them into the diffusion generation model for training after data preprocessing, and finally generating new data to form the seeds for fuzz testing. The schematic diagram of the diffusion generation model is as Figure 2 shown, showing the forward and reverse processes of the diffusion generation model. Finally Figure 4 It is the effect comparison diagram of the path coverage rate and branch coverage rate of fuzz testing and normal fuzz testing on the open-source program libxml using the method of the present invention, proving the effectiveness of this method.

[0070] Although the embodiments of the present invention have been shown and described, for those of ordinary skill in the art, it can be understood that various changes, modifications, substitutions, and variations can be made to these embodiments without departing from the principles and spirit of the present invention. The scope of the present invention is defined by the appended claims and their equivalents.

Claims

1. A fuzz testing seed generation method based on a diffusion generative model, characterized in that, It includes the following steps: Step 1, collect the training set: First, perform parallel fuzz testing on the target program, and use all files that trigger new paths and unique crashes as the training data set; Step 2, data preprocessing: Preprocess the training data set collected in Step 1 to generate data samples that meet model training requirements; Step 3, build a diffusion generation model: The model includes a training process and a sampling process. In the training process, noise is added to the data samples step by step, and in the sampling process, the noise-added data is iteratively denoised backward; Step 4, generate the initial seed: Continuously train the model in Step 3. After the training is completed, a weight file will be generated. Use the weight file to generate a new binary file as the initial seed for fuzz testing; Step 5, perform fuzz testing using the newly generated seed: Use the binary file generated in Step 4 as the initial seed to start fuzz testing.

2. The method for generating fuzzing test seeds based on a diffusion generative model according to claim 1, wherein: The parallel fuzz testing of the target program in Step 1 is as follows: Use the existing seed as the input for fuzz testing, and perform parallel fuzz testing on the target program for a predetermined time. The existing seed is a conventional input file.

3. A method for generating fuzzing test seeds based on a diffusion generative model according to claim 1, characterized in that, The specific steps of the data preprocessing in Step 2 are as follows: Step 2.1, read files in binary: Read all files in the training data set in binary form to obtain binary strings; Step 2.2, Base64 encoding: Use Base64 to encode the binary strings and convert the encoding results into values within the digital range (0-64) to form the transformed samples; Step 2.3, normalize the samples: Perform a normalization operation on the transformed samples to make the sample data at the same order of magnitude; Step 2.4, conversion matrix: Store the normalized data in a matrix and input it into the model for training.

4. A fuzz testing seed generation method based on a diffusion generative model according to claim 1, characterized in that, The training process in Step 3 is: For the data samples preprocessed in Step 2, starting from the real data distribution, gradually add noise to each data sample; Add a little noise at each step until the data becomes a pure noise distribution; The sampling process is: Train a neural network to simulate the reverse denoising process, and through multiple steps of iteration, gradually restore from the noise samples to the noise-free samples.

5. A fuzzing test seed generation method based on a diffusion generative model according to claim 1, characterized in that, The continuous training of the model in Step 4 and the generation of a weight file after the training is completed are specifically as follows: During the entire training process, the model continuously learns the denoising ability, and after the training is completed, the learned ability is saved to the final weight file for generating new data in the subsequent steps.

6. A method for generating fuzzing test seeds based on a diffusion generative model according to claim 1, characterized in that, The use of the weight file to generate a new binary file as the initial seed for fuzz testing in Step 4 is specifically as follows: Randomly sample a pure noise data, use the denoising ability learned from the weight file to denoise the randomly sampled data to generate new matrix data, and then convert the matrix data into a binary file as the initial seed for fuzz testing.

7. A fuzz testing seed generation method based on a diffusion generative model according to claim 3, characterized in that, In Step 2.3, the normalization formula is: In the formula, X represents the normalized data, x is the value of a single data, min is the minimum value of the data in the column, and max is the maximum value of the data in the column.

8. A fuzz testing seed generation method based on a diffusion generative model according to claim 4, characterized in that, The neural network model formula is: where α t is a coefficient related to the time step, ∈ θ (x t , t) is the noise prediction network obtained through training, x t is the data at the t-th step in this process, x t-1 is the data of the previous step, that is, after denoising the data at the t-th step.